1,351 to 1,375 of 2,378 Observability Jobs in London

Engineering Director

Location
Greater London, England, United Kingdom
design decisions across your area, engaging directly with engineers on complex technical challenges. Build engineering practices that strengthen testing, code review, CI/CD, observability and deployment confidence. Own the operational health of your services, using incidents and customer feedback to drive lasting improvements. Hire, coach and develop engineering managers ...

Lead UI Developer – Single Dealer Platform (TypeScript / RxJS / React)

Location
Greater London, England, United Kingdom
edge. Enforce automated testing (unit, integration, and UI E2E) and CI/CD with quality gates. Instrument analytics, logging, and front end observability (e.g., Web Vitals, error tracking). Ensure secure front end practices (XSS/CSRF protection, content security policy, secrets handling, auth flows with OAuth/OIDC/ ...

Sr Manager, Enterprise Tools & Platforms

Location
City Of London, England, United Kingdom
technology, operations, engineering, business, or a related field. Preferred experience Atlassian and ITIL certifications or equivalent practical expertise. Experience with identity and access integrations, observability, testing platforms, enterprise architecture repositories, configuration management, and software license-management processes. Experience leading platform consolidation, business-process migration, merger, acquisition, divestiture, separation, or large ...

Enterprise Architect Microsoft Dynamics 365 CRM & Power Platform

Hiring Organisation
Bearing Point
Location
London, UK
Employment Type
Full-time
driven, batch and real-time integration patterns. Guide the use of Azure Integration Services, APIs, middleware and enterprise integration platforms. Set expectations for resilience, observability, performance, failure handling and ownership. Ensure clear system-of-record, master-data and data-responsibility boundaries. Data, AI & Analytics ArchitectureDefine data architecture across Dataverse, enterprise ...

Software Engineer

Location
Greater London, England, United Kingdom
Python, C/C++, Go, Typescript, or Java. CI/CD and DevOps: Knowledge of building pipelines, automated testing and deployment strategies. Monitoring and observability: Familiarity with tools like Grafana, Prometheus or other observability platforms. Containerisation and orchestration: Docker, Kubernetes. We are open to a wide range of backgrounds ...

Staff AI Engineer, Payments Intelligence

Location
Greater London, England, United Kingdom
70+ languages, adapting to context and intent. Evaluation, safety, and guardrails — define how the unit measures agent quality and safety, with rigorous evaluation, observability, and guardrails so agents behave reliably in a compliance‐sensitive, multi‐market environment. Intelligence and data products — architect the client‐intelligence systems that turn payment data … systems at scale. Model serving and LLM inference, with the cloud infrastructure to run AI workloads in production (AWS or similar). Evaluation and observability for AI — hands‐on with tools like LangFuse, LangSmith, Braintrust, or MLflow — plus solid automated testing and CI/CD. Proven technical leadership — mentoring engineers ...

Lead AI Software Engineer

Location
Greater London, England, United Kingdom
aligned to business requirements. Working closely with Technical Leads, Architects, Product Owners, Platform Engineers, and delivery teams, you will drive implementation quality, testing, observability, operational readiness, and continuous improvement throughout the software development lifecycle. What you will do Lead build execution within a squad, platform capability, or engineering domain. Translate … generated and engineer‐written code to ensure correctness, maintainability, security, and alignment with specifications. Drive engineering excellence through automated testing, contract testing, regression testing, observability, and production verification. Support CI/CD processes, deployment readiness, operational handover, runbooks, and service ownership. Ensure AI‐generated outputs are explainable, traceable, secure ...

Senior Platform Engineer - AI Native SaaS Platform

Location
City Of London, England, United Kingdom
someone who can design, build and operate scalable systems, not just manage cloud infrastructure. You'll work across infrastructure, automation, CI/CD, observability and production reliability, while writing high-quality software in Go and helping improve the way engineering teams build and run services at scale. What will … making builds, tests and deployments faster and safer Build automation and internal tooling that removes manual toil for engineering teams Improve observability across metrics, logging, tracing and production monitoring Define and manage SLOs across latency, availability and error rates Lead incident response, triage complex production issues and ship long-term ...

Senior Forward Deployment Engineer

Hiring Organisation
Luxoft
Location
London, UK
Employment Type
Full-time
analysis, upgrading Java and NPM runtimes, modernizing Spring and legacy middleware applications, improving CI/CD pipelines, containerizing applications, automating deployments, and introducing standard observability and resilience patterns. The Expert FDE is expected to lead complex engagements, work directly with development and client stakeholders, define the technical remediation approach, implement … testing, release, resilience, and legacy technology challenges with development teams. Assess application code, dependencies, runtime environment, test coverage, deployment architecture, CI/CD pipelines, observability, and operational risks. Write, debug, review, and enhance production-quality code and configuration throughout engagements. Define and implement practical modernization and remediation plans with clear ...

Senior Backend Engineer | AI Platform

Location
Greater London, England, United Kingdom
high degree of autonomy and ownership, as you'll be responsible for designing scalable solutions that empower multiple engineering teams while ensuring reliability, observability, and cost efficiency. What are we looking for: 5+ years of experience in Software Engineering, Backend Engineering, or Platform Engineering. Strong experience building and maintaining backend … LangChain, LangGraph, CrewAI, or similar. Experience working with cloud platforms such as Google Cloud Platform (preferred), AWS, or Azure. Strong understanding of system reliability, observability, monitoring, and incident management. Experience with Infrastructure as Code and cloud-native architectures. Previous experience working within a Platform Engineering team is a strong plus. ...

Lead AI Software Engineer London, United Kingdom Value Stream Engineering Posted 12 hours ago

Location
Greater London, England, United Kingdom
aligned to business requirements. Working closely with Technical Leads, Architects, Product Owners, Platform Engineers, and delivery teams, you will drive implementation quality, testing, observability, operational readiness, and continuous improvement throughout the software development lifecycle.## **What you will do*** Lead build execution within a squad, platform capability, or engineering domain.* Translate … generated and engineer-written code to ensure correctness, maintainability, security, and alignment with specifications.* Drive engineering excellence through automated testing, contract testing, regression testing, observability, and production verification.* Support CI/CD processes, deployment readiness, operational handover, runbooks, and service ownership.* Ensure AI-generated outputs are explainable, traceable, secure ...

Lead AI Software Engineer

Hiring Organisation
Vodafone
Location
London, UK
Employment Type
Full-time
aligned to business requirements. Working closely with Technical Leads, Architects, Product Owners, Platform Engineers, and delivery teams, you will drive implementation quality, testing, observability, operational readiness, and continuous improvement throughout the software development lifecycle. What you will doLead build execution within a squad, platform capability, or engineering domain. Translate specifications … generated and engineer-written code to ensure correctness, maintainability, security, and alignment with specifications. Drive engineering excellence through automated testing, contract testing, regression testing, observability, and production verification. Support CI/CD processes, deployment readiness, operational handover, runbooks, and service ownership. Ensure AI-generated outputs are explainable, traceable, secure ...

Staff Platform Engineer - AI Native SaaS Platform

Location
City Of London, England, United Kingdom
role for someone who can design and build scalable systems, not just manage cloud infrastructure. You'll work across infrastructure, automation, CI/CD, observability and production reliability, while writing high-quality software in Go and helping set the technical standard for how engineering teams build and operate services …/CD pipelines, making builds, tests and deployments faster and safer Build automation and internal tooling that removes manual toil for engineering teams Improve observability across metrics, logging, tracing and production monitoring Define and manage SLOs across latency, availability and error rates Lead incident response, resolve complex production issues ...

Sr Lead Software Engineer - Market Risk Technology SRE

Location
Greater London, England, United Kingdom
business issues, and serves as a culture carrier and site reliability adoption champion for your team Collaborates with others to create and implement observability and reliability designs for complex systems which are robust, stable, and do not incur additional toil or technical debt Uses enterprise-authorized AI capabilities within … culture and principles and a track record of demonstrating how to implement site reliability within an application or platform Advanced knowledge and experience in observability such as white and black box monitoring, service level objectives, alerting, and telemetry collection Demonstrated experience using enterprise-authorized AI capabilities within the work environment ...

Foundation Engineering - SRE Platforms - Site Reliability Engineer - Associate - London

Hiring Organisation
Goldman Sachs
Location
London, UK
Employment Type
Full-time
build, run and continuously improve this business-critical service. The role combines software engineering, systems engineering and production expertise to improve the reliability, scalability, observability, incident response and operational efficiency of the CTL platform. Required Skills/ExperienceStrong programming ability in one or more modern languages – Java or Go strongly … building maintainable automation beyond simple scripts. Good understanding of networking, messaging, distributed systems, data structures, algorithms and software design fundamentals. Hands-on experience with observability tooling, including metrics, logging, tracing and dashboarding platforms such as Prometheus, Grafana, ELK or OpenTelemetry. Proven ability to investigate production issues, identify root causes ...

Staff Software Engineer-AI

Location
Greater London, England, United Kingdom
technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging … technologies in production environments Strong experience designing and implementing application programming interfaces, distributed systems, event-driven architectures, data pipelines, PostgreSQL, MongoDB, Redis, vector databases, observability, and automated deployment pipelines Demonstrated ability to influence technical direction while remaining close to the codebase, mentoring engineers through design reviews, code reviews, pairing, debugging ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
Senior DevOps Engineer to take ownership of the cloud infrastructure and DevOps practices powering our BIM Platform, working across Azure, Kubernetes, CI/CD, observability, security and developer tooling. The Role This is a hands‐on senior engineering role with broad ownership across our cloud and platform infrastructure. … Code using Terraform, ARM templates and Helm Design and improve CI/CD pipelines using GitHub Actions, enabling fast, safe and repeatable deployments Build observability across our infrastructure and services through monitoring, logging, alerting and distributed tracing Improve platform reliability, scalability and performance through capacity planning, autoscaling and resource optimisation ...

Lead Engineer, Site Reliability Engineering

Location
Greater London, England, United Kingdom
Write automation to scale systems sustainably, prevent service issues, or when they occur, quickly recover service. Partner with development teams to improve system reliability, observability, and release velocity. Participate in on-call rotations, incident response, postmortems, and root cause analysis and resolution. Be a vocal advocate of strong/sound … Infrastructure as Code using Terraform.* Hands on Experience with one of the following cloud platforms: Azure, AWS, or GCP.* Knowledge on Docker and Kubernetes.* Observability tools like Datadog, Dynatrace or similar.* Implement and maintain CI/CD pipelines.* Incident response and running blameless post-mortems.* Proficient in Git workflows.## ## ...

DevOps / SRE Engineer (London)

Location
Greater London, England, United Kingdom
DevOps/SRE Engineer to join our growing Infrastructure team. In this role, you’ll help shape and strengthen the reliability, scalability and observability of our cloud‐native platform. You’ll work across the business to improve how we build, deploy and monitor our systems , while playing a key role … platforms Solid experience with Terraform and IaC automation Experience participating in or managing production incidents and on‐call Strong grasp of monitoring, alerting, and observability principles Ability to diagnose and fix complex distributed systems issues Demonstrated use of GenAI tools (ChatGPT, GitHub Copilot, Claude) in engineering workflows Excellent communication ...

Foundation Engineering - SRE Platforms - Site Reliability Engineer – Associate - London

Location
Greater London, England, United Kingdom
build, run and continuously improve this business-critical service. The role combines software engineering, systems engineering and production expertise to improve the reliability, scalability, observability, incident response and operational efficiency of the CTL platform. Required Skills/Experience Strong programming ability in one or more modern languages – Java … building maintainable automation beyond simple scripts. Good understanding of networking, messaging, distributed systems, data structures, algorithms and software design fundamentals. Hands‐on experience with observability tooling, including metrics, logging, tracing and dashboarding platforms such as Prometheus, Grafana, ELK or OpenTelemetry. Proven ability to investigate production issues, identify root causes ...

Principal Software Architect (UK)

Location
Greater London, England, United Kingdom
containerization frameworks (Kubernetes, Docker, Helm). Experience with infrastructure-as-code, configuration management (Ansible, Terraform), and network automation protocols (NetConf, YANG). Distributed Systems & Observability: Deep familiarity with modern distributed systems tooling, including RESTful APIs, gRPC, message brokers (Kafka, NATS), and data serialization formats (JSON, YAML). Demonstrated expertise … system observability, telemetry, and monitoring stacks (Prometheus, Grafana). High-Performance Networking: Knowledge of high‐performance packet processing techniques (e.g., DPDK, eBPF) and low‐latency systems tuning. Proficiency in network protocols and architecture, including TCP/IP, UDP, routing, and network performance optimization. Preferred Qualifications Master's degree in computer ...

Lead Agentic AI Architect

Hiring Organisation
CBSbutler Holdings Limited
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
Up to £710 per day + inside ir35
memory, orchestration and tool integration Production-grade LLM and RAG architectures Intelligent search, knowledge assistants, workflow automation and decision-support systems Enterprise LLMOps, MLOps, observability and AI governance AI security, guardrails, evaluation and Responsible AI Cloud AI platforms and hybrid/multi-cloud architecture You'll also provide technical leadership … Hybrid/Multi-Cloud Engineering & Platform Python | Java | Scala | TypeScript | SQL | Kubernetes | Docker | Terraform | Bicep | CI/CD | Infrastructure as Code AI Governance & Observability LLMOps | MLOps | Model Governance | AI Security & Guardrails | LangSmith | Arize | Monitoring & Evaluation Frameworks Qualifications Relevant qualifications may include: Master's degree in Computer Science, AI, Data Science ...

Analytics Services Platform Engineer

Location
Greater London, England, United Kingdom
developer experience Collaborating with research, data and engineering teams to accelerate time‐to‐insight through modern analytics solutions Driving improvements in automation, observability and resilience across analytics services Evaluating and adopting emerging technologies such as AI assistants, data mesh and cloud‐native analytics solutions Defining SLAs, KPIs and monitoring strategies … using Terraform or Ansible Deep understanding of AWS analytics technologies including EMR, MSK, Athena, Redshift, Glue and MWAA Experience with CI/CD and observability tools such as Jenkins, ArgoCD, Prometheus, Grafana and OpenTelemetry Strong problem‐solving skills and a systematic approach to diagnosing and resolving issues Highly Desirable Skills ...

Software Engineering Manager

Hiring Organisation
Halian Technology Limited
Location
Central London, London, United Kingdom
Employment Type
Permanent
making sound architectural and design decisions. Champion software quality, security, performance, and operational excellence. Encourage modern engineering practices, including CI/CD, automated testing, observability, and cloud-native development. Stakeholder Management Build strong relationships with business and technology stakeholders. Communicate progress, risks, and dependencies effectively. Align engineering activities with organisational … with highly available, transaction-heavy platforms. Technology Environment .NET Microservices architecture RESTful APIs Kubernetes and Docker AWS, Azure, or GCP CI/CD tooling Observability and monitoring platforms There is a 2 - 3 stage interview process, with interview slots now available with a range of benefits and a bonus ...

Analytics Services Platform Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
usability, scalability and the developer experienceCollaborating with research, data and engineering teams to accelerate time-to-insight through modern analytics solutionsDriving improvements in automation, observability and resilience across analytics servicesEvaluating and adopting emerging technologies such as AI assistants, data mesh and cloud-native analytics solutionsDefining SLAs, KPIs and monitoring strategies … code, using Terraform or AnsibleDeep understanding of AWS analytics technologies including EMR, MSK, Athena, Redshift, Glue and MWAAExperience with CI/CD and observability tools such as Jenkins, ArgoCD, Prometheus, Grafana and OpenTelemetryStrong problem-solving skills and a systematic approach to diagnosing and resolving issuesHighly desirable skillsExperience with streaming frameworks ...