1,401 to 1,425 of 2,378 Observability Jobs in London

Backend Engineer, Core BI Platform & Automation

Location
Greater London, England, United Kingdom
infrastructure powering BI platforms used across the DoorDash ecosystem. You will automate the platform lifecycle, implement integrations with BI platforms, and improve availability, observability and security. You will also apply AI tooling to deliver reliable, scalable solutions and mentor the team on best practices. #J-18808-Ljbffr ...

Platform Engineering Leader: Cloud, DevOps & Security

Location
Greater London, England, United Kingdom
consistent engineering standards to deliver secure, reliable software. In this hybrid London role, you will partner with Security, Architecture and Engineering teams, drive resilience, observability, and incident response, and mentor engineers while managing #J-18808-Ljbffr ...

Senior Quality Engineer - Champion Quality in Cloud Apps

Location
Greater London, England, United Kingdom
perform exploratory testing, and collaborate with Product and Development to build the right software for customers, embracing shift-left quality practices. You will enhance observability, apply risk-based testing, and contribute to automation while growing a culture of quality across teams. #J-18808-Ljbffr ...

Senior Data Platform SRE — Hybrid, Obs & Reliability Leader

Location
Greater London, England, United Kingdom
Senior Site Reliability Engineer to embed within the Data Engineering team. You will own reliability, performance, and operability of the data platform, building observability from the ground up and leading incident response for data outages. You'll balance feature velocity with system stability, applying SRE practices and capacity planning ...

Senior SET: Platform Quality Engineer (Hybrid, London)

Location
Greater London, England, United Kingdom
architecture level, automate validation across environments, and manage delivery risk in distributed systems. You will define standards for testing distributed systems, enable observability-driven debugging, automate validation of availability and latency, and contribute to security posture. #J-18808-Ljbffr ...

Senior Data Platform Engineer | ELT & Analytics

Location
Greater London, England, United Kingdom
senior member, you will influence platform architecture, data products, and engineering standards while collaborating with analytics teams to deliver reliable production solutions and observability across key datasets. #J-18808-Ljbffr ...

SRE Lead: Data, Cloud & Developer Experience

Location
Greater London, England, United Kingdom
Blackstone is seeking a Site Reliability Engineer to lead the adoption of SRE practices, improve observability, and ensure reliable services across the firm. The role involves instrumentation, monitoring, and automation to reduce toil and incident impact. The successful candidate will collaborate with development and operations teams to design resilient systems ...

Site Reliability Engineer - Live Ops & Cloud Resilience

Location
Greater London, England, United Kingdom
experienced Site Reliability Engineer to design, build, and operate resilient, secure platforms underpinning our digital and live operations. You’ll focus on reliability, observability, automation, and disaster recovery across hybrid environments, collaborating with engineering, operations, and project stakeholders. The role emphasizes improving service availability, incident response, and continuous improvement ...

Platform Engineering Lead: GPU Infra & Kubernetes

Location
Greater London, England, United Kingdom
Volta is seeking a hands-on Platform Engineering leader to own reliability, observability, and security across a Kubernetes-native compute platform. You will lead a team, set technical direction, and translate product requirements into scalable platform features while collaborating with security and cross-team leads. This role requires 5+ years ...

AI-Ops SRE & Operations Leader

Location
Greater London, England, United Kingdom
experienced SRE Manager to lead the reliability function for production services used by internal and external customers. You will drive AI-Ops adoption, automation, observability and incident response. You will own end-to-end incident and problem management, coach team leads, ensure RCAs and post-mortems are completed, and balance ...

Kubernetes SRE: Secure, Scalable Sandbox Infra

Location
Greater London, England, United Kingdom
Senior SRE to scale Kubernetes workloads and guarantee tenant isolation. You will own the infra that keeps agent testing safe and always-on, improving observability and efficiency. You will join a highly concurrent system, manage on-call rotations, and drive cost and performance optimizations while shaping security practices across deployments. ...

Senior Backend Architect for Core Platform

Location
Greater London, England, United Kingdom
operate core backend services enabling secure C2 operations across Nova Cloud deployments. The role focuses on service architecture, APIs, data models, authn/authz, observability, and developer experience, supporting multiple Nova modules and product teams. Strong emphasis on reliability and security. #J-18808-Ljbffr ...

Senior Backend Engineer, Scalable ChatGPT Infra

Location
Greater London, England, United Kingdom
safe, scalable capabilities. The role emphasizes reliability, performance, and ownership of production behavior, with opportunities to lead architectural improvements and contribute to system-wide observability and on-call duties. #J-18808-Ljbffr ...

AI-First Analytics Engineer: Data Platforms & Insights

Location
Greater London, England, United Kingdom
assets in Snowflake, dbt and Looker, using AI-assisted development to focus on intent and clarity. You will review SQL, extend data contracts and observability, integrate data from Amplitude, Segment and Google Ads, and enable self-service in BI tools for business users. #J-18808-Ljbffr ...

GenAI Architect: Enterprise AI on AWS & Bedrock

Location
Greater London, England, United Kingdom
delivery of enterprise-scale AI solutions on AWS, leveraging Bedrock and RAG techniques. You will shape reusable patterns for prompt orchestration, agentic workflows, and observability while working with security, compliance and engineering teams. The role emphasizes governance, model risk management, data privacy, and production-grade deployment in highly regulated environments. ...

Senior Data Analyst: AI/ML UX & Hybrid Data Pipelines

Location
City Of London, England, United Kingdom
three on-site days in City of London. Responsibilities include building scalable data pipelines (BigQuery, Dataflow/Apache Beam, Airflow), ensuring data quality and observability (Looker, Monte Carlo), and collaborating with product engineering and data science teams to plan data tracking and ingestion tasks. #J-18808-Ljbffr ...

Senior UI Platform Engineer: Reusable Frontend Foundations

Location
Greater London, England, United Kingdom
TypeScript packages, and shape onboarding and access experiences across teams. The role focuses on platform engineering for internal web apps, with emphasis on authentication, observability, and developer workflows. Base salary ranges £120,000–£155,000 for London, plus a full compensation package. #J-18808-Ljbffr ...

Senior Cloud Platform Engineer: Self-Service & Automation

Location
Greater London, England, United Kingdom
technical leadership role requires guiding architectural decisions while remaining practical and collaborative. The role emphasises platform engineering as a product, with focus on automation, observability, and governance to accelerate software #J-18808-Ljbffr ...

Generative AI Testing Engineer - Python & Validation

Location
Greater London, England, United Kingdom
system validation in production environments. The role focuses on architecture design and hands-on development of evaluation pipelines, synthetic data generation, and observability layers to deliver robust quality #J-18808-Ljbffr ...

Engineering Director - AI-Driven Research Analytics

Location
Greater London, England, United Kingdom
enable AI-assisted experiences for researchers and institutions. You will oversee workforce planning, architectural governance, and roadmaps, while maintaining high standards for reliability, observability and security. #J-18808-Ljbffr ...

Platform Engineer – Scale GPU Infra for AI Platform

Location
Greater London, England, United Kingdom
infrastructure powering large-scale GPU compute, ensuring reliability, perf and fast developer feedback. You’ll work across Kubernetes, containers, cloud providers and observability tools to keep the platform humming and accelerate experimentation. This is a hands-on role with real ownership and impact. #J-18808-Ljbffr ...

Platform SRE Lead: On-Call, Reliability & Cloud Ops

Location
Greater London, England, United Kingdom
platform. Open to mid-level or senior SREs, the role is ops-heavy with focus on health of live systems, Kubernetes, cloud infra, deployments, observability, and collaboration with product teams. Hybrid work with 3 days in the office. #J-18808-Ljbffr ...

Analytics Engineer: AI-Driven Data & Trust

Location
Greater London, England, United Kingdom
collaborate with senior engineers to raise our analytics depth. You'll work with dbt, Snowflake, Looker and modern data platforms, own data quality and observability, and enable self-service BI for the business. #J-18808-Ljbffr ...

Senior AI Engineer – GenAI, RAG & Agents (Hybrid)

Location
Greater London, England, United Kingdom
including Generative AI, RAG, and AI agents, collaborating with architecture, security, and infrastructure teams. You will lead the adoption of best practices, model evaluation, observability, and governance across SDLC while translating #J-18808-Ljbffr ...

Senior Network Reliability Engineer - Azure Networking

Location
City Of London, England, United Kingdom
hybrid network estate across Azure and on-premises, with a hands-on, incident-driven focus. You will lead complex networking challenges, drive automation and observability, and contribute to a more Azure-native architecture from our Wimbledon office. You will partner with Cloud, Security and Engineering teams, mentor peers, and help ...