3,776 to 3,800 of 4,209 Observability Jobs

AI-Ops SRE & Operations Leader

Location
Greater London, England, United Kingdom
experienced SRE Manager to lead the reliability function for production services used by internal and external customers. You will drive AI-Ops adoption, automation, observability and incident response. You will own end-to-end incident and problem management, coach team leads, ensure RCAs and post-mortems are completed, and balance ...

AI Platform & SRE Leader — Scale & Govern AI

Location
Greater London, England, United Kingdom
Platform & Site Reliability Engineering Managing Consultant to help clients design, build and scale secure, reliable AI platforms. You will lead platform engineering, SRE, observability and guardrails, moving AI from experimentation to enterprise-scale delivery. You will partner with CIO/CTO stakeholders, shape platform strategies, and guide multidisciplinary delivery teams ...

Site Reliability Engineer - Live Ops & Cloud Resilience

Location
Greater London, England, United Kingdom
experienced Site Reliability Engineer to design, build, and operate resilient, secure platforms underpinning our digital and live operations. You’ll focus on reliability, observability, automation, and disaster recovery across hybrid environments, collaborating with engineering, operations, and project stakeholders. The role emphasizes improving service availability, incident response, and continuous improvement ...

Platform Engineering Lead: GPU Infra & Kubernetes

Location
Greater London, England, United Kingdom
Volta is seeking a hands-on Platform Engineering leader to own reliability, observability, and security across a Kubernetes-native compute platform. You will lead a team, set technical direction, and translate product requirements into scalable platform features while collaborating with security and cross-team leads. This role requires 5+ years ...

Production Reliability Lead

Location
Norwich, England, United Kingdom
team to move from firefighting to proactive reliability engineering. This hands-on role requires leading SRE/DevOps practices, defining SLIs/SLOs, improving observability, and ensuring clear ownership with blameless accountability. #J-18808-Ljbffr ...

Senior AI Engineer - Real-Time Personalization & AI Systems

Location
Manchester, England, United Kingdom
players. As a senior engineer, you’ll own the design and delivery of real-time AI systems, focusing on low latency, reliability, and observability, while collaborating with product managers and data analysts. #J-18808-Ljbffr ...

Staff Data Platform Engineer — Real-Time Pipelines & Infra

Location
United Kingdom
datasets that underpin our products, delivering scalable real-time and batch data pipelines for geospatial data. You will drive architectural decisions and champion reliability, observability, and end-to-end ownership across cloud-based infrastructure. The Staff Engineer role involves mentoring engineers, aligning priorities, and collaborating with product and domain stakeholders ...

VP, SRE: Architect Resilient, Scalable Systems

Location
Birmingham, England, United Kingdom
Vice President in Site Reliability Engineering (SRE) within Core Engineering at the Birmingham location. You will lead reliability efforts across distributed systems, driving SLOs, observability, and incident response while shaping scalable, automated platforms. This role emphasizes design reviews, reliability improvements, and reducing toil through automation, with a focus ...

Kubernetes SRE: Secure, Scalable Sandbox Infra

Location
Greater London, England, United Kingdom
Senior SRE to scale Kubernetes workloads and guarantee tenant isolation. You will own the infra that keeps agent testing safe and always-on, improving observability and efficiency. You will join a highly concurrent system, manage on-call rotations, and drive cost and performance optimizations while shaping security practices across deployments. ...

Senior Backend Architect for Core Platform

Location
Greater London, England, United Kingdom
operate core backend services enabling secure C2 operations across Nova Cloud deployments. The role focuses on service architecture, APIs, data models, authn/authz, observability, and developer experience, supporting multiple Nova modules and product teams. Strong emphasis on reliability and security. #J-18808-Ljbffr ...

Cloud Native SWE II: Kubernetes & OpenTelemetry

Location
Manchester, England, United Kingdom
platform. You will work with a team of engineers defining what it means to be cloud native, contributing to components and integrations across Kubernetes, observability and automation tooling. Under guidance from senior engineers, you will deepen your knowledge of the Cloud Native ecosystem, help bring Couchbase everywhere ...

Field-Ready AI Engineer: Deploy & Scale Agentic AI

Location
Crawley, England, United Kingdom
production across diverse platforms. The role focuses on building production architectures, APIs, CI/CD pipelines and security practices, with emphasis on reliability, observability and scalable performance in scientific and enterprise contexts. #J-18808-Ljbffr ...

Operations Team Lead — Production Reliability & Scale

Location
Cambridge, England, United Kingdom
seeking an Operations Team Lead to own production and scale systems. You will lead operational excellence across live customer-facing platforms, ensuring reliability, observability, and proactive improvements. This hands-on role involves shaping processes, guiding incidents, building the team, and moving from firefighting to sustainable reliability engineering. The ideal candidate ...

Senior Cloud Platform Engineer: Self-Service & Automation

Location
Greater London, England, United Kingdom
technical leadership role requires guiding architectural decisions while remaining practical and collaborative. The role emphasises platform engineering as a product, with focus on automation, observability, and governance to accelerate software #J-18808-Ljbffr ...

Engineering Director - AI-Driven Research Analytics

Location
Greater London, England, United Kingdom
enable AI-assisted experiences for researchers and institutions. You will oversee workforce planning, architectural governance, and roadmaps, while maintaining high standards for reliability, observability and security. #J-18808-Ljbffr ...

Generative AI Testing Engineer - Python & Validation

Location
Greater London, England, United Kingdom
system validation in production environments. The role focuses on architecture design and hands-on development of evaluation pipelines, synthetic data generation, and observability layers to deliver robust quality #J-18808-Ljbffr ...

Senior Backend Engineer, Scalable ChatGPT Infra

Location
Greater London, England, United Kingdom
safe, scalable capabilities. The role emphasizes reliability, performance, and ownership of production behavior, with opportunities to lead architectural improvements and contribute to system-wide observability and on-call duties. #J-18808-Ljbffr ...

Senior IT Operations & Incident Lead

Location
Gloucester, England, United Kingdom
Santander UK (TSB) is seeking a Senior Operations Analyst to monitor core services and drive observability improvements in Edinburgh on a 12-month fixed-term contract. You will triage incidents, coordinate recovery, and work with IT Operations & supplier teams to minimise disruption. Ideal candidates bring hands-on monitoring across applications ...

Senior Platform Engineer: Backend & Developer Tools

Location
City Of London, England, United Kingdom
platform capabilities that improve how teams build, deploy, and operate software. The ideal candidate will create reusable services, frameworks, and tooling, spanning API enablement, observability, and governance. This is a high-ownership position within a fast-moving engineering environment. #J-18808-Ljbffr ...

AI-First Analytics Engineer: Data Platforms & Insights

Location
Greater London, England, United Kingdom
assets in Snowflake, dbt and Looker, using AI-assisted development to focus on intent and clarity. You will review SQL, extend data contracts and observability, integrate data from Amplitude, Segment and Google Ads, and enable self-service in BI tools for business users. #J-18808-Ljbffr ...

Senior UI Platform Engineer: Reusable Frontend Foundations

Location
Greater London, England, United Kingdom
TypeScript packages, and shape onboarding and access experiences across teams. The role focuses on platform engineering for internal web apps, with emphasis on authentication, observability, and developer workflows. Base salary ranges £120,000–£155,000 for London, plus a full compensation package. #J-18808-Ljbffr ...

Senior Data Analyst: AI/ML UX & Hybrid Data Pipelines

Location
City Of London, England, United Kingdom
three on-site days in City of London. Responsibilities include building scalable data pipelines (BigQuery, Dataflow/Apache Beam, Airflow), ensuring data quality and observability (Looker, Monte Carlo), and collaborating with product engineering and data science teams to plan data tracking and ingestion tasks. #J-18808-Ljbffr ...

GenAI Architect: Enterprise AI on AWS & Bedrock

Location
Greater London, England, United Kingdom
delivery of enterprise-scale AI solutions on AWS, leveraging Bedrock and RAG techniques. You will shape reusable patterns for prompt orchestration, agentic workflows, and observability while working with security, compliance and engineering teams. The role emphasizes governance, model risk management, data privacy, and production-grade deployment in highly regulated environments. ...

Principal Engineer — Transformation & Cloud-Native Banking

Location
Greater London, England, United Kingdom
core banking, shaping architecture and cloud-native platforms for safe, frequent releases. You will collaborate across squads, drive best practices in CI/CD, observability, and security, and embed AI-enabled engineering techniques with strong governance. #J-18808-Ljbffr ...

Enterprise Account Executive

Location
United Kingdom
advantage of all structured and unstructured data — securing and protecting private information more effectively — Elastic’s complete, cloud-based solutions for search, security, and observability help organizations deliver on the promise of AI. What is The Role: Elastic, the Data Analytics, AI, Search, Observability & Security company, is seeking a dynamic … high-level technical curiosity and the gravitas of a value-based seller. Because Elastic isn't just a "product" but a versatile search, observability, and security platform, the role requires moving beyond simple transactions to solving complex data problems. As a dynamic Enterprise Account Executive, you will be an integral ...