3,426 to 3,450 of 4,168 Observability Jobs

MLOps & Infra Engineer

Location
Greater London, England, United Kingdom
software on HPC platforms, enabling distributed systems across the team. You’ll work with Kubernetes, Terraform, and CI/CD practices, shaping cloud infrastructure, observability, and ML workflow orchestration within a flexible hybrid UK setup. #J-18808-Ljbffr ...

Staff Cloud Native Engineer: Kubernetes AI Infra Leader

Location
United Kingdom
with core networking components on GPU-backed infrastructure. You will extend control plane capabilities, own significant components end-to-end, and raise reliability and observability across platforms. Strong Go skills and hands-on Kubernetes internals are essential for collaboration with platform teams. #J-18808-Ljbffr ...

Distinguished Engineer (Head of Service Reliability Engineering) - HMRC - SCS1

Location
Manchester, England, United Kingdom
Head of Service Reliability Engineering, you will lead a distinct, independent capability spanning services, products and platforms. You will make reliability, operability and observability integral to engineering from the outset. With enterprise-wide reach and influence, you will set the direction for SRE, raise service maturity and build a lasting … adoption of SRE practices by using SLOs, error budgets and reliability metrics to drive measurable improvements in service performance and operational decision-making. Observability: Establishes and governs enterprise observability capabilities, using telemetry, dependency mapping and modern monitoring platforms to improve operational insight and decision quality. Resilience and Recovery: Designs ...

Principal Engineer - Member Experience Platform

Location
Skipton, England, United Kingdom
Quality), and bar‐raising across squads: you shorten lead times, increase deployment frequency, hold change‐failure rate low, and improve MTTR through release‐linked observability - turning fast, safe flow into the default way of working. Operating at platform scale, you define cross‐cutting architecture and delivery standards (API/event … contracts, resilience, observability, language/dependency baselines) and drive adoption through the Golden Path: policy‐as‐code CI/CD, progressive delivery (feature flags, canary/blue‐green), automated rollback/forward‐fix, ephemer... data‐ready environments, and guardrails that make security and compliance by design. You partner with Platform ...

Senior AWS Platform Engineer - Secure Cloud Automation

Location
United Kingdom
architects and consultants to deliver end-to-end platform solutions, implement secure-by-design principles, and mentor early-career colleagues while promoting automation and observability across projects. #J-18808-Ljbffr ...

Senior Backend Engineer — Scalable AI Media Platforms

Location
United Kingdom
work with DevOps on platform infrastructure. Occasional frontend touchpoints help expose backend capabilities. Responsibilities include API and system design, data processing, model integration, observability, and performance optimization to ensure low latency and high #J-18808-Ljbffr ...

Senior Security Platform Architect (SCA & Backend)

Location
Cambridge, England, United Kingdom
design and implement backend services, Python APIs, and workflow components to enable tool onboarding, analysis execution, results processing, and delivery, while improving scalability and observability across the platform. #J-18808-Ljbffr ...

Senior Platform Engineer - AWS Cloud, Automation & Self-Service

Location
Greater London, England, United Kingdom
tooling to boost developer productivity and automate key operational processes. You will lead complex platform initiatives, mentor engineers, and drive best practices for reliability, observability and governance, influencing senior stakeholders on technology choices. #J-18808-Ljbffr ...

Lead AI Platform Engineer - Scale Inference & Security

Location
City of Edinburgh, Scotland, United Kingdom
models and inference services. You will own the platform layer above GPU infrastructure, deploy scalable inference services with containers and Kubernetes, improve observability, and help define secure, resilient operating standards for enterprise environments. #J-18808-Ljbffr ...

GenAI Full-Stack Engineer (Python/TypeScript) – Hybrid

Location
Belfast City District, Northern Ireland, United Kingdom
/Azure cloud architecture, delivering production GenAI systems used across global operations. This hybrid role focuses on feature implementation, system integration, cloud architecture and observability tooling, ensuring secure coding and scalable deployments. #J-18808-Ljbffr ...

Hybrid AI-Driven SRE & Reliability Engineer

Location
United Kingdom
bet365 Group is seeking a Site Reliability Engineer to enhance system reliability, observability, and performance. You will treat reliability as a software problem, protecting uptime and driving improvements across critical systems. Responsibilities include building tools, dashboards, and automation, contributing to live incident resolution and post-mortems, and mentoring teammates ...

Senior Backend Engineer – AI-Driven FinTech Platform

Location
Greater London, England, United Kingdom
that supports AI-driven advisory workflows and regulatory compliance. You will contribute to event-driven architectures, robust data models and reliable integrations, ensuring correctness, observability and scalable operations for complex financial services platforms. #J-18808-Ljbffr ...

Senior AI/ML Solutions Architect (GenAI & MLOps)

Location
Greater London, England, United Kingdom
Field Engineering team to design production‐grade AI solutions on the Databricks platform. You will drive GenAI initiatives, RAG architectures, agentic systems, AI observability, and NLQ of structured data, while mentoring peers and influencing the platform roadmap. Some travel may be required. #J-18808-Ljbffr ...

Latency‐Focused Cloud Infrastructure Engineer

Location
Greater London, England, United Kingdom
latency-sensitive trading environment, focused on performance, resilience, and safe change management. You will work on IaC (Terraform/Terragrunt), CI/CD, and observability, collaborating with exchanges and internal stakeholders to support continuous trading operations. #J-18808-Ljbffr ...

Engineering Manager - Lead Scalable Betting Platform

Location
Greater London, England, United Kingdom
drive delivery with accountability. The role requires hands-on Java Spring and Kafka expertise, cloud-native development (AWS), and strong focus on reliability, observability, and AI-enabled tooling. #J-18808-Ljbffr ...

Lead GenAI & LLM Platform Engineer

Location
Greater London, England, United Kingdom
knowledge with strong software engineering. You will lead deployment of scalable AI systems, integrate LLMs for enterprise planning, and develop API services and observability to ensure robust performance and cost efficiency. #J-18808-Ljbffr ...

Senior Cloud Infrastructure Engineer | Scale a Global Platform

Location
Greater London, England, United Kingdom
cloud-native core and payments tech to modernize banking. Senior Software Engineers in Infrastructure design and deploy scalable platform tooling, focusing on multi-cloud, observability, and data systems. You will contribute to a cloud-agnostic core platform, build automation, and integrate with open-source tools, while ensuring high reliability ...

AI Platform Architect & Infra Lead

Location
City of Edinburgh, Scotland, United Kingdom
inference services, while owning the platform layer above managed GPU infrastructure. You will deploy scalable inference services using containers, Kubernetes, and AI gateways, improve observability, and enforce security, performance and availability. #J-18808-Ljbffr ...

Principal Cloud Engineer | AI/ML FinOps & Terraform

Location
Greater London, England, United Kingdom
modular microservices, and deploy AI/ML models to optimize cloud spend. As a principal-level IC, you will own data modeling, orchestration, and observability while collaborating with FinOps analysts, cloud engineers, and platform leads to elevate the entire cost-management practice. #J-18808-Ljbffr ...

Senior Backend Engineer: Chaos & Reliability (Remote)

Location
Greater London, England, United Kingdom
guide product direction, and collaborate with friendly colleagues who live our FAITH values. You’ll design and run chaos experiments, improve load testing and observability, and introduce new tooling to boost reliability. Global teams collaborate on scalable solutions. #J-18808-Ljbffr ...

Senior DevOps Engineer - Platform & Automation Champion

Location
Ashford, England, United Kingdom
England is looking for an experienced professional in DevOps and Platform Engineering. You'll play a crucial role in modernizing our payment platforms, enhancing observability, and leading automation efforts across various engineering teams. The ideal candidate will have over 5 years of hands-on experience, a strong technical background ...

Remote Data Engineer for AI Data Platform

Location
United Kingdom
with a strong emphasis on privacy, security, and robust data practices. As part of a collaborative team, you’ll design data pipelines, APIs, and observability, applying IaC and container orchestration tools to keep our platform scalable, reliable, and self-serve for internal teams. #J-18808-Ljbffr ...

SRE & Operations Leader - Reliability at Scale

Location
United Kingdom
services used by internal and external customers. You will drive reliability improvements, advance automation and AI-Ops capabilities, and lead a team focused on observability, incident response, and continuous improvement. You will line manage team leaders, shape strategic direction, ensure incident management and RCAs are completed, and collaborate with ...

AWS DevOps Engineer – EKS, Terraform, GitOps

Location
Greater London, England, United Kingdom
with developers and SREs in an automation-first culture to deliver reliable, scalable systems. The role covers on-call support, incident response, cost optimization, observability, and secure workload access across AWS services, with emphasis on security and reliability. #J-18808-Ljbffr ...

Senior Backend Engineer (Go) — Revenue & Billing

Location
Greater London, England, United Kingdom
data flows that turn usage into invoices. You will build in Go, integrate with ERP and the Data Platform, and ensure correctness and observability as revenue scales. Joining a cross-functional team with Finance and Go To Market, you will shape the Revenue team's foundations, design robust tests ...