1,251 to 1,275 of 2,378 Observability Jobs in London

Senior DevOps Engineer - Remote, FinOps & Compliance

Location
Greater London, England, United Kingdom
contribute to secure multi-account governance, IAM/SOC 2 alignment, and cost management while supporting a distributed engineering team. You will shape reliability, observability, and automation using Terraform, CDK, or Pulumi, in a fully remote setup with a global team and a strong focus on security and operational excellence. ...

Senior Trust & Safety Engineer — Scale trust & reviews

Location
Greater London, England, United Kingdom
design-time decisions. You should have strong TypeScript and Node.js skills, be comfortable with Next.js and an event-driven mindset, and apply Datadog for observability alongside SQL and GCP basics. #J-18808-Ljbffr ...

Web App Developer - React/Python, Azure Cloud, Hybrid

Location
Greater London, England, United Kingdom
Vite frontends, contributing to reliability, performance, and developer experience. In this hybrid London role, you’ll collaborate with engineering teams on CI/CD, observability, security and scalable cloud deployments, with a strong focus on delivering robust software and reusable engineering patterns. #J-18808-Ljbffr ...

Senior Applied AI Engineer — Enterprise AI Platform

Location
Greater London, England, United Kingdom
lead the delivery of enterprise AI products and platform services. You will design, build, and productionise AI systems, govern them through testing and observability, and mentor fellow engineers to raise engineering standards. You will work with Azure AI Foundry, Accio, Nexus, and other components, shaping how the team operates ...

Engineering Lead for Acquisition & AI Growth

Location
Greater London, England, United Kingdom
raise standards while shaping how AI tools are used across the team. You’ll own end-to-end delivery, drive CI/CD and observability improvements, and balance reliability with cost across AWS services. Hybrid work is offered in a dynamic, AI-driven environment. #J-18808-Ljbffr ...

Senior Python Backend Engineer – Fintech API Systems

Location
Greater London, England, United Kingdom
Python, FastAPI, SQLAlchemy, and related technologies. Collaborate with cross-functional teams, own complex problems, and deliver production-ready code with strong emphasis on maintainability, observability, and reliability. #J-18808-Ljbffr ...

Senior Data Platform Engineer - Lakehouse & Catalog

Location
Greater London, England, United Kingdom
Engineer to join the Data Lake and Catalog team in London. You will design, build, and operate scalable data lake services, focusing on reliability, observability, and end-to-end platform quality. You will mentor engineers, partner with cross-functional teams, and help evolve the lakehouse architecture using Hive, Spark ...

Senior Platform Engineer: Cloud Infra Leader (Remote)

Location
Greater London, England, United Kingdom
shared platform layer across B2B and B2C products in a high-growth fintech environment. You will build and improve scalable, secure cloud infrastructure, enhance observability and CI/CD, and collaborate with product and engineering teams to enable rapid, reliable releases. #J-18808-Ljbffr ...

Senior MLOps Product Manager — AI Infra Platform

Location
Greater London, England, United Kingdom
owns strategy, roadmap, and delivery of the ML platform and cloud services. You will collaborate across GPU compute, training, inference, orchestration, model lifecycle, tooling, observability, and platform integrations with engineering, SRE, security, commercial teams, and customers. This role focuses on making it easier for customers to build, train, deploy ...

Senior HPC Network Engineer — RDMA, Leaf-Spine, Equity

Location
Greater London, England, United Kingdom
architecture through day-2 operations, including RDMA fabrics, leaf-spine design, and per-tenant isolation. You’ll automate provisioning with Python and Ansible, build observability dashboards, and mentor colleagues while expanding the data-centre and office networks. #J-18808-Ljbffr ...

Associate Site Reliability Engineer, SRE Platforms

Location
Greater London, England, United Kingdom
help build, run and continuously improve the Consolidated Trade Ledger platform in London. This role combines software and systems engineering to boost reliability, observability and incident response for a cloud-native service. The candidate will implement automation, maintain canary and blue/green deployment approaches, and contribute to scalable design ...

Python Engineer - AI-Driven Energy Platform & Modelling

Location
Greater London, England, United Kingdom
systems from day one, owning APIs, data pipelines, AWS infrastructure, and tooling. You’ll work across the stack, shipping features, hardening pipelines, and improving observability, with autonomy and a bias for practical design. Experience with AI tools and LLM-powered applications is a plus. #J-18808-Ljbffr ...

Senior Platform Engineer – Production Reliability & On-Call

Location
Greater London, England, United Kingdom
Platform/SRE role in London. You will own production reliability, participate in on‐call and incident response, and drive improvements across observability, deployments, and operational tooling. This is an ops‐heavy, hands‐on position suitable for mid‐level to senior SREs who enjoy ownership and direct impact on real ...

AI Platform Engineer for Production ML & Kubernetes

Location
Greater London, England, United Kingdom
failures in distributed systems, and guide customers through complex Kubernetes and GPU workloads in production. We seek engineers who can diagnose performance bottlenecks, improve observability, and build runbooks. This hybrid London role requires in-office presence at least 2 days/week and supports large-scale AI #J-18808-Ljbffr ...

Senior SRE Lead - Resiliency, AI-Driven Incidents

Location
Greater London, England, United Kingdom
across the SDLC to minimize incidents and maximize service levels. You will collaborate with cross-functional partners to define SLOs, apply best practices in observability, and mentor teams while owning critical incident response and architectural decisions. #J-18808-Ljbffr ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
party intelligence platform, reliability, infrastructure, and delivery systems. What you’ll do Own platform reliability, deployment workflows, and infrastructure improvements for Xapien systems. Improve observability, operational safety, and delivery speed for engineering teams. Support secure, scalable infrastructure for an AI-powered intelligence product. Requirements Senior DevOps, platform, or infrastructure engineering ...

Senior Data Engineer: Scale Pipelines & AI-Driven Data

Location
Greater London, England, United Kingdom
hands-on role, you’ll mentor peers through code reviews and knowledge sharing while promoting best practices and an emphasis on CI/CD, observability, and IaC. The team operates in a hybrid model across the UK. #J-18808-Ljbffr ...

Senior AI Engineer for Production-Grade LLMs

Location
Greater London, England, United Kingdom
features inside cross-functional Innovation Squads. The role focuses on production-grade AI systems, retrieval pipelines, agent workflows, and enterprise integration, with emphasis on observability and guardrails. You will collaborate with host-function stakeholders, mentor capability-building, and ensure solutions fit real workflows, not assumptions. Strong CI/ ...

Senior Full-Stack Engineer — Backend-Heavy with AI Impact

Location
Greater London, England, United Kingdom
deliver measurable impact at scale. This is a remote, internationally oriented role with opportunities across 175 countries. You’ll contribute to architectural direction, champion observability, and foster engineering excellence while balancing short-term delivery with long-term #J-18808-Ljbffr ...

Senior Real-Time Full-Stack Engineer (Voice & AI)

Location
Greater London, England, United Kingdom
front ends, Scala and Python services, and cloud infrastructure. You will own realtime voice integrations, streaming pipelines, and multi-tenant architecture while ensuring safety, observability, and robust go-lives across partner integrations. The role emphasizes senior ownership, low-latency delivery, and collaboration with product, design, and operations to ship iteratively ...

Backend Engineer — AI-Powered Expense Platform

Location
Greater London, England, United Kingdom
will ship backend and frontend features, build AI-backed agents, and improve extraction, policy, and audit capabilities at scale. The role emphasizes production quality, observability, and cross-team collaboration across time zones. Ideal candidates have 5+ years in production software, strong CS fundamentals, and experience with Java/Spring Boot ...

Lead Backend Engineer — AI-Driven Insurance Platform

Location
Greater London, England, United Kingdom
real-time insurance platform. You will guide several engineers, shape backend architecture, and drive end-to-end delivery with a strong emphasis on simplicity, observability, and reliability. You will collaborate across squads, mentor peers, and advocate AI-first approaches to deliver high-impact solutions in a fast-paced, hybrid London ...

Platform Engineer

Location
City Of London, England, United Kingdom
platform, working across IAM, Kubernetes, networking, logging, monitoring, and cloud architecture. Key experience: AWS Platform Engineering Kubernetes/EKS IAM & Cloud Security CloudTrail, GuardDuty & Observability Cloud Architecture & Infrastructure Documentation Security or compliance focused engineering projects London, 2 days onsite per week Up to £450/day, Outside IR35 #J ...

Hybrid SDET: AI-Driven QA Automation & CI/CD

Location
Greater London, England, United Kingdom
automation across UI, API, and data layers. You will work closely with the Head of QA to improve CI/CD quality gates and observability while reducing manual testing effort. You will contribute hands-on code, review pull requests, and advance our testing strategy by determining what to test ...

Senior Dedicated Support Engineer - 24/7 Incident Response

Location
Greater London, England, United Kingdom
incident resolution, automation, and cross-team collaboration to ensure high-quality service experiences. The position requires strong coding in Python, experience with AWS and observability tools, and willingness to join a 24x7 on-call rotation. Bachelor's degree is preferred, with 5–7 years of relevant work experience. #J ...