1,976 to 2,000 of 2,372 Observability Jobs in London

Head of Developer Experience (DevEx)

Location
Greater London, England, United Kingdom
infrastructure team, with strong AWS knowledge (ECS Fargate, S3, RDS, Lambda, EventBridge/SNS/SQS) and hands-on Terraform experience. CI/CD & Observability: Hands-on with GitHub Actions, Terraform Cloud and Cloudflare, and comfortable using observability tools to pinpoint performance and reliability issues. AI Fluency: specificallywith agentic coding ...

Data Engineer

Hiring Organisation
Ashdown Group
Location
London, UK
Employment Type
Full-time
work from home 2 days per week. This is a high-impact role focused on improving data quality, reducing incidents, and building scalable observability across a modern enterprise data platform. You'll help ensure data across the organisation is accurate, reliable, and trusted for critical business decision-making. … platforms. You'll have excellent SQL and Python skills coupled with experience working in modern cloud-based data environments. Hands-on experience with data observability tools such as Grafana, Monte Carlo, or Acceldata, and data governance/quality platforms like Informatica, Collibra or Microsoft Purview is highly desirable. Experience within ...

Director, Cloud Infrastructure

Location
Greater London, England, United Kingdom
Cloud Infrastructure to lead that work. This person will own the platform foundations Sanity engineers build on every day: cloud infrastructure, Kubernetes, networking, routing, observability, CI/CD, deployment paths, incident response, and the standards that make production ownership work across product teams. The scale is real. Content Lake alone … edge, gateway, caching, object storage, and GCP infrastructure. Some of the work is already in motion: moving Varnish and Mead onto Fastly, tightening observability, finishing our developer on‐call rollout, and making production readiness a normal part of shipping. This is a leadership role for someone who can still ...

Software Engineer, Backend (BI Platform)

Location
Greater London, England, United Kingdom
infrastructure our BI platforms run on. This is a hands-on software and platform engineering role: services and integrations, identity and access automation, observability, reliability and developer experience. BI is the domain you serve (primarily Looker and Sigma, and the systems around them), but the problems are backend and platform … Build the integrations between BI platforms and the wider ecosystem: warehouses, identity providers, orchestration, CI/CD and internal developer tooling Improve the availability, observability, performance, scalability and security of the systems you own, and carry them in production Use AI tooling as part of how you engineer, and where ...

Infrastructure & Developer Platform Lead

Location
Greater London, England, United Kingdom
cloud architecture strategies into implementation plans and engineering roadmaps. Build and maintain reliable, secure and self‐service engineering environments.Drive automation across provisioning, deployment, testing, observability and operational processes. Ensure infrastructure meets agreed requirements for security, operability, compliance and resilience.Improve developer productivity through reusable platform services, engineering tooling and automation. Support … native software delivery through platform capabilities, governance, automation and operational controls.Drive reliability, security, observability and operational readiness across platform services. Manage technical risks, infrastructure incidents, dependencies and operational blockers.Provide technical leadership, mentoring and engineering guidance while remaining actively involved in platform engineering decisions. Build platform capability, operational knowledge and engineering ...

Senior Vice President, Production Services Application Support

Hiring Organisation
The Bank of New York Mellon
Location
London, UK
Employment Type
Full-time
escalation management processes to reduce repeat issues, control gaps, and operational risk while improving platform stability. Champion an automation-first mindset, leveraging AI, observability, and modern engineering practices to reduce manual effort and improve service quality. Provide strategic leadership across EFX platforms, ensuring strong understanding of front-to-back trade … stability, reduction of recurring issues, and strong service recovery execution in high-pressure environments. Demonstrated automation-first and AI-enabled mindset, with experience driving observability, tooling, operational analytics, and process automation initiatives. Strong ownership mentality and execution focus, with the ability to take strategic initiatives from concept through delivery while ...

Infrastructure & Developer Platform Lead

Hiring Organisation
Vodafone
Location
London, UK
Employment Type
Full-time
architecture strategies into implementation plans and engineering roadmaps. Build and maintain reliable, secure and self-service engineering environments. Drive automation across provisioning, deployment, testing, observability and operational processes. Ensure infrastructure meets agreed requirements for security, operability, compliance and resilience. Improve developer productivity through reusable platform services, engineering tooling and automation. … Support AI-native software delivery through platform capabilities, governance, automation and operational controls. Drive reliability, security, observability and operational readiness across platform services. Manage technical risks, infrastructure incidents, dependencies and operational blockers. Provide technical leadership, mentoring and engineering guidance while remaining actively involved in platform engineering decisions. Build platform capability ...

AI Platform & Site Reliability Engineering Consultant

Hiring Organisation
Akkodis
Location
City of London, London, United Kingdom
Employment Type
Permanent
Salary
£88000 - £96000/annum
include: Define and embed SRE engagement models aligned to modern engineering and traditional ITSM/ITIL practices Establish SLIs, SLOs, and Error Budgets Shape observability strategies using metrics, logs, and traces Design incident response models and post-incident learning loops Reduce toil through automation and engineering excellence Deliver SRE capability … Looking For Extensive experience in SRE, cloud operations, or DevOps Proven consulting or advisory background Experience with AWS, Azure, or GCP Strong observability and incident management expertise Ability to obtain UK SC clearance Modis International Ltd acts as an employment agency for permanent recruitment and an employment business ...

Senior Engineering Manager, Developer Experience

Location
Greater London, England, United Kingdom
About the team The Developer Experience team owns the internal platform that every the company engineer touches daily: CI/CD pipelines, observability tooling, our developer portal, and an emerging AI platform. It's a high-visibility role: the work you lead directly shapes the productivity of hundreds of engineers … looks like at the company as we scale. What you'll do Lead and develop a growing team of 5+ highly motivated engineers across observability, CI/CD, developer portal (Backstage), and FinOps tooling — setting clear priorities and establishing strong ways of working. Own and evolve the technical roadmap across ...

Senior Director Release Management London, England, United Kingdom Research and Development

Location
Greater London, England, United Kingdom
Engineering, DevOps, Security, Product, Quality, Cloud Operations and engineering leadership, the Head of Release Management will drive the adoption of automated readiness checks, environment observability, telemetry-led insights, predictive risk management and modern deployment practices. A key part of the role is building a release management capability that acts … across complex product, platform or service environments. Strong knowledge of CI/CD, progressive delivery, feature flags, deployment automation, policy-as-code, automated compliance, observability, SRE principles and modern engineering practices. Experience defining release approaches for planned, emergency, security, platform, service and out-of-band releases. Experience with environment strategy ...

Lead AI Engineer IRC302971

Location
Greater London, England, United Kingdom
clients transform their digital platforms and products. As an AI Engineer, you’ll own the development of the core agentic frameworks, evaluation pipelines, and observability tooling that let it operate with safety, trust, and intelligence at scale. You won’t just be integrating AI into a product … CrewAI is a plus. Production Python engineering : Clean, modular, testable, maintainable code — you care about system reliability as much as model output. Evaluation & observability : Practical experience instrumenting tracing (LangSmith, Arize) and building CI/CD pipelines built specifically for LLMs. Engineering foundations : Distributed systems, scalable data pipelines (Kafka ...

Software Engineer III - Fullstack (Java, React, Python and AI) Engineer

Location
Greater London, England, United Kingdom
with product, marketing, and partners to translate requirements into well-designed technical solutions Improve engineering excellence through code reviews, test automation, CI/CD, observability, performance tuning, and operational best practices Ensure solutions meet security, privacy, and compliance expectations including consent management, data minimization, and access controls Required Qualifications, Capabilities … with campaign management and attribution systems Background in AI-enabled content generation, experimentation, or workflow automation Skills in test automation, CI/CD, and observability tools Experience optimizing performance and scalability Familiarity with consent management and data minimization practices Ability to drive innovation in personalization and measurement Experience working ...

Network Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
responsibilities of the role include: Owning the operational performance, stability, availability and security posture of network and network security platforms end to endEnhancing observability across telemetry, alerting and performance monitoring using tools such as Grafana, Prometheus and ThousandEyes to enable proactive operationsImplementing and maintaining scalable, resilient datacentre network infrastructure built … looking for? The ideal candidate will have the following experience: Strong background as a Network Engineer in enterprise or large-scale environmentsExperience applying SRE, observability and automation principles to networking, using technologies such as Python, Prometheus, Grafana, OpenTelemetry, Ansible and JenkinsExperience with Cisco and Arista switching and routing, alongside network ...

Senior UI Platform Engineer, London

Location
Greater London, England, United Kingdom
developer workflows· Help define standards for application structure, testing, deployment, and supportability· Partner with other engineering teams on shared concerns such as authentication, authorization, observability, and frontend/backend integration patterns· Contribute to internal platform applications such as onboarding, access administration, and service discovery experiences· Evaluate and maintain third-party … internal tooling· Experience with Vite, Next.js, or both· Experience with UI/component libraries such as Ant Design, AG Grid, or similar· Experience with observability tooling such as Datadog, OpenTelemetry, or Grafana· Experience with Microsoft Entra ID/Azure AD or similar identity platforms· Experience publishing and maintaining internal ...

Lead Engineer - Network Connectivity

Location
Greater London, England, United Kingdom
enterprise connectivity platforms, including WAN, LAN, Wi‐Fi, SD‐WAN, SASE and cloud connectivity services. The role provides technical leadership across network engineering, automation, observability, operational resilience, standards governance and service lifecycle management for a large‐scale multi‐site enterprise environment. It combines deep technical expertise with platform ownership, service … tooling, configuration management and operational telemetry Cloud networking and hybrid connectivity across Azure, GCP and enterprise environments Experience designing resilient network services for automation, observability and supportability Technical governance, standards development and engineering best practice Ability to translate business and operational requirements into scalable engineering solutions Strong stakeholder management ...

Full-Stack Software Engineer, Model Development Platform London, United Kingdom Sunnyvale, California USA

Location
Greater London, England, United Kingdom
data models that power them. You will take ownership of significant technical areas, from understanding user needs and shaping solutions through to implementation, deployment, observability, and production support. This is a hands-on engineering role with meaningful technical leadership responsibilities. You will influence technical decisions within the team, establish effective … APIs, and backend services that perform reliably as usage and complexity grow. Apply sound security practices, identify bottlenecks and failure modes, and use testing, observability, and performance analysis to continuously improve the systems you own. Shape system design and technical direction Lead the design of significant features and services, making ...

Lead Platform Engineer

Location
Greater London, England, United Kingdom
quantum resources. Architect capability matching, task orchestration, fallback mechanisms, policy enforcement and backend adapters across the platform. Set engineering standards for reliability, security, testing, observability and extensibility while remaining highly hands‐on with implementation. Lead integrations with Dell/AMD infrastructure, Databricks, QPU systems and customer environments. Coordinate technical contributors … Strong experience designing APIs, containerised systems and event‐driven architectures. Experience with Kubernetes, Slurm or comparable workload orchestration technologies. Strong understanding of security and observability within distributed platforms. Proven ability to design stable interfaces across heterogeneous compute or hardware resources. The technical depth to set architectural direction while remaining highly ...

Network Engineer

Location
Greater London, England, United Kingdom
responsibilities of the role include: Owning the operational performance, stability, availability and security posture of network and network security platforms end to end Enhancing observability across telemetry, alerting and performance monitoring using tools such as Grafana, Prometheus and ThousandEyes to enable proactive operations Implementing and maintaining scalable, resilient datacentre network … ideal candidate will have the following experience: Strong background as a Network Engineer in enterprise or large-scale environments Experience applying SRE, observability and automation principles to networking, using technologies such as Python, Prometheus, Grafana, OpenTelemetry, Ansible and Jenkins Experience with Cisco and Arista switching and routing, alongside network security ...

Full-Stack Software Engineer, Model Development Platform

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
data models that power them. You will take ownership of significant technical areas, from understanding user needs and shaping solutions through to implementation, deployment, observability, and production support. This is a hands-on engineering role with meaningful technical leadership responsibilities. You will influence technical decisions within the team, establish effective … APIs, and backend services that perform reliably as usage and complexity grow. Apply sound security practices, identify bottlenecks and failure modes, and use testing, observability, and performance analysis to continuously improve the systems you own. Shape system design and technical directionLead the design of significant features and services, making thoughtful ...

Analytics Engineer

Location
Greater London, England, United Kingdom
work, and we focus less on hand-writing code and more on framing problems, specifying intent, and owning the trust infrastructure (testing, contracts, observability) that the business relies on. About the Role We're looking for an Analytics Engineer who's genuinely excited about how this craft is changing. … build data models with the support of AI-assisted workflows, and you'll own a growing share of the testing, data contracts, and observability that let the rest of the business rely on our data with confidence. You'll work closely with senior engineers and stakeholders, with plenty of room ...

Senior Data Engineer - Data Science Platform

Hiring Organisation
ASOS
Location
London, UK
Employment Type
Full-time
Spark and Databricks. Owning and evolving platform components that help engineers and data scientists build, test, deploy and monitor data products. Improving platform reliability, observability, data quality and operational excellence. Creating reusable libraries, tooling and engineering patterns that enable teams to deliver data products more efficiently. Partnering with Data Scientists … Scala. Experience designing data architectures that balance scalability, reliability and cost efficiency. Experience implementing modern engineering practices, including CI/CD, automated testing, observability and Infrastructure as Code. Demonstrated ability to solve complex engineering problems and improve platform capabilities that help other teams work more effectively. Experience leading the design ...

Principal Software Engineer (Hybrid) New London, England, United Kingdom

Location
Greater London, England, United Kingdom
other teams shift faster and more reliably Taken AI and agentic systems from prototype to production, with a strong understanding of infrastructure, authentication, and observability Delivered on high-stakes consulting engagements across multiple language paradigms, stacks, ecosystems, and client industries Built high-quality, maintainable software collaboratively, incrementally, and through … legacy systems with short and long-term business needs Led and delivered solutions to architecture-level problems including scalability, security, reliability, performance, maintainability, and observability Facilitated alignment across technical and non-technical stakeholders to move initiatives forward through ambiguity and complexity Provided mentorship and team support at scale, while sharing ...

Senior Pre-Sales Consultant, Enterprise Payments

Location
City Of London, England, United Kingdom
workshops, both virtually and on-site. Lead and contribute to RFI and RFP responses, providing detailed technical input across areas including: Integration Security Resiliency Observability Compliance Identify risks, gaps, constraints, and assumptions while clearly articulating solution trade-offs. Plan, design, and support Proof of Concept (POC) activities, enabling customers … identity frameworks, including OpenID Connect and OAuth Cloud platforms such as: AWS Microsoft Azure Google Cloud Platform (GCP) IBM Cloud Oracle Cloud Monitoring and observability tools, including Splunk and SIEM solutions Pre-Sales and Customer-Facing Experience Proven experience in a customer-facing technical role, solution architecture role ...

Senior Java Developer IRC303859

Location
Greater London, England, United Kingdom
core engineering principle. You’ll be comfortable designing and delivering robust, scalable and resilient distributed systems, with a strong focus on reliability, observability and operational excellence. Exposure to agentic coding practices where it adds value, would be hugely beneficial You’ll also need: Strong experience building robust, resilient and defensively … coded distributed systems, ideally leveraging technologies such as Kafka and an understanding of designing and operating RESTful APIs at scale, including versioning, observability, monitoring, and forward, backward compatibility Experience with event driven architectures, both orchestrated and choreographed is expected Experience working effectively in teams that value autonomy and empowerment, where ...

Senior Full Stack Engineer - Vialto Rewards -

Hiring Organisation
Vialto Partners
Location
London, UK
Employment Type
Full-time
over tax and payroll guidance. Evaluate, integrate, and operationalize AI APIs, models, and toolchains within application architectures, ensuring responsible AI usage, cost management, and observability—especially where AI touches financially material outputs. Mentor and support other engineers through code reviews, pairing, and technical guidance on payroll, compensation, and equity domain … architecture, performance and accessibility Backend: C#, ASP.NET, WebAPI, RESTful design, authentication and authorization patterns, high-integrity calculation services Cloud: Azure PaaS services, observability, secure networking, identity, deployment automation Domain: Compensation Collection workflows, Shadow Payroll calculation and reporting, Equity event processing and tax treatment, cross-border payroll integrations Engineering practices ...