25 of 25 Reliability Engineer Jobs in London

Senior Network Reliability Engineer

Location
Wimbledon, England, United Kingdom
Senior Network Reliability Engineer Permanent Position Hybrid Role from our Wimbledon Office with On Call Work Domestic & General is growing fast across 12 countries, protecting the appliances that keep households running and cutting waste by extending product life. Recognised by Great Place to Work and the Inclusive Employers … Standard, we're building a modern, technology lead business and we're looking for a Senior Network Reliability Engineer to drive the performance, reliability and security of our hybrid, enterprise network estate. About the Role This is a hands-on senior engineering position with responsibility ...

Product Reliability Engineer – APM

Location
Greater London, England, United Kingdom
## Product Reliability Engineer - APMLondon, UK · Full-time#### About The PositionCoralogix is a modern, full-stack observability platform transforming how businesses process and understand their data. Our unique architecture powers in-stream analytics without reliance on expensive indexing or hot storage. We specialize in comprehensive monitoring of logs … such as APM, RUM, SIEM, Kubernetes monitoring, and more, enhancing operational efficiency and reducing observability spending by up to 70%.We seek a **Product Reliability Engineer** who ensures that the Coralogix APM Product and Process exceed the quality and reliability standards, establish a competitive edge, and prevent ...

Senior Hybrid Azure Network Reliability Engineer

Location
Wimbledon, England, United Kingdom
Domestic & General Group is seeking a Senior Network Reliability Engineer for a permanent hybrid role based around our Wimbledon office. You will own the reliability, performance and security of our hybrid Azure network, lead major incident resolution, and drive automation and observability across our estate. ...

Remote Senior Cloud Reliability Engineer - AWS & Kubernetes

Location
Greater London, England, United Kingdom
Salve.Inno Consulting is seeking a Senior Cloud Reliability Engineer to own highly available, cloud-native production environments on AWS and Kubernetes. You will handle on-call incidents, lead root-cause analyses, and drive permanent improvements, using Terraform and GitOps to automate infrastructure. The role emphasizes production readiness, security ...

APM Platform Reliability Engineer - London

Location
Greater London, England, United Kingdom
Coralogix is hiring a Product Reliability Engineer in London to ensure the APM product and processes meet high reliability standards. You will help reduce engineering interruptions, improve customer satisfaction, and drive product quality through benchmarking and knowledge sharing. The role requires PromQL experience, SaaS metrics background ...

Principal Cloud Reliability Engineer Multi-Cloud

Location
Greater London, England, United Kingdom
Veson Nautical is seeking a Principal Site Reliability Engineer to design, build, and operate scalable cloud infrastructure across AWS and GCP. The role focuses on growing the GCP footprint, with cross-region and cross-cloud alignment, in a hybrid London-based environment. You will lead greenfield projects, improve … existing estates, and influence architectural direction while mentoring peers and collaborating with development teams to ensure reliability and cost efficiency. #J-18808-Ljbffr ...

Senior M&E Reliability Engineer – EMEA Travel

Location
Greater London, England, United Kingdom
Nscale, the GPU cloud for AI, seeks a Reliability Engineer (Mechanical & Electrical) to support its EMEA data centre estate. You’ll be the on-site technical escalation point for M&E systems and help set maintenance baselines. This UK role requires regular travel across Europe/Nordics, with … focus on high-density AI cooling and power. You’ll perform design reviews, audits, RCA/CAPA, and coach site engineers to improve reliability and energy efficiency. #J-18808-Ljbffr ...

Reliability Engineer SME (M&E)

Location
Greater London, England, United Kingdom
work. If you join our team, you’ll be contributing to building the technology that powers the future. About the Role (Job Purpose) The Reliability Engineer (Mechanical & Electrical) provides cross-discipline M&E engineering expertise and hands‐on technical support across Nscale's EMEA data centre estate — owned … supporting high-density AI compute. As GPU rack densities and power draw increase, cooling performance and power resilience become two of the most critical reliability factors across the estate — and neither can sensibly be engineered in isolation from the other. What You’ll be Doing (Responsibilities) M&E Technical ...

SRE Engineer – FinTech Reliability, Observability & Cloud

Location
Greater London, England, United Kingdom
Hamilton Barnes Associates Limited is seeking a Site Reliability Engineer to work at the intersection of software engineering and infrastructure. You'll develop internal platforms, tooling, and automation across Linux, distributed systems, and cloud-native technologies to improve reliability and operational efficiency in a global production environment. ...

Senior SRE Engineer: Reliability, Cloud & Automation

Location
Greater London, England, United Kingdom
London Stock Exchange Group is looking for a Senior Engineer in Site Reliability who will join a driven team focused on system availability, performance, and scalability. Responsibilities include maintaining service level objectives, writing automation for system resilience, and partnering with development teams. Required qualifications include a Bachelor … computer science, experience in Object Oriented programming and cloud systems, and DevOps familiarity. The role is pivotal in ensuring 24/7 system reliability and promoting engineering best practices. #J-18808-Ljbffr ...

AI Platform Reliability Engineer

Location
Greater London, England, United Kingdom
Elavon, a global payments leader, is seeking an AI DevOps Engineer to own the deployment, operations, and reliability of the AI‐SDLC platform as it moves to a centralized hosted service. You will build and maintain infrastructure, CI/CD pipelines, and observability, ensure enterprise SLAs and security … support multi‐team adoption across products. Candidates bring strong Terraform, AWS, Docker/Kubernetes, OpenTelemetry, and scripting skills, with a focus on reliability and cost efficiency. #J-18808-Ljbffr ...

Remote Senior Production Reliability Engineer

Location
Greater London, England, United Kingdom
Unlimited Inc. in London is seeking a Senior Technical Support Specialist to own the health, reliability, and trustworthiness of production agentic AI systems. This is a senior, incident-driven role that operates at the intersection of AI behavior, workflows, and enterprise integrations. You will diagnose complex multi-agent failures ...

Senior Backend Engineer: Chaos & Reliability (Remote)

Location
Greater London, England, United Kingdom
Camunda seeks a Senior Software Engineer, Backend, to own automated reliability testing and chaos engineering for Camunda 8. You will break things in safe environments to strengthen the platform, guide product direction, and collaborate with friendly colleagues who live our FAITH values. You’ll design and run chaos ...

Senior Software Engineer II — Reliability & Observability

Location
Greater London, England, United Kingdom
company is hiring a Senior Software Engineer II to join the OPX team within the Developer Experience organization. You will design and build automated reliability and self-healing systems at scale, delivering platform tooling that engineers across the company adopt for their services. You will own incident management ...

Senior DevOps Engineer: Platform & Reliability Lead

Location
Greater London, England, United Kingdom
Auriga is seeking a Senior DevOps Engineer to architect, lead and scale the platform engineering and operational backbone of our solutions. You will define and maintain roadmaps for infrastructure, CI/CD, observability, reliability and security, across cloud, on-premises, hybrid and customer-site environments. The role requires ...

Senior Infrastructure Engineer, AI-Driven Reliability

Location
Greater London, England, United Kingdom
United States Digital Space LLC is seeking an Infrastructure Engineer to join our hybrid team in London or New York. You’ll design scalable, reliable systems across AWS, GCP, and Azure, and automate critical tasks with Python or Go. You’ll own on-call, incident response, and platform reliability ...

Infra Engineer: AI-Powered Reliability & Scale

Location
Greater London, England, United Kingdom
WRITER is hiring an Infrastructure Engineer to own the reliability and performance of our enterprise-grade AI platform. This hybrid role can be based in London or New York, reporting to the Director of Engineering. You will build scalable infrastructure, automate across the stack, and lead incident response ...

GPU Infrastructure Engineer — Scale & Reliability for Production

Location
Greater London, England, United Kingdom
OpenAI in London seeks a software engineer to build and operate large-scale production systems powering ChatGPT. You’ll develop tooling for fleet health, capacity planning, automation, and incident response, collaborating with infra, research, and product teams to improve reliability and compute utilization. The role suits engineers ...

Senior Dependency Tracing Engineer - Cloud Reliability

Location
Greater London, England, United Kingdom
United States Digital Space LLC seeks an experienced software engineer to join our team building next-generation technologies. You will work on high-impact projects that affect billions of users, spanning distributed systems, data storage, and security, with opportunities to switch teams as our fast-paced business grows. ...

DevEx Platform Engineer — Build Tooling & Reliability

Location
Greater London, England, United Kingdom
seeking a Developer Experience engineer in London to own end-to-end service delivery for internal developer platforms. You will work with US/UK teams on tooling, observability, and workflows to empower engineers across the firm. Role requires hands-on experience with Kubernetes, CI/CD, and infrastructure ...

Senior API Enterprise Engineer — Scale, Governance & Reliability

Location
Greater London, England, United Kingdom
DeepL, a global AI translation leader, seeks an experienced engineer to drive enterprise API adoption and governance. You will own features that enable large organizations to use APIs securely, with strong reliability and scalable backend design. You will collaborate across teams to ensure data residency, RBAC, and compliance … while delivering four-nines reliability and robust testing. Hybrid work and global collaboration are core to the role. #J-18808-Ljbffr ...

Senior ML Engineer - Financial AI & Model Reliability

Location
Greater London, England, United Kingdom
will build rigorous benchmarks, fine-tune open models for financial guidance, and investigate why models fail on numbers and layered rules to improve reliability and auditability. #J-18808-Ljbffr ...

Azure DevOps Engineer — Cloud Automation & Reliability

Location
Greater London, England, United Kingdom
United States Digital Space LLC is seeking a Microsoft DevOps Engineer to design, deploy and support secure, automated Azure platforms. You will own CI/CD pipelines, IaC, cloud operations and monitoring, collaborating with multiple teams to accelerate reliable technology delivery. The role emphasizes automated deployment, security by design ...

Senior Software Engineer — Hybrid, Cloud, Reliability & Security

Location
Greater London, England, United Kingdom
United Kingdom is seeking a Senior Software Engineer to own and improve critical digital services, collaborate with engineering, QA and client teams, and deliver resilient, high‐performance solutions. The role blends hands‐on development with service improvement in a hybrid working arrangement and may require UK security clearance. ...

Staff Data Platform Engineer — Scale + Reliability

Location
Greater London, England, United Kingdom
A cloud infrastructure company in Greater London is looking for a Staff+ level engineer. The role involves defining and leading the database strategy, influencing production infrastructure design, and working with various data technologies such as ...