26 to 50 of 70 Senior Reliability Engineer Jobs in the UK

Remote Senior Production Reliability Engineer

Location
Greater London, England, United Kingdom
Unlimited Inc. in London is seeking a Senior Technical Support Specialist to own the health, reliability, and trustworthiness of production agentic AI systems. This is a senior, incident-driven role that operates at the intersection of AI behavior, workflows, and enterprise integrations. You will diagnose complex multi ...

Senior AI Infra & LLM Reliability Engineer

Location
Glasgow, Scotland, United Kingdom
J.P. Morgan is seeking a Senior Lead Software Engineer to shape reliable AI production systems. You will build and operate large language model serving infrastructure, with cloud and Kubernetes deployments, deep observability, and cost-aware tuning. Own reliability, performance, and security across end-to-end LLM endpoints ...

Senior Platform Engineer Reliability & Automation (Hybrid)

Location
Wales, United Kingdom
Centerprise International Limited in the United Kingdom is seeking a Principal Platform Engineer to own day-to-day reliability, operability and continual improvement of the Centerprise platform estate. You will lead a small engineering team, establish standards and ensure platform services are stable and efficient to support customers. … This senior, hands-on role blends technical leadership with hands-on design, operation and driving automation across platform services. #J-18808-Ljbffr ...

Senior Kubernetes & Platform Reliability Engineer (Remote)

Location
Greater London, England, United Kingdom
Range is seeking a senior infrastructure and operations engineer to own and evolve platform reliability. You will design, operate, and maintain a Kubernetes-based infrastructure, build reliable monitoring and alerting pipelines, and ensure system stability under real-world load and failure conditions. This hands-on role requires deep ...

Senior Software Engineer II — Reliability & Observability

Location
Greater London, England, United Kingdom
company is hiring a Senior Software Engineer II to join the OPX team within the Developer Experience organization. You will design and build automated reliability and self-healing systems at scale, delivering platform tooling that engineers across the company adopt for their services. You will own incident ...

Senior Platform Reliability Engineer – Onsite Birmingham

Location
Birmingham, England, United Kingdom
Infused Solutions Ltd. is seeking a Senior Full Stack Developer to join our Birmingham-based team onsite. You will own and improve large-scale SaaS platforms, focusing on stability, performance, and engineering excellence across the product. You’ll resolve complex production issues, profile and optimise across multi-layer stacks ...

Senior Data Centre Ops & Reliability Engineer

Location
Newcastle upon Tyne, England, United Kingdom
Pulsant in the United Kingdom is seeking a Senior Data Centre Services Engineer to support critical infrastructure and mentor the Data Centre Services team. You will deputise for the Data Centre Manager and help ensure high-availability services that power digital workloads for clients. Based in Newcastle with ...

Senior Software Engineer, Reliability & Large-Scale Infra

Location
Greater London, England, United Kingdom
Google is seeking software engineers to join the Dependency Tracing Team within Platform Reliability Engineering to help minimize outages and improve cloud infrastructure reliability. The role emphasizes building scalable systems, collaborating across teams, and driving reliability across Google Cloud Platform. As part of Google’s technical workforce ...

Senior Cloud Reliability & Operations Engineer

Location
Belfast City District, Northern Ireland, United Kingdom
Oracle Cloud is seeking a Senior Technical Operations Engineer to operate production environments, including systems and databases, and to drive reliability and performance for critical workloads. You will guide junior engineers, participate in large-scale incident bridges, and help build new processes and procedures to keep Oracle ...

Senior Cloud Reliability Engineer - Hybrid Role

Location
Southampton, England, United Kingdom
NiCE Ltd. is seeking an experienced DevOps/SRE to manage production systems, automate platform infrastructure, and improve reliability across a suite of distributed applications. You will monitor health, tune performance, and partner with development teams to streamline releases in a hybrid work model. Candidates should have ...

Senior Cloud Reliability Engineer - Hybrid Role

Location
Greater London, England, United Kingdom
NiCE Ltd. is seeking an experienced DevOps/SRE to manage production systems, automate platform infrastructure, and improve reliability across a suite of distributed applications. You will monitor health, tune performance, and partner with development teams to streamline releases in a hybrid work model. Candidates should have ...

Senior AI/ML Platform Reliability Engineer

Location
Auchentibber, Scotland, United Kingdom
JPMorganChase is seeking a Software Engineer III within the AI/ML Data Platforms group to design and deliver trusted, scalable AI/ML platforms, with a focus on reliability and security. As part of the Reliability Engineering team, you will own non-functional requirements, build tooling ...

Senior Go Engineer: Scale, Automate & Improve Reliability

Location
Greater London, England, United Kingdom
Jobtailor is seeking a senior software engineer in London to drive the evolution of large-scale systems using Go, AWS, and Azure. You will contribute to platform reliability, observability, and scalable delivery while mentoring junior engineers and collaborating across teams. The role emphasizes CI/CD, automation ...

Senior SRE: Hybrid Cloud Reliability Engineer

Location
Horsell, England, United Kingdom
Capgemini is seeking Site Reliability Engineers to grow their careers in a hybrid setup, delivering and maintaining services for UK clients. You’ll work within Pods, focusing on reusable platform blueprints, GitOps, and multi-cloud operations while engaging in on-call rotations and on-site client engagements. You will ...

Senior Performance QA Engineer: Scale & Reliability

Location
Greater London, England, United Kingdom
Chip UK is seeking a Senior QA Engineer specializing in Performance & Scalability to own performance across 30+ microservices and legacy dependencies. You will define pass/fail thresholds, embed performance in delivery plans, and lead performance initiatives beyond gate testing. You’ll design tests, automate, and monitor across ...

Senior Full Stack Engineer — Reliability & On-Call Lead

Location
Leeds, England, United Kingdom
Burendo is seeking an experienced Full Stack Engineer focused on operating and improving production systems. You will join an L2/L3 managed service team supporting business-critical applications across cloud platforms, participating in on-call rotations and leading technical investigations. The role emphasizes production stability, incident response ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Location
Glasgow, Scotland, United Kingdom
reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands‐on with cloud and Kubernetes-based deployments, deep observability, and cost‐aware performance tuning. … enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. As a Senior Lead Software Engineer at JPMorganChase within the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site ...

Senior Software Engineer, Dependency Tracing & Reliability

Location
City of Westminster, England, United Kingdom
Google's Dependency Tracing team seeks a Senior Software Engineer to design and optimize large-scale distributed systems in a cloud context. You will translate complex requirements into robust architectures, mentor juniors, and drive quality through rigorous testing. You will collaborate with cross-functional teams to minimize outages ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. … enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. As a Senior Lead Software Engineer at JPMorganChase within the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Location
Auchentibber, Scotland, United Kingdom
reliably in production at scale. In this role, you’ll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You’ll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. … enjoy solving hard production problems and making platforms measurably better, you’ll find meaningful impact and growth here. As a Senior Lead Software Engineer at JPMorganChase within the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site ...

Senior Data Center Facilities Engineer & Reliability Lead

Location
United Kingdom
Oracle is seeking a senior data center engineering SME in the United Kingdom to lead maintenance and repair of critical infrastructure, from liquid to chip, across global sites. You will drive change management, site governance, and reliability improvements with high autonomy. You will set standards for design, operations ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Milton, Cambridgeshire, UK
reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. … enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. As a Senior Lead Software Engineer at our client within the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Paisley, Renfrewshire, UK
reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. … enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. As a Senior Lead Software Engineer at our client within the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management ...

Senior Software Engineer, Testing & AI-Driven Reliability

Location
United Kingdom
Identity’s OneLogin team seeks a Senior Software Engineer who owns design, build, test, deploy and operate responsibilities with a bias for action. You will shape testing strategy and automation, and contribute to AI-augmented development to maximize confidence and release safety at scale. You will work across … backend, APIs, cloud, and operations, improving system reliability and reducing operational burden while maintaining product velocity in a hands-on role. #J-18808-Ljbffr ...

Senior Data & MLOps Engineer - AI Reliability Platform

Location
Greater London, England, United Kingdom
CoreWeave is recruiting a Senior Data & MLOps Engineer to design and scale the GPU Intelligence Platform infrastructure, building pipelines for data, features, and model training while delivering insights for system health and optimization. You will transition prototypes to production across a fleet, focusing on scalable distributed services, separating ...