26 to 50 of 71 Reliability Engineer Jobs in the UK

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. … enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. As a Lead Software Engineer at JPMorgan Chase in the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability ...

AI Platform Reliability Engineer

Location
Greater London, England, United Kingdom
Elavon, a global payments leader, is seeking an AI DevOps Engineer to own the deployment, operations, and reliability of the AI‐SDLC platform as it moves to a centralized hosted service. You will build and maintain infrastructure, CI/CD pipelines, and observability, ensure enterprise SLAs and security … support multi‐team adoption across products. Candidates bring strong Terraform, AWS, Docker/Kubernetes, OpenTelemetry, and scripting skills, with a focus on reliability and cost efficiency. #J-18808-Ljbffr ...

AI/ML Platform Reliability Engineer III

Location
Glasgow, Scotland, United Kingdom
JPMorganChase is seeking a Software Engineer III within the AI/ML Data Platforms organization to enhance reliability and scalability of AI/ML platforms. You will design, build, and maintain production-grade code and tooling, and collaborate across teams to deliver high-impact AI capabilities with operational ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. … enjoy solving hard production problems and making platforms measurably better, you'll find meaningful impact and growth here. As a Senior Lead Software Engineer at JPMorganChase within the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability ...

Senior Platform Engineer Reliability & Automation (Hybrid)

Location
Wales, United Kingdom
Centerprise International Limited in the United Kingdom is seeking a Principal Platform Engineer to own day-to-day reliability, operability and continual improvement of the Centerprise platform estate. You will lead a small engineering team, establish standards and ensure platform services are stable and efficient to support customers. ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Location
Glasgow, Scotland, United Kingdom
reliably in production at scale. In this role, you’ll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You’ll work hands-on with cloud and Kubernetes-based deployments, deep observability, and cost-aware performance tuning. … enjoy solving hard production problems and making platforms measurably better, you’ll find meaningful impact and growth here. As a Senior Lead Software Engineer at JPMorganChase within the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability ...

Staff Software Engineer (Reliability & Platform)

Location
United Kingdom
while remaining capable across the stack (backend, APIs, cloud, and operations). You will play a key role in maintaining product velocity, improving system reliability, and reducing operational burden. This is a hands‐on role that includes on‐call responsibilities and requires a strong commitment to modern engineering practices … design, build, test, deploy, and operate — rather than handing work off between silos. In this role, you will focus on improving production operability and reliability by debugging complex issues, strengthening observability, and eliminating recurring problems at the source. Responsibilities Own production operability by debugging complex issues, improving system visibility ...

Remote Senior Platform Reliability Engineer

Location
United Kingdom
Megaport is seeking a Senior Platform Engineer to champion DevOps and SRE culture, ensuring secure, maintainable, and highly available systems across multiple time zones. You will engage with stakeholders to gather requirements and demonstrate solutions, while hands-on engineering and continuous learning drive customer success and company goals. … will work on improving production reliability across an SRE‐scoped team, promote best practices, and contribute to the evolution of our #J-18808-Ljbffr ...

Remote Senior Platform Reliability Engineer

Location
United Kingdom
Megaport is seeking a Senior Platform Engineer to champion DevOps and SRE practices across our global team. You will work hands-on to keep systems secure, reliable and scalable, coordinating with stakeholders in multiple timezones. You will engage in on-call rotation, incident reviews, and evolve solutions through peer ...

Senior Reliability Engineer – Platform & Observability

Location
United Kingdom
Identity’s OneLogin team is seeking Senior Software Engineers who own production systems and drive reliability, observability, and operability across the stack. You’ll tackle complex issues, design for resilience, and apply AI-assisted development approaches to accelerate debugging and root-cause analysis. The role emphasizes ...

Remote Senior Production Reliability Engineer

Location
Greater London, England, United Kingdom
Unlimited Inc. in London is seeking a Senior Technical Support Specialist to own the health, reliability, and trustworthiness of production agentic AI systems. This is a senior, incident-driven role that operates at the intersection of AI behavior, workflows, and enterprise integrations. You will diagnose complex multi-agent failures ...

Platform Reliability Engineer — Full-Stack (Python + React)

Location
East Midlands, England, United Kingdom
Sperry is hiring a versatile Full Stack Engineer to build and run the platform behind our inspection products. You will work across the stack—Python services and AWS infrastructure at the back, React at the front—moving between projects as priorities shift, and you will own the reliability ...

DRAM Test & Reliability Engineer — Equity

Location
Bristol, England, United Kingdom
Fractile is seeking a DRAM Product Engineer to design, implement, and support production-ready test solutions for a high-volume board integrating multiple AI accelerators and PCIe Gen6 interconnect. You will develop robust tests, diagnose DRAM issues, and collaborate across design, architecture, and manufacturing to drive memory test innovation. ...

SNR Reliability Maintenance Engineer

Hiring Organisation
Amazon TA
Location
Derby, England, United Kingdom
Reliability Maintenance Engineering (RME) team is central to Amazon's commitment to innovation. As Amazon evolves and adapts, this team makes sure that the tools and technologies we use do as well. As a Senior RME Technician, you'll help us stay one step ahead, adopting the latest technologies … tasks and schedule additional servicing when required - Supervise technicians on shift to support their development and act as the first point of contact for Reliability Maintenance Engineers - Solve issues in equipment to reduce downtime for operations so they can process packages as quickly as possible - Support in finding ways ...

Full Stack Engineer, Platform Reliability

Location
East Midlands, England, United Kingdom
Role Summary We are hiring a versatile Full Stack Engineer to build and run the platform behind Sperry's inspection products. You will work across the stack – Python services and AWS infrastructure at the back, React at the front – moving between projects as priorities shift, and you will … reliability and resilience of what we ship alongside the features you add to it. We have put build and run in the same seat on purpose. The engineers who keep a platform dependable are the ones closest to how it is actually used, and that puts you in front ...

Software Engineer III - AI ML Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Bournemouth, Dorset, South West, United Kingdom
Employment Type
Permanent
artificial intelligence and machine learning - and we need engineers like you to help us do it reliably, securely, and at scale. As a Software Engineer III at JPMorganChase within the AI/ML Data Platforms organization, you will be a key member of the Reliability Engineering team, contributing … design and delivery of trusted, market-leading technology products. You will apply your technical expertise and problem-solving skills to enhance the reliability and scalability of AI/ML platforms, build reusable services and tooling, and partner across teams to unblock high-impact AI use cases. This ...

Software Engineer III - AI/ML Platform Reliability

Location
Glasgow, Scotland, United Kingdom
artificial intelligence and machine learning - and we need engineers like you to help us do it reliably, securely, and at scale. As a Software Engineer III at JPMorganChase within the AI/ML Data Platforms organization, you will be a key member of the Reliability Engineering team, contributing … design and delivery of trusted, market-leading technology products. You will apply your technical expertise and problem-solving skills to enhance the reliability and scalability of AI/ML platforms, build reusable services and tooling, and partner across teams to unblock high-impact AI use cases. This ...

Senior Backend Engineer: Chaos & Reliability (Remote)

Location
Greater London, England, United Kingdom
Camunda seeks a Senior Software Engineer, Backend, to own automated reliability testing and chaos engineering for Camunda 8. You will break things in safe environments to strengthen the platform, guide product direction, and collaborate with friendly colleagues who live our FAITH values. You’ll design and run chaos ...

Senior Software Engineer II — Reliability & Observability

Location
Greater London, England, United Kingdom
company is hiring a Senior Software Engineer II to join the OPX team within the Developer Experience organization. You will design and build automated reliability and self-healing systems at scale, delivering platform tooling that engineers across the company adopt for their services. You will own incident management ...

Senior DevOps Engineer: Platform & Reliability Lead

Location
Greater London, England, United Kingdom
Auriga is seeking a Senior DevOps Engineer to architect, lead and scale the platform engineering and operational backbone of our solutions. You will define and maintain roadmaps for infrastructure, CI/CD, observability, reliability and security, across cloud, on-premises, hybrid and customer-site environments. The role requires ...

Senior Reliability & Maintenance Engineer — Lead Technician

Location
United Kingdom
Senior Reliability Maintenance Engineering Technician, RME Job ID: 10428314 | Amazon UK Services Ltd. - A10 Our Reliability Maintenance Engineering (RME) team is central to Amazon’s commitment to innovation. As Amazon evolves and adapts, this team makes sure that the tools and technologies we use do as well. … tasks and schedule additional servicing when required Supervise technicians on shift to support their development and act as the first point of contact for Reliability Maintenance Engineers Solve issues in equipment to reduce downtime for operations so they can process packages as quickly as possible Support in finding ways ...

Network Engineer - Network Source of Truth, Automation & Network Reliability

Location
Cambridge, England, United Kingdom
Role: Senior Network Engineer ( Network Source of Truth, Automation & Network Reliability ) Employment: Contract - Inside IR35 Location: Cambridge,UK - Hybrid Required Technical Skills Network Source of Truth, IRM, IPAM & DCIM Hands-on experience with one or more of: NetBox, Nautobot, IP Fabric, Auvik, SolarWinds, Device42 or Infoblox. Core Networking ...

Senior Infrastructure Engineer, AI-Driven Reliability

Location
Greater London, England, United Kingdom
United States Digital Space LLC is seeking an Infrastructure Engineer to join our hybrid team in London or New York. You’ll design scalable, reliable systems across AWS, GCP, and Azure, and automate critical tasks with Python or Go. You’ll own on-call, incident response, and platform reliability ...

Infra Engineer: AI-Powered Reliability & Scale

Location
Greater London, England, United Kingdom
WRITER is hiring an Infrastructure Engineer to own the reliability and performance of our enterprise-grade AI platform. This hybrid role can be based in London or New York, reporting to the Director of Engineering. You will build scalable infrastructure, automate across the stack, and lead incident response ...

Senior AI Infra Engineer – LLM Ops & Reliability

Location
Auchentibber, Scotland, United Kingdom
JPMorganChase is seeking a Senior Lead Software Engineer to shape reliable AI infrastructure at scale. You will own LLM serving stacks, optimize performance and costs, and lead incident response across cloud and on-prem GPU clusters. The role emphasizes secure software delivery, observability, and strong engineering fundamentals. You will … collaborate with engineering to deliver scalable AI platforms, drive reliability, and build reusable patterns for robust production deployments, with a focus #J-18808-Ljbffr ...