76 to 100 of 134 Reliability Engineer Jobs

Software Engineer III - AI/ML Platform Reliability

Location
Auchentibber, Scotland, United Kingdom
artificial intelligence and machine learning - and we need engineers like you to help us do it reliably, securely, and at scale. As a Software Engineer III at JPMorganChase within the AI/ML Data Platforms organization, you will be a key member of the Reliability Engineering team, contributing … design and delivery of trusted, market-leading technology products. You will apply your technical expertise and problem-solving skills to enhance the reliability and scalability of AI/ML platforms, build reusable services and tooling, and partner across teams to unblock high-impact AI use cases. This ...

Data Reliability Engineer: Postgres & ClickHouse

Location
Greater London, England, United Kingdom
Fuse Energy, LLC is seeking a backend engineer with a focus on data to own and manage our database infrastructure. The ideal candidate will have over 3 years of experience, proficient in SQL and Python, and be skilled in designing data models that reflect real-world business logic. … will be responsible for building scalable data pipelines and ensuring the reliability and performance of our data systems. We offer competitive salaries and benefits, including a biannual bonus scheme and tailored tech support. #J-18808-Ljbffr ...

Senior Data Center Facilities Engineer & Reliability Lead

Location
United Kingdom
Kingdom to lead maintenance and repair of critical infrastructure, from liquid to chip, across global sites. You will drive change management, site governance, and reliability improvements with high autonomy. You will set standards for design, operations and commissioning, escalate complex issues, and coordinate cross-functional teams with vendors ...

Waste Site SCADA Engineer – Reliability & OT Upgrades

Location
Phillack, England, United Kingdom
Thames Water is seeking a Waste Site SCADA Systems Engineer to maintain and enhance site SCADA across our extensive wastewater operations. You will support AVEVA System Platform/Wonderware, FactoryTalk View and Iconics, ensuring reliability and security across hundreds of sites. The role involves diagnosing faults, delivering upgrades ...

Senior SRE Engineer — Cloud Reliability & Automation

Location
City of Westminster, England, United Kingdom
Google London, UK is seeking a Software Engineer III in Site Reliability Engineering for the GCE AI team. This mid-level role focuses on building reliable, scalable systems, code development, and mentoring junior team members. The position emphasizes deep expertise in distributed systems, problem solving, and collaboration across ...

Senior Platform Reliability Engineer – Onsite Birmingham

Location
Birmingham, England, United Kingdom
Infused Solutions Ltd. is seeking a Senior Full Stack Developer to join our Birmingham-based team onsite. You will own and improve large-scale SaaS platforms, focusing on stability, performance, and engineering excellence across the ...

Full Stack Engineer, Platform Reliability

Location
East Midlands, England, United Kingdom
Role Summary We are hiring a versatile Full Stack Engineer to build and run the platform behind Sperry's inspection products. You will work across the stack – Python services and AWS infrastructure at the back, React at the front – moving between projects as priorities shift, and you will … reliability and resilience of what we ship alongside the features you add to it. We have put build and run in the same seat on purpose. The engineers who keep a platform dependable are the ones closest to how it is actually used, and that puts you in front ...

Senior Software Engineer II — Reliability & Observability

Location
Greater London, England, United Kingdom
company is hiring a Senior Software Engineer II to join the OPX team within the Developer Experience organization. You will design and build automated reliability and self-healing systems at scale, delivering platform tooling that engineers across the company adopt for their services. You will own incident management ...

Service Reliability Engineer (GVMC)

Hiring Organisation
Sky Group
Location
London, United Kingdom
Salary
£ 70 K
Provide concise operational updates, maintain accurate records and contribute to improved procedures and knowledge articles. Work as part of a 24/7 service reliability function and remain calm and effective during major incidents. Essential knowledge and experience A strong telecommunications background, preferably gained in a network operations, service … Diameter, SS7/SIGTRAN, roaming or MVNO operations. Exposure to cloud-native network functions, automation or modern observability tools. The team Service Reliability provides the 24/7 operational contact point for Group Networks. The team monitors services, coordinates incidents, carries out approved recovery actions and works with engineering ...

Senior Backend Engineer: Chaos & Reliability (Remote)

Location
Greater London, England, United Kingdom
Camunda seeks a Senior Software Engineer, Backend, to own automated reliability testing and chaos engineering for Camunda 8. You will break things in safe environments to strengthen the platform, guide product direction, and collaborate with friendly colleagues who live our FAITH values. You’ll design and run chaos ...

Senior DevOps Engineer: Platform & Reliability Lead

Location
Greater London, England, United Kingdom
Auriga is seeking a Senior DevOps Engineer to architect, lead and scale the platform engineering and operational backbone of our solutions. You will define and maintain roadmaps for infrastructure, CI/CD, observability, reliability and security, across cloud, on-premises, hybrid and customer-site environments. The role requires ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, United Kingdom
Employment Type
Permanent
Salary
GBP Annual
reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms click apply for full job details ...

Senior AI Infra Engineer – LLM Ops & Reliability

Location
Auchentibber, Scotland, United Kingdom
JPMorganChase is seeking a Senior Lead Software Engineer to shape reliable AI infrastructure at scale. You will own LLM serving stacks, optimize performance and costs, and lead incident response across cloud and on-prem GPU clusters. The role emphasizes secure software delivery, observability, and strong engineering fundamentals. You will … collaborate with engineering to deliver scalable AI platforms, drive reliability, and build reusable patterns for robust production deployments, with a focus #J-18808-Ljbffr ...

Infra Engineer: AI-Powered Reliability & Scale

Location
Greater London, England, United Kingdom
WRITER is hiring an Infrastructure Engineer to own the reliability and performance of our enterprise-grade AI platform. This hybrid role can be based in London or New York, reporting to the Director of Engineering. You will build scalable infrastructure, automate across the stack, and lead incident response ...

Senior Software Engineer, Cloud & Reliability

Location
England, United Kingdom
JPMorgan Chase is seeking a Software Engineer III for its Corporate and Investment Bank Payments Technology – Account Services to strengthen reliability, performance, and automation of mission-critical systems. You will design robust software, build automation, monitor health, and push for scalable, secure solutions in a fast-paced environment. ...

Engineer R&D - Glass Core PKG - Product and Process Reliability

Hiring Organisation
AT&S Austria Technologie & Systemtechnik AG
Location
Leoben, Steiermark, Austria
Employment Type
Permanent
Salary
EUR Annual
drive to make a difference. To enhance our successful R&D Glass Core PKG Teamin Leoben, Austria , we are looking for a passionate Engineer R&D - Product and Process Reliability Aufgaben In your role, you will drive the development and reliability enhancement of next-generation glass core … technology readiness for future market demands. Develop mechanical defect-free process for glass core substrate by working closely with materials suppliers Lead the reliability design, evaluation, and qualification of glass core substrate, including TGV structure and coating configuration Conduct reliability test plans covering thermal cycling, mechanical stress, warpage ...

GPU Infrastructure Engineer — Scale & Reliability for Production

Location
Greater London, England, United Kingdom
OpenAI in London seeks a software engineer to build and operate large-scale production systems powering ChatGPT. You’ll develop tooling for fleet health, capacity planning, automation, and incident response, collaborating with infra, research, and product teams to improve reliability and compute utilization. The role suits engineers ...

Senior Platform Engineer – Production Reliability & On-Call

Location
Greater London, England, United Kingdom
Heidi is hiring for a Platform/SRE role in London. You will own production reliability, participate in on‐call and incident response, and drive improvements across observability, deployments, and operational tooling. This is an ops‐heavy, hands‐on position suitable for mid‐level to senior SREs who enjoy ...

Senior Go Engineer: Scale, Automate & Improve Reliability

Location
Greater London, England, United Kingdom
Jobtailor is seeking a senior software engineer in London to drive the evolution of large-scale systems using Go, AWS, and Azure. You will contribute to platform reliability, observability, and scalable delivery while mentoring junior engineers and collaborating across teams. The role emphasizes CI/CD, automation ...

Network Engineer - Network Source of Truth, Automation & Network Reliability

Location
Cambridge, England, United Kingdom
Role: Senior Network Engineer ( Network Source of Truth, Automation & Network Reliability ) Employment: Contract - Inside IR35 Location: Cambridge,UK - Hybrid Required Technical Skills Network Source of Truth, IRM, IPAM & DCIM Hands-on experience with one or more of: NetBox, Nautobot, IP Fabric, Auvik, SolarWinds, Device42 or Infoblox. Core Networking ...

Azure SRE Engineer: Cloud Reliability & CI/CD Expert

Location
York and North Yorkshire, England, United Kingdom
Hamilton Barnes Associates Limited is seeking an SRE Engineer to join the Cloud Services Group in the UK. You will support the reliability, operation, security, and evolution of customer Azure platforms, designing and implementing systems to improve availability, scalability and performance. Collaborate with development and operations teams ...

Senior Automation Engineer , EMA2 Reliability Maintenance Engineering

Hiring Organisation
Amazon
Location
Mansfield, Nottinghamshire, United Kingdom
Salary
£ 60 K
Here at Amazon we are looking to hire an experienced Senior Automation Engineer to join the team at our Fulfillment Center (FC) in EMA2, Mansfield.The Senior Automation Engineer will ensure that Safety comes first in all Facilities efforts. This position will be responsible for troubleshooting, design/implement … ability to multi-task and deliver results in a dynamic environment.Key job responsibilitiesThe following roles and responsibilities are required for a successful Senior Automation Engineer:• Prepare specifications and technical detail to fully define performance on equipment, material and services• Perform PLC control level issue diagnosis• Follow change management process ...

Network Engineer – Observability, Automation & Reliability

Location
Greater London, England, United Kingdom
Research is seeking an experienced Network Engineer to drive reliability, performance, resilience and security of our network and security infra from our London HQ. You will own end-to-end operations, enhance observability, and champion automation across Cisco and Arista environments. The role requires hands-on experience with ...

Senior Systems Engineer: Automation & Reliability

Location
Greater London, England, United Kingdom
LexisNexis Legal & Professional is seeking a Systems Engineer to support enterprise systems through design, maintenance and change management. You will troubleshoot issues, automate tasks, and collaborate with development, support, and vendor teams to deliver reliable technology solutions. The role focuses on system stability, incident response, and ongoing improvements across ...

Senior Cloud Reliability & Operations Engineer

Location
Belfast City District, Northern Ireland, United Kingdom
Oracle Cloud is seeking a Senior Technical Operations Engineer to operate production environments, including systems and databases, and to drive reliability and performance for critical workloads. You will guide junior engineers, participate in large-scale incident bridges, and help build new processes and procedures to keep Oracle Cloud ...