501 to 525 of 1,104 Site Reliability Engineering Jobs in the UK

Service Delivery Manager – HPC/Super Computing

Location
United Kingdom
Microsoft technology with solid overview of the Microsoft cloud servicesProject Management/Prosci Change Management/ITIL certification is a plusUnderstanding DevOps, Site Reliability Engineering, Continuous Improvement is a plusSupercomputing/HPC experience is a plus.Ability to meet Microsoft, customer and/or government security screening requirements ...

Senior Network SRE: Cloud Reliability & IaC

Location
Greater London, England, United Kingdom
Miro is seeking a Senior Network Site Reliability Engineer to help strengthen reliability, availability, and scalability of our production environment. You will focus on cloud automation, IaC, and governance across our AWS infra, contributing to highly available services for millions of users. You will own automation, observability ...

Senior SRE - Platform Reliability & Observability Lead

Location
Cambridge, England, United Kingdom
Bango plc is seeking a Senior Site Reliability Engineer to lead reliability, performance and continuous improvement of the Bango Platform. You will … reliability, observability and incident response across infrastructure and delivery pipelines, serving as a technical centre of gravity for the SRE function. You will shape the bench across the third-line support and PICE teams, grow engineers, and lead cross-cutting work end-to-end while remaining an individual contributor ...

Lead Data Engineer - Python, Databricks

Location
Glasgow, Scotland, United Kingdom
efficiency. Contribute to a team culture of diversity, opportunity, inclusion, and respect. Required Qualifications, Capabilities, And Skills Formal training or certification on data engineering concepts and advanced applied experience. Hands-on experience developing and maintaining data pipelines using Python. Proficiency with Databricks for large-scale data processing and analytics. … Experience working with enterprise-level datasets in a large, complex organization. Familiarity with cloud-based data platforms and modern data architecture patterns. Experience in site reliability engineering or infrastructure deployment roles. Knowledge of data governance, access control, and security best practices. ABOUT US J.P. Morgan ...

Managing Engineer - Observability, Pipeline & Analytics (Hybrid)

Location
Belfast City District, Northern Ireland, United Kingdom
large‐scale telemetry pipelines, modern analytics platforms, cloud‐native observability solutions, and engineering automation practices. The successful candidate will partner closely with Observability, SRE, Infrastructure, Security, Data, and Product Engineering organizations to deliver a unified telemetry strategy that improves reliability, accelerates troubleshooting, enables AI‐driven operations, reduces … scale operational and platform data. Lead architecture reviews and provide technical direction for onboarding applications, infrastructure platforms, cloud services, and emerging technologies. Partner with SRE, Platform Engineering, Security, Services, and Product teams to improve reliability, observability maturity, and operational insights. Implement automation‐first self‐service operational practices using ...

Senior DevOps Engineer

Location
United Kingdom
container image scanning. Implement and monitor infrastructure and application security controls. Support the organisation's ongoing compliance and certification requirements. Reliability & SRE Establish and maintain observability across distributed systems. Develop proactive monitoring, alerting and performance-tuning strategies. Help maintain service-level objectives and platform availability. Investigate and resolve infrastructure … with development teams to embed DevSecOps practices throughout the software lifecycle. What We're Looking For We're looking for a seasoned DevOps or SRE professional who combines strong hands-on technical skills with a pragmatic, collaborative approach . You'll ideally have: Proven professional experience in DevOps, SRE ...

Senior DevOps Engineer

Hiring Organisation
MarkIT Placements
Location
Didcot, Oxfordshire, South East, United Kingdom
Employment Type
Permanent
container image scanning. Implement and monitor infrastructure and application security controls. Support the organisation's ongoing compliance and certification requirements. Reliability & SRE Establish and maintain observability across distributed systems. Develop proactive monitoring, alerting and performance-tuning strategies. Help maintain service-level objectives and platform availability. Investigate and resolve infrastructure … with development teams to embed DevSecOps practices throughout the software lifecycle. What We're Looking For We're looking for a seasoned DevOps or SRE professional who combines strong hands-on technical skills with a pragmatic, collaborative approach . You'll ideally have: Proven professional experience in DevOps, SRE ...

Lead Data Engineer - Python, Databricks

Location
Glasgow, Scotland, United Kingdom
efficiency. Contribute to a team culture of diversity, opportunity, inclusion, and respect. Required qualifications, capabilities, and skills Formal training or certification on data engineering concepts and advanced applied experience. Hands-on experience developing and maintaining data pipelines using Python. Proficiency with Databricks for large-scale data processing and analytics. … Experience working with enterprise-level datasets in a large, complex organization. Familiarity with cloud-based data platforms and modern data architecture patterns. Experience in site reliability engineering or infrastructure deployment roles. Knowledge of data governance, access control, and security best practices. #J-18808-Ljbffr ...

AI-Powered Observability Tech Lead

Location
Greater London, England, United Kingdom
Collaboration Technology Group in London seeks a Technical Leader to drive architectural vision and implement an AI-powered Production Intelligence platform. You will blend Site Reliability Engineering with agentic AI to improve monitoring, incident response, and auto-remediation across global SaaS infrastructure. You will mentor engineers, shape ...

Senior Product Manager for AI Observability

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
with AI Evaluation PM (previous role), Model Risk, GSSR, Legal and Compliance to align telemetry with governance frameworks. Work closely with Engineering and SRE teams to drive observability improvements and reliability engineering for AI systems. Optimisation & Insights Identify cost inefficiencies across model and MCP usage, and drive … support auditability, compliance and explainability requirements. Skills & Competencies Required Experience in product management with a strong foundation in observability, telemetry, data platforms, monitoring, or SRE/DevOps-driven products. Understanding of LLMs, embeddings, vector search, MCP tools, and AI inference workflows. Deep familiarity with logging, tracing, metrics, and event-based ...

Cyber Platform Engineer

Hiring Organisation
Hays Specialist Recruitment Limited
Location
Sheffield, South Yorkshire, United Kingdom
Employment Type
Full-Time
Salary
£550.00 - £650.00 per day
with a strong software development background who enjoy building production-grade systems and are comfortable operating those systems in cloud environments using DevOps and SRE practices. This is not a traditional infrastructure-focused DevOps or Platform Engineering position. Key responsibilities: Software Engineering (Primary Focus): Design, build, test … provisioning. Work with cloud networking concepts and secure service integrations. Support engineering platforms used by Cyber Security and Data Engineering teams. DevOps & SRE Practices: Build and maintain CI/CD pipelines. Implement monitoring, alerting and observability. Troubleshoot production issues and participate in operational support. Continuously improve reliability ...

Solutions Architect, Studios — Pre-Sales & Architecture

Location
Greater London, England, United Kingdom
solutions across client contracts, new business opportunities, and procurement activity. The role emphasizes pre-sales engagement, proofs of concept, and collaboration with DevOps and Site Reliability Engineering teams to ensure automatable, observable, and secure solutions. The position reports into IMG Studios and involves working across commercial, technical ...

Manager - DE - Technology Consulting - FS

Location
Greater London, England, United Kingdom
want it to go. Join EY and help to build a better working world. Rank: Manager | Team: Financial Services Organisation, Technology Consulting, Digital Engineering Role Summary: Hands-on Digital Engineering Manager responsible for leading software engineering delivery across EY FSO client engagements. The role combines engineering … such as Kafka, API management, data quality and operational data analysis Production support Monitoring, logging, incident management, RCA, performance tuning, runbooks, ServiceNow/Jira, SRE-aligned operational readiness Experience delivering technology change for banking, insurance, wealth, asset management, risk, finance, payments or market infrastructure clients Consulting capabilities Stakeholder management, workshops ...

Senior SRE Engineer: Reliability, Cloud & Automation

Location
Greater London, England, United Kingdom
London Stock Exchange Group is looking for a Senior Engineer in Site Reliability who will join a driven team focused on system availability, performance, and scalability. Responsibilities include maintaining service level objectives, writing automation for system resilience, and partnering with development teams. Required qualifications include a Bachelor … computer science, experience in Object Oriented programming and cloud systems, and DevOps familiarity. The role is pivotal in ensuring 24/7 system reliability and promoting engineering best practices. #J-18808-Ljbffr ...

Research Engineer, Safety Oversight, DeepMind

Location
Greater London, England, United Kingdom
technical products. Experience in the domain area of generative AI and Large Language Models (LLMs). Preferred qualifications: Master’s degree or PhD in Engineering, Computer Science, or a related technical field. 3 years of experience developing code, running experiments and analyses collaboratively with coding agents. Experience building large … Software Engineer, Full Stack, Google AdsGoogle-2w agoLondon, UKFull-time14DetailsG### Software Engineer III, Full Stack, Publisher InventoryGoogle-2w agoLondon, UKFull-time15DetailsG### Software Engineer III, Site Reliability Engineering, Traffic Network Load BalancingGoogle-2w agoLondon, UKFull-time15Details## Explore related hubsCountry hubUnited Kingdom JobsCompany pageGoogle JobsSalary pageSoftware Engineer SalaryVisa pageSkilled ...

AI-Powered Production Intelligence Tech Lead

Location
Greater London, England, United Kingdom
Cisco is seeking a Technical Leader to drive architectural vision for an AI-powered Production Intelligence platform. You will blend Site Reliability Engineering with agentic AI to improve monitoring, diagnosis, and auto-remediation across global SaaS infrastructure. You will lead architecture, mentor engineers, and partner across teams ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
impact***** ****Build and maintain infrastructure-as-code (Terraform or similar), CI/CD pipelines, and Kubernetes platforms as products consumed by engineering teams********SRE & reliability***** ****Define and drive SLOs, error budgets, and observability standards (metrics, logging, tracing) across the platform***** ****Take part in post-incident reviews and help … roadmap for the platform function, balancing reliability investment against delivery****## ****What we're looking for********Must have***** ****A background in Operations or SRE running highly available, redundant production platforms — you understand failure domains, graceful degradation, and what "five nines" costs***** ****Deep hands-on experience with AWS (VPC design ...

Senior SRE: GCP, Kubernetes & Automation Leader

Location
Greater London, England, United Kingdom
Brevan Howard CFD LTD is seeking a Senior Site Reliability Engineer (SRE) to enhance the reliability, scalability, and performance of its core platform. The successful candidate will provide operational support and lead infrastructure projects. This role requires hands-on experience with Google Cloud Platform (GCP) and Kubernetes ...

Site Reliability Engineer- Spacetime UK

Location
Greater London, England, United Kingdom
Role Overview This isn't a "keep the lights on" SRE role. This is a strategic, high-impact opportunity to build the nervous system for a platform that transforms how networks of satellites, ground stations, and fleets are interconnected and orchestrated. You will be building the core observability stack that … cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade observability stack from the ground up, this role ...

Lead SRE: AWS & Python for Scalable Reliability

Location
Glasgow, Scotland, United Kingdom
JPMorganChase in the United Kingdom is seeking a Lead Site Reliability Engineer to drive reliability at scale across critical production systems. You will partner with software engineering and product teams to embed reliability into the software development lifecycle and deliver resilient services for millions ...

Cloud SRE Lead: Platform Reliability & Automation

Location
Halifax, England, United Kingdom
Lloyds Banking Group is seeking a Lead Site Reliability Engineer to strengthen reliability across Azure and Google Cloud Platform. You will lead a team of SREs, set engineering standards and drive improvements in observability, incident response and platform reliability. Collaborate with Product Owners, Engineering Leads ...

SRE & Cloud Engineer — Multi-Cloud, Kubernetes, Terraform

Location
Glasgow, Scotland, United Kingdom
Dianaduggan is seeking an experienced Site Reliability Engineer (SRE) to join a high-performing engineering team in Glasgow as a Cloud Engineer. You will enhance cloud reliability, scalability, automation and operational excellence across enterprise cloud platforms. Based in Glasgow, the role is on-site ...

Senior Platform Engineer

Location
Abingdon, England, United Kingdom
deliver sustainable fusion energy and maximise its scientific and economic impact. The Computing Division supports this through digital capabilities spanning research, simulation, data, engineering, business systems and communications. The Software Operations Group (SORG) provides the infrastructure and expertise needed for secure software development and operations, promoting DevSecOps and MLOps … improving processes and supporting the growth of technical capabilities. Drive continuous improvement and operational excellence across UKAEA, managing stakeholders, promoting secure‐by‐design and SRE practices, and maintaining high standards of safety and quality. What You’ll Bring Technical qualification or equivalent experience in a scientific, engineering or technical ...

Senior SRE: Observability, Automation & Reliability

Location
Cambridge, England, United Kingdom
Altium is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, and performance of large-scale SaaS platforms across regions. You will automate operations, improve observability, and contribute to incident management in collaboration with DevOps and engineering teams. Join a team that champions … principles, and scalable deployments, driving reliability improvements and faster incident resolution across the cloud platform. #J-18808-Ljbffr ...

Senior Cloud Engineer

Hiring Organisation
Pacific Life
Location
London, UK
Employment Type
Full-time
operational excellence, security, reliability, and lifecycle management of cloud applications across multiple regions. The role sits at the intersection of cloud architecture, SRE, security, and developer experience, helping to raise the overall standard of cloud usage across PL Re. While experience in platform provisioning is required, this role does … hours on‐call rotas as required. Support the full lifecycle of application workloads, including maintenance, patching, backups, upgrades, and controlled decommissioning. Reliability, SRE & Observability Drive improvements in reliability, resilience, performance, and efficiency through strong SRE practices. Own and enhance observability and monitoring, ensuring meaningful alerting, clear operational dashboards ...