551 to 575 of 1,104 Site Reliability Engineering Jobs in the UK

MLOps Engineer

Location
Greater London, England, United Kingdom
2025. Learn more at www.coreweave.com. We're proud to be a Living Wage accredited Employer. What You'll Do CoreWeave’s Physical AI Platform Engineering team builds and scales the data and workflow backbone powering advanced engineering simulation and AI workflows. Our ambition is to become the super … engineers on production‐grade ML practices. Who You Are 5–6+ years of professional experience in MLOps, ML platform engineering, ML infrastructure, or SRE/DevOps for production machine learning systems. Proven experience building, operating, and automating production ML pipelines covering experiment tracking, model registries, artifact versioning, dataset management ...

Observability AVP: SRE & Cloud Telemetry Lead

Location
Greater London, England, United Kingdom
Citigroup Inc. is seeking a Site Reliability Engineer (SRE) to lead hands-on deployment of observability principles in a large-scale environment. The role focuses on reliability, performance, and resiliency across applications and services. You will migrate monitoring tools to Google Cloud Observability and Grafana using OpenTelemetry … author reusable deployment solutions, and onboard engineering teams to new telemetry standards. This is a senior technical role within Production Management. #J-18808-Ljbffr ...

SRE Engineer – FinTech Reliability, Observability & Cloud

Location
Greater London, England, United Kingdom
Hamilton Barnes Associates Limited is seeking a Site Reliability Engineer to work at the intersection of software engineering and infrastructure. You'll develop internal platforms, tooling, and automation across Linux, distributed systems, and cloud-native technologies to improve reliability and operational efficiency in a global production ...

Head of Quality Enablement

Location
United Kingdom
first-class discipline. Resilience and chaos engineering across critical services. The evolution of IRIS's future Scalability & Resiliency practice and eventual SRE capability. Quality signals as a product input through the integration of observability, telemetry, and customer experience data. What We’re Looking For Essential Experience Experience leading quality … aligned pods. Exposure to performance, resilience, or chaos engineering at a meaningful level, enough to credibly grow the role into those areas. Full SRE leadership experience is welcome but not required at hire. Leadership Style & Approach We are looking for someone who: Leads with quality, not headcount. Holds ...

Lead SRE: AWS & Python for Scalable Reliability

Location
United Kingdom
JPMorganChase is seeking a Lead Site Reliability Engineer to ensure the availability, performance, and resilience of production systems serving millions of customers worldwide. You will partner with software engineering and product teams to embed reliability into the SDLC, drive incident response, build automation, mentor junior engineers … champion proactive reliability improvements and observability across platforms. #J-18808-Ljbffr ...

Sr Lead Infrastructure Engineer- Devops/AWS

Location
Auchentibber, Scotland, United Kingdom
looking for a hands-on DevOps/SRE engineer ready to take their career to new heights. Join the ranks of top talent at one of the world's most influential companies. As a Lead DevOps/SRE Engineer - Vice President at JPMorgan Chase within the International Private Bank … engineering Understanding of production security and change-management controls Strong communication skills and the ability to set operational standards across a team Formal SRE experience in a regulated or high-availability environment Master's degree in Computer Science, Engineering, or a related technical field (or equivalent applied experience ...

Sr Lead Infrastructure Engineer- Devops/AWS

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
DESCRIPTION We're looking for a hands-on DevOps/SRE engineer ready to take their career to new heights. Join the ranks of top talent at one of the world's most influential companies. As a Lead DevOps/SRE Engineer - Vice President at JPMorgan Chase within the International … engineering Understanding of production security and change-management controls Strong communication skills and the ability to set operational standards across a team Formal SRE experience in a regulated or high-availability environment Master's degree in Computer Science, Engineering, or a related technical field (or equivalent applied experience ...

DevOps / SRE Engineer (London)

Location
Greater London, England, United Kingdom
Bumper continues to scale across the UK and Europe — and we’re excited to welcome a passionate and experienced DevOps/SRE Engineer , based in our London office, to help us take our platform reliability and engineering excellence to the next level. Our Head Office is in Sheffield … with the vision of being the leading automotive payment and insights platform! A bit about the role... We’re looking for a DevOps/SRE Engineer to join our growing Infrastructure team. In this role, you’ll help shape and strengthen the reliability, scalability and observability of our cloud ...

DevOps / SRE Engineer (Sheffield)

Location
Sheffield, England, United Kingdom
Bumper continues to scale across the UK and Europe — and we’re excited to welcome a passionate and experienced DevOps/SRE Engineer , based in our Sheffield office, to help us take our platform reliability and engineering excellence to the next level. Our Head Office is in Sheffield … with the vision of being the leading automotive payment and insights platform! A bit about the role... We’re looking for a DevOps/SRE Engineer to join our growing Infrastructure team. In this role, you’ll help shape and strengthen the reliability, scalability and observability of our cloud ...

Lead AI Infra SRE: Scale, Reliability & Mentorship

Location
Gloucester, England, United Kingdom
Radiant is pursuing an experienced Infrastructure Site Reliability Engineer to run and evolve our AI-native infrastructure stack in the UK. You’ll cover bare-metal, virtualization, and orchestration layers while mentoring teammates and improving automation to support AI/HPC workloads. You will configure and operate resilient ...

Cloud and Infrastructure Operations Manager (AWS)

Hiring Organisation
Radius Payment Solutions
Location
Chester, United Kingdom
focus on our AWS environment.Reporting directly to the CIO, this is a key leadership role within our IT Operations department, with responsibility for the reliability, performance, security and scalability of our core infrastructure services. You'll lead people across TechOps, SysOps, Networks and Voice, whilst remaining closely involved … infrastructure security principles and governance.Excellent stakeholder management and communication skills.Highly DesirableExposure to Azure environments.Enterprise networking expertise across LAN, WAN, VPN and cloud networking.Experience with SRE or reliability engineering practices.ITIL-based service management experience.Experience supporting modern development teams with infrastructure and platform services.Experience with voice or contact-centre technologies ...

Remote SRE: Cloud Reliability & CI/CD Expert

Location
Dacorum, England, United Kingdom
Haven is seeking a hands-on Site Reliability Engineer to join our Product Technology function. This remote-first role involves working with engineering teams to design, implement and support scalable, reliable systems across CI/CD, observability and database reliability. You will shape deployment strategies, own incident ...

AI Native DevOps Platform Engineer

Location
Greater London, England, United Kingdom
capabilities and platform standardisation. Improve release processes and operational excellence across teams. Reliability & Observability Implement monitoring, logging, tracing, and alerting solutions. Establish platform SRE principles and operational standards. Proactively identify and resolve reliability, security, and performance issues. Lead incident response and continuous improvement initiatives. Security & Governance Embed security … platform services and automation capabilities. Support model serving infrastructure, and AI observability tooling. Role Requirements Significant experience as a Platform Engineer, DevOps Engineer, SRE, or Cloud Engineer. Strong expertise with Azure cloud technologies. Deep understanding of cloud‐native platforms and container orchestration. Experience building CI/CD pipelines using Azure ...

Senior DevOps Analyst

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
Azure/GCP.Implement Infrastructure as Code (IaC) using Terraform, CloudFormation, or ARM templates. Build and manage containerized applications using Docker and Kubernetes. Ensure system reliability, scalability, security, and performance. Implement monitoring, logging, and alerting solutions. Collaborate with development, QA, and security teams. Troubleshoot production issues and lead incident response. … stack, Splunk. Version Control: Git (GitHub/GitLab/Bitbucket).Security: IAM, secrets management, vulnerability scanning. Experience & Qualifications8+ years of experience in DevOps/Site Reliability/Infrastructure Engineering. Strong understanding of DevOps and CI/CD best practices. Experience supporting high-availability, production systems. Experience in Agile ...

Senior SRE: Front‐Office Trading Reliability

Location
Greater London, England, United Kingdom
Engineer within its Trading Technology group in London. You will embed with the software engineering team that builds front‐office trading platforms, shaping SRE patterns, observability, and resilience across globally distributed systems and AI‐accelerated incident response. You will partner with traders and senior stakeholders, lead incident responses, implement … reliable code, and drive modern SRE practices across CI/CD, telemetry, and automated #J-18808-Ljbffr ...

Managed Service Operations - Head of Practice

Location
Greater London, England, United Kingdom
service operations including incident, problem, change, event, monitoring, resilience, continuity, capacity, and on‐call models. Strong understanding of ITIL practices blended with modern DevOps, SRE, Agile and platform‐engineering approaches. Broad technical awareness across cloud platforms, application architectures, data platforms, networks, observability tooling, security‐by‐design, and automation. Ability … engineering, service readiness, and live‐service best practices. Key experiences Running and growing operational or engineering teams in a Managed Service, SRE, DevOps, or live‐service environment—with responsibility for hiring, coaching, development and performance. Leading high‐pressure operational functions including incident management, problem resolution, major incident ...

Principal Recruitment Consultant

Hiring Organisation
Harrison Clarke
Location
City of London, London, United Kingdom
class engineering teams. Our mission is simple: align exceptional engineers with groundbreaking businesses that are shaping the future. We operate across Cloud (DevOps, SRE, Platform Engineering, DevSecOps, Performance Engineering) and Data & AI (Machine Learning, MLOps, Data Engineering, AI Research, and related domains). Through our portfolio … recruitment consultant with experience placing technical talent into startups or technology companies Strong domain knowledge or demonstrable interest in Cloud, Data, AI, DevOps/SRE, Platform Engineering, Machine Learning or related fields Excellent stakeholder management skills, comfortable engaging at founder and senior engineering leadership level Entrepreneurial attitude with ...

Senior SRE ServiceNow Platform Reliability

Location
England, United Kingdom
JPMorgan Chase is seeking a Site Reliability Engineer III to own end-to-end reliability for the ServiceNow platform. … will manage incident response, SLOs, and platform performance, and drive efficiency through automation and improved telemetry across the stack. Ideal candidates combine 3+ years SRE experience with strong scripting, JVM tuning, and database optimization skills in a large enterprise environment. #J-18808-Ljbffr ...

Senior Platform Engineering Manager at Prolific

Location
United Kingdom
human data infrastructure that's reshaping the landscape of AI development. As our Platform Engineering Manager, you will lead our Cloud Platform and SRE teams, taking charge of the technical foundation and reliability that allows Prolific to scale securely and efficiently. Why Prolific At Prolific, our product development … infrastructure and platform ensuring robust management of our services like Kubernetes (GKE) clusters. Focus on continuous improvement, operational excellence, and innovation. Champion SRE Culture: Own availability and embed SRE principles across the organization, including defining SLOs, SLAs, error budgets, and enhancing observability and incident remediation. Platform & Developer Experience ...

Senior SRE Technical Lead — Reliability & Observability

Location
Greater London, England, United Kingdom
leading global financial markets infrastructure provider is seeking a Technical Lead SRE in Greater London. In this role, you will enhance the reliability engineering capabilities, collaborating with various teams to establish observability standards and ensure operational excellence. The ideal candidate will have over 10 years of experience … SRE or related fields, strong AWS and Kubernetes skills, and a proven track record in building resilient platforms. Join us to make a significant impact in financial markets infrastructure. #J-18808-Ljbffr ...

DevOps Engineer

Hiring Organisation
NSD
Location
Newcastle Upon Tyne, Tyne and Wear, North East, United Kingdom
Employment Type
Permanent
Salary
£70,000
DevOps best practices, and providing hands-on technical support across development and delivery teams. We're looking for someone who enjoys solving challenging engineering problems and getting stuck into the technical detail! Due to the sensitive nature of the work, SC Clearance eligibility is required. DEVOPS ENGINEER ESSENTIAL EXPERIENCE … Experience in DevOps or Platform engineering role, building and maintaining AWS infrastructure. * Large scale AWS environments * Container technologies e.g. Kubernetes, docker * CI/CD e.g. Jenkins, Github * Infrastructure as code e.g. Terraform, Cloud Formation * Monitoring e.g. Prometheus, Grafana * Strong communication and stakeholder management skills * Eligible for SC Clearance DEVOPS ...

Director of Platform Engineering

Location
United Kingdom
Director of Platform Engineering Location: UK preferred, office preferred Contract: Permanent full-time Start Date: ASAP About the Role As Director of Platform Engineering, you are responsible for building the shared foundations that make Hawk‐Eye’s technology organisation easier to trust, easier to scale, and easier … calibre Platform Engineering function with strong technical credibility and operational judgement. Hire, develop, and align engineering leaders and specialists across platform, cloud, SRE, and data disciplines. Foster a culture of accountability, clarity, calm execution, and continuous improvement. Act as a senior internal leader who can bring structure ...

Senior SRE, Observability & Cloud Reliability

Location
Greater London, England, United Kingdom
company is seeking a Principal Site Reliability Engineer, Infrastructure Observability to guide a team of SREs focused on observability, reliability, and scalable cloud/on‐prem solutions. … role requires hands‐on expertise and collaboration with diverse partners to drive measurable improvements. The ideal candidate has extensive cloud experience, DevOps/SRE leadership, and strong automation skills, with a track record in designing resilient systems and implementing effective monitoring and #J-18808-Ljbffr ...

Remote SRE: Cloud Reliability & AI-Driven Ops

Location
Hemel Hempstead, England, United Kingdom
Haven is seeking a hands-on Site Reliability Engineer to join our Product Technology team. This remote-first role involves shaping CI/CD, observability, and incident response while collaborating with engineers and tech leads to ensure reliable, scalable platforms for guests and colleagues. You’ll tackle infrastructure ...

Staff Platform Engineer - AI Native SaaS Platform

Location
City Of London, England, United Kingdom
2025. Their platform is redefining the sector, and with revenues nearly 10x since the start of last year, they're continuing to expand their Engineering team to match the ambition of their product and customers. The product is real-time, data-rich and AI-native, creating complex engineering … infrastructure as code Strong understanding of networking, IAM, managed services and secure cloud architecture Experience owning observability, logging, APM or distributed tracing tooling An SRE mindset across SLOs, reliability, incident response and toil reduction A track record of technical leadership, mentoring and delivering complex projects through others Strong systems ...