151 to 175 of 538 Site Reliability Engineering Jobs in London

Site Reliability Engineer

Location
Greater London, England, United Kingdom
benefits packages, technology talks by our experts, a beautiful modern office, daily catered lunches, and more. As a Site Reliability Engineer (SRE), you will work at the intersection of production operations and software development as you improve, manage, and monitor production-critical infrastructure and data pipelines. At Voleon … make a real difference: your contributions will make our critical systems more reliable, lower operational risk, and increase the efficiency of our engineering effort. Responsibilities Improve fault-tolerance and maintainability of code in proprietary data pipelines and trading systems Diagnose and fix bugs in code Lead complex deployments Automate ...

SRE Director — AI-Driven Reliability & Scale

Location
Greater London, England, United Kingdom
EPAM Systems in London, United Kingdom, is seeking a Director of Site Reliability Engineering to lead a global SRE organization in a hybrid work setting. The role focuses on reliability, operational excellence, and governance across mission‐critical platforms, with an emphasis on AI‐enabled automation … improving engineering standards. The successful candidate will drive resilience, define KPIs, and collaborate across product, platform, operations, and security teams to embed reliability #J-18808-Ljbffr ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
client, a world leading systematic multi strat hedgefund is looking for a Site Reliability Engineer to work within their ETF Trading systems. This role combines building and owning the observability of the ETF trading platform as well as coordinating across exchange connectivity quant platforms and risk team. This … risk teams as new venues are onboarded. Automate exchange go-live operations, data pipelines, and desk workflows - leveraging AI tooling Requirements: Experience in production SRE or platform engineering in a latency sensitive financial environment Deep experience in Linux systems cloud and infrastructure as code Experience with monitoring stacks ...

Mid-Level DevOps Engineer

Location
Greater London, England, United Kingdom
engineers and technical leadership to ensure our production systems remain reliable, secure, and performant as the company grows. In short, you will help the engineering team keep our production systems running smoothly while contributing to infrastructure improvements and automation initiatives. What You Bring 3–5 years of DevOps, Site Reliability Engineering, or Infrastructure Engineering experience Experience working with AWS and cloud-based infrastructure Hands-on experience with Kubernetes and Docker in production environments Familiarity with networking concepts including VPCs, VPNs, and secure infrastructure design Experience building or maintaining CI/CD pipelines using CircleCI, Jenkins ...

Product Support Engineer

Location
Greater London, England, United Kingdom
base continue to scale, they're looking for a Product Support Engineer who can work across the boundaries of Production Support, Software Engineering, SRE and DevOps . This is a broad, hands-on engineering role. You'll take ownership when things go wrong in production, investigate issues across … debugging and modifying backend code to resolve issues and implement smaller fixes and improvements. Working hands-on with AWS or GCP and collaborating with SRE, DevOps and Platform Engineering. Improving monitoring, alerting and observability to identify problems before they impact customers. Conducting Root Cause Analysis and turning findings into permanent ...

Platform Engineer

Location
Greater London, England, United Kingdom
capital providers and quota share partners make us nimble. Our breadth of expertise and capabilities deliver outstanding market returns. The role The Analytics Product Engineering team at The Fidelis Partnership (TFP) builds and manages a bespoke analytical platform that powers our partnership-driven business model. We combine actuarial expertise … with advanced technology to deliver innovative, scalable solutions. The Platform Engineer partners with the Analytics product engineering team to design, build and operate platform capabilities that enable our market leading high-performance, distributed computing products to run reliably, securely and efficiently at scale. This includes infrastructure automation, runtime orchestration ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City, London, United Kingdom
Employment Type
Permanent
Salary
GBP 85,000 Annual
Site Reliability Engineer Up to £85,000 + Benefits Central London Hybrid (2/3 days a week in the office) Build, Scale & Improve the Reliability of a Fast-Growing SaaS Platform We're partnering with a fast-growing SaaS company that's going through an exciting … period of growth and investing heavily in its engineering and platform capabilities click apply for full job details ...

Senior Data & MLOps Engineer

Location
Greater London, England, United Kingdom
proud to be a Living Wage accredited Employer. What You’ll Do The Data Science team is focused on developing an advanced reliability platform. This system covers various aspects of data processing and analysis, including data intake, deriving meaningful metrics, identifying unusual patterns, predicting potential issues, finding slow processes … level performance analysis systems. Experience developing agentic or LLM‐powered reasoning systems for diagnostics or operational intelligence. Background in reliability engineering or SRE practices. Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences ...

Director of Platform Engineering

Location
Greater London, England, United Kingdom
incident triage, anomaly detection, runbook automation, knowledge search, root cause support or service desk workflows. Knowledge of observability and reliability practices, including SRE principles, service‐level objectives, alert tuning, capacity planning and production readiness reviews. Experience operating SaaS products for financial services, enterprise technology or other regulated customers with … requirements. Experience with container security, policy‐as‐code, image scanning, secrets management, role‐based access control and Kubernetes security hardening. A background in DevOps, SRE, Cloud Operations or Platform Engineering, with a track record of improving automation, reliability and operational maturity. Health Insurance and Dental Health Cover ...

Senior Network SRE: Automation, Reliability & Observability

Location
Greater London, England, United Kingdom
leading IT solutions provider in London is seeking a Senior Network Site Reliability Engineer (SRE) with extensive experience in network engineering and automation. The ideal candidate will design and maintain high-availability network solutions while employing SRE principles for improved reliability and performance. A strong background ...

Senior Site Reliability Engineer - Scale & Automation

Location
Greater London, England, United Kingdom
Google is seeking a Site Reliability Engineer to design, build, and operate scalable, fault-tolerant systems across Google Cloud services. You’ll blend software and systems engineering to ensure reliability, uptime, and rapid improvement while reducing toil through automation. You will work on large-scale distributed ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
## Site Reliability EngineerApplyremote type: Hybridlocations: London, United Kingdomtime type: Full timeposted on: Posted Todayjob requisition id: 2020336## **Meet the Team**Cisco's Webex Engineering Group is redefining the future of collaboration. We're building a world where people connect effortlessly to enjoy modern, uncompromised collaboration across … Adaptable & Problem-Solver**: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance.* **Ownership & Quality**: Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments and full reproducibility, including ...

Senior Infrastructure Engineer

Hiring Organisation
Community Fibre Limited
Location
London, UK
Employment Type
Full-time
building new servers/systems as well as ensuring that all our existing ones are maintained, reliable and resilient. The role covers backend systems engineering, infrastructure, and site reliability engineering within the Network Technology group. You should be hungry for hands-on experience, have a desire … WSGIDeep understanding of key service health metricsKnowledge of current security best practices including encryption and systems hardeningExperience implementing technologies with a focus on operability, reliability, and scalability. Network aptitude. The skill to take a new technology, research it, test it in the lab and deploy it live. Excellent understanding ...

Service Reliability Engineer - London

Hiring Organisation
Fitch Group
Location
Greater London, United Kingdom
Employment Type
Full Time
career in technology and data at Fitch? Visit: https://careers.fitch.group/content/Technology-and-Data/About the Team Fitch Group SRE provides Service Reliability Engineering expertise to Fitch’s development organizations. This squad joins Core Engineering and other SRE groups as part … escalation point and participate in the L3 on‐call rotation. You May be a Good Fit if: You have deep, hands-on experience in SRE, DevOps, or Platform Engineering across both AWS and Azure, with a strong track record operating Docker and Kubernetes in production environments. You’re highly ...

ML Compute SRE Lead: Scale, Uptime & Automation

Location
City of Westminster, England, United Kingdom
Google London is seeking a Systems Engineering Manager for Site Reliability Engineering in ML Compute. You will lead a multi-disciplinary team, own uptime, and shape reliability strategy for large-scale services. You will mentor engineers, drive end-to-end availability, and collaborate with cross ...

Senior Site Reliability Engineer - AI-Driven Reliability

Location
Greater London, England, United Kingdom
JPMorgan Chase is seeking a Site Reliability Engineer to join the International Consumer Bank group in the UK. The role focuses on building reliable, scalable digital banking services and leading initiatives to reduce operational toil through automation. You will collaborate with product and platform teams to implement observability ...

Head of Production Management- J.P. Morgan Personal Investing

Location
Greater London, England, United Kingdom
powered solutions and intelligent automation to reduce manual intervention, fast-track resolution, and continuously improve operational efficiency. Champion an automation-first, shift-left SRE culture leveraging shared tooling and automation to ensure consistency, reduce duplication, and maintain alignment with firmwide standards. Oversee capacity management and planning, ensuring infrastructure scales … management standards — change, incident, capacity, and automation — across multiple engineering teams operating in a you-build-it-you-run-it model, underpinned by SRE principles and disaster recovery planning. Composure, decisiveness, and authority during incidents, vendor failure, or regulatory escalation, with a proven ability to protect business lines under ...

Senior SRE Engineer — Cloud Reliability & Automation

Location
City of Westminster, England, United Kingdom
Google London, UK is seeking a Software Engineer III in Site Reliability Engineering for the GCE AI team. This mid-level role focuses on building reliable, scalable systems, code development, and mentoring junior team members. The position emphasizes deep expertise in distributed systems, problem solving, and collaboration ...

AI Native DevOps Platform Engineer

Location
Greater London, England, United Kingdom
capabilities and platform standardisation. Improve release processes and operational excellence across teams. Reliability & Observability Implement monitoring, logging, tracing, and alerting solutions. Establish platform SRE principles and operational standards. Proactively identify and resolve reliability, security, and performance issues. Lead incident response and continuous improvement initiatives. Security & Governance Embed security … capabilities and platform standardisation. Improve release processes and operational excellence across teams. Reliability & Observability Implement monitoring, logging, tracing, and alerting solutions. Establish platform SRE principles and operational standards. Proactively identify and resolve reliability, security, and performance issues. Lead incident response and continuous improvement initiatives. Security & Governance Embed security ...

Site Reliability Engineer: Scale Global Services

Location
Greater London, England, United Kingdom
Apple Services Engineering is seeking a Site Reliability Engineer to own and improve the massive-scale infrastructure powering App Store, Apple TV, Apple Music, and more. You’ll work across Linux-based systems, build automated tooling, and collaborate with developers to deliver reliable services. Responsibilities include troubleshooting ...

Site Reliability Engineer,

Location
Greater London, England, United Kingdom
BIG. Operating at our scale, across multiple geographically dispersed data centers and servicing hundreds of millions of users presents unique challenges. As an SRE at Apple, you'll need to solve these problems using data, teamwork, and your own expertise. SREs at Apple own the full infrastructure stack; from device … close partnership with our development teams and aim to design & build new services together. We're passionate about software and automation in SRE and develop a variety of tooling and infrastructure. Our services run on mixed & hybrid platforms. Responsibilities Create outstanding customer experience, and help developers write better code faster ...

SC Cleared Devops/Site Reliability Engineer

Location
Greater London, England, United Kingdom
leading global IT and business consulting firm, deeply Embedded in UK Defence and National Security, to identify an experienced SC-cleared DevOps/SRE Engineer to join their defence team. You'll be building and running the CI/CD pipelines, automation and observability tooling behind systems that support national … minimum of 5 years Due to the secure nature of the role, candidates must be sole UK nationals Hands-on DevOps/SRE experience, CI/CD pipelines, automation, monitoring/observability Comfortable working within secure, regulated environments Desirable: Defence, Intelligence, or wider public sector background Experience with DevSecOps principles ...

Senior Solutions Engineer

Location
Greater London, England, United Kingdom
relationship that will provide continuous value to our customers. Finally, you will have the opportunity to work cross-functionally with our Product Management and Engineering teams to share your knowledge and experiences to ultimately improve our business and our customers' success. We seek talent who wants to leverage their … technical validations during the Proof of Value phase Be successful working with all levels of an organization, from executives down to individual developers and Site Reliability Engineers Deliver product and technical demonstrations of the Sumo Logic service Work cross functionally with Product Management and Engineering to improve ...

Senior Site Reliability Engineer — Hybrid Kubernetes & GitOps

Location
Greater London, England, United Kingdom
Cisco’s Webex Engineering Group in London is seeking a Senior Site Reliability Engineer to own the design, deployment, and operation of Kubernetes-based microservices, delivering reliability and scalable deployments in a hybrid work environment. You will drive GitOps workflows with Argo CD, use Helm ...

Platform Engineer

Location
Greater London, England, United Kingdom
Full-Time Working Pattern: Monday to Friday Start Date: ASAP About the role As a Platform Engineer, you will be responsible for the management, reliability, performance, and continuous improvement of our Kubernetes infrastructure. You’ll play a key role in operating and enhancing our Kubernetes environments, containerised workloads … troubleshooting complex application and infrastructure issues. Knowledge of networking principles including DNS, load balancing, firewalls, and TLS/SSL. Experience working within DevOps or Site Reliability Engineering practices. Experience with Amazon EKS. Exposure to GitHub Actions, GitLab CI, Azure DevOps, or Jenkins. Knowledge of security best practices ...