176 to 200 of 593 Site Reliability Engineering Jobs in London

Director of Platform Engineering

Location
Greater London, England, United Kingdom
incident triage, anomaly detection, runbook automation, knowledge search, root cause support or service desk workflows. Knowledge of observability and reliability practices, including SRE principles, service‐level objectives, alert tuning, capacity planning and production readiness reviews. Experience operating SaaS products for financial services, enterprise technology or other regulated customers with … requirements. Experience with container security, policy‐as‐code, image scanning, secrets management, role‐based access control and Kubernetes security hardening. A background in DevOps, SRE, Cloud Operations or Platform Engineering, with a track record of improving automation, reliability and operational maturity. Health Insurance and Dental Health Cover ...

Site Reliability Engineer I

Location
Greater London, England, United Kingdom
drive real change. Constantly grow as you work hard for a mission that matters at a company where you matter. Your Impact As an SRE contributor in Axon's Real Time Operations organization, you are passionate about delivering solutions to the real-time problems our mission-critical cloud native services … encounter. You are also obsessed about achieving the high quality and reliability our customers demand. You will work closely not only with your peers, but also the RTO engineering teams, allowing your technical deliverables to reach the entire engineering organization, enabling product teams to continuously deliver features ...

TechOps Engineer

Hiring Organisation
Addition
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£400.00 per day
environment, taking ownership of incident management, security remediation, CI/CD and automation. The Opportunity As TechOps Engineer, you’ll be responsible for the reliability, security and performance of complex UAT and production environments hosted on Azure. You’ll act as a senior escalation point … automation, scripting and CI/CD. • Participating in a 24/7 on-call rotation. About You You’ll be an experienced TechOps, Infrastructure, SRE or Production Operations professional with a strong technical background and the ability to support complex, business-critical platforms. You will also have: • At least eight ...

Senior Site Reliability Engineer - Kubernetes & Automation

Location
Greater London, England, United Kingdom
FactSet is seeking a Lead Site Reliability Engineer to ensure the reliability, scalability, and performance of our systems. You will design robust infrastructure, automate processes, and drive engineering best practices in a hybrid London-based team. You will monitor production, respond to incidents, define SLOs/… SLIs, and partner with development to bake reliability in from the ground up. Strong Kubernetes and cloud experience are essential. #J-18808-Ljbffr ...

DevOps Engineer, Studios

Location
Greater London, England, United Kingdom
infrastructure and delivery pipelines across our digital and broadcast platforms. The successful candidate will play a key role in enabling continuous delivery, improving system reliability, and supporting high-profile clients and live event services, ensuring optimal performance and resilience across all environments. Key Responsibilities and Accountabilities Design, build … software delivery across multiple teams. Automate infrastructure provisioning using Infrastructure as Code tools such as Terraform, CloudFormation, or similar. Monitor system performance, availability, and reliability using observability tools such as Prometheus, Grafana, and ELK stack. Ensure high availability and disaster recovery strategies are in place and tested regularly. Collaborate ...

Site Reliability Engineer: Scale Global Services

Location
Greater London, England, United Kingdom
Apple Services Engineering is seeking a Site Reliability Engineer to own and improve the massive-scale infrastructure powering App Store, Apple TV, Apple Music, and more. You’ll work across Linux-based systems, build automated tooling, and collaborate with developers to deliver reliable services. Responsibilities include troubleshooting ...

cloud engineer in AI solutions

Location
Greater London, England, United Kingdom
implement intelligent, cloud-native solutions using Google Cloud and Gemini Enterprise for major UK organisations. They work across cloud architecture, AI, and data engineering to build scalable platforms and agentic solutions that drive business transformation. Задачи Create and deploy agentic and AI solution components using Gemini Enterprise Agent Platform … Google Cloud products and services, such as GKE, BigQuery, Cloud Run, and Cloud Storage Have hands-on experience in enterprise technology delivery across software engineering, cloud development, or data engineering Have hands-on experience building and deploying AI solutions using Gemini Enterprise, Gemini Enterprise Agent Platform (previously Vertex ...

Senior Engineering & Architecture Opportunities

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
used worldwide. We are recruiting for individual contributors and leadership roles across a varied technical tech. You'll work on challenging problems across software engineering where performance, data quality, and reliability are non-negotiable, while leveraging modern cloud (AWS) and AI-assisted development tooling to accelerate delivery without … global finance. We welcome professionals with experience in one or more of the following areas: Software EngineeringCloud & Platform EngineeringSite Reliability Engineering (SRE)AI & Machine LearningArchitecture & Distributed SystemsBring your curiosity, expertise, and ambition—and help build what's next for global investment and market intelligence. ABOUT US: LSEG (London ...

SRE Lead: CloudOps, IaC & 24/7 Reliability

Location
Greater London, England, United Kingdom
IQVIA is seeking an experienced Site Reliability Engineering Lead to head a CloudOps team delivering mission-critical cloud services for a UK Public Sector client. The role blends hands-on technical work with leadership to ensure high availability, reliability, and scalability of AWS-based platforms. ...

Senior Sales Engineer

Hiring Organisation
Sumo Logic
Location
London, UK
Employment Type
Full-time
relationship that will provide continuous value to our customers. Finally, you will have the opportunity to work cross-functionally with our Product Management and Engineering teams to share your knowledge and experiences to ultimately improve our business and our customers' success. We seek talent who wants to leverage their … technical validations during the Proof of Value phase Be successful working with all levels of an organization, from executives down to individual developers and Site Reliability EngineersDeliver product and technical demonstrations of the Sumo Logic serviceWork cross functionally with Product Management and Engineering to improve the Sumo ...

Senior Sales Engineer London, England, United Kingdom

Location
Greater London, England, United Kingdom
relationship that will provide continuous value to our customers. Finally, you will have the opportunity to work cross-functionally with our Product Management and Engineering teams to share your knowledge and experiences to ultimately improve our business and our customers’ success. We seek talent who wants to leverage their … technical validations during the Proof of Value phase Be successful working with all levels of an organization, from executives down to individual developers and Site Reliability Engineers Deliver product and technical demonstrations of the Sumo Logic service Work cross functionally with Product Management and Engineering to improve ...

Site Reliability Engineer, London

Location
Greater London, England, United Kingdom
BIG. Operating at our scale, across multiple geographically dispersed data centers and servicing hundreds of millions of users presents unique challenges. As an SRE at Apple, you'll need to solve these problems using data, teamwork, and your own expertise. SREs at Apple own the full infrastructure stack; from device … close partnership with our development teams and aim to design & build new services together. We're passionate about software and automation in SRE and develop a variety of tooling and infrastructure. Our services run on mixed & hybrid platforms. Responsibilities Create outstanding customer experience, and help developers write better code faster ...

Google Cloud Engineer

Location
Greater London, England, United Kingdom
some of the UK’s largest and most strategically important organisations. We work at the cutting edge of cloud architecture, AI, and data engineering to build scalable platforms and agentic that drive real business transformation. Google Cloud Engineer Responsibilities Hands-on creation and deployment of agentic and AI solution … range of Google Cloud products and services (e.g. GKE, BigQuery, Cloud Run, Cloud Storage, etc.) Hands-on experience in enterprise technology delivery spanning software engineering, cloud development, or data engineering Hands-on experience building and deploying AI solutions using Gemini Enterprise, Gemini Enterprise Agent Platform (previously Vertex ...

Senior SRE - Reliability & Automation Lead (Flexible Hours)

Location
City Of London, England, United Kingdom
Elsevier Limited is seeking a Senior Site Reliability Engineer to join embedded teams enhancing reliability, scalability, and performance of critical platforms. You will drive automation, reduce toil, and collaborate with engineering to build resilient systems that deliver strong customer experiences. You will lead reliability initiatives ...

DevOps Engineer

Location
Greater London, England, United Kingdom
troubleshooting complex application and infrastructure issues. Knowledge of networking principles including DNS, load balancing, firewalls, and TLS/SSL. Experience working within DevOps or Site Reliability Engineering practices. Experience with Amazon EKS. Exposure to GitHub Actions, GitLab CI, Azure DevOps, or Jenkins. Knowledge of security best practices … availability, and capacity. Investigate and resolve platform, infrastructure, and application‐related incidents. Work closely with Development, DevOps, and Application Support teams to improve system reliability and deployment processes. Collaborating with developers work closely with software engineers and operations teams to understand business requirements and translate into technical solutions Contribute ...

DevOps Engineer

Location
Greater London, England, United Kingdom
SBOM generation, dependency and artifact provenance, signing/verification of build artifacts, and vetting of third‐party actions/images used in pipelines. Help engineering teams to define infrastructure and deployment standards, and act as a resource for developers adopting IaC and CI/CD best practices. Maintain clear … documentation of infrastructure, runbooks, and security procedures. What you’ll bring 3–5 years of experience in a DevOps, Site Reliability Engineering, or Cloud Infrastructure role. Strong hands‐on experience with the AWS stack (e.g., EC2, VPC, IAM, S3, Lambda, ECS/EKS, CloudWatch, Route 53). ...

Platform Engineer

Location
Greater London, England, United Kingdom
Full-Time Working Pattern: Monday to Friday Start Date: ASAP About the role As a Platform Engineer, you will be responsible for the management, reliability, performance, and continuous improvement of our Kubernetes infrastructure. You’ll play a key role in operating and enhancing our Kubernetes environments, containerised workloads … troubleshooting complex application and infrastructure issues. Knowledge of networking principles including DNS, load balancing, firewalls, and TLS/SSL. Experience working within DevOps or Site Reliability Engineering practices. Experience with Amazon EKS. Exposure to GitHub Actions, GitLab CI, Azure DevOps, or Jenkins. Knowledge of security best practices ...

Site Reliability Engineer: Automation & Observability

Location
Greater London, England, United Kingdom
Apple Inc. is seeking a Site Reliability Engineer in London to join the Apple Services Engineering team. You will help sustain large-scale services powering the App Store, Apple Music, TV, Podcasts and Books for users worldwide. As an SRE, you’ll work with Linux, open source tools and internal software to manage configuration, deployment, logging and monitoring. You’ll collaborate with development teams ...

AI Solutions Support Lead

Hiring Organisation
Axis Capital
Location
London, UK
Employment Type
Full-time
ability to drive efficiency, improve decision making, and enhance customer and employee experiences. As AI capabilities mature and move into production, maintaining their effectiveness, reliability, fairness, and business value becomes critical. The AI Solutions Support Lead is responsible for the ongoing support, maintenance, monitoring, and enhancement of AI solutions … ensure production AI solutions remain stable, performant, compliant, and aligned to evolving business needs. This role acts as the central coordination point between AI engineering, Data Science & AI Delivery, Data Engineering, BTS, business stakeholders, and governance teams to ensure AI solutions continue to deliver measurable value throughout their ...

Technical Site Reliability Engineer

Hiring Organisation
Anduril Industries
Location
London, UK
Employment Type
Full-time
operate together in future contested multi-domain environments. You'll join a small, multinational team of engineers spanning multiple disciplines such as wargaming, game engineering, HPC simulations, LLM agents and VR environments; all to give our warfighters and researchers the leverage to explore faster, test more ideas, and better … current, and trustworthy. A failed scenario run or a silent regression after a software release costs operators and engineers' real time. As our founding Site Reliability Engineer, you will design, build, and operate the infrastructure that makes this possible. You'll work at the intersection of hardware, simulation ...

Technical Site Reliability Engineer

Location
Greater London, England, United Kingdom
operate together in future contested multi-domain environments. You'll join a small, multinational team of engineers spanning multiple disciplines such as wargaming, game engineering, HPC simulations, LLM agents and VR environments; all to give our warfighters and researchers the leverage to explore faster, test more ideas, and better … current, and trustworthy. A failed scenario run or a silent regression after a software release costs operators and engineers’ real time. As our founding Site Reliability Engineer, you will design, build, and operate the infrastructure that makes this possible. You'll work at the intersection of hardware, simulation ...

Site Reliability Engineer - Live Ops & Cloud Resilience

Location
Greater London, England, United Kingdom
seeking an experienced Site Reliability Engineer to design, build, and operate resilient, secure platforms underpinning our digital and live operations. You’ll focus on reliability, observability, automation, and disaster recovery across hybrid environments, collaborating with engineering, operations, and project stakeholders. The role emphasizes improving service availability ...

Site Reliability Engineer I: Cloud-Native & Observability

Location
Greater London, England, United Kingdom
Axon is seeking an experienced Site Reliability Engineer for its Real Time Operations in London. You will contribute to building reliable cloud-native services, collaborate with RTO engineering teams, and enable product teams to scale features globally. We value engineers who write clean code, champion reliability, and create self-service tooling for rapid provisioning and incident response. This role is based in London with on-site work expectations. #J-18808-Ljbffr ...

Data Platform SRE

Location
Greater London, England, United Kingdom
bar on everything they ship. You’ll act as the reliability voice inside the Data Engineering team — not a separate, siloed SRE function — balancing feature velocity against system stability, and bringing SRE practice (SLOs, error budgets, blameless post‐mortems, capacity planning) to a team that has historically … disaster recovery and business continuity planning for data systems, including backup strategy, failover testing, and documented runbooks. Contribute to the broader platform/infrastructure SRE community at Outerlimit, sharing tooling and practices across teams. Support compliance and audit requirements (e.g. SOC 2, data governance) as they relate to data platform ...

Site Reliability Engineer — Production & Incident Response

Location
Greater London, England, United Kingdom
慨正橡扯 is looking for a Site Reliability Engineer to manage incident response and oversee the offshore Production Support team … London. This hybrid position will require you to ensure system reliability and maintain high observability standards. Ideal candidates will bring significant experience in SRE, embracing automation tools to enhance performance. You will be instrumental in protecting client interests and contributing to the company during a vital stage of growth. ...