76 to 100 of 538 Site Reliability Engineering Jobs in London

Site Reliability Engineer

Hiring Organisation
Anson Mccade
Location
london, south east england, united kingdom
Troubleshooting problems across the application, infrastructure and platform stack Working with cloud infrastructure and containerised/microservices environments Contributing to the wider DevOps and SRE community Helping introduce new technologies and approaches that improve reliability and operational efficiency What were looking for Youll ideally have experience across a number … opportunity to work on technically challenging projects where your engineering skills can have a genuine impact. You'll be part of a growing SRE/DevOps community, working with modern technologies across cloud, automation, monitoring and distributed systems, while having the opportunity to develop your skills and work ...

DevOps Engineer

Location
Greater London, England, United Kingdom
adjacent cloud foundation and infrastructure initiatives.## **Key Responsibilities*** Build and maintain reusable IaC modules and automation using Terraform.* Develop self-service capabilities for engineering teams to provision and manage approved AWS resources.* Provision and operate AWS services including EC2, EKS, RDS, S3, MSK, ElastiCache, IAM, and CloudWatch.* Build … prem environments, and other cloud platforms we use.* Implement and maintain secure IAM, encryption, network security, secrets management, and compliance controls.* Improve the reliability, scalability, security, and operational supportability of shared platforms.* Automate manual processes and improve the consistency of infrastructure and application deployments.* Support development teams with infrastructure ...

Senior Platform Engineer (12 Month FTC)

Location
Greater London, England, United Kingdom
leadership in shaping, evolving, and scaling our clients cloud platform. The role will establish robust, reusable platform capabilities and self-service solutions that enable engineering teams to deliver software faster, more reliably, and with a consistently high developer experience. The Principal Platform Engineer will operate across architecture, engineering … management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. Establishing meaningful service ...

Principal Platform Engineer (12 Month FTC)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
leadership in shaping, evolving, and scaling our clients cloud platform. The role will establish robust, reusable platform capabilities and self-service solutions that enable engineering teams to deliver software faster, more reliably, and with a consistently high developer experience. The Principal Platform Engineer will operate across architecture, engineering … management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. Establishing meaningful service ...

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
Site Reliability Engineer (SRE) - Assistant Vice President is a technical professional responsible for the hands-on execution, technical implementation, and deployment of SRE and observability principles in a complex, critical, and large-scale multi-disciplinary environment. In this role, you will apply a deep understanding of multiple technology … authority, authoring reusable deployment solutions, configuring telemetry collectors, and providing direct technical onboarding support to application teams. We are seeking a passionate and experienced SRE to join our Production Management team. In this role, you will be instrumental in executing our strategy for end-to-end observability and resiliency, collaborating ...

AI Technical Platform Leader

Location
Greater London, England, United Kingdom
business stakeholders to mold and implement strategy. The role applies broad technical knowledge with depth in generative AI, agentic systems, enterprise platforms and engineering governance to ensure that AI platforms at WTW are optimally configured to achieve company vision and imperatives. It has end-to-end ownership of implementation … operations in a large, global enterprise: Software or platform engineering for global, enterprise-scaled solutions Management of Site reliability engineering (SRE) programs for mission‐critical systems Creation and management of DevOps practices for automated, consistent, and secure solution deployment in regulated environments. Literacy in global compliance ...

Senior SRE Engineer - ASE Traffic & Secure Services Network

Location
Greater London, England, United Kingdom
Senior SRE Engineer - ASE Traffic & Secure Services Network Shanghai, Shanghai, China Software and Services At Apple, we build systems that power services used by hundreds of millions of people around the world, and every second counts. The Services Engineering organization is at the heart of this mission, ensuring … platforms are performant, secure, and always available. We're seeking a technically strong Site Reliability Engineer (SRE) to join our growing London team, focused on the future of traffic management, load balancing, and secure networking infrastructure.You’ll play a key role in shaping the next generation ...

site reliability engineer

Location
Greater London, England, United Kingdom
engineering, cloud, and AI-enabled transformation services, focusing on complex software product development and digital platform engineering. Задачи Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices Define … deliver systemic improvements across production environments Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services Drive resilience strategies with highly available architectures and disaster recovery readiness Champion an automation-first ...

Site Reliability Engineer - Frontend

Hiring Organisation
Capital On Tap
Location
London, UK
Employment Type
Full-time
just getting started! ðLondon, Old Street | ð 2 Days in OfficeSRE at Capital On Tap ðAt Capital On Tap, we run a hybrid embedded SRE model - We aim to work closely with the teams to provide them the best support. As a Site Reliability Engineer (SRE) you will … apply. Interview process ðFirst stage: 30 minute intro and values call with Talent PartnerSecond stage: 60 minute CV overview and technical chat with the SRE team lead and the SRE & Platform Engineering ManagerThird stage: 75 minute technical exercise & questions with the SRE lead Final stage: 30 minute chat with ...

Senior Software Engineer

Location
Greater London, England, United Kingdom
contribute to modern digital capabilities, drive continuous improvement, and support the delivery of future-ready solutions. This is an opportunity to shape high-quality engineering outcomes, embrace innovation and AI-enabled ways of working, and create lasting value in a complex, enterprise-scale environment.Hybrid working:The places that … principles, secure software development practices, and security-focused engineering approaches.Experience working within regulated Financial Services environments.Understanding of Site Reliability Engineering (SRE) concepts, operational resilience, and service reliability practices.Relevant cloud, DevOps, or engineering certifications.We are a Disability Confident Employer:Capgemini is proud ...

Senior Platform Engineer

Location
Greater London, England, United Kingdom
employ over 130 colleagues across Jersey, Geneva, London, Singapore, New York and Shanghai. Purpose and Overview of Role This is a hands‐on senior engineering role within the Platform Engineering team, which forms part of Technology Operations. Platform Engineering is responsible for building and operating the infrastructure … platforms and developer tooling that enable our engineering and quantitative research teams to deliver software reliably, securely and at scale. The role will contribute to the design, build, automation and operation of a hybrid production platform across AWS and on‐premises environments, with a particular focus on the HashiCorp ...

Devops SRE

Location
Greater London, England, United Kingdom
Cloud Engineering team is seeking a seasoned and passionate Senior Cloud Engineer with deep hands‐on development and cloud engineering expertise. In this role, you will serve as a key technical contributor within a cloud‐focused engineering team, working on one of the Group’s flagship initiatives … best practices and business goals. Required Skills & Experience Core Cloud & DevOps Competencies Extensive experience in DevOps or Site Reliability Engineering (SRE) roles across consumer or SaaS environments. Strong expertise in deploying and managing production‐grade Kubernetes clusters and containerised services. Hands‐on experience with Kubernetes ...

Platform Engineer

Location
Greater London, England, United Kingdom
Platform Engineer Department: Technology Employment Type: Permanent - Full Time Location: London Reporting To: Segun Ikuesan Description This is a hands‐on engineering role within the Platform Engineering team, which forms part of Technology Operations. Platform Engineering is responsible for building and operating the infrastructure, platforms and developer … tooling that enable our engineering and quantitative research teams to deliver software reliably, securely and at scale. The role will contribute to the design, build, automation and operation of a hybrid production platform across AWS and on‐premises environments, with a particular focus on the HashiCorp platform, including Nomad ...

Venue & Studio Deployments System Engineer, Event Productions

Hiring Organisation
Amazon
Location
London, UK
Employment Type
Full-time
Bachelor's degree in Systems Engineering, Computer Science, or related field or relevant work experience- Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration- Experience working with Linux- Experience in systems engineering- Experience ...

Staff Software Engineer, AI Reliability Engineering

Hiring Organisation
Humanloop
Location
London, UK
Employment Type
Full-time
systems. About the RoleClaude has your back. AIRE has Claude's. Help us keep Claude reliable for everyone who depends on it. AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths -- every hop from the SDK through … comes from people who've built product stacks, scaled databases, run massive distributed systems, and everything in between. Strong candidates may alsoHave been an SRE, Production Engineer, or in similar reliability-focused roles on large scale systemsHave experience operating large-scale model serving or training infrastructure (>1000 GPUs).Have ...

AI Technical Platform Leader

Hiring Organisation
WTW
Location
London, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
business stakeholders to mold and implement strategy. The role applies broad technical knowledge with depth in generative AI, agentic systems, enterprise platforms and engineering governance to ensure that AI platforms at WTW are optimally configured to achieve company vision and imperatives. It has end-to-end ownership of implementation … operations in a large, global enterprise: * Software or platform engineering for global, enterprise-scaled solutions * Management of Site reliability engineering (SRE) programs for mission-critical systems * Creation and management of DevOps practices for automated, consistent, and secure solution deployment in regulated environments. * Literacy in global compliance ...

AWS Cloud Architect, Technology Consulting

Location
Greater London, England, United Kingdom
security and automation technologies. You will help design scalable, secure and resilient cloud solutions for our clients, enabling digital transformation through modern architecture, platform engineering and DevOps practices. You should bring practical experience of cloud technologies and infrastructure modernisation, alongside an understanding of how AI-enabled tooling and modern … capabilities Strong understanding of cloud networking, identity, security and resilience patterns Experience supporting cloud engineering, DevOps or Site Reliability Engineering (SRE) teams Cloud architecture certifications in additional public cloud platforms (Azure or GCP) Experience using AI or agentic techniques within the role through GitHub Copilot, Claude ...

Systems Engineer, Cryptography, Access and Identity Services

Hiring Organisation
AmazonWebServices
Location
London, UK
Employment Type
Full-time
availability environment, building and operating critical Cryptography, Access and Identity services for our customers. This exciting role is designed for someone with a strong engineering background and a passion for driving efficiency, quality, and process improvements within our service operations. As a Systems Engineer at Amazon you will utilize … supported in the workplace and at home, there's nothing we can't achieve. Basic qualifications- Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration- Experience working with Linux- Experience in systems engineering- Experience ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
+ 3 years in Computer Science, Software Engineering, or a related technical field. Proven record as a Technical Lead or Lead SRE/Software Engineer delivering distributed, high-availability SaaS platforms at scale. Strong proficiency in Python, Go, Java, or C++ with experience designing microservices, APIs, and production automation. … Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA. Preferred Qualifications AI & Agentic Systems: Hands-on experience building LLM pipelines, AI Agents, Model Context ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
welcome; join us and let’s build what’s next - together! Role Overview We are seeking a ML Ops Engineer to join our Platform Engineering team at Anaplan. In this role, you will design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our cutting-edge AI-infused … tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Your Skills Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ML infrastructure. Proven track record of deploying, scaling, and operationalising machine learning models ...

Systems Engineer, Database Services (AWS)

Location
Greater London, England, United Kingdom
supported in the workplace and at home, there’s nothing we can’t achieve. Basic Qualifications Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration Experience working with Linux Experience in any of the following: Python, Java, Perl ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
Site Reliability Engineer Reports to: Labs Team Lead Looper Insights Remote-first, with regular visits to our Byfleet and Hounslow data centres The company Looper Insights builds analytics products that help the world’s leading media and entertainment companies understand how their content is performing across digital platforms … collaborate closely to produce a valuable service for an industry about which we are all passionate. The role We’re looking for a Site Reliability Engineer to keep the global LooperBox fleet running, the physical backbone behind every piece of data Looper Insights produces. LooperBoxes sit in front ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
such as Elasticsearch Exploring and adopting new technologies, including emerging AI tools and technologies We’re looking for someone with a strong background in SRE, DevOps, production engineering or infrastructure engineering, combined with solid Oracle experience. You don't need to tick every single box. If you have … strong core SRE/DevOps background and particularly good Oracle experience, we'd still be keen to hear from you. Security Clearance Due to the nature of the work, you must be a UK national and either already hold SC clearance or be willing and eligible to obtain it. What ...

Sr. Observability Engineer – Kings Cross, London

Location
Greater London, England, United Kingdom
data for swift root cause identification. Drive post-incident reviews and implement long-term solutions to enhance system resilience.* Collaborate & Influence: Partner with Development, SRE, and Infrastructure leaders to embed observability into the entire technology lifecycle. Influence and drive the adoption of observability best practices across the global organization. Champion … this.**Job Requirements:**Essential Qualifications* Experience: 5-7+ years of hands-on experience in an Observability, Site Reliability Engineering (SRE), or DevOps role, with a proven track record of leading complex projects.* Technical Leadership: Demonstrated experience in architecting and designing large-scale monitoring and observability solutions. ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
Title Site Reliability Engineer Location Remote with 1 day per week in London Salary up to £75,000 base The Company We are working with an established technology business that has just secured its first external round of funding after more than 10 years of organic growth. … hands on role where your judgement will be trusted and your input valued. What They’re Looking For Around 3 to 5 years in SRE, DevOps, or production infrastructure roles Strong AWS experience across services such as EC2, IAM, EKS, RDS, and networking Infrastructure as Code experience, ideally Terraform Familiarity ...