51 to 75 of 572 Site Reliability Engineering Jobs in London

Senior Site Reliability Engineer

Location
City Of London, England, United Kingdom
internal workflows that help our people deliver better outcomes for customers, faster. About the role: As a Senior Site Reliability Engineer (SRE), you will play a key role in ensuring the reliability, scalability, and performance of our critical platforms and services. You will lead complex reliability … during incidents. Makes contributions during post-mortems and RCAs. Participates in disaster recovery tests. Implements automation and executes code in production environments. Contributes to SRE knowledge documentation. Supports the deployment, monitoring, and reliability of services integrating AI tools. Design for Reliability Can support architecture and senior engineers ...

Software Engineer, Model Deployment- ChatGPT Engineering

Location
Greater London, England, United Kingdom
Software Engineer, Model Deployment- ChatGPT Engineering Applied AI Engineering - London, UK About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. We develop the systems and tools that make it possible to introduce new models, manage production deployments, respond … operational issues, and use infrastructure effectively at scale. Our work spans distributed systems, platform engineering, infrastructure automation, and developer experience. We partner closely with research, infrastructure, and product teams to make model deployment more reliable, more efficient, and easier to manage. About the Role We are looking ...

Product Engineering Environment Lead

Location
Greater London, England, United Kingdom
Lead Location: UK (Hybrid with occasional travel) Contract: Interim 6 Months Rate: Inside IR35 Level: Senior Manager/Head of Function Reporting to: Product Engineering Director Essential Experience - Please Read Before Applying We are seeking a technically credible transformation leader who can operate at the intersection of engineering … role is likely to suit candidates from backgrounds such as: DevOps Leadership Platform Engineering Leadership Environment Management Site Reliability Engineering (SRE) Engineering Enablement Technology Operations Software Delivery Transformation Telecommunications Technology Large-scale Digital Engineering Organisations This role is NOT primarily looking for: A hands ...

Site Reliability Manager - Environment Strategy

Location
Greater London, England, United Kingdom
looking for a Site Reliability Manager to join our team in London, United Kingdom in a hybrid working mode. In this role, you will lead a team focused on environment strategy, automation, patch governance and operational reliability for AWS-based platforms. Your responsibilities include setting roadmaps, driving … while ensuring strong technical standards, compliance and resilience across all production and non-production systems. Responsibilities Define and own the vision and roadmap for site reliability and environment strategy Lead, mentor and develop a team of DevOps and environment engineers Set and enforce standards for environment provisioning, lifecycle ...

AWS DevOps Engineer

Location
Greater London, England, United Kingdom
Cost Explorer, Compute Optimizer, and EKS workload right‐sizing. Skills: 3+ years of hands‐on experience in DevOps, Site Reliability Engineering (SRE), infrastructure engineering, or closely related roles (focused on AWS/EKS environments). Proven track record building and maintaining production‐grade CI/… incidents/post‐mortems into actionable learning for personal and team growth. Proven ability to collaborate effectively with cross‐functional stakeholders, including engineering, SRE, security, and product teams. Maintain clear technical documentation, runbooks, and operational playbooks for production systems. Chinese proficiency is preferred We Offer: Experience a dynamic ...

Site Reliability Engineer , Cryptography, Access and Identity Services

Hiring Organisation
AmazonWebServices
Location
London, UK
Employment Type
Full-time
availability environment, building and operating critical Cryptography, Access and Identity services for our customers. This exciting role is designed for someone with a strong engineering background and a passion for driving efficiency, quality, and process improvements within our service operations. We are an operations team, but a key focus … supported in the workplace and at home, there's nothing we can't achieve. Basic qualifications- Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration- Experience working with Linux- Experience in systems engineering- Experience ...

Site Reliability Engineer - NS London

Hiring Organisation
BAE SYSTEMS
Location
London, UK
Employment Type
Full-time
maintained. This role blends operational product support with software engineering to create applications to understand the overall health of our systems. The SRE team sits within a wider programme at the core of the customer mission. The role holder: As an SRE, fundamentally you will be doing work that … human labour, with the objective of limiting traditional manual operations work (incident tickets, on-call etc.) to no more than half of the SRE team's time (and aiming for considerably less). You will have an enthusiasm to learn and experiment, to develop tools to understand application health ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
trading technology stack is undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self‐healing workflows, and resilience engineering. Use enterprise‐authorized AI capabilities within the work environment to accelerate major‐incident triage ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
trading technology stack is undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage ...

Security Engineer (Site Reliability Engineering) - SC Cleared

Hiring Organisation
Sanderson Government and Defence
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£545 - £590 per day + Inside IR35
Length: 6-18 Months Rate: £545-£590 per day (Inside IR35) Positions Available: 2 About the Role We are seeking two experienced Security Engineers (SRE) to join a specialist consultancy delivering cyber security services across a portfolio of government projects and digital transformation programmes. This role is ideal for security … best practices. Drive continual improvements across security engineering and platform security functions. Essential Skills & Experience Strong experience as a Security Engineer, Security-focused SRE, Platform Security Engineer, or similar role. Advanced knowledge of Enterprise Security Architecture principles. Hands-on experiencewithSplunk, including: Security monitoring Dashboard development Alerting and reporting ...

Lead Site Reliability Engineer London

Location
Greater London, England, United Kingdom
core engineering concerns - AI is the primary lever for doing that at scale, not a bolt on. You will be our dedicated SRE, working alongside Systems Engineering and embedded with Product and Engineering across roughly 120 engineers. Teams own their services and their own on-call. … runbooks an agent can execute rather than prose that describes, and diagnostics good enough that an engineer or an agent reaches resolution without an SRE in the room. You will be the primary advocate for reliability with Product and Engineering, making the case with data rather than assertion ...

Site Reliability Engineer – 11863CF

Location
Greater London, England, United Kingdom
11863CF £500 – 580 per day Site Reliability Engineer – DV Cleared Our client is urgently looking for an experienced Site Reliability Engineer to join their team on a contract basis, initially for 6 months with a view to extend. Please note, the role is OUTSIDE of IR35. … must hold live UKIC DV. The role is on-site 4 days per week in Gloucester. Site Reliability Engineer – Key Skills: Must hold a live UKIC DV Modern configuration management tools (such as Ansible, Chef or similar) Experience working with Terraform Docker containers & container orchestration tools (such ...

Site Reliability Engineer / Production Support

Location
Greater London, England, United Kingdom
fastest growing fintech in 2025. The momentum is real. THE OPPORTUNITY Monument’s production environment is the heartbeat of a licensed bank, and the SRE role is the single point of ownership when incidents occur. You will directly oversee the offshore Production Support team, run on-call and incident response … eliminate it. Quality-driven - you care about alert quality, observability standards, and reliability patterns that prevent problems at source. WHAT YOU BRING Strong SRE or production support experience with accountability for incident response in a production environment. Deep understanding of observability tools, alerting, logging, and distributed systems debugging. Experience ...

Lead Engineer, Site Reliability Engineering

Location
Greater London, England, United Kingdom
## Our TeamWe are evolving our Reliability Engineering team to move beyond support and operations. As a Senior Engineer in Site Reliability, you will be part of a diverse and inclusive organization that has full ownership of the availability, performance, and scalability … purpose. Write automation to scale systems sustainably, prevent service issues, or when they occur, quickly recover service. Partner with development teams to improve system reliability, observability, and release velocity. Participate in on-call rotations, incident response, postmortems, and root cause analysis and resolution. Be a vocal advocate of strong ...

Site Reliability Engineer - Sports Betting

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Site Reliability Engineer – Fintech/Linux100,000Quant Capital is urgently looking for a Site Reliability Engineer to join our high profile client. Our client is a data company specialised in the aggregation and global distribution of real time sports betting odds. They use unique online data … trading environment for their clients. They have grown massively and recently were voted in the top 50 fintech firms globally. Day to Day the Site Reliability Engineer will: Ensuring the 24/7 availability of systems, infrastructure and our real-time, low latency data services. Designing, building ...

Site Reliability Engineer (Chinese speaking, £90k, Financial Services,Technology, London)

Location
Greater London, England, United Kingdom
have an exciting opportunity for a Site Reliability Engineer role in a Singapore‐based financial services company. The company is a fast‐growing global trading platform looking for a Site Reliability Engineer to bridge development and operations using software engineering. Job Requirements A bachelor's degree … computer science or a closely related field is required. Three or more years of DevOps, Site Reliability, and Infrastructure Support experience. Linux technical abilities (CentOS preferred) including shell scripting, Docker, and Ansible. Familiarity with Python (preferred) and/or Go or Ruby. Experience with cloud infrastructure such ...

Lead SRE - Chase UK

Hiring Organisation
JP Morgan Chase
Location
London, UK
Employment Type
Full-time
building the bank of the future from the ground up, offering you the chance to join us and make a significant impact. As a Site Reliability Engineer at JPMorgan Chase within the International Consumer Bank, you will play a crucial role in this initiative, dedicated to delivering … oriented and possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with ...

Lead SRE - Chase UK

Location
Greater London, England, United Kingdom
building the bank of the future from the ground up, offering you the chance to join us and make a significant impact. As a Site Reliability Engineer at JPMorgan Chase within the International Consumer Bank, you will play a crucial role in this initiative, dedicated to delivering … oriented and possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with ...

Principal Site Reliability Engineer

Location
Greater London, England, United Kingdom
driven workflows, trusted data, and seamless collaboration, to deliver the insight and context needed for confident, competitive decision-making. The Opportunity As a Principal Site Reliability Engineer at Veson Nautical, you will design, build, monitor, and support the cloud infrastructure that underpins our rapidly growing SaaS platform. This … observable, and easier to operate. You will have significant influence over the architectural direction of the platform. The Team You'll join a global Site Reliability Engineering team with members in the United States and the United Kingdom. This is a senior individual contributor role without direct ...

SRE Engineer

Hiring Organisation
Ricoh
Location
London, UK
Employment Type
Full-time
role in driving reliability, performance, and operational excellence across Ricoh's hybrid cloud and on‐prem environments. You will help shape SRE practices, support incident and problem management, embed automation, and ensure infrastructure operations meet the highest security and compliance standards, including ISO 27001.This is a hands‐on technical … capacity, and scalabilityLeading and contributing to root‐cause analysis and major incident reviewsSupporting a blameless post‐mortem culture with clear action trackingDefining and implementing SRE practices, tooling, and engineering standardsDriving infrastructure‐as‐code and automation across Azure and on‐premImproving image bakery pipelines for secure, repeatable server buildsEmbedding observability ...

SRE Engineer

Location
Greater London, England, United Kingdom
role in driving reliability, performance, and operational excellence across Ricoh’s hybrid cloud and on‐prem environments. You will help shape SRE practices, support incident and problem management, embed automation, and ensure infrastructure operations meet the highest security and compliance standards, including ISO 27001. This is a hands … Leading and contributing to root‐cause analysis and major incident reviews Supporting a blameless post‐mortem culture with clear action tracking Defining and implementing SRE practices, tooling, and engineering standards Driving infrastructure‐as‐code and automation across Azure and on‐prem Improving image bakery pipelines for secure, repeatable server ...

Site Reliability Engineer (remote working)

Location
Greater London, England, United Kingdom
opportunity for someone who enjoys working at a technical level but also wants genuine ownership of projects and the opportunity to influence how SRE and DevOps are delivered across a large enterprise environment. The team is continuing to develop its SRE capability, with a significant pipeline of projects focused … resilience of cloud-based services Reducing manual intervention across deployment and release processes Automating repetitive operational tasks and reducing engineering toil Helping develop SRE practices around SLOs, error budgets and reliability Supporting and improving Kubernetes environments used by engineering teams Working with Azure DevOps and Infrastructure ...

DevOps Engineer

Location
Harrow, England, United Kingdom
adjacent cloud foundation and infrastructure initiatives. Key Responsibilities Build and maintain reusable IaC modules and automation using Terraform. Develop self-service capabilities for engineering teams to provision and manage approved AWS resources. Provision and operate AWS services including EC2, EKS, RDS, S3, MSK, ElastiCache, IAM, and CloudWatch. Build … prem environments, and other cloud platforms we use. Implement and maintain secure IAM, encryption, network security, secrets management, and compliance controls. Improve the reliability, scalability, security, and operational supportability of shared platforms. Automate manual processes and improve the consistency of infrastructure and application deployments. Support development teams with infrastructure ...

Senior Site Reliability Engineer

Hiring Organisation
Pathfinder Business Solutions Ltd
Location
City, London, United Kingdom
Employment Type
Permanent
Salary
GBP 90,000 Annual
Senior Site Reliability Engineer (SRE) London, hybrid, one day a week onsite Up to £90,000 plus 10% cash allowance, 10% non contributory pension and discretionary bonus We're looking for a Senior Site Reliability Engineer with hands on Azure, Kubernetes and Terraform experience to join ...

Senior Platform Engineer (12 Month FTC)

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
leadership in shaping, evolving, and scaling our clients cloud platform. The role will establish robust, reusable platform capabilities and self-service solutions that enable engineering teams to deliver software faster, more reliably, and with a consistently high developer experience. The Principal Platform Engineer will operate across architecture, engineering … management. Experience establishing repeatable, standardised engineering workflows. Reliability & Performance Engineering Designing platforms for resilience, fault tolerance, observability, and operational excellence. Applying SRE principles and practices to improve availability and reduce operational risk. Performance analysis, capacity planning, scalability engineering, and proactive reliability improvement. Establishing meaningful service ...