301 to 325 of 580 Site Reliability Engineering Jobs in London

Platform Engineering Manager (Cloud Foundations) London, UK

Location
Greater London, England, United Kingdom
Platform Engineering Manager (Cloud Foundations) London, UK Our vision is to give everyone the belief they can make their move. We aim to make moving simpler, by giving everyone the best place to turn to and return to for access to the tools, expertise, trust, and belief to make … delivery in line with expectations. Align cloud platform strategy and delivery plans with business goals, partnering with technical product manager, DX, DBA, security, SRE, and data teams. Coach and mentor engineers to improve skills, confidence, and impact. Sets clear, achievable development goals and provides actionable feedback. Aligns individual growth plans ...

AVP, Observability & SRE Engineer

Location
Greater London, England, United Kingdom
Citi is seeking a Site Reliability Engineer - Assistant Vice President in London to drive end‐to‐end observability, migrate legacy monitoring to Google Cloud Observability and Grafana, and implement OpenTelemetry instrumentation across OpenShift/Kubernetes environments. The role emphasizes hands‐on deployment, automation (Ansible/Terraform), and collaboration ...

Senior Database Engineer

Location
Greater London, England, United Kingdom
here the tooling and judgement to write data access that holds up in production. You’ll join a small, agile team of DBAs and site reliability engineers in a hands‐on, senior individual contributor role. You won't be working a ticket queue. You’ll choose the work … that moves the needle on latency, cost and reliability, and you’ll be trusted to call out what needs re-architecting. What You’ll Do Own query and index performance across our transactional SQL Server estate, using execution plans, Query Store, wait statistics, real production telemetry or other data ...

Senior Database Engineer

Hiring Organisation
Capital On Tap
Location
London, UK
Employment Type
Full-time
here the tooling and judgement to write data access that holds up in production. You'll join a small, agile team of DBAs and site reliability engineers in a hands-on, senior individual contributor role. You won't be working a ticket queue. You'll choose the work … that moves the needle on latency, cost and reliability, and you'll be trusted to call out what needs re-architecting. What You'll DoOwn query and index performance across our transactional SQL Server estate, using execution plans, Query Store, wait statistics, real production telemetry or other data driven ...

Senior Platform Engineer - AI Native SaaS Platform

Location
City Of London, England, United Kingdom
2025. Their platform is redefining the sector, and with revenues nearly 10x since the start of last year, they're continuing to expand their Engineering team to match the ambition of their product and customers. The product is real-time, data-rich and AI-native, creating complex engineering … infrastructure as code Strong understanding of networking, IAM, managed services and secure cloud architecture Experience owning observability, logging, APM or distributed tracing tooling An SRE mindset across SLOs, reliability, incident response and toil reduction Clear communication skills and the ability to influence how engineers build and operate services This ...

Site Reliability Engineer - Private Cloud Compute

Location
Greater London, England, United Kingdom
approach to cloud intelligence, extending the security and privacy of Apple devices into the cloud to unlock even more intelligence for our users. This SRE team is responsible for the availability and automation of the critical systems and services that enable PCC to deliver cloud intelligence without compromising user privacy. … future of privacy-preserving cloud infrastructure at scale, this is the opportunity for you! Description We're looking for a hardworking and passionate SRE Engineer to join this amazing team. You will be an accomplished builder and problem-solver, eager to tackle challenging technical problems. You have a deep understanding ...

Lead Observability Engineer

Hiring Organisation
Tria
Location
London, United Kingdom
Employment Type
Contract
days per week (Sheffield as an alternative) Rate: £tbd/day inside IR35 Duration: 6 months+ Are you a Senior Observability Engineer/SRE Lead, with demonstrable experience of assessing and defining observability and monitoring roadmaps within enterprise scale environments? If so, apply now for this new contract opportunity. … Lead Observability Engineer/SRE Lead will be required to assess a complex hybrid estate, understand how services, platforms, infrastructure and networks should be monitored, and work across multiple internal teams, partners and suppliers to build a consolidated view of existing telemetry, monitoring and alerting capabilities. As well as short ...

Lead SRE: Build Reliable, Scalable Systems

Location
Greater London, England, United Kingdom
FactSet is seeking a Lead Site Reliability Engineer to ensure the reliability, scalability, and performance of our systems in a hybrid UK environment. You will partner with development and operations teams to automate workflows, improve observability, and drive reliability initiatives across production services. The role involves ...

Senior Site Reliability Engineer - Cloud & Observability

Location
Greater London, England, United Kingdom
leading global financial markets company is seeking a Senior Engineer in Site Reliability. This role involves maintaining service level objectives, enhancing system reliability, and automating to ensure scalability. With a strong focus on cloud platforms, particularly Azure, candidates should have extensive experience in scripting and infrastructure as code. ...

Senior Platform Engineer – Remote UK (Cloud/SRE)

Location
Greater London, England, United Kingdom
Hudl in London, United Kingdom, is seeking a Senior Engineer to join our Platform Engineering team. You’ll work on site reliability, cloud infrastructure, observability and production operations to keep Hudl’s platform highly available, scalable and secure. You’ll lead with technical excellence, mentor engineers ...

Senior SRE – Remote-First, Cloud-Native Reliability

Location
Greater London, England, United Kingdom
McNally Recruitment Ltd in London seeks a Senior Site Reliability Engineer. The role is largely remote with one day per week in the London office, focused on improving availability, performance, and incident response for cloud-native services. You will work with feature teams to meet service level objectives ...

Staff Cloud SRE for AI/ML Platform & GPU Compute

Location
Greater London, England, United Kingdom
Wayve is seeking a founding Staff Cloud Site Reliability Engineer to shape the reliability of large-scale AI systems and GPU compute infrastructure. … will build and scale the reliability foundations of our AI cloud platform, including model development and GPU compute environments. This is a founding SRE role situated at the intersection of AI research, cloud infrastructure, and operations. You will define frameworks, automate standards, and ensure scalable, secure deployments with ...

Principal DevOps Engineer

Location
Greater London, England, United Kingdom
define standards and reference architectures, lead complex initiatives across Windows/Linux platforms (with deep expertise in Windows clustering), and own end-to-end reliability—from Infrastructure as Code to release engineering and observability. What you will be doing: Act as technical lead for DevOps/Platform/… platform roadmaps What you will bring to the role: Bachelors degree in Computer Science or related field 7+ years (or equivalent) in DevOps/SRE/Infrastructure Engineering , including leadership in complex environments Expert-level experience designing and operating Windows Server HA and clustering (Failover Clustering and related components ...

Software Engineer — Observability Instrumentation

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
ideas take time to evolve. Together we're building a world-class platform to amplify our teams' most powerful ideas. As part of our engineering team, you'll shape the platforms and tools that drive high-impact research - designing systems that scale, accelerate discovery and support innovation across … engineering teams to debug and operate servicesDevOps tooling experience using tools such as Terraform, ArgoCD, Helm or JenkinsInterest in AI engineering and SRE practices to improve incident response and RCADesirable but not essential experience includes: Prometheus/PromQL, VictoriaMetrics, OpenSearch, Grafana or similar observability backendsAuto-instrumentation, distributed tracing ...

Secure Cloud Platform Engineer | DevSecOps & SRE

Location
Greater London, England, United Kingdom
Envitia in the UK is hiring a DevOps/Platform Engineer to design, build and operate secure cloud capabilities within classified government environments. You will work across AWS and cloud-native technologies, delivering reliable, scalable ...

Test Environment Manager

Location
Greater London, England, United Kingdom
This role ensures that test environments are reliable, scalable, automated, observable, and aligned with modern SDLC and DevOps practices. The TEM drives technical excellence, SRE-inspired culture, and continuous improvement across development, QA, and operations teams. Key Operational Responsibilities: Environment Automation & Lifecycle Management Design and implement Infrastructure as Code … based on reliability KPIs. Promote a culture of shared ownership, blameless problem-solving, and strong cross-team collaboration among development, QA, DevOps, and SRE teams. Capacity Planning & Scalability Forecast environment capacity requirements based on usage trends, test cycles, and upcoming projects. Ensure infrastructure elasticity and scalability to meet future ...

Senior DevOps & Cloud Engineer: AWS, CI/CD, Kubernetes

Location
Greater London, England, United Kingdom
Morgan’s Tek is seeking a seasoned DevOps/SRE engineer to own cloud infrastructure, pipelines, and monitoring that keep our grocery platform running. You’ll shape scalable AWS-based systems and reliable CI/CD pipelines. You’ll implement containerised services with Docker and Kubernetes, bake in security ...

Machine Learning Engineer (Applied AI ML)

Location
Greater London, England, United Kingdom
capabilities and results to technical and non-technical audiences Document approaches, techniques, and processes to comply with industry regulation Collaborate with cloud and SRE teams and take a leading role in designing and delivering production architectures Act as an individual contributor; optional management responsibility may be available depending on experience ...

Senior Software Engineer - Database team

Location
Greater London, England, United Kingdom
tooling and judgement to write data access that holds up in production. You’ll join a small, agile team of Database Engineers and Site Reliability Engineers in a hands on, senior individual contributor role. You won’t be working a ticket queue. You’ll choose the work that … moves the needle on latency, cost and reliability and you’ll be trusted to call out what needs re-architecting. What You’ll Do Evangelise advanced database technologies to the software engineering teams. Research, demo and promote technologies such as graph databases and zero-downtime data migrations ...

Senior/Principal Product Manager - Operations

Location
Greater London, England, United Kingdom
work across CMDB, asset and service inventory, internal APIs, workflow automation, third-party integrations, and platform administration, partnering closely with engineering, SRE, infrastructure, networking, security, and operations. This is a highly technical product role focused on reducing operational friction, improving self-service, and creating reliable, scalable internal systems. … requirements, specs, user stories, data models, and priorities. Partner with engineering to define solutions across APIs, integrations, automation, and platform tooling. Work with SRE, infrastructure, security, support, and other internal users to identify pain points and improve automation and self-service. Establish reliable sources of truth for infrastructure, services ...

Principal Engineer I — Prepurchase Platform

Location
City Of London, England, United Kingdom
Summary: JOB DESCRIPTION –Principal Engineer I, Prepurchase Platform Location:London, UK Division:TicketmasterUK Limited LineManager:VP Engineering, Accounts, Identity and Pre-Purchase Contract Terms:Full-time permanent, 40h/per week THE TEAM The Prepurchase Platform group owns the backend services that power every fan's path … strategies. Collaborate with the Enterprise Architecture team on architectural decisions and evolve engineering standards across the Prepurchase domain. Partner with product, security, and SRE teams to align technical decisions with business priorities. Drive observability improvements - ensuring services are instrumented for monitoring, alerting, and rapid incident response. Identifyandeliminatesingle points ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
Senior DevOps Engineer About the team The Core Engineering teams build the systems and tooling that enable Wolt, Deliveroo, and DoorDash engineers to develop, ship, and operate software at scale. We treat infrastructure, CI/CD, and developer tooling as software products, engineered with the same care as customer … Redis, CI/CD such as GitHub Actions) and the ability to weigh trade-offs between them for a given problem. Solid understanding of SRE principles, including incident response, fault-tolerant architecture, and capacity planning. Experience unblocking and mentoring other engineers (e.g. code/design reviews, knowledge sharing, or driving ...

Site Reliability Engineer Infrastructure View role →

Location
Greater London, England, United Kingdom
highly concurrent, agent-driven system. Drive cost and performance optimization across compute-heavy scanning workloads. What we're looking for 4+ years in SRE/infrastructure roles operating production Kubernetes at scale. Strong grasp of network and workload isolation (namespaces, gVisor/Firecracker-style sandboxing, or similar). Experience with ...

SRE Engineer - Global Trading Platform | Cloud + Automation

Location
Greater London, England, United Kingdom
Mandarin Recruitment is looking for a Site Reliability Engineer to join a fast-growing Singapore-based financial services company. The ideal candidate will have a bachelor's degree in computer science and at least three years of experience in DevOps, ideally focusing on infrastructure support. Successful applicants will ...

Senior Platform Engineer

Hiring Organisation
Workable Software Limited
Location
South East London, London, United Kingdom
Employment Type
Permanent, Work From Home
Engineer to take ownership of the platform and infrastructure behind it. As a small, ambitious team, you'll work directly with our founder and engineering team and have significant influence over how our next-generation architecture is designed and built. What You'll Do Design & Scale our Platform: Build … Requirements Degree in a related technical field: Computer science, software engineering or similar. 7+ years of professional experience: in Platform Engineering, Infrastructure, SRE, DevOps or distributed backend engineering. Strong AWS experience , ideally including EKS/ECS, RDS/PostgreSQL, S3 and Lambda. Distributed Systems Experience: Strong understanding ...