101 to 125 of 157 Site Reliability Engineer Jobs in London

Site Reliability Engineer, Infrastructure - ThousandEyes

Location
City Of London, England, United Kingdom
deeply integrated across the Cisco technology portfolio, delivering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios. Our distributed Site Reliability Engineering team of approximately nine engineers owns the availability, latency, performance, efficiency, monitoring, emergency response, and capacity planning of the platform while partnering … operational on-call rotation. Hands-on experience with infrastructure-as-code tooling and codebases, preferably Terraform. Hands-on experienceleveraging AIas a force multiplier of SRE activities, such as automati ng toil away and improving operational efficiency. Professional experience administering and troubleshooting GNU/Linux systems, including system libraries, file systems ...

Site Reliability Engineer - Private Cloud Compute

Location
Greater London, England, United Kingdom
approach to cloud intelligence, extending the security and privacy of Apple devices into the cloud to unlock even more intelligence for our users. This SRE team is responsible for the availability and automation of the critical systems and services that enable PCC to deliver cloud intelligence without compromising user privacy. … future of privacy-preserving cloud infrastructure at scale, this is the opportunity for you! Description We're looking for a hardworking and passionate SRE Engineer to join this amazing team. You will be an accomplished builder and problem-solver, eager to tackle challenging technical problems. You have a deep ...

Staff Platform Site Reliability Engineer

Hiring Organisation
Index Exchange
Location
London, United Kingdom
Salary
£ 80 K
users—it just works.Mentor and raise the bar. Coach engineers, foster a culture of engineering excellence, and collaborate across Cloud Platform Operations, SRE, Network, Security, and Software Engineering teams.What You BringYou've built and shipped platform infrastructure at scale—not just operated someone else's. You think in systems … ambiguity, and you care as much about the developer experience of your platform as you do about its architecture.Must Have8+ years in platform engineering, SRE, infrastructure engineering, or DevOps.Deep experience with Linux internals: kernel tuning, network stack, system observability, security.Strong Kubernetes expertise: cluster lifecycle, networking, storage, RBAC, multi-cluster—across ...

Site Reliability Engineer – Privacy‐First Cloud at Scale

Location
Greater London, England, United Kingdom
Apple Inc. in London, England invites a hardworking SRE Engineer to join the Private Cloud Compute team. You will tackle challenging problems, own responsibilities for high-availability systems and contribute to global-scale infrastructure with privacy-centric principles. This role emphasizes automation, performance optimization and secure, scalable service delivery. ...

Technical Lead - Site Reliability Engineering

Location
Greater London, England, United Kingdom
Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Markets and Risk Intelligence division.As a **Technical Lead SRE**, you will be a senior hands‐on technical person help shape the foundations of reliability across both new and existing platforms. You will collaborate … person who is passionate about reliability engineering and who bring a continuous improvement approach to everything they do!Lead the establishment of SRE foundations for new projects building environments, monitoring, alerting, and ensuring operational readiness from day one.Collaborate with Architecture and Engineering teams to embed reliability, scalability, security ...

Systems Engineering Manager, Site Reliability Engineering, ML Compute

Hiring Organisation
Google
Location
London, United Kingdom
Salary
£ 80 K
technical field involving coding (e.g., physics or mathematics).Track record of mentoring technical leads.Proven success leading and influencing multiple technical teams.Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google's services—both our internally … critical and our externally-visible systems—have reliability, uptime appropriate to users' needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work ...

Senior SRE Engineer — Cloud Reliability & Automation

Location
City of Westminster, England, United Kingdom
Google London, UK is seeking a Software Engineer III in Site Reliability Engineering for the GCE AI team. This mid-level role focuses on building reliable, scalable systems, code development, and mentoring junior team members. The position emphasizes deep expertise in distributed systems, problem solving, and collaboration ...

Senior Site Reliability Engineer - Cloud & Observability

Location
Greater London, England, United Kingdom
leading global financial markets company is seeking a Senior Engineer in Site Reliability. This role involves maintaining service level objectives, enhancing system reliability, and automating to ensure scalability. With a strong focus on cloud platforms, particularly Azure, candidates should have extensive experience in scripting and infrastructure ...

Principal Platform Engineer (SRE/Cloud)

Hiring Organisation
Beamery
Location
London, United Kingdom
Salary
£ 120 K
trust, empathy & honesty ensuring our workforce is able to bring their full selves to work.ABOUT THE ROLEPrincipal Platform Engineers at Beamery solve the toughest reliability, scalability and infrastructure problems with the highest impact. Together they collaborate to set the standards for how Engineering will build, run and operate services … across the whole engineering organisationWHO ARE WE LOOKING FOR We are seeking a hands-on technical leader with deep Site Reliability Engineering (SRE) and Cloud expertise who can set direction across the engineering organisation. Key skills/experience:A proven track record of designing and delivering scalable, reliable ...

Principal Platform Engineer (SRE/Cloud)

Location
Greater London, England, United Kingdom
honesty ensuring our workforce is able to bring their full selves to work. ABOUT THE ROLE Principal Platform Engineers at Beamery solve the toughest reliability, scalability and infrastructure problems with the highest impact. Together they collaborate to set the standards for how Engineering will build, run and operate services … whole engineering organisation WHO ARE WE LOOKING FOR? We are seeking a hands-on technical leader with deep Site Reliability Engineering (SRE) and Cloud expertise who can set direction across the engineering organisation. Key skills/experience: A proven track record of designing and delivering scalable, reliable cloud ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
Title Platform Engineer Location Leeds or London, UK (3 days per week on-site; travel & accommodation expenses covered) Position Type Permanent Salary £80,000 - £85,000 GBP per annum (Flexible for exceptional candidates) Security Clearance Active SC Clearance Required Notice Period Immediate Joiners or up to 2 Weeks … About the Role We are seeking a Senior Platform Engineer to design, build, and maintain efficient, scalable, and reliable platform solutions. In this role, you will optimize AWS resources, drive automation across CI/CD pipelines, and collaborate directly with development and strategy teams to elevate engineering standards across ...

Senior AWS Site Reliability Engineer

Hiring Organisation
Spectrum IT Recruitment Limited
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£70,000
Datadog, PagerDuty, or Rundeck Experience using configuration management platforms like Ansible, Puppet, or Chef Professional certifications in cloud DevOps, such as AWS Certified DevOps Engineer or Google Cloud Professional DevOps Engineer, or similar credentials Do You Have What It Takes? 3-6 years of hands-on experience … similar role, with a strong emphasis on systems engineering, automation, and service reliability Proficient in at least one programming language such as Python, Go, Java, or C#, along with scripting skills in Bash or PowerShell Solid grasp of cloud platforms like AWS, including an understanding of how core services ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
Take on ambiguous reliability, scalability, and efficiency challenges and drive solutions across SRE and development teams Build and run large-scale, massively distributed, fault-tolerant systems supporting the Genesis platform Optimize existing systems and build infrastructure Eliminate toil through automation to improve uptime and rate of change Cultivate … culture of reliability throughout the organization Guide technical decisions balancing system health with product priorities Ensure long-term health, maintainability, and reliability of services Perform capacity planning and performance analysis Proactively prevent incidents Work across teams to build robust, reusable solutions Requirements Strong software engineering skills in Python ...

Site Reliability Engineer- Spacetime UK

Location
Greater London, England, United Kingdom
Role Overview This isn't a "keep the lights on" SRE role. This is a strategic, high-impact opportunity to build the nervous system for a platform that transforms how networks of satellites, ground stations, and fleets are interconnected and orchestrated. You will be building the core observability stack that … cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade observability stack from the ground up, this role ...

Site Reliability Engineer II — Commercial Cloud

Location
Greater London, England, United Kingdom
CrowdStrike is seeking an Engineer II for the TechOps SRE team focused on our Commercial Cloud. You will be a deeply technical, hands-on engineer building automation and tooling to ensure mission-critical services run reliably across thousands of servers. You will work with Linux engineering, on-call ...

Senior or Staff Software Engineer, SRE/ Platform Team

Location
Greater London, England, United Kingdom
Senior or Staff Software Engineer, SRE/Platform Team OneSignal is a leading omnichannel customer engagement solution, powering personalized customer journeys across mobile and web push notifications, in-app messaging, SMS, and email. On a mission to democratize customer engagement, we enable businesses to keep their 1.5B monthly active … Go. This potent combination of high performance with efficient resource utilization has given us an incredible competitive edge. We are seeking a Platform Engineer to join our team and help us scale by managing and developing the next generation of our infrastructure. While we currently maintain a 99.95 % uptime ...

Site Reliability Engineer

Hiring Organisation
GoCardless
Location
London, United Kingdom
Salary
£ 70 K
About usGoCardless is a global bank payment company. Over 100,000 businesses, from start-ups to household names, use GoCardless to collect, manage and send bank payments through Direct Debit, real-time payments and open ...

Site Reliability Engineer

Hiring Organisation
Wheely
Location
London, United Kingdom
Salary
£ 80 K
About WheelyWheely is redefining premium transportation across major cities in Europe, the US, and the Middle East. We blend cutting-edge technology with the craft of five-star chauffeuring to deliver an experience trusted by ...

SRE Engineer

Location
Greater London, England, United Kingdom
client is looking for a SRE Engineer combining software and IT engineering principles to build and maintain reliable, scalable, and high-performing systems. Job Responsibilities: Scope technical projects and break them down into user stories and tasks within an engineering team Directly contribute to the design and coding ...

Software & Site Reliability Engineer for AI Platform

Location
Greater London, England, United Kingdom
Jobtailor seeks a senior software engineer to build and enhance enterprise capabilities using AI agent visibility and governance. You will own reliability, improve observability, and strengthen infrastructure while collaborating with stakeholders to streamline customer workflows. You will lead initiatives around incident response and automation, contributing to the technical ...

Senior Site Reliability Engineer

Hiring Organisation
CISCO Systems
Location
London, United Kingdom
Salary
£ 70 K
risk.Adaptable & Problem-Solver: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance.Ownership & Quality: Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments and full reproducibility, including backup ...

Senior Site Reliability Engineer

Location
City Of London, England, United Kingdom
Adaptable & Problem-Solver : Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality : Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments and full reproducibility, including ...

Site Reliability Engineer

Hiring Organisation
CISCO Systems
Location
London, UK
Employment Type
Full-time
Adaptable & Problem-Solver: Address complex challenges across configuration, policy, observability, and data services. Apply a data-driven approach using Prometheus and Grafana to improve reliability and performance. Ownership & Quality: Own end-to-end configuration quality, enforcing governance with Open Policy Agent. Ensure secure, compliant deployments and full reproducibility, including ...

Staff Site Reliability Engineer

Location
Greater London, England, United Kingdom
that usable, and MystraAI is the agentic layer we are building on top of it. This is a Staff-level role that owns the reliability, performance, security and integrity of that infrastructure end-to-end — and sets the technical direction that other teams build on. You will lead … source level rather than as a black box — and ideally have contributed code upstream. Reliability engineering for data platforms. You bring true SRE discipline — SLOs, observability, capacity planning and incident response — to analytical data systems and pipelines. Data-as-a-Service productisation. You think in terms of data ...

SRE Engineer: Build Reliable, Scalable Systems

Location
Greater London, England, United Kingdom
leading recruitment agency is seeking an SRE Engineer to combine software and IT engineering principles for building reliable systems. Responsibilities include scoping projects, designing software, and automating infrastructure management. Ideal candidates will have proficiency in programming languages like Python or Go, experience with Terraform, and familiarity with CI/ ...