101 to 125 of 262 Site Reliability Engineering Jobs in London

Senior Site Reliability Engineer – Cloud, MLOps & HPC

Hiring Organisation
Jobleads-UK
Location
City of Westminster, England, United Kingdom
Roche Holding AG in the United Kingdom seeks a Senior Site Reliability Engineer to design resilient, cloud-based platforms for MLOps and HPC workloads at global scale. You will work with Data & Digital Catalyst teams to craft robust systems and accelerate scientific discovery. You will architect IaC using ...

Director, Head of Technology Resilience and Production Operations

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
vision to become ‘the worlds most trusted Bank’ and ensuring continued regulatory confidence* The role holder will work with the Head of Digital Engineering Services and Solutions Department Head to deliver transformation through a reliable, robust, sustainable, scalable and efficient operating model by leveraging best practices* The role holder … Resilience related projects and programmes measuring the effectiveness of these services delivered* This is a Leadership position and an integral part of the Digital Engineering Solutions and Services Leadership team maintain compliance and regulatory obligations.**NUMBER OF DIRECT REPORTS**TBC - Team Size circa 25**KEY RESPONSIBILITIES****Planning & Strategy ...

Software Engineer III, Full Stack, Publisher Inventory

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
control and optimize ad placement, policy compliance, and publisher revenue. Leverage Google's state-of-the-art AI tooling and LLM APIs to improve engineering velocity and build AI-augmented product features. Collaborate with cross-functional partners including PMs, UX, and cross-sites with other engineering partners … deliver seamless end-to-end features. Write robust, well-tested, and high-performance code, ensuring high quality and reliability across our platform.### Skills Required* Kotlin* C++* TypeScript* Dart* Angular* Java* APIs* SDKs* LLM APIs* Generative AI### Tags:UKLondonAngular DeveloperFlutter DeveloperShare Job:Application planning## What to evaluate before applying### Visa ...

SRE Engineering Manager: Traffic Steering & DNS

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
United States Digital Space LLC is seeking a Site Reliability Engineer to lead a seasoned team ensuring the availability and performance of mission-critical services. You will automate responses, oversee incident management, and drive reliability across globally distributed systems. Join a culture of curiosity and collaboration, mentoring ...

Site Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
iGaming company based in London. Estimated Benefits Health Insurance Pension Stock Options Benefits estimated based on industry standards We’re hiring a Site Reliability Engineer to join our London team This is a fantastic opportunity for someone passionate about reliability, scalability and automation. You’ll be pivotal ...

Cloud Platform Engineer (Senior / Lead)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
identity, network, workloads and data for both human and machine/agent identities; secure by default, least privilege, secrets management and continuous compliance. Apply SRE practices—SLOs/SLIs, observability, capacity planning, resilience and blameless incident management—to keep the platform reliable and cost‐efficient. Partner with data engineering … ability (e.g. Python, Go) and strong observability, reliability and cost‐optimisation practices. Desirable requirements: Experience working as a Site Reliability Engineer (SRE) with SLOs/SLIs, error budgets and incident management. A third top‐tier cloud certification, or specialist security/Kubernetes certifications (e.g. CKA/ ...

Principal Platform Engineer

Hiring Organisation
Sanderson Recruitment
Location
City of London, London, United Kingdom
Employment Type
Permanent
large-scale distributed systems and database platforms? We're looking for a hands-on technical leader to help shape the future of our platform engineering capability. This is an opportunity to lead complex engineering initiatives, define technical strategy, and act as a subject matter expert across AWS infrastructure … technical authority for distributed database and persistence technologies Required Experience 8+ years' experience in Platform Engineering, Infrastructure Engineering, DevOps, SRE or Software Engineering Expert-level AWS infrastructure experience Strong Infrastructure as Code expertise with Terraform Strong Linux systems administration and networking knowledge Experience designing and operating distributed ...

Senior Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
diversified range of trading strategies. We employ over 130 colleagues in Jersey, Geneva, London, Singapore, New York and Shanghai. This is a hands‐on engineering role within the Platform Engineering team, part of Technology Operations. Platform Engineering is responsible for building and operating the infrastructure, platforms … comply with all organisational, statutory and regulatory policies and procedures. Experience, Knowledge & Skills Five or more years of experience in platform engineering, DevOps, SRE, infrastructure engineering or a closely related role. Strong experience operating production or production‐like infrastructure, ideally across hybrid cloud and on‐premises environments. Hands ...

Service Management Lead (Cloud Services)

Hiring Organisation
17918
Location
London, United Kingdom
opportunity to take ownership of critical cloud platform services, driving service excellence across modern cloud technologies while working closely with engineering, SRE and senior business stakeholders. Excellent Day Rate Available. Key Skills Technology Service Management or Cloud Operations leadership Kubernetes (GKE and/or AKS) API platforms including Apigee … Azure APIM ITIL Service Management Major Incident & Service Governance Platform Engineering & SRE Stakeholder & Executive Management Service Improvement, SLAs, SLOs & KPIs If you're experienced in leading cloud services within complex enterprise environments and enjoy driving operational excellence, we'd like to hear from you. For more information ...

Senior AWS SRE Lead — Cloud Reliability & Automation

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
company in the UK is seeking a Lead Site Reliability Engineer to strengthen reliability, availability and performance of global digital platforms. The role emphasizes AWS expertise, automation, monitoring, and incident management within a hybrid working model. You will mentor engineers, collaborate with cross‐functional teams, and drive … best practices in cloud engineering, security, and reliability across multiple regions. #J-18808-Ljbffr ...

AI-Driven Senior SRE Leader for Large-Scale Systems

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
United States Digital Space LLC is looking for a seasoned engineer to join their Site Reliability Engineering team in London, UK. This pivotal role focuses on ensuring reliable and robust Google Ads services while utilizing AI-driven automation. Candidates should have a Bachelor's degree in Computer … Science and substantial software development experience. The position requires expertise in problem-solving, innovative engineering solutions, and mentoring skills to guide a team of engineers within a dynamic environment. #J-18808-Ljbffr ...

Developer Enablement, Technical Architect – Release on Demand (SVP)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
effort to make RoD elastically scalable and enterprise-grade reliable. Design for high availability, graceful degradation, and zero-downtime deployments. Define SLOs, own the SRE practice for the platform, and be accountable when things need fixing.* **Define the Observability Strategy.** Establish a comprehensive observability framework — distributed tracing, structured logging, metrics … development language. Designing, building and consuming domain specific RESTful services & APIs* Proficiency with relational and/or NoSQL databases: PostgreSQL, MongoDB or Couchbase* Demonstrated SRE or platform engineering experience — SLOs, incident management, reliability engineering at scale* Experience defining and implementing observability strategies: distributed tracing, structured logging, metrics ...

Senior Site Reliability Engineer

Hiring Organisation
Staffworx Limited
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
team is expanding and hiring engineers now The role: Build, operate and maintain high-performance, scalable, reliable services across UK Government deployments Own production reliability: monitoring, alerting, configuration management and upgrades Lead automation to reduce manual operations, using modern platforms including LLM/AI tooling Deploy new products into … need: UK SC clearance, or eligibility to obtain it (active SC strongly preferred; no visa sponsorship) 1-5 years in infrastructure engineering or SRE, building and deploying production systems - not just debugging Hands-on Kubernetes/Docker in production; Terraform, Ansible or CI/CD pipeline experience Proficiency ...

Site Reliability Engineer- Spacetime UK

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Role Overview This isn't a "keep the lights on" SRE role. This is a strategic, high-impact opportunity to build the nervous system for a platform that transforms how networks of satellites, ground stations, and fleets are interconnected and orchestrated. You will be building the core observability stack that … cloud-native tools to a robust, scalable, and insightful platform built on best-in-class technologies (Prometheus, OpenTelemetry, etc.). If you are an SRE who thrives on platform-building challenges and wants to be relied upon to build a production-grade observability stack from the ground up, this role ...

Site Reliability Engineer

Hiring Organisation
iXceed Solutions
Location
City of London, London, United Kingdom
backups, secrets, and platform operations. ✅ Monitor platform health using dashboards, logging, and alerting tools. ━━━━━━━━━━━━━━━━━━━━━━ 🔹 Required Skills ✔ Experience in DevOps, Infrastructure, Platform Engineering, or SRE ✔ Strong Python and Shell scripting ✔ Hands-on experience with Azure, AWS, or GCP ✔ Kubernetes, Docker, and Linux administration ✔ CI/CD pipelines (GitHub Actions, GitLab ...

Site Reliability Engineer

Hiring Organisation
NEEV LIMITED
Location
London, United Kingdom
Employment Type
Contract
NOTE: VISA SPONSORSHIP IS NOT PROVIDED Role: SRE Location: London, UK(5 Days/Week Onsite) Type: Contract InsideIR35/Permanent Exp: Minimum 8+ Years Skills: SRE experience with Python-based applications (not Java) Exposure to cloud technologies Familiarity with Athena ecosystem or similar (SecDB, Quartz) banking and risk domain … exposure SRE Role description We need an experienced SRE to focus predominantly on automation, optimization, and process re-engineering using AI for the Market Risk Platform. Success is measured by capacity created 9toil eliminated, fewer manual steps, faster recovery, safer/faster changes) not by being the primary ...

Kubernetes SRE: DevOps, Observability & Reliability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Cisco's Webex Engineering Group in London is seeking a Site Reliability Engineer to own the design, deployment, and operation of Kubernetes-based microservices. You will build reusable configurations and advance a robust, scalable platform. In this role, you will implement GitOps workflows with Argo CD, manage … canary releases and secret management with Vault, and monitor reliability with Prometheus and Grafana, collaborating across teams in a hybrid London office. #J-18808-Ljbffr ...

Grafana Observability Engineer (6-month contract)

Hiring Organisation
17918
Location
London, United Kingdom
version control. Develop KQL queries, dashboards and visualisations. Define monitoring standards, SLIs and SLOs where appropriate. Produce documentation and share knowledge across the engineering team. Essential Skills & Experience Significant experience in Observability, Platform Engineering or Site Reliability Engineering. Strong hands-on experience with Grafana administration, dashboards ...

Performance & Observability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Analyse application, database, and infrastructure performance to identify bottlenecks and inefficiencies. Develop performance benchmarks and SLIs to measure service responsiveness and stability. Collaborate with SRE and DevOps teams to optimise CI/CD pipelines for performance improvements. Collaborate with wider IT teams to Implement caching strategies, query optimisation, and autoscaling … translate observability insights into business impact for stakeholders. A continuous improvement mindset, focused on reducing toil and improving efficiency. Experience working in a DevOps, SRE, or Platform Engineering environment. Team Information Technology Working Pattern Full time Location London Contract Permanent Diversity & Inclusion We are committed to attracting people from ...

Senior DevOps Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
DevOps Platform Engineer London We are currently supporting a leading organisation looking to hire an experienced Senior DevOps Platform Engineer to join their cloud engineering function. This is an excellent opportunity for a highly skilled DevOps professional to work on complex cloud-native environments, designing and automating scalable platforms … Implement GitOps delivery models using tools such as ArgoCD and Helm Develop reusable automation frameworks for infrastructure provisioning, configuration management and deployments Improve platform reliability, scalability and security through automation and engineering best practices Support cloud migrations from traditional data centre environments into modern cloud-native platforms Implement ...

Front-Office SRE Lead: Observability & AI-Driven Reliability

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
JPMorgan Chase & Co. in London is seeking a Lead Site Reliability Engineer to shape next‐gen SRE patterns, observability, and reliability across globally distributed trading systems. You will partner with front‐office traders, contribute production code (Java/Python/Kotlin), drive incident response, and lead … assisted reliability initiatives while collaborating with infrastructure, cloud, and security teams. #J-18808-Ljbffr ...

Senior Azure SRE & Cloud Reliability Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Willis Re in London seeks an experienced Site Reliability Engineer to design and operate enterprise Azure platforms. You will automate infrastructure, define reliability targets, and drive resilience across multi-region deployments. The role requires deep Azure knowledge, Terraform/IaC discipline, and strong DevOps practices, including incident ...

Senior Azure SRE: Cloud Reliability & Automation

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Willis Re (UK) Limited is expanding its Cloud & Infrastructure team to design, deploy and support enterprise-scale Azure platforms. The role focuses on reliability, automation, and governance … reduce operational overhead. The successful candidate will lead reliability engineering, implement observability, and drive resilient architecture with Terraform, Azure DevOps, and strong SRE practices. UK-based, with flexible same-team collaboration across regions. #J-18808-Ljbffr ...

Architect & Delivery Lead (68018)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Mandatory Skills Broad and deep technical knowledge across infrastructure, cloud, networking, security, IAM, SRE, and operations. Key Responsibilities Define and own the end‐to‐end technical architecture and target operating model across all six service towers Lead the three‐phase transformation programme (Stabilise → Optimise → Transform), managing timelines, milestones, risks … dependencies Make architectural decisions across IAM, Cloud, SRE, Network, Data, and Security — ensuring coherence, reusability, and alignment with business objectives Establish and chair the Architecture & Engineering Governance board, providing technical assurance across all workstreams Own the programme roadmap, resource plan, and financial model — tracking cost savings, team reduction trajectory ...

SRE Platform Engineer Lead — High-Availability Cloud

Hiring Organisation
Jobleads-UK
Location
City Of London, England, United Kingdom
Goldman Sachs Group, Inc. seeks a Lead Site Reliability Platform Engineer (SRE) in London. This role emphasizes system reliability, performance, and collaboration across teams. Your mission is to design secure, scalable systems on AWS, streamline software delivery, and mentor technical staff. The ideal candidate has over … years of experience in SRE or DevOps and is skilled in AWS and infrastructure as code. You'll lead initiatives to enhance operational efficiency and contribute significantly to security and system integration. #J-18808-Ljbffr ...