19 of 19 Permanent Site Reliability Engineering Jobs in the City of Westminster

Systems Engineering Manager, Site Reliability Engineering, ML Compute

Location
City of Westminster, England, United Kingdom
. Track record of mentoring technical leads. Proven success leading and influencing multiple technical teams. About the job Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google's services—both … internally critical and our externally-visible systems—have reliability, uptime appropriate to users' needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating ...

Software Engineer III, Site Reliability Engineering, GCE AI

Location
City of Westminster, England, United Kingdom
Science or Engineering. 2 years of experience designing, analyzing, and troubleshooting large-scale distributed systems. About the job Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both … internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure ...

Production Engineering Manager

Location
City of Westminster, England, United Kingdom
Meta is seeking a Production Engineering Manager to lead a team responsible for the reliability, scalability, and operational excellence of Meta's production infrastructure and services. In this role, you will manage a team of production engineers who own the full lifecycle of systems — from capacity planning … performance optimization to incident response and automation. You will drive technical strategy, champion AI-augmented workflows, and partner closely with software engineering, infrastructure, and product teams to ensure Meta's services operate at global scale with high availability and efficiency.Production Engineering Manager Responsibilities:Manage a team of production ...

DevSecOps Engineer

Location
City of Westminster, England, United Kingdom
operating the software delivery infrastructure required to develop and deploy advanced autonomous systems for defence applications. This role sits at the intersection of software engineering, platform engineering, cyber security, and defence systems engineering. The DevSecOps Engineer works alongside autonomy, software, systems, integration, and test engineers to create secure … delivery pipelines that enable teams to rapidly develop, integrate, test, and deploy mission critical software. The ideal candidate has a strong software and platform engineering background combined with significant experience operating within UK defence environments. They have a strong understanding of the UK Ministry of Defence/NATO approach ...

Software Engineer Lead - Site Reliability

Location
City of Westminster, England, United Kingdom
strategic partners and third parties., As a Lead DevOps Engineer, you will be a hands-on contributor and technical lead focused on improving the engineering foundations of the Customer Digital Platform. You will work across internal squads and with strategic partners, third parties, service, architecture and security teams … GitHub/GitHub Actions, Azure DevOps, Terraform and automated testing. Improve deployment safety, release readiness and operational readiness for customer-facing digital services. Apply SRE principles pragmatically to improve availability, recoverability, monitoring and incident learning. Strengthen monitoring, logging, tracing, alerting and service-health dashboards across digitally connected workloads. Reduce single ...

Lead Site Reliability Engineer

Location
Westminster, West End, United Kingdom
trading technology stack is undergoing a multi year convergence and modernization journey. You will play a pivotal role in shaping our next generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self healing workflows, and resilience engineering. Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage ...

Lead SRE - Chase UK

Location
Westminster, West End, United Kingdom
building the bank of the future from the ground up, offering you the chance to join us and make a significant impact. As a Site Reliability Engineer at JPMorgan Chase within the International Consumer Bank, you will play a crucial role in this initiative, dedicated to delivering … oriented and possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with ...

Full-Stack Engineer - ML Platform, Applied AI - Vice President

Location
City of Westminster, England, United Kingdom
Investment Banking, you will design and deliver production architectures for AI-powered products and services. You will work at the intersection of software engineering and applied research to translate innovative ideas into scalable, enterprise-grade solutions. You will collaborate closely with cloud and site reliability engineering … distributed, multi-threaded, and scalable applications Build, test, and deploy highly secure automated pipelines for cloud systems, desktop applications, and ML solutions Apply software engineering and computer science best practices Develop and deploy business-critical, data-intensive applications Leverage foundational libraries and services for reuse across teams Utilize MLOps ...

Product Associate - SRE Team - Chase UK

Location
Westminster, West End, United Kingdom
oriented and possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with … engineering teams to ensure services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams ...

ML Compute SRE Lead: Scale, Uptime & Automation

Location
City of Westminster, England, United Kingdom
Google London is seeking a Systems Engineering Manager for Site Reliability Engineering in ML Compute. You will lead a multi-disciplinary team, own uptime, and shape reliability strategy for large-scale services. You will mentor engineers, drive end-to-end availability, and collaborate with cross ...

Senior SRE Engineer — Cloud Reliability & Automation

Location
City of Westminster, England, United Kingdom
Google London, UK is seeking a Software Engineer III in Site Reliability Engineering for the GCE AI team. This mid-level role focuses on building reliable, scalable systems, code development, and mentoring junior team members. The position emphasizes deep expertise in distributed systems, problem solving, and collaboration ...

Head of Production Management- J.P. Morgan Personal Investing

Location
Westminster, West End, United Kingdom
powered solutions and intelligent automation to reduce manual intervention, fast-track resolution, and continuously improve operational efficiency. Champion an automation-first, shift-left SRE cultureleveraging shared tooling and automation to ensure consistency, reduce duplication, and maintain alignment with firmwide standards. Oversee capacity management and planning, ensuring infrastructure scales to meet … management standards change, incident, capacity, and automation across multiple engineering teams operating in a you-build-it-you-run-it model, underpinned by SRE principles and disaster recovery planning. Composure, decisiveness, and authority during incidents, vendor failure, or regulatory escalation, with a proven ability to protect business lines under ...

Senior Applied AI Engineer (Manager) TC

Location
City of Westminster, England, United Kingdom
will work across a diverse portfolio of clients spanning Financial Services, the Public Sector, and the Private Sector. Our Applied AI Engineering teams deliver production-grade AI systems in regulated financial institutions as well as government, health, infrastructure, consumer, industrial and energy organisations. This cross-sector model gives … serverless, IAM and network security. Data engineering depth (Spark/Databricks; ETL/ELT); cloud-native data + AI architectures. Enterprise integration and SRE principles (SLIs/SLOs, runbooks, rollback). Consulting leadership: stakeholder, budget and risk management; team leadership., Graph/big-data stacks; streaming; cloud architect certifications ...

SRE & Reliability Lead — AI-Ops, Observability, & On-Call

Location
City of Westminster, England, United Kingdom
ICIS, part of RELX Group, seeks an SRE Manager to lead the reliability function for production services. You will drive automation, AI-Ops capabilities, and a team focused on observability, incident response, and continuous improvement across internal and external customers. You will shape strategy, manage and develop leaders ...

Senior DevOps Lead, Site Reliability for Cloud Platform

Location
City of Westminster, England, United Kingdom
Standard Life in City of Westminster, UK, is seeking a Lead DevOps Engineer to strengthen the Customer Digital Platform by delivering immediate engineering improvements, focusing on Azure API platform integration, cloud infrastructure …/CD automation, observability and deployment safety. You will lead hands-on coding and technical direction across squads, partners and security teams, applying SRE and DORA-like metrics, improving monitoring, incident learning and recovery capabilities, and #J-18808-Ljbffr ...

Business Support Engineer

Location
City of Westminster, England, United Kingdom
looking for an engineer to play a key role in providing technical, engineering support to Meta’s Advertising partners, Customers and clients globally. You will have the opportunity to work together with a global team of Business Support Engineers who are expert in Meta’s Adtech to provide proactive … broad range of partners across the globe to integrate Meta’s Business Products into their offering.Business Support Engineer Responsibilities:Provide proactive and reactive engineering support for partners, independently managing complex outages to ensure high partner satisfactionTroubleshoot large-scale distributed systems and partner integrations, championing operational excellence and engineering ...

BXTI, Technical Product Manager, Data, Cloud and Developer Experience - AVP

Location
City of Westminster, England, United Kingdom
features that give leadership and security teams visibility into and control over AI platform usage, including cost attribution, chargeback, and usage governance. Partner with engineering to deliver scalable, secure architecture that meets Blackstone's enterprise and regulatory requirements. Define APIs … integration points that enable agent workflows, tool access, and system connectivity across Copilot, AWS Bedrock, and embedded applications. Work closely with Platform Engineering, SRE, Security, and Data teams to understand cross-platform pain points and translate them into prioritized product requirements. Build relationships with product managers across business-unit ...

Operations and SRE Manager

Location
City of Westminster, England, United Kingdom
SRE Manager, you will lead the operational reliability function for ICIS, ensuring stable, resilient, and well-supported production services for both internal and external customers. You will be responsible for driving reliability improvements, advancing automation and AI-Ops capabilities, and leading a team focused on observability, incident response … RCAs, post-mortems and improvement actions are owned, tracked and completed. Strengthen operational process adherence, ensuring responsibilities are clear and delegation is effective. Drive SRE practices across observability, automation, disaster recovery, design for reliability, on-call readiness and production support. Protect service levels by ensuring engineering effort ...

Deal Architect

Location
Westminster, West End, United Kingdom
with the customers business outcomes. Key Responsibilities 1. Deal & Solution Architecture Leadership Lead the architecture of complex, multi-tower solutions covering cloud, data, digital engineering, service management, platforms, and application modernization. Translate high-level business requirements into integrated technical, service, and commercial solution frameworks. Govern solution integrity across technology … value engineering mindset Technical Multi-cloud architectures (Azure/AWS) Data, integration, and automation frameworks Modern engineering practices (CI/CD, DevOps, SRE) Security & compliance frameworks Behavioural Executive presence and gravitas Structured problem-solving Collaborative leadership High resilience under bid pressure Ability to navigate ambiguity and derive clarity ...