1 to 25 of 538 Site Reliability Engineering Jobs in London

Cloud Operating Model - Managing Consultant

Location
Greater London, England, United Kingdom
build and scale secure, reliable and operationally effective AI platforms. You will combine expertise in platform engineering, Site Reliability Engineering (SRE), observability and intelligent operations to help organisations move from isolated AI experimentation to production-grade, enterprise-scale AI services.You will work with technology, engineering … observability, platform automation and operational guardrails. Enable reliable and repeatable delivery of AI services from experimentation through to production.• Reliability Engineering & SRE: Establish SRE practices including SLIs, SLOs, error budgets, capacity planning, resilience engineering and reliability governance. Help clients shift from reactive operations to data-driven ...

Head Of Infrastructure and Cloud - Internal Applicants Only

Location
Greater London, England, United Kingdom
transition from traditional infrastructure management to a platform-centric, product-led operating model, integrating platform engineering, DevOps, Site Reliability Engineering (SRE), and Network Operations (NOC) to enable scalable, automated, and resilient technology services. To place the interests of customers at the centre of all activities … YBIYRI) model with shared accountability for service delivery and operational outcomes. Establish and integrate Site Reliability Engineering (SRE) practices, defining and managing service‐level objectives (SLOs), error budgets, and proactive reliability engineering across critical services. Ensure end‐to‐end service reliability and resilience, including ...

Head Of Infrastructure and Cloud

Hiring Organisation
Arbuthnot Latham
Location
London, UK
Employment Type
Full-time
transition from traditional infrastructure management to a platform-centric, product-led operating model, integrating platform engineering, DevOps, Site Reliability Engineering (SRE), and Network Operations (NOC) to enable scalable, automated, and resilient technology services. To place the interests of customers at the centre of all activities … YBIYRI) model with shared accountability for service delivery and operational outcomes. Establish and integrate Site Reliability Engineering (SRE) practices, defining and managing service-level objectives (SLOs), error budgets, and proactive reliability engineering across critical services. Ensure end-to-end service reliability and resilience, including ...

Systems Engineering Manager, Site Reliability Engineering, ML Compute

Location
City of Westminster, England, United Kingdom
. Track record of mentoring technical leads. Proven success leading and influencing multiple technical teams. About the job Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google's services—both … internally critical and our externally-visible systems—have reliability, uptime appropriate to users' needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating ...

Software Engineer III, Site Reliability Engineering, GCE AI

Location
City of Westminster, England, United Kingdom
Science or Engineering. 2 years of experience designing, analyzing, and troubleshooting large-scale distributed systems. About the job Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that Google Cloud's services—both … internally critical and our externally-visible systems—have reliability, uptime appropriate to customer's needs and a fast rate of improvement. Additionally SRE’s will keep an ever-watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure ...

Engineer - Site Reliability

Location
Greater London, England, United Kingdom
early‐session coverage from London that ensures continuous, high‐availability operations across Cboe's real‐time low‐latency trading platforms. The London‐based SRE provides technical support to Cboe Trade Desk and Operations Support Center staff across time zones, and works closely with Software Engineering, Systems Engineering … spoken — is required. This role demands clear, precise, and unambiguous communication at all times. As the operational bridge between Cboe's European and APAC SRE teams and its US‐based leadership, the ability to communicate with clarity across time zones, cultures, and technical disciplines is fundamental to the success ...

Vice President, Site Reliability Engineering

Hiring Organisation
The Bank of New York Mellon
Location
London, UK
Employment Type
Full-time
seeking a future team member for the role of Vice President - Site Reliability Engineer to join our team. This role is located in London. Role SummaryBNY is seeking a Vice President - Site Reliability Engineer to design, build, deploy, and scale resilient, automated, and centrally managed engineering … driven automation, and modern software delivery practices. Experience supporting distributed systems, cloud-native platforms, or container-based architectures. Knowledge of Agile, DevOps, and SRE operating models, including continuous improvement and blameless post-incident practices. Ability to influence engineering standards and drive adoption of common tooling and automation patterns across ...

Vice President, Site Reliability Engineering

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
hackajob is partnering directly with BNY Mellon to hire for this role. Were seeking a future team member for the role of Vice President - Site Reliability Engineer to join our team. This role is located in London. Role Summary BNY is seeking a Vice President - Site Reliability … driven automation, and modern software delivery practices. Experience supporting distributed systems, cloud-native platforms, or container-based architectures. Knowledge of Agile, DevOps, and SRE operating models, including continuous improvement and blameless post-incident practices. Ability to influence engineering standards and drive adoption of common tooling and automation patterns across ...

Director of Site Reliability Engineering

Hiring Organisation
Hackajob Ltd
Location
london, south east england, united kingdom
your profession to new heights by contributing to revolutionary projects. You've discovered the perfect environment to have a major impact. As a Principal Site Reliability Engineer at JPMorgan Chase within the Corporate Technology and Enterprise Technology Team, you draw upon your advanced knowledge to identify new opportunities … lifecycle of software development for the firm. You will have the opportunity to manage, design, and implement infrastructure components to improve reliability and ensure operational efficiency. Job Responsibilities Identifies and solves problems of high complexity and drives improvements as outcomes Uses enterprise-authorized AI capabilities within the work environment ...

Senior DevSecOps Engineer

Location
Greater London, England, United Kingdom
repeatable, and secure delivery of autonomy and mission software • Embed security throughout the software development lifecycle, integrating security controls, testing, evidence, and assurance into engineering workflows • Apply UK MOD Secure by Design principles and help engineering teams meet cyber security and technical assurance responsibilities • Collaborate with software, autonomy … controls including static analysis, dependency scanning, container scanning, secrets detection, software composition analysis, vulnerability management, and policy enforcement • Champion a DevSecOps culture emphasizing security, reliability, deployability, and operational performance ownership Requirements BS or MS in Computer Science, Software Engineering, Cyber Security, Electrical Engineering, Systems Engineering ...

Technical Lead - Site Reliability Engineering

Location
Greater London, England, United Kingdom
Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Markets and Risk Intelligence division.As a **Technical Lead SRE**, you will be a senior hands‐on technical person help shape the foundations of reliability across both new and existing platforms. You will collaborate … person who is passionate about reliability engineering and who bring a continuous improvement approach to everything they do!Lead the establishment of SRE foundations for new projects building environments, monitoring, alerting, and ensuring operational readiness from day one.Collaborate with Architecture and Engineering teams to embed reliability ...

Director of Site Reliability Engineering

Location
Greater London, England, United Kingdom
influence engineering standards, enhance operational frameworks, and foster a culture of continuous improvement across mission‐critical environments. Responsibilities Lead and scale a global SRE organization, focusing on engineering excellence and team empowerment Collaborate with product, platform, operations, and security teams to embed reliability within SDLC practices Define … deliver systemic improvements across production environments Establish observability strategies with standardized tooling for metrics, logs, and tracing to support distributed systems Adopt and enforce SRE practices, including SLIs, SLOs, SLAs, and error budgets across services Drive resilience strategies with highly available architectures and disaster recovery readiness Champion an automation‐first ...

Principal Cloud SRE / Cloud SME - LSEG Workspace

Hiring Organisation
London Stock Exchange Group
Location
London, UK
Employment Type
Full-time
build technology that helps people make informed decisions in global financial markets. We are looking for a Principal Cloud Site Reliability Engineer (SRE) with strong cloud platform experience to join us in evolving the reliability, scalability, and operational health of our LSEG Workspace platform. Workspace … supporting teams to deliver safely and efficiently. Collaboration & Knowledge SharingWe work as partners across engineering and product teams. You will share cloud and SRE knowledge, support less experienced engineers, and help create clear, reusable patterns that improve platform consistency. We value mentoring, documentation, and open technical discussion as much ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
City of London, London, United Kingdom
Employment Type
Permanent, Part Time, Work From Home
Salary
£80,000
support economic growth across the UK. The Department for Business, Innovation, Science and Trade (BIST), in partnership with Inspire People, is seeking a Senior SRE Squad Lead with experience leading and developing engineers, strong DevOps and Site Reliability Engineering expertise, cloud platform experience, infrastructure-as-code capability … UK. BIST's Digital, Data and Technology (DDaT) directorate develops and operates the tools and services that enable this mission. As a Senior SRE Squad Lead, you will play a key role in leading engineers while remaining hands-on in the design, delivery and continuous improvement of reliable, secure ...

Lead Site Reliability Engineer

Hiring Organisation
Inspire People
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£67,547 - £83,778 per annum
support economic growth across the UK. The Department for Business, Innovation, Science and Trade (BIST), in partnership with Inspire People, is seeking a Senior SRE Squad Lead with experience leading and developing engineers, strong DevOps and Site Reliability Engineering expertise, cloud platform experience, infrastructure-as-code capability … UK. BIST's Digital, Data and Technology (DDaT) directorate develops and operates the tools and services that enable this mission. As a Senior SRE Squad Lead, you will play a key role in leading engineers while remaining hands-on in the design, delivery and continuous improvement of reliable, secure ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
london, south east england, united kingdom
systems run reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability … JPMorgan Chase in the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI. You will own the reliability, performance, and cost-efficiency of the LLM inference platform end to end. ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
period of growth and investing heavily in its engineering and platform capabilities. They're looking for an experienced Site Reliability Engineer (SRE) to join the team and play a key role in building highly reliable, scalable, and observable infrastructure. This is a hands-on role focused … experience Help improve platform resilience, scalability, and disaster recovery capabilities Contribute to capacity planning and performance optimisation as the platform scales Establish and champion SRE best practices across the wider engineering function What We're Looking For Proven commercial experience working as an SRE, DevOps Engineer, Platform Engineer ...

Lead Site Reliability Engineer

Location
City of Westminster, England, United Kingdom
defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Infrastructure Platforms team, you hold a leadership role in your team … ability to expand and collaborate across different levels and stakeholder groups Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity. Ability to evaluate AI-assisted operational recommendations ...

Site Reliability Engineer

Hiring Organisation
Bristow Holland Ltd
Location
London, United Kingdom
Employment Type
Permanent
Salary
£55000 - £60000/annum - Offering 100% Work from home
exciting global technology organisation is looking for a Site Reliability Engineer (SRE) to join its growing engineering team. This is a fully remote position, offering the opportunity to work on large-scale, business-critical platforms used by customers around the world. The role would suit an experienced … Site Reliability, DevOps, Platform or Cloud Engineer with strong hands-on experience across Kubernetes and Microsoft Azure who enjoys solving complex production problems, improving reliability and automating manual processes. You will work closely with Development and DevOps teams, helping to design, build, operate and scale highly available ...

Lead Site Reliability Engineer - Chief Technology Office

Hiring Organisation
Hackajob Ltd
Location
london, south east england, united kingdom
defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Chief Technology Office, youwill solve complex and broad business problems, demonstrate … technical lead for medium to large-sized products, and provide advice and mentoring to other engineers. Job responsibilities Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team Leads initiatives to improve the reliability and stability of your team's applications ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
Senior Site Reliability EngineerApplylocations: London (82)time type: Full timeposted on: Posted Todayjob requisition id: JR101516**Senior Site Reliability Engineer (SRE) - GCP/Kubernetes****About the Role**We are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small … agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance of our core platform with a high degree of autonomy and ownership.The successful candidate will split their time between providing expert operational support for our critical systems and leading exciting new infrastructure ...

Site Reliability Engineer

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Site Reliability Engineer – FintechQuant Capital is urgently looking for a Site Reliability Engineer to join or well known Fintech50 client who produces software disrupting the wealth management market. My client is a market leading SAAS provider to financial advisory business nationwide. They are currently in growth … stable solutions. Familiarity with development languages, such as .NET, Java or PythonRedisDocker/KubernetesDatabase experiencesThis role suits a senior Engineer from a DevOps or SRE background who is real technologist interested in the latets tooling and technologies that support software development and infrasturtcure. The firm has a corporate feel ...

principal engineer- international technology & Starbucks digital solutions

Location
Greater London, England, United Kingdom
technical excellence across EMEA while aligning to global technology strategy and leading the Starbucks Digital Solutions technical direction for International markets. It will set engineering direction, raise standards and guide decisions across internally developed and third-party platforms that matter most to our business, customers, partners, baristas and shareholders.As … Management organisations in a product-led operating model.• Knowledge of modern engineering practices including Platform Engineering, Site Reliability Engineering (SRE), AI-assisted development and Developer Experience (DevEx).What else should you know?• We have a flexible working policy. Meaning 50% of the time we collaborate ...

Site Reliability Engineer, Studios

Location
Uxbridge, England, United Kingdom
rotations, to support live operations and critical systems. Occasional travel may be required depending on project and client needs. IMG is looking for a Site Reliability Engineer to help design, build, operate, and continuously improve resilient, secure, and highly available platforms that underpin our digital, cloud, and broadcast … adjacent services. This role is suited to someone who combines strong infrastructure and software engineering capability with an operational mindset, and who can help embed reliability engineering practices across systems that support live, business‐critical environments. The successful candidate will play a key role in improving service ...

Site Reliability Engineer - Fintech

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Site Reliability Engineer – FintechQuant Capital is urgently looking for a Site Reliability Engineer to join or well known Fintech50 client who produces software disrupting the wealth management market. My client is a market leading SAAS provider to financial advisory business nationwide. They are currently in growth … with development languages, such as .NET, Java or Python·Redis·Docker/Kubernetes·Database experiencesThis role suits a senior Engineer from a DevOps or SRE background who is real technologist interested in the latets tooling and technologies that support software development and infrasturtcure. The firm has a corporate feel ...