26 to 50 of 572 Site Reliability Engineering Jobs in London

Senior Machine Learning Engineer (MLOps)

Hiring Organisation
ASOS
Location
London, UK
Employment Type
Full-time
monitor machine learning solutions in a reliable, scalable and cost-effective way. We're looking for a Senior Machine Learning Engineer - with strong Engineering experience - who enjoys solving complex engineering challenges at scale. This role is ideal for someone with a strong software engineering/MLOps, distributed … recognise that expertise can be developed through a variety of backgrounds, including Backend Engineering, Platform Engineering, Site Reliability Engineering (SRE), Cloud Engineering, MLOps or Machine Learning Engineering. We'd love to see experience in several of the following: Strong software engineering fundamentals with ...

Site Reliability Engineer, Studios

Location
Uxbridge, England, United Kingdom
rotations, to support live operations and critical systems. Occasional travel may be required depending on project and client needs. IMG is looking for a Site Reliability Engineer to help design, build, operate, and continuously improve resilient, secure, and highly available platforms that underpin our digital, cloud, and broadcast … adjacent services. This role is suited to someone who combines strong infrastructure and software engineering capability with an operational mindset, and who can help embed reliability engineering practices across systems that support live, business‐critical environments. The successful candidate will play a key role in improving service ...

Systems Engineering Manager

Hiring Organisation
Hackajob Ltd
Location
Westminster, Greater London, UK
right qualifications and skills for this job Find out below, and hit apply to be considered. Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that our client's services—both our internally … critical and our externally-visible systems—have reliability, uptime appropriate to users' needs and a fast rate of improvement. Additionally SRE's will keep an ever-watchful eye on our systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work ...

Systems Engineering Manager

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
client is a global technology company. Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems. SRE ensures that our client's servicesboth our internally critical and our externally-visible systemshave reliability, uptime appropriate … systems capacity and performance. Much of our software development focuses on optimizing existing systems, building infrastructure and eliminating work through automation. On the SRE team, youll have the opportunity to manage the complex challenges of scale which are unique to our client, while using your expertise in coding, algorithms, complexity ...

Site Reliability Engineer - Fintech

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Site Reliability Engineer – FintechQuant Capital is urgently looking for a Site Reliability Engineer to join or well known Fintech50 client who produces software disrupting the wealth management market. My client is a market leading SAAS provider to financial advisory business nationwide. They are currently in growth … with development languages, such as .NET, Java or Python·Redis·Docker/Kubernetes·Database experiencesThis role suits a senior Engineer from a DevOps or SRE background who is real technologist interested in the latets tooling and technologies that support software development and infrasturtcure. The firm has a corporate feel ...

Site Reliability Engineer - Fintech / Linux

Hiring Organisation
Quant Capital
Location
London, UK
Employment Type
Full-time
Site Reliability Engineer – Fintech/Linux Site Reliability Engineer – Fintech/Linux85,000 Plus BonusQuant Capital is urgently looking for a Site Reliability Engineer to join our high profile client. Our client is a major global financial exchange, driven by technology. They … trading environment for their clients. They have grown massively and recently were voted in the top 50 fintech firms globally. Day to Day the Site Reliability Engineer will: Analyzing and optimizing trading platform Monitoring development activities, change management tickets Monitoring U.S. production, disaster recovery, and certification systems ...

Site Reliability Engineer, Studios

Location
Greater London, England, United Kingdom
rotations, to support live operations and critical systems. Occasional travel may be required depending on project and client needs. IMG is looking for a Site Reliability Engineer to help design, build, operate, and continuously improve resilient, secure, and highly available platforms that underpin our digital, cloud, and broadcast … adjacent services. This role is suited to someone who combines strong infrastructure and software engineering capability with an operational mindset, and who can help embed reliability engineering practices across systems that support live, business-critical environments. The successful candidate will play a key role in improving service ...

Production Engineering Manager

Location
City of Westminster, England, United Kingdom
Meta is seeking a Production Engineering Manager to lead a team responsible for the reliability, scalability, and operational excellence of Meta's production infrastructure and services. In this role, you will manage a team of production engineers who own the full lifecycle of systems — from capacity planning … performance optimization to incident response and automation. You will drive technical strategy, champion AI-augmented workflows, and partner closely with software engineering, infrastructure, and product teams to ensure Meta's services operate at global scale with high availability and efficiency.Production Engineering Manager Responsibilities:Manage a team of production ...

Jobshare - Sr Lead Software Engineer - Site Reliability Engineer, Python & Infrastructure management - Part time/Jobshare

Location
Greater London, England, United Kingdom
domains, and advise others on the technical and business issues facing them. You will will set the vision, strategy, and operating model for our SRE transformation - enabling our business-aligned support teams to deliver higher reliability, stronger resilience, and a measurably better end-user experience across the board. … responsibilities Defines the SRE vision, north-star outcomes, and multi-year roadmap for the Production Management team, aligned to both CIB and JPM Global Technology priorities. Establishes the SRE operating model across global regions (ways of working, intake, prioritization, engagement with engineering teams and production support). Partners with ...

Software Engineer, GPU Infrastructure- ChatGPT Engineering

Location
Greater London, England, United Kingdom
About the Team ChatGPT Engineering builds and operates the compute platform powering one of the world's largest AI products. Every ChatGPT conversation relies on massive GPU clusters serving inference workloads with high reliability, efficiency, and performance. As our GPU fleet continues to grow, we're investing … production infrastructure, preferably GPU clusters or other compute-intensive distributed systems. Have a background in Production Engineering, Site Reliability Engineering (SRE), Infrastructure Engineering, or Platform Engineering. Have built software that automates operational workflows rather than relying on manual processes. Have experience with Kubernetes, Linux systems ...

Lead Product Manager AIOPs

Hiring Organisation
S&P Global
Location
London, UK
Employment Type
Full-time
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across … Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI.Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning priorities ...

Lead Product Manager AIOPs

Location
Greater London, England, United Kingdom
responsible for S&P Global's enterprise AIOps platform and strategy, driving the modernization of IT Operations and Site Reliability Engineering (SRE) through intelligent observability, event intelligence, automation, and AI-driven insights. DTS Platform & Tools – Service Enablement: We serve as thought leaders in AIOps, partnering across … Operations, SRE, engineering, infrastructure, service management, and application teams to solve enterprise operational challenges. Our mission is to improve reliability, reduce operational complexity, optimize technology investments, and enable more proactive and resilient technology operations by applying AI. Responsibilities and Impact: Own and execute the AIOps product roadmap, aligning ...

Site Reliability Engineering (SRE) / Observability Technical Lead

Hiring Organisation
NTT DATA
Location
London, UK
Employment Type
Full-time
team you'll be working with: We are seeking an experienced Site Reliability Engineer (SRE)/Observability Technical Lead to join our team and drive the strategy and execution of observability and reliability projects across our clients. The ideal candidate will have deep expertise in Application Performance … will guide the design, implementation, and continuous improvement of observability solutions, ensuring system reliability, performance, and scalability while fostering best practices in SRE and DevOps. What you'll be doing: Lead the strategic development and management of observability and reliability frameworks across the organization, ensuring alignment with business ...

Senior DevSecOps Engineer

Location
Greater London, England, United Kingdom
operating the software delivery infrastructure required to develop and deploy advanced autonomous systems for defence applications. This role sits at the intersection of software engineering, platform engineering, cyber security, and defence systems engineering. The DevSecOps Engineer works alongside autonomy, software, systems, integration, and test engineers to create secure … delivery pipelines that enable teams to rapidly develop, integrate, test, and deploy mission critical software. The ideal candidate has a strong software and platform engineering background combined with significant experience operating within UK defence environments. They have a strong understanding of the UK Ministry of Defence/NATO approach ...

Senior DevSecOps Engineer

Location
City Of London, England, United Kingdom
operating the software delivery infrastructure required to develop and deploy advanced autonomous systems for defence applications. This role sits at the intersection of software engineering, platform engineering, cyber security, and defence systems engineering. The DevSecOps Engineer works alongside autonomy, software, systems, integration, and test engineers to create secure … delivery pipelines that enable teams to rapidly develop, integrate, test, and deploy mission critical software. The ideal candidate has a strong software and platform engineering background combined with significant experience operating within UK defence environments. They have a strong understanding of the UK Ministry of Defence/NATO approach ...

Foundation Engineering - SRE Platforms - Site Reliability Engineer - Associate - London

Hiring Organisation
Goldman Sachs
Location
London, UK
Employment Type
Full-time
Role OverviewGoldman Sachs has embarked on one of its most ambitious engineering programs: the Consolidated Trade Ledger (CTL), a ground-up reimagining of the front-to-back architecture that underpins every trade the firm executes. CTL is a flagship initiative jointly sponsored by Global Markets and Engineering leadership … flags or chaos/game days. Experience developing AI-assisted operations capabilities, such as alert enrichment, anomaly detection, triage support or runbook automation. Key SRE CompetenciesReliability engineering: SLOs, SLIs, error budgets, incident reduction and service health. Production excellence: monitoring, alerting, observability, runbooks, handoffs and operational readiness. Engineering mindset ...

Senior Network Engineer- IP

Location
Greater London, England, United Kingdom
will take the lead on complex, high-impact fault resolution spanning multiple platforms and services, acting as a senior technical escalation point. Applying SRE principles and deep technical knowledge, you will drive improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing … automation and IaC tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building ...

Junior SRE – Endpoint Focus

Location
Greater London, England, United Kingdom
technical components. Automation & Continuous Improvement: Proactively isolate recurring operational issues, eliminating manual workflow friction through shell scripting, automated provisioning design, and strategic process enhancements. SRE Transition: Partner closely with Senior SRE team members to progressively absorb production infrastructure tasks, system monitoring duties, and core site reliability principles. Qualifications … performance. Professional Trajectory: A clear, defined motivation to evolve technically and professionally into a Production Engineering or Site Reliability Engineering (SRE) role. Commute Compliance: Willingness and ability to work regularly from our client’s modern office facilities located in the Moorgate area of London (minimum ...

SRE Engineer

Hiring Organisation
TEKsystems
Location
London, UK
Site Reliability Engineer to join a client SRE team. As a Site Reliability Engineer, you will drive adoption of SRE best practice across our cloud estate. By utilising both your soft skills and technical experience, you will work with teams to ensure our standards and governance … Site Reliability Engineer, you will play a pivotal role in ensuring the reliability and performance of our applications and infrastructure. The SRE team will put you in the position to work with application teams across the department on developing reliable and secure solutions to provide to citizens ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
processes and platform lifecycle management Experience establishing repeatable, standardised engineering workflows Experience designing for resilience, fault tolerance, observability, and operational excellence Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives, monitoring, alerting … resume keywords AWS Cloud Platform Architecture Infrastructure as Code CI/CD Automation Platform-as-a-Product Principles Site Reliability Engineering (SRE) Practices ATS Optimization Keywords Hard Skills Cloud Infrastructure Distributed Systems Networking Security Performance Analysis Capacity Planning Scalability Engineering Operational Excellence Monitoring and Alerting Fault ...

Security Operations Manager

Location
Greater London, England, United Kingdom
outsourced support, and set up how alerts are handled and escalated. You’ll also lead how we respond to incidents, working closely with our engineering and reliability colleagues, since a lot of this sits alongside how they already keep … services running. This is a deeply collaborative role. In particular you’ll work hand in hand with our site reliability engineering (SRE) team, since much of security monitoring and response builds on the same tooling and ways of working they already use to keep our services running. ...

Staff Software Engineer, AI Reliability Engineering

Location
Greater London, England, United Kingdom
Staff Software Engineer, AI Reliability Engineering London, UK About Anthropic Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed … from people who've built product stacks, scaled databases, run massive distributed systems, and everything in between. Strong candidates may also Have been an SRE, Production Engineer, or in similar reliability-focused roles on large scale systems Have experience operating large-scale model serving or training infrastructure (>1000 GPUs ...

Senior Site Reliability Engineer

Hiring Organisation
Brevan Howard
Location
London, UK
Employment Type
Full-time
Senior Site Reliability Engineer (SRE) - GCP/KubernetesAbout the RoleWe are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance … Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post ...

Oracle OSS Stack Lead

Hiring Organisation
Vodafone
Location
London, UK
Employment Type
Full-time
serve as the key customer contact for critical OSS-related incidents and strategic programmes while driving DevOps and Site Reliability Engineering (SRE) transformation initiatives. The role offers the opportunity to influence operational excellence, automation strategy, stakeholder engagement, and technology roadmaps supporting business priorities. Based in Newbury … times. Govern production readiness reviews, release planning, deployment activities, and change management controls. Drive automation across deployment, monitoring, and operational processes using DevOps and SRE practices. Support the adoption of CI/CD capabilities using Jenkins, GitHub, and Azure DevOps. Define observability standards using monitoring tools such as Dynatrace, Splunk ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
internal workflows that help our people deliver better outcomes for customers, faster.**About the role:** As a **Senior Site Reliability Engineer (SRE)**, you will play a key role in ensuring the reliability, scalability, and performance of our critical platforms and services. You will lead complex reliability … during incidents.* Makes contributions during post-mortems and RCAs.* Participates in disaster recovery tests.* Implements automation and executes code in production environments.* Contributes to SRE knowledge documentation.* Supports the deployment, monitoring, and reliability of services integrating AI tools.**Design for Reliability** * Can support architecture and senior engineers ...