226 to 250 of 1,104 Site Reliability Engineering Jobs in the UK

AWS Platform Engineer

Location
Langley Mill, England, United Kingdom
generation of analytics, AI and digital services. This is a rare opportunity to shape a strategic AWS platform from the ground up, establishing the engineering foundations, automation, security controls and operational standards that future teams will build upon. Working within our Platform Engineering team, you’ll design … implement reusable AWS capabilities that enable engineers across Microlise to deploy securely, operate confidently and innovate at scale. If you’re passionate about platform engineering, Infrastructure as Code, cloud automation and creating exceptional developer experiences, we’d love to hear from you. What will you be doing? Build ...

AI Technical Platform Leader

Hiring Organisation
WTW
Location
London, United Kingdom
Employment Type
Full-Time
Salary
Competitive salary
business stakeholders to mold and implement strategy. The role applies broad technical knowledge with depth in generative AI, agentic systems, enterprise platforms and engineering governance to ensure that AI platforms at WTW are optimally configured to achieve company vision and imperatives. It has end-to-end ownership of implementation … operations in a large, global enterprise: * Software or platform engineering for global, enterprise-scaled solutions * Management of Site reliability engineering (SRE) programs for mission-critical systems * Creation and management of DevOps practices for automated, consistent, and secure solution deployment in regulated environments. * Literacy in global compliance ...

Head of Cyber, Platforms & IT

Location
Chester, England, United Kingdom
enforce secure‐by‐design principles, including cybersecurity standards, cloud architecture guardrails and operational controls Lead DevOps and Site Reliability Engineering (SRE) maturity, embedding monitoring, observability, automated testing and structured incident response Drive adoption of automation and AI‐enabled tooling to improve anomaly detection, incident management, vulnerability management … , cybersecurity or reliability teams in complex SaaS environments Deep expertise in cloud‐native architecture (Azure, AWS or equivalent) Strong understanding of DevOps, SRE principles and SaaS production operations Experience implementing secure‐by‐design frameworks and managing cybersecurity governance Experience owning uptime, reliability and systemic operational performance Strong ...

Site Reliability Engineer - NS London

Hiring Organisation
BAE SYSTEMS
Location
London, United Kingdom
maintained. This role blends operational product support with software engineering to create applications to understand the overall health of our systems. The SRE team sits within a wider programme at the core of the customer mission.The role holder:As an SRE, fundamentally you will be doing work that … human labour, with the objective of limiting traditional manual operations work (incident tickets, on-call etc.) to no more than half of the SRE team's time (and aiming for considerably less). You will have an enthusiasm to learn and experiment, to develop tools to understand application health ...

AWS Cloud Architect, Technology Consulting- London, Leeds, Manchester or Newcastle

Hiring Organisation
Momentum Worldwide
Location
London, United Kingdom
security and automation technologies. You will help design scalable, secure and resilient cloud solutions for our clients, enabling digital transformation through modern architecture, platform engineering and DevOps practices.You should bring practical experience of cloud technologies and infrastructure modernisation, alongside an understanding of how AI-enabled tooling and modern engineering … capabilities- Strong understanding of cloud networking, identity, security and resilience patterns- Experience supporting cloud engineering, DevOps or Site Reliability Engineering (SRE) teams- Cloud architecture certifications in additional public cloud platforms (Azure or GCP)- Experience using AI or agentic techniques within the role through GitHub Copilot, Claude ...

AWS Cloud Architect, Technology Consulting

Hiring Organisation
Momentum Worldwide
Location
London, United Kingdom
Salary
£ 70 K
security and automation technologies. You will help design scalable, secure and resilient cloud solutions for our clients, enabling digital transformation through modern architecture, platform engineering and DevOps practices.You should bring practical experience of cloud technologies and infrastructure modernisation, alongside an understanding of how AI-enabled tooling and modern engineering … capabilities- Strong understanding of cloud networking, identity, security and resilience patterns- Experience supporting cloud engineering, DevOps or Site Reliability Engineering (SRE) teams- Cloud architecture certifications in additional public cloud platforms (Azure or GCP)- Experience using AI or agentic techniques within the role through GitHub Copilot, Claude ...

Systems Engineer, Cryptography, Access and Identity Services

Hiring Organisation
AmazonWebServices
Location
London, UK
Employment Type
Full-time
availability environment, building and operating critical Cryptography, Access and Identity services for our customers. This exciting role is designed for someone with a strong engineering background and a passion for driving efficiency, quality, and process improvements within our service operations. As a Systems Engineer at Amazon you will utilize … supported in the workplace and at home, there's nothing we can't achieve. Basic qualifications- Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration- Experience working with Linux- Experience in systems engineering- Experience ...

Senior – Principal Consultant - Architecture

Location
West of England, England, United Kingdom
secure and resilient supportable cloud‐native solutions Define infrastructure, platform and application architectures Support migration and modernisation initiatives across legacy and cloud environments Establish engineering standards, patterns and automation approaches Provide assurance on technical designs and implementation practices Drive adoption of Infrastructure as Code, automation and modern engineering … Terraform Bicep CloudFormation Infrastructure as Code Configuration management CI/CD pipelines GitHub Azure DevOps GitLab Knowledge of: Knowledge of DevOps, automation and modern engineering practices Platform engineering Site Reliability Engineering principles Data & Analytics Experience with technologies such as: Databricks Microsoft Fabric Azure Data Factory ...

Software Engineering Tech Lead (SRE + AI)

Location
Greater London, England, United Kingdom
+ 3 years in Computer Science, Software Engineering, or a related technical field. Proven record as a Technical Lead or Lead SRE/Software Engineer delivering distributed, high-availability SaaS platforms at scale. Strong proficiency in Python, Go, Java, or C++ with experience designing microservices, APIs, and production automation. … Deep experience with Kubernetes, Docker, and container orchestration in large-scale multi-cluster environments. Proven background in SRE practices: SLI/SLO design, observability platforms (metrics/logs/traces), incident management, and automated RCA. Preferred Qualifications AI & Agentic Systems: Hands‐on experience building LLM pipelines, AI Agents, Model Context ...

VP, SRE: Architect Resilient, Scalable Systems

Location
Birmingham, England, United Kingdom
Goldman Sachs is seeking a Vice President in Site Reliability Engineering (SRE) within Core Engineering at the Birmingham location. You will lead reliability efforts across distributed systems, driving SLOs, observability, and incident response while shaping scalable, automated platforms. This role emphasizes design reviews, reliability ...

Site Reliability Engineer

Hiring Organisation
Twinstream Limited
Location
Cheltenham, Gloucestershire, United Kingdom
Employment Type
Permanent
Salary
GBP 95,000 Annual
Site Reliability Engineer Up to £95,000 DOE Fully Remote Initially Cheltenham Future Hybrid Working Build resilient systems. Solve complex challenges. Make a real impact. Are you a Site Reliability Engineer or DevOps professional who enjoys getting stuck into complex infrastructure challenges? Do you want … work with talented engineers on meaningful projects where reliability, security and t click apply for full job details ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
welcome; join us and let’s build what’s next - together! Role Overview We are seeking a ML Ops Engineer to join our Platform Engineering team at Anaplan. In this role, you will design, scale, and maintain high-performance MLOps and LLMOps infrastructure supporting our cutting-edge AI-infused … tools like Prometheus, Grafana, OpenTelemetry, and Weights & Biases or MLflow. Your Skills Hands-on production experience in DevOps, Site Reliability Engineering (SRE), or Platform Engineering, with some experience dedicated to AI/ML infrastructure. Proven track record of deploying, scaling, and operationalising machine learning models ...

Staff Cloud SRE – AI/ML Platform & GPU Compute London, United Kingdom on-site

Location
Greater London, England, United Kingdom
other to deliver impact. Make Wayve the experience that defines your career! The role This is a rare opportunity to be a founding Staff SRE shaping the reliability of large-scale AI systems and GPU compute infrastructure from the ground up. As a Staff Cloud Site Reliability … Compute platform (large-scale, multi-tenant GPU fleets and scheduling systems driving model training and inference at scale). This is a founding Cloud SRE role. You won’t inherit a mature SRE function, you’ll help create it. You will define the frameworks, automation, and operational standards that ensure ...

Principal AWS Cloud Architect, Technology Consulting- London, Leeds, Manchester or Newcastle

Hiring Organisation
Momentum Worldwide
Location
London, United Kingdom
cloud solutions that accelerate digital transformation and operational excellence. You will help clients realise the full value of cloud adoption through modern architecture, platform engineering, DevOps practices and infrastructure automation, delivering sustainable outcomes in complex and regulated environments.This is a permanent within our Technology Solutions function. At Credera … models- Strong understanding of cloud networking, identity, security and resilience patterns- Experience supporting cloud engineering, DevOps or Site Reliability Engineering (SRE) transformations- Cloud architecture certifications in additional public cloud platforms (Azure or GCP)- Experience using AI or agentic techniques within the role through GitHub Copilot, Claude ...

Systems Engineer, Database Services (AWS)

Location
Greater London, England, United Kingdom
supported in the workplace and at home, there’s nothing we can’t achieve. Basic Qualifications Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration Experience working with Linux Experience in any of the following: Python, Java, Perl ...

Lead Site Reliability / DevOps Engineer

Hiring Organisation
JP Morgan Chase
Location
Glasgow, UK
Employment Type
Full-time
defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Software Engineer at JPMorgan Chase within the Commercial & Investment Bank, you hold a leadership role in your team, demonstrate strong knowledge across … engineers, act as a technical lead for medium to large-sized products, and provide advice and mentoring to other engineers. Job responsibilitiesDemonstrates and champions site reliability culture and practices and exerts technical influence throughout your teamLeads initiatives to improve the reliability and stability of your team ...

Site Reliability Engineer

Location
Greater London, England, United Kingdom
Site Reliability Engineer Reports to: Labs Team Lead Looper Insights Remote-first, with regular visits to our Byfleet and Hounslow data centres The company Looper Insights builds analytics products that help the world’s leading media and entertainment companies understand how their content is performing across digital platforms … collaborate closely to produce a valuable service for an industry about which we are all passionate. The role We’re looking for a Site Reliability Engineer to keep the global LooperBox fleet running, the physical backbone behind every piece of data Looper Insights produces. LooperBoxes sit in front ...

Staff Release Engineer

Location
West of England, England, United Kingdom
title: Staff Release Engineer Location: London or Bristol (Including Hybrid) Salary: £93,600 - £117,000 Team: Release Engineering Reporting To: Software Engineering manager This role is based in the UK and requires an existing right to work in the UK. At this time, we are not able … practices that reliably move software from source to production through centralized developer tooling across all Kaluza engineering domains. Drive Engineering Productivity & SRE Practices: Strategize and embed automated deployment & release mechanisms, testing frameworks, Site Reliability Engineering principles and AI tooling to drastically reduce toil and accelerate ...

Lead Site Reliability / DevOps Engineer

Location
Glasgow, Scotland, United Kingdom
defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Software Engineer at JPMorgan Chase within the Commercial & Investment Bank, youhold a leadership role in your team, demonstrate strong knowledge across multiple … technical lead for medium to large-sized products, and provide advice and mentoring to other engineers. Job responsibilities Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team Leads initiatives to improve the reliability and stability of your team's applications ...

Platform Engineer (Mid-level)

Location
Wakefield, England, United Kingdom
About the Role We are seeking a Platform Engineer (Mid-level) to join our Platform Engineering function. This hands-on technical role focuses on the platforms, infrastructure, automation, and tooling that support our development teams and business operations. The successful candidate will help maintain, improve, and modernise our platform … game development, or creative production environments. Experience migrating services from on-premises infrastructure to cloud platforms. Understanding of Site Reliability Engineering (SRE) principles. #J-18808-Ljbffr ...

Sr. Observability Engineer – Kings Cross, London

Location
Greater London, England, United Kingdom
data for swift root cause identification. Drive post-incident reviews and implement long-term solutions to enhance system resilience.* Collaborate & Influence: Partner with Development, SRE, and Infrastructure leaders to embed observability into the entire technology lifecycle. Influence and drive the adoption of observability best practices across the global organization. Champion … this.**Job Requirements:**Essential Qualifications* Experience: 5-7+ years of hands-on experience in an Observability, Site Reliability Engineering (SRE), or DevOps role, with a proven track record of leading complex projects.* Technical Leadership: Demonstrated experience in architecting and designing large-scale monitoring and observability solutions. ...

Military Data Centre Engineering Operations (DCEO), Amazon Web Services (AWS)

Location
Greater London, England, United Kingdom
Military Data Centre Engineering Operations (DCEO), Amazon Web Services (AWS) Job ID: 10509900 | Amazon Data Services UK Limited This role focuses on those who have military experience interested in working in the private sector. Amazon Web Services (AWS) is seeking a Critical Facilities Technician to join our Data Center … Engineering Operations (DCEO) team. In this role you will support the operation, monitoring and maintenance of the electrical and mechanical infrastructure that powers AWS cloud services. AWS data centers operate 24/7 and rely on highly reliable power and cooling systems. As a technician in this environment ...

Principal Site Reliability Engineer, Infrastructure Observability

Location
Greater London, England, United Kingdom
toolchain and systems, code build and deployment, incident response, and 24x7 monitoring and support. The candidate will also have extensive experience operating within a SRE function within a complex, distributed environment. They will have a demonstrated ability to work horizontally and vertically within an organization with diverse partners and sponsor … learning through blameless post-mortems to improve the shared goal of reliability across services Transform operations teams by facilitating internal change to adopt SRE standard methodologies across the organization and driving strategic growth in this area within Global Technology Analyzes incidents impacting technology availability for high-level trends across ...

Software Engineers

Location
Douglas, Isle of Man, United Kingdom
third-party platforms. Supporting and optimising cloud infrastructure and platform services within Azure environments. Driving operational excellence through monitoring, observability, automation and service reliability practices. Developing and maintaining automated testing frameworks and quality engineering standards. Supporting DevOps, CI/CD and infrastructure automation initiatives. Contributing to technology modernisation … platform upgrades and transformation programmes. Collaborating with business stakeholders, project teams and technology colleagues to deliver successful outcomes. Championing engineering best practices, innovation and continuous improvement. Key Skills and Experience Software engineering and application development. Microsoft technologies including Azure, .NET, C#, Power Platform and associated services. Solution architecture ...

Product Associate - SRE Team - Chase UK

Location
London, United Kingdom
oriented and possess an interest in the financial sector and focus on addressing our customer needs. We work in teams focused on improving the reliability, resilience, observability, and operability of customer-facing digital banking services. We build automation, define measurable reliability practices, reduce operational friction, and partner with … engineering teams to ensure services are designed, delivered, and operated with reliability in mind. Job responsibilities Support the product strategy and delivery of reliability capabilities, including standards, observability, incident practices, automation, and developer experience improvements. Partner with engineers, site reliability engineers, and cross-functional teams ...