51 to 75 of 909 Site Reliability Engineering Jobs in England

Cloud and Platform Engineer-Consultant-AI and Digital Factory

Location
Manchester, England, United Kingdom
more of the following disciplines: cloud engineering, platform engineering, infrastructure engineering, DevOps, and Site Reliability Engineering (SRE). As an engineer in our Cloud Team, you will design, innovate, optimize, build, and run the infrastructure and platforms our clients depend on. Depending on your … service tooling that improve developer experience• Work with clients and internal teams to shape new engineering opportunities and grow a strong DevOps/SRE/platform culture• Provide operational support: monitoring, alerting, troubleshooting, and production issue resolution• Conduct systems tests for security, performance, resilience, and availability• Share your knowledge ...

Cloud and Platform Engineer-Consultant-AI and Digital Factory

Location
Greater London, England, United Kingdom
more of the following disciplines: cloud engineering, platform engineering, infrastructure engineering, DevOps, and Site Reliability Engineering (SRE). As an engineer in our Cloud Team, you will design, innovate, optimize, build, and run the infrastructure and platforms our clients depend on. Depending on your … service tooling that improve developer experience• Work with clients and internal teams to shape new engineering opportunities and grow a strong DevOps/SRE/platform culture• Provide operational support: monitoring, alerting, troubleshooting, and production issue resolution• Conduct systems tests for security, performance, resilience, and availability• Share your knowledge ...

Cloud and Platform Engineer-Consultant-AI and Digital Factory

Location
Newcastle upon Tyne, England, United Kingdom
more of the following disciplines: cloud engineering, platform engineering, infrastructure engineering, DevOps, and Site Reliability Engineering (SRE). As an engineer in our Cloud Team, you will design, innovate, optimize, build, and run the infrastructure and platforms our clients depend on. Depending on your … service tooling that improve developer experience• Work with clients and internal teams to shape new engineering opportunities and grow a strong DevOps/SRE/platform culture• Provide operational support: monitoring, alerting, troubleshooting, and production issue resolution• Conduct systems tests for security, performance, resilience, and availability• Share your knowledge ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Milton, Cambridgeshire, UK
systems run reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability … client in the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI. You will own the reliability, performance, and cost-efficiency of the LLM inference platform end to end. ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Milton, Cambridgeshire, UK
systems run reliably in production at scale. In this role, you'll build and operate large language model serving infrastructure, bringing strong engineering fundamentals and site reliability practices to cutting-edge AI platforms. You'll work hands-on with cloud and Kubernetes-based deployments, deep observability … client within the AI and Machine Learning Platform team, you will build and scale AI infrastructure that modernizes traditional infrastructure management and site reliability engineering through applied AI. You will own the reliability, performance, and cost-efficiency of the large language model inference platform ...

Site Reliability & Observability Engineer – Datadog / Azure

Hiring Organisation
MYO Talent
Location
West Midlands, United Kingdom
Employment Type
Full-Time
Salary
£450.00 - £600.00 per day
Site Reliability & Observability Engineer/Datadog – Synthetic Monitoring, APM, RUM, Log Management, SLO’s, Alerting/Azure/Azure DevOps/Cloudflare/6-month contract/Hybrid – West Midlands/Remote/£450 – 600 per day Inside IR35. One of our leading clients is seeking a Lead … with distributed systems, microservices, and cloud-native architectures. Desirable: Datadog certifications. Azure certifications. Cloudflare administration experience. Background in Site Reliability Engineering (SRE) or Platform Engineering leadership roles. ...

Site Reliability & Observability Engineer Datadog / Azure

Hiring Organisation
MYO Talent
Location
Birmingham, West Midlands, United Kingdom
Employment Type
Contract
Contract Rate
From £450 to £600 per day Inside IR35
Site Reliability & Observability Engineer/Datadog Synthetic Monitoring, APM, RUM, Log Management, SLOs, Alerting/Azure/Azure DevOps/Cloudflare/6-month contract/Hybrid West Midlands/Remote/£450 600 per day Inside IR35. One of our leading clients is seeking a Lead Site … with distributed systems, microservices, and cloud-native architectures. Desirable: Datadog certifications. Azure certifications. Cloudflare administration experience. Background in Site Reliability Engineering (SRE) or Platform Engineering leadership roles. ...

Site Reliability Engineer

Hiring Organisation
REVYBE IT RECRUITMENT LIMITED
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
period of growth and investing heavily in its engineering and platform capabilities. They're looking for an experienced Site Reliability Engineer (SRE) to join the team and play a key role in building highly reliable, scalable, and observable infrastructure. This is a hands-on role focused … experience Help improve platform resilience, scalability, and disaster recovery capabilities Contribute to capacity planning and performance optimisation as the platform scales Establish and champion SRE best practices across the wider engineering function What We're Looking For Proven commercial experience working as an SRE, DevOps Engineer, Platform Engineer ...

Lead Site Reliability Engineer, Athena Core

Location
Greater London, England, United Kingdom
defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at JPMorgan Chase within the Commercial and Investment Banking, Markets Technology – Athena Core team, you hold … programming languages such as: Python, Java/Spring Boot, .Net Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity. Ability to evaluate AI-assisted operational recommendations for correctness ...

Head of Infrastructure and Cyber Security

Location
Greater London, England, United Kingdom
systems through modern tech paradigms such as Infrastructure as Code (IaC), automated CI/CD pipelines, FinOps, and Site Reliability Engineering (SRE). You will establish clear technical roadmaps, professional standards and investment priorities that support Hackney’s wider digital transformation. Cyber security will be central … improved commercial arrangements. – Strong supplier, procurement, budget and risk‐management experience. – Experience with modern engineering practices including Site Reliability Engineering (SRE), Infrastructure as Code (IaC), DevSecOps, and cloud FinOps optimization. This is an opportunity to lead a strategically important service, strengthen Hackney’s cyber resilience ...

Site Reliability Engineer

Location
Manchester, England, United Kingdom
Site Reliability Engineer, you will enhance system reliability, observability and performance through a strong engineering approach and assist with incident resolution and best practices. Full-time Closes 05/08/2026 You will have strong software engineering skills, approaching system reliability and observability … policy. Preferred Skills and Experience Knowledge and experience of modern software development techniques and lifecycles. Excellent knowledge of Site Reliability Engineering (SRE) principles, including the creation and management of effective Service Level Indicators (SLI's) and Service Level Objectives (SLO's) for reliability and customer satisfaction. ...

SRE Technical Lead

Hiring Organisation
83zero Limited
Location
Wokingham, Berkshire, South East, United Kingdom
Employment Type
Permanent, Work From Home
SRE Technical Lead Location: Hybrid (UK - office, client site and home-based working) Salary: Up to £100,000 + 5% bonus We're looking for an experienced SRE Technical Lead to take ownership of the reliability, availability and operational excellence of critical platforms within complex, multi-vendor environments. … clearance. Be a sole UK national. Unfortunately, candidates who do not meet both of these essential requirements cannot be considered. The Role As the SRE Technical Lead, you will: Define and drive the SRE strategy, standards, SLAs, SLOs and error budgets. Embed reliability engineering principles into platform ...

Site Reliability Engineer

Location
West of England, England, United Kingdom
Site Reliability/Platform Engineer We are recruiting for a growing technology infrastructure business that is building a new operational capability to support large-scale, high-performance computing environments. This is an excellent opportunity for an experienced Site Reliability Engineer or Platform Engineer who enjoys automating … that infrastructure more reliable. You'll be joining a growing organisation where you'll have considerable autonomy and the opportunity to help shape the SRE and operational automation capability rather than simply inherit an established environment. Salary: TBC Location/hybrid working: TBC #J-18808-Ljbffr ...

Field Applications Engineer

Location
Bolton, England, United Kingdom
work experience, education level, skill set, and/or location. This is a current vacant position. Job Purpose The Field Applications Engineer supports the reliability, availability, and performance of battery-buffered DC fast charging and energy storage systems deployed in the field. Working closely with Field Service, Engineering … deployment and validation of corrective actions and measure their effectiveness after release. Collaborate across electrical, mechanical, firmware, software, systems, test, operations, field service, and SRE teams to drive issues through root cause and permanent resolution. Document technical investigations, root causes, corrective actions, test results, and lessons learned to continuously improve ...

Site Reliability Engineer

Location
Newcastle upon Tyne, England, United Kingdom
procedures for incident response and operational tasks. Collaborate with cross-functional teams to review and provide feedback on technical designs, ensuring alignment with SRE principles. Participate in on-call rotations and handle critical incidents with confidence and expertise. Continuously improve documentation for systems and services, contributing to a knowledge-sharing … incident management processes like Prometheus, Grafana, New Relic, DataDog, Splunk, Cloudwatch, Sumologic etc. Extensive understanding of networking and security concepts. Bonus Points For Specialized SRE observability experience with New Relic or DataDog. Familiarity with OpenTelemetry, AIOps, MLOps, or SecOps. Location: Newcastle, UK - In-Office (at least 4 days per week ...

Software Engineer/ SRE (Linux)

Hiring Organisation
Visa
Location
Basingstoke, Hampshire, UK
Employment Type
Full-time
work that matters - to you, to your community, and to the world. Progress starts with you. Job Description Site Reliability Engineering (SRE) is essential to Visa's Cloud platform strategy. In this role, you'll ensure our development platform and tools let engineers focus on innovation instead … full coverage. Hands-on expertise is required, especially with major DevTools like GitHub, Jenkins, Jira, and Artifactory. We seek a Software Engineer + SRE hybrid engineer. The ideal candidate deeply understands at least one major DevTool, quickly resolves tool-related issues in collaboration with developers, and applies systems thinking ...

Site Reliability Engineer

Hiring Organisation
E-Solutions IT Services UK Ltd
Location
Leeds, West Yorkshire, United Kingdom
Employment Type
Full-Time
Salary
£280.00 - £300.00 per day
onsite Mode of employment: Inside IR 35 Contract About the Role: We are looking for a passionate and experienced Site Reliability Engineer (SRE) to join our Cloud Platform team. The ideal candidate will have hands-on experience managing large-scale Kubernetes clusters on public cloud environments … strong understanding of modern SRE and DevOps practices. You will be responsible for ensuring high availability, reliability, scalability, and performance of our cloud-native infrastructure and CI/CD systems. Key Responsibilities: • Manage, monitor, and optimize large-scale Kubernetes clusters hosted on public cloud platforms (Azure ...

Manager, Site Reliability Engineering

Location
Harrogate, England, United Kingdom
combine to deliver a unique set of products and services that help people, businesses and governments realize their greatest potential. Title and Summary Manager, Site Reliability Engineering About Mastercard Mastercard is a global technology company in the payments industry, dedicated to enabling an inclusive digital economy that … process: Over 90% of UK salaries More than 70% of household bills Almost all state benefit payments The Opportunity We are seeking a Manager, Site Reliability Engineer to support the Production Manager in overseeing a team of individual contributors who support the Card Transaction Services and Faster Payment ...

Senior Site Reliability Engineer

Location
Nottingham, England, United Kingdom
Site Reliability Engineering capabilities to strengthen reliability, observability, security, and operational excellence across our Risk Intelligence division. As a Senior SRE, you will be a senior hands‐on technical person help shape the foundations of reliability across both new and existing platforms. You will collaborate … requires a highly proactive, hard-working expert with strong leadership presence and ownership of platform reliability outcomes. Key Responsibilities Lead the establishment of SRE foundations for new projects building environments, monitoring, alerting, and ensuring operational readiness from day one. Define, implement, and champion observability standards, tooling, and guidelines across ...

Monitoring & Observability Engineer

Location
Reading, England, United Kingdom
monitor the health and performance of our live cryogenic systems, respond to operational incidents and continuously improve our observability capabilities. You'll collaborate across engineering, operations, software and reliability teams to develop dashboards, refine alerting strategies and automate operational responses that improve reliability and reduce downtime. What … Experience monitoring high-availability technical systems. Understanding of incident management, root cause analysis, reliability engineering or Site Reliability Engineering (SRE) principles. Experience with HTTP/REST APIs, Git, containerisation, Kubernetes or Infrastructure as Code. Continuous improvement, ownership or leadership experience. Why Join OQC You will ...

Senior Software Engineer

Location
Manchester, England, United Kingdom
work-life balance and offer a range of working patterns, including full-time, part-time, and compressed hours. Hybrid working, which combines working on-site and from home, may be more limited due to the nature of the work. However, some homeworking may be available depending on business needs. … existing systems, establish and promote best practices, and deliver high-quality software solutions. Drawing on your expertise in a range of software engineering methodologies, you’ll introduce fresh ideas and innovative approaches that make a real impact at the core of our mission: keeping the UK safe, both ...

Senior Software Engineer Ref. 3839

Location
Cheltenham, England, United Kingdom
work-life balance and offer a range of working patterns, including full-time, part-time, and compressed hours. Hybrid working, which combines working on-site and from home, may be more limited due to the nature of the work. However, some homeworking may be available depending on business needs. … existing systems, establish and promote best practices, and deliver high-quality software solutions. Drawing on your expertise in a range of software engineering methodologies, you’ll introduce fresh ideas and innovative approaches that make a real impact at the core of our mission: keeping the UK safe, both ...

Azure CloudOps Engineer

Location
Greater London, England, United Kingdom
Infrastructure as Code, enhancing observability and AIOps capabilities, and driving automation across both application and infrastructure lifecycles. This role combines Cloud Engineering, DevOps, SRE, and AIOps practices, leveraging automation, AI-assisted operations, and self-healing capabilities to improve platform reliability, operational efficiency, and service availability. Key Responsibilities: Design … cloud and hybrid environments. Essential Skills & Experience 3+ years' experience in Cloud Engineering, DevOps, Platform Engineering, Site Reliability Engineering (SRE), or Cloud Operations roles. Strong experience with Microsoft Azure services and cloud infrastructure environments. Strong experience with Terraform is preferred. Experience with Helm, CloudFormation ...

Site Reliability Engineer

Hiring Organisation
Proactive Appointments
Location
Gloucester, Gloucestershire, United Kingdom
Employment Type
Full-Time
Salary
£500.00 - £580.00 per day
Site Reliability Engineer – DV Cleared Our client is urgently looking for an experienced Site Reliability Engineer to join their team on a contract basis, initially for 6 months with a view to extend. Please note, the role is OUTSIDE of IR35. You must hold live UKIC … DV. The role is on-site 4 days per week in Gloucester. Site Reliability Engineer – Key Skills: Must hold a live UKIC DV Modern configuration management tools (such as Ansible, Chef or similar) Experience working with Terraform Docker containers & container orchestration tools (such as Kubernetes, OpenShift ...

Senior SRE (AWS)

Hiring Organisation
VIQU IT
Location
Wavendon, Bedfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 65,000 - 75,000 Annual
Senior Site Reliability Engineer (AWS focused) Up to £75,000 + bonus + on call allowance Milton Keynes (2 days on site a week) VIQU have partnered with a well-established B2B SaaS company who are going through a significant platform transformation. and so are hiring … Senior Site Reliability Engineer to build stability, respond to live incidents, and assist with system upkeep. The role will also play a key part in on implementing and adopting new tooling and processes surrounding the wider transformation. This is a genuine opportunity to own and operate ...