101 to 125 of 909 Site Reliability Engineering Jobs in England

Senior DevSecOps Engineer

Location
City Of London, England, United Kingdom
operating the software delivery infrastructure required to develop and deploy advanced autonomous systems for defence applications. This role sits at the intersection of software engineering, platform engineering, cyber security, and defence systems engineering. The DevSecOps Engineer works alongside autonomy, software, systems, integration, and test engineers to create secure … delivery pipelines that enable teams to rapidly develop, integrate, test, and deploy mission critical software. The ideal candidate has a strong software and platform engineering background combined with significant experience operating within UK defence environments. They have a strong understanding of the UK Ministry of Defence/NATO approach ...

Senior Network Engineer- IP

Location
Ipswich, England, United Kingdom
will take the lead on complex, high-impact fault resolution spanning multiple platforms and services, acting as a senior technical escalation point. Applying SRE principles and deep technical knowledge, you will drive improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing … automation and IaC tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building ...

Senior Network Engineer- IP

Location
Birmingham, England, United Kingdom
will take the lead on complex, high-impact fault resolution spanning multiple platforms and services, acting as a senior technical escalation point. Applying SRE principles and deep technical knowledge, you will drive improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing … automation and IaC tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building ...

Senior Network Engineer- IP

Location
Greater London, England, United Kingdom
will take the lead on complex, high-impact fault resolution spanning multiple platforms and services, acting as a senior technical escalation point. Applying SRE principles and deep technical knowledge, you will drive improvements in service availability and reliability through end-to-end business ownership – implementing flawless network change, championing … automation and IaC tools (e.g. Ansible, Terraform, Netconf/YANG) to manage network infrastructure at scale and reduce operational toil. Proven ability to apply SRE principles – automation, observability and toil reduction – to improve service availability, with proficiency in a programming or scripting language such as Python. Strong proficiency in building ...

Junior SRE – Endpoint Focus

Location
Greater London, England, United Kingdom
technical components. Automation & Continuous Improvement: Proactively isolate recurring operational issues, eliminating manual workflow friction through shell scripting, automated provisioning design, and strategic process enhancements. SRE Transition: Partner closely with Senior SRE team members to progressively absorb production infrastructure tasks, system monitoring duties, and core site reliability principles. Qualifications … performance. Professional Trajectory: A clear, defined motivation to evolve technically and professionally into a Production Engineering or Site Reliability Engineering (SRE) role. Commute Compliance: Willingness and ability to work regularly from our client’s modern office facilities located in the Moorgate area of London (minimum ...

Senior Software Engineer

Hiring Organisation
Visa
Location
Basingstoke, Hampshire, UK
Employment Type
Full-time
work that matters - to you, to your community, and to the world. Progress starts with you. Job Description Build AI-driven automation that powers reliability at global scale. Work on real engineering problems, writing production-grade software that improves resilience, automation, and operational efficiency across critical payment infrastructure. … Advanced Degree (e.g. Masters, MBA, JD, MD) What We're Looking For Core Requirements 1-5 years of software engineering experience, or equivalent SRE-DevOps experience with strong software development skills Strong programming skills in Python, Java, or Go Solid understanding of data structures, algorithms, system design, and distributed ...

Principal Platform Engineer

Location
Greater London, England, United Kingdom
processes and platform lifecycle management Experience establishing repeatable, standardised engineering workflows Experience designing for resilience, fault tolerance, observability, and operational excellence Experience applying SRE principles and practices Experience in performance analysis, capacity planning, scalability engineering, and proactive reliability improvement Experience establishing service-level objectives, monitoring, alerting … resume keywords AWS Cloud Platform Architecture Infrastructure as Code CI/CD Automation Platform-as-a-Product Principles Site Reliability Engineering (SRE) Practices ATS Optimization Keywords Hard Skills Cloud Infrastructure Distributed Systems Networking Security Performance Analysis Capacity Planning Scalability Engineering Operational Excellence Monitoring and Alerting Fault ...

Security Operations Manager

Location
Greater London, England, United Kingdom
outsourced support, and set up how alerts are handled and escalated. You’ll also lead how we respond to incidents, working closely with our engineering and reliability colleagues, since a lot of this sits alongside how they already keep … services running. This is a deeply collaborative role. In particular you’ll work hand in hand with our site reliability engineering (SRE) team, since much of security monitoring and response builds on the same tooling and ways of working they already use to keep our services running. ...

Staff Software Engineer, AI Reliability Engineering

Location
Greater London, England, United Kingdom
Staff Software Engineer, AI Reliability Engineering London, UK About Anthropic Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed … from people who've built product stacks, scaled databases, run massive distributed systems, and everything in between. Strong candidates may also Have been an SRE, Production Engineer, or in similar reliability-focused roles on large scale systems Have experience operating large-scale model serving or training infrastructure (>1000 GPUs ...

Senior Site Reliability Engineer

Hiring Organisation
Brevan Howard
Location
London, UK
Employment Type
Full-time
Senior Site Reliability Engineer (SRE) - GCP/KubernetesAbout the RoleWe are seeking an experienced and highly motivated Senior Site Reliability Engineer (SRE) to join our small, agile engineering team. This role offers the unique opportunity to drive the reliability, scalability, and performance … Kubernetes application deployment. Monitoring & Observability: Implement and manage robust monitoring, alerting, and logging solutions to ensure clear system visibility and proactive issue identification. Reliability & Performance: Define, measure, and enforce Service Level Objectives (SLOs) and Service Level Indicators (SLIs). Participate in on-call rotation (if applicable) and lead post ...

Site Reliability Engineer (DV Security Clearance)

Location
Manchester, England, United Kingdom
seeking an experienced and motivated Site Reliability Engineer (SRE) to join a high-performing team supporting multiple data product and platform groups. This role is focused on improving the reliability, scalability, observability, deployment, and operational support of critical data-driven platforms and services operating within complex production … environments. The successful candidate will work closely with engineering, platform, and operational support teams to strengthen monitoring and alerting capabilities, improve logging and traceability, troubleshoot incidents, support deployments, and automate operational processes wherever possible. The environment includes Kubernetes, Helm, the ELK stack, and a broad range of modern Site ...

Oracle OSS Stack Lead

Location
Newbury, England, United Kingdom
serve as the key customer contact for critical OSS-related incidents and strategic programmes while driving DevOps and Site Reliability Engineering (SRE) transformation initiatives. The role offers the opportunity to influence operational excellence, automation strategy, stakeholder engagement, and technology roadmaps supporting business priorities. Based in Newbury … times. Govern production readiness reviews, release planning, deployment activities, and change management controls. Drive automation across deployment, monitoring, and operational processes using DevOps and SRE practices. Support the adoption of CI/CD capabilities using Jenkins, GitHub, and Azure DevOps. Define observability standards using monitoring tools such as Dynatrace, Splunk ...

Oracle OSS Stack Lead

Hiring Organisation
Vodafone
Location
London, UK
Employment Type
Full-time
serve as the key customer contact for critical OSS-related incidents and strategic programmes while driving DevOps and Site Reliability Engineering (SRE) transformation initiatives. The role offers the opportunity to influence operational excellence, automation strategy, stakeholder engagement, and technology roadmaps supporting business priorities. Based in Newbury … times. Govern production readiness reviews, release planning, deployment activities, and change management controls. Drive automation across deployment, monitoring, and operational processes using DevOps and SRE practices. Support the adoption of CI/CD capabilities using Jenkins, GitHub, and Azure DevOps. Define observability standards using monitoring tools such as Dynatrace, Splunk ...

Senior Site Reliability Engineer

Location
Greater London, England, United Kingdom
internal workflows that help our people deliver better outcomes for customers, faster.**About the role:** As a **Senior Site Reliability Engineer (SRE)**, you will play a key role in ensuring the reliability, scalability, and performance of our critical platforms and services. You will lead complex reliability … during incidents.* Makes contributions during post-mortems and RCAs.* Participates in disaster recovery tests.* Implements automation and executes code in production environments.* Contributes to SRE knowledge documentation.* Supports the deployment, monitoring, and reliability of services integrating AI tools.**Design for Reliability** * Can support architecture and senior engineers ...

Senior Site Reliability Engineer

Location
City Of London, England, United Kingdom
internal workflows that help our people deliver better outcomes for customers, faster. About the role: As a Senior Site Reliability Engineer (SRE), you will play a key role in ensuring the reliability, scalability, and performance of our critical platforms and services. You will lead complex reliability … during incidents. Makes contributions during post-mortems and RCAs. Participates in disaster recovery tests. Implements automation and executes code in production environments. Contributes to SRE knowledge documentation. Supports the deployment, monitoring, and reliability of services integrating AI tools. Design for Reliability Can support architecture and senior engineers ...

Software Engineer, Model Deployment- ChatGPT Engineering

Location
Greater London, England, United Kingdom
Software Engineer, Model Deployment- ChatGPT Engineering Applied AI Engineering - London, UK About the Team ChatGPT relies on a large and growing GPU fleet to serve inference workloads reliably and efficiently. We develop the systems and tools that make it possible to introduce new models, manage production deployments, respond … operational issues, and use infrastructure effectively at scale. Our work spans distributed systems, platform engineering, infrastructure automation, and developer experience. We partner closely with research, infrastructure, and product teams to make model deployment more reliable, more efficient, and easier to manage. About the Role We are looking ...

SRE/Infra | Quant Trading

Location
Greater London, England, United Kingdom
Senior SRE/Platform Engineer | Elite Quant Hedge Fund | London Paragon Alpha is partnering with a leading ~$70BN AUM quantitative trading firm following an exceptional 2025 and continuing to invest aggressively across its global technology organisation. As the business scales, they're expanding a world-class Infrastructure/Site … just engineering metrics – they're a competitive advantage. This is a genuinely high-impact engineering role sitting at the intersection of SRE, Platform Engineering, and Software Engineering , where you'll build the tooling and infrastructure that enables researchers and traders to operate at scale. ...

Product Engineering Environment Lead

Location
Greater London, England, United Kingdom
Lead Location: UK (Hybrid with occasional travel) Contract: Interim 6 Months Rate: Inside IR35 Level: Senior Manager/Head of Function Reporting to: Product Engineering Director Essential Experience - Please Read Before Applying We are seeking a technically credible transformation leader who can operate at the intersection of engineering … role is likely to suit candidates from backgrounds such as: DevOps Leadership Platform Engineering Leadership Environment Management Site Reliability Engineering (SRE) Engineering Enablement Technology Operations Software Delivery Transformation Telecommunications Technology Large-scale Digital Engineering Organisations This role is NOT primarily looking for: A hands ...

Site Reliability Manager - Environment Strategy

Location
Greater London, England, United Kingdom
looking for a Site Reliability Manager to join our team in London, United Kingdom in a hybrid working mode. In this role, you will lead a team focused on environment strategy, automation, patch governance and operational reliability for AWS-based platforms. Your responsibilities include setting roadmaps, driving … while ensuring strong technical standards, compliance and resilience across all production and non-production systems. Responsibilities Define and own the vision and roadmap for site reliability and environment strategy Lead, mentor and develop a team of DevOps and environment engineers Set and enforce standards for environment provisioning, lifecycle ...

AWS DevOps Engineer

Location
Greater London, England, United Kingdom
Cost Explorer, Compute Optimizer, and EKS workload right‐sizing. Skills: 3+ years of hands‐on experience in DevOps, Site Reliability Engineering (SRE), infrastructure engineering, or closely related roles (focused on AWS/EKS environments). Proven track record building and maintaining production‐grade CI/… incidents/post‐mortems into actionable learning for personal and team growth. Proven ability to collaborate effectively with cross‐functional stakeholders, including engineering, SRE, security, and product teams. Maintain clear technical documentation, runbooks, and operational playbooks for production systems. We Offer: Experience a dynamic and team‐orientated work environment ...

Site Reliability Engineer

Location
Leeds, England, United Kingdom
Description At evoke, Site Reliability Engineering (SRE) is all about delivering exceptional customer experiences through reliable, scalable and high-performing technology. Operating at the heart of our betting and gaming platforms, our SRE team combines observability, automation and engineering excellence to ensure our systems perform when … systems can scale effectively to meet changing customer demand. Drive automation initiatives that improve operational efficiency, reduce manual effort and enhance service reliability. Promote SRE best practices throughout the organisation, influencing teams through data-driven recommendations and continuous improvement initiatives. Who we are looking for We are committed to responsible ...

Site Reliability Engineer , Cryptography, Access and Identity Services

Hiring Organisation
AmazonWebServices
Location
London, UK
Employment Type
Full-time
availability environment, building and operating critical Cryptography, Access and Identity services for our customers. This exciting role is designed for someone with a strong engineering background and a passion for driving efficiency, quality, and process improvements within our service operations. We are an operations team, but a key focus … supported in the workplace and at home, there's nothing we can't achieve. Basic qualifications- Experience in site reliability engineering (SRE), systems engineering, systems administration, DevOps, security administration, or network administration- Experience working with Linux- Experience in systems engineering- Experience ...

Lead Site Reliability Engineer

Hiring Organisation
Hackajob Ltd
Location
Milton, Cambridgeshire, UK
defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability. As a Lead Site Reliability Engineer at our client within the Commercial & Investment Bank, you hold a leadership role in your team … technical lead for medium to large-sized products, and provide advice and mentoring to other engineers. Job responsibilities Demonstrates and champions site reliability culture and practices and exerts technical influence throughout your team Leads initiatives to improve the reliability and stability of your team's applications ...

Software Engineer Lead - Site Reliability

Location
Telford, England, United Kingdom
proactive, self-starting engineer who enjoys getting things done and improving the reliability of live digital services, Standard Life could be the place for you. We’re looking for a Lead DevOps Engineer to join our Digital Engineering team. This role is focused on making immediate, practical improvements … GitHub/GitHub Actions, Azure DevOps, Terraform and automated testing. Improve deployment safety, release readiness and operational readiness for customer-facing digital services. Apply SRE principles pragmatically to improve availability, recoverability, monitoring and incident learning. Strengthen monitoring, logging, tracing, alerting and service-health dashboards across digitally connected workloads. Reduce single ...

Lead Site Reliability Engineer

Location
Greater London, England, United Kingdom
trading technology stack is undergoing a multi‐year convergence and modernization journey. You will play a pivotal role in shaping our next‐generation SRE patterns, reliability frameworks, observability strategy, and performance engineering capabilities across globally distributed systems. This role is ideal for an SRE specialist who thrives … codebase (Java, Kotlin, Python) to implement reliability improvements, performance optimisations, bug fixes, and automation. Lead the design and rollout of modern SRE patterns across trading systems, including automated remediation, self‐healing workflows, and resilience engineering. Use enterprise‐authorized AI capabilities within the work environment to accelerate major‐incident triage ...