1 to 25 of 29 Permanent Site Reliability Engineering Jobs in the East of England

Site Reliability Engineer

Location
Cambridge, England, United Kingdom
learn more, visit http://www.darktrace.com. **Job D****escription****:**## **About the Role**We’re looking for a **Site Reliability Engineer (SRE)** to bring deep expertise in a key reliability domain and help shape the future of our platform reliability strategy.SRE sits at the heart … your area of specialism**, working across teams to embed best practices, solve complex reliability challenges, and improve system resilience at scale.Unlike a generalist SRE, this role focuses on a **core domain of expertise**—such as **observability, performance engineering, data infrastructure reliability, security-focused SRE, or network reliability ...

Site Reliability Engineer

Location
Cambridge, England, United Kingdom
world’s largest content providers, including Amazon, Google and Microsoft trust Bango technology to reach subscribers everywhere. Bango, where people subscribe. Role As a Site Reliability Engineer at Bango, you own the reliability, performance and continuous improvement of the Bango Platform end-to-end — from the infrastructure … constructive, detailed feedback. Call out areas where Bango can improve service or reduce cost. Essentials 3+ years' experience in a Cloud, Platform, DevOps or SRE role in a commercial environment. Production experience with a major cloud provider (AWS, Azure or GCP). Strong Linux administration and troubleshooting (process management, memory ...

Senior / Lead Site Reliability Engineer

Location
Watford, England, United Kingdom
customer-facing systems during both normal operation and peak lottery events. The role combines hands-on engineering, incident leadership, and ownership of the SRE improvement backlog and reporting, working across platform, product, and operational teams. Objectives of the role Own reliability outcomes across services using SLOs, SLIs … performance optimisation: Latency reduction Throughput scaling Cost efficiency (AWS utilisation and associated log costs, observability license consumption) Backlog ownership & reporting Own and prioritise the SRE backlog, balancing: Reliability improvements Technical debt Automation opportunities to reduce/offload toil Produce structured reporting covering: SLO performance Incident trends and MTTR Platform ...

Site Reliability Engineer: Observability & Platform Resilience

Location
Cambridge, England, United Kingdom
Darktrace Ltd is looking for a Site Reliability Engineer in Cambridge, UK, to enhance their platform reliability strategy. In this pivotal role, you will work closely with Platform Engineering and DevSecOps to implement best practices and solve complex reliability challenges. The ideal candidate will have … substantial experience in Site Reliability Engineering or a related field, with a strong background in programming, cloud platforms, and performance engineering. The role offers a dynamic work environment with competitive benefits including private medical insurance, life insurance, and additional holiday days. #J-18808-Ljbffr ...

Cloud Platform Tech Lead

Hiring Organisation
Danaher
Location
Cambridgeshire, United Kingdom
Employment Type
Full Time
Azure supporting specific business and technology requirements. Working closely with Enterprise Architecture, Cyber Security, Product Engineering, and Site Reliability Engineering (SRE), the Platform Technology Lead delivers cloud platform capabilities, infrastructure automation, self-service services, and operational improvements that enable faster, safer, and more reliable technology delivery. … driving record It would be a plus if you also possess previous experience in: Microsoft Azure and Azure Landing Zones. Platform Engineering and SRE operating models. GitHub Actions and Azure DevOps. Observability and monitoring platforms. Abcam, a Danaher operating company, offers a broad array of comprehensive, competitive benefit programs ...

Senior Site Reliability Engineer

Location
Bishop's Stortford, England, United Kingdom
Want your engineering skills to enable real scientific breakthroughs? Were looking for a Senior Site Reliability Engineer to join our highly skilled IT Operations team at EMBL-EBI. At EMBL-EBI, our IT & Technical Services department underpins groundbreaking research that improves human and planetary health. As part … small but highly skilled IT Operations team, youll play a critical role in ensuring the availability, reliability and efficiency of services that support researchers across the globe. In this hands‐on role, youll ensure the reliability, performance, and availability of the services that thousands of scientists rely ...

Site Reliability Engineer – Platform & Automation

Location
Cambridge, England, United Kingdom
Bango plc is seeking a Site Reliability Engineer to own reliability, performance and continuous improvement of the Bango Platform end-to-end. You’ll merge platform engineering, automation and delivery ownership with proactive incident response and customer impact management, shaping the automation and security posture across … stack. As a member of the Managed Services & Support team, you’ll collaborate with NOC, Partner Support, TSM and other engineering groups to ensure secure, #J-18808-Ljbffr ...

Senior Site Reliability Engineer: Lead Resilience & Incidents

Location
Watford, England, United Kingdom
Allwyn UK in Watford seeks a Senior/Lead Site Reliability Engineer to drive reliability across the digital estate, ensuring high availability and performance of customer-facing systems during normal operation and peak lottery events. You will own SLOs/SLIs, push automation with Terraform, mentor engineers ...

Software Development Team Lead

Location
Cambridge, England, United Kingdom
cross-functional engineers. This team includes back-end developers (C#/.NET, Node.js), front-end developers (HTML, CSS, JavaScript, Vue.js), test automation engineers, and Site Reliability Engineers (AWS). Responsibilities Lead and develop a cross-functional software development team, including back-end, front-end, test automation, and Site Reliability Engineers. Shape software architecture and guide the implementation across the full software development lifecycle with focus on performance, security, and maintainability. Act as a hands-on technical leader while supporting design decisions, coding, and ensuring best practices are followed. Own the technical direction of the product roadmap ...

Senior Platform Engineer

Location
Welwyn Garden City, England, United Kingdom
Senior Platform Engineer, you will lead the design, evolution, and reliability of the core platform that underpins our engineering ecosystem. You will set technical direction, define standards, and drive best practices that enable product teams to deliver securely, efficiently, and at scale. Your role goes beyond implementation … tools, technologies, and approaches to keep the platform modern, efficient, and competitive. What we would like from you Strong experience in platform engineering, SRE, or DevOps within a distributed cloud environment. Deep expertise in Kubernetes and containerised workloads, ideally in managed environments such as AKS. Proven experience designing ...

Senior SRE — Research Infrastructure & Reliability

Location
Bishop's Stortford, England, United Kingdom
EMBL-EBI is seeking a Senior Site Reliability Engineer to join the IT Operations team. You will ensure reliability, performance, and availability of global research services, tackling infrastructure challenges and driving automation and resilience. Hybrid working with on-site presence required. You will work on identity ...

Staff DevOps Engineer

Location
Cambridge, England, United Kingdom
Mentor team members and drive knowledge sharing. Participate in project management and on‐going strategic planning. Qualifications 6+ years of experience in DevOps/SRE environment. 6+ years of experience working with Linux and Windows systems. Strong understanding and knowledge of Kubernetes. Strong understanding and knowledge of cloud infrastructure components ...

Head of Cloud

Location
Norwich, England, United Kingdom
embedded finance solutions, trusted by 90,000+ businesses worldwide. We build modern, scalable platforms that power payments, data, and AI-driven experiences. Join an engineering culture that values ownership, clean and testable code, and continuous improvement. The Opportunity We are seeking a Head of Cloud to serve … infrastructure-as-code mindset. Strategic Thinker: Ability to influence directors and align technical initiatives with commercial objectives. Reliability & Resilience: Expert understanding of SRE principles, building for failure, and high-availability systems. Nice to Have Experience with service mesh and advanced traffic management. Expertise in data persistence (RDS, Aurora, DynamoDB ...

Senior Private Cloud Engineer

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
great opportunity to build and operate a greenfield private cloud platform based on OpenStack, delivering scalable and reliable infrastructure services for Arm's engineering teams. The platform underpins large-scale engineering workloads and is central to the organisation's strategy. The platform is to be built using modern … Experience: Experience with Linux systems. Hands-on experience with Terraform/Ansible or other IaC tools. Experience of working in a DevOps/SRE environment. Programming experience in Python or similar. Understanding and ability to debug compute, networking, storage in large production environments."Nice To Have" Skills and Experience: Exposure ...

Staff Private Cloud Engineer

Location
Cambridge, England, United Kingdom
lead the design, build and operation of a greenfield multi-tenant private cloud platform based on OpenStack, delivering scalable and reliable infrastructure services for engineering teams. The platform underpins large-scale engineering workloads and is central to the organisation’s infrastructure strategy. The platform is to be built … storage or compute. Hands-on GitOps driven and CI/CD pipelines (Jenkins, ArgoCD, FluxCD, etc). Experience of working in a DevOps/SRE environment. “Nice To Have” Skills and Experience: Experience managing hybrid platforms across multiple data centres and clouds. Hands-On experience with Kubernetes and cloud-native ...

Platform Engineer - Essex

Location
Chelmsford, England, United Kingdom
Platform Engineer , you will design, build and continuously improve our Internal Developer Platform (IDP) and cloud infrastructure. You’ll work at the intersection of engineering, infrastructure, security and developer experience — creating the tools, automation and standards that allow our teams to ship high-quality software faster, safer and more … that are secure, scalable, resilient and cost-efficient . Monitor platform health and performance, identify bottlenecks and continuously optimise our infrastructure. Partner closely with SRE, Security and application engineering teams to establish standards, patterns and best practices. Automate repetitive operational tasks and eliminate toil through tooling, automation and effective ...

Lead Private Cloud Engineer — OpenStack Platform

Location
Cambridge, England, United Kingdom
experience across compute, networking and storage layers. The role involves leading incident response, improving scalability and performance, and mentoring engineers in a DevOps/SRE culture within a hybrid working framework. #J-18808-Ljbffr ...

Senior Network Engineer

Hiring Organisation
Infoplus Technologies UK Limited
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£450.00 - £500.00 per day
with enterprise platforms such as ServiceNow . The successful candidate will also contribute towards expanding network automation and developing a strategic roadmap towards Network SRE practices . Key Responsibilities Review and improve the existing network data landscape, data quality and integrations. Develop and enhance NetBox for enterprise network inventory … documentation. Perform network device baselining and configuration validation. Improve network monitoring, data quality and operational reliability. Define KPIs, SLIs and SLOs to support Network SRE practices. Identify opportunities for proactive monitoring, fault prevention and self-healing automation. Work closely with Operations, Security, Architecture, Platform and ITSM teams. Required Technical Skills ...

Operations Team Lead (Production & Reliability)

Location
Basildon, England, United Kingdom
improving. This is a hands‐on role. You’ll shape process, lead incidents, build the team, and move us from reactive firefighting to proactive reliability engineering. What You’ll Own Production Stability and availability of all live systems Operational readiness for new releases Safe production access and change coordination … Raise the bar on operational discipline You’re responsible for both system performance and team performance. What We’re Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under ...

Operations Team Lead (Production & Reliability)

Location
Norwich, England, United Kingdom
improving. This is a hands‐on role. You’ll shape process, lead incidents, build the team, and move us from reactive firefighting to proactive reliability engineering. What You’ll Own Production Stability and availability of all live systems Operational readiness for new releases Safe production access and change coordination … Raise the bar on operational discipline You’re responsible for both system performance and team performance. What We’re Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under ...

Operations Team Lead (Production & Reliability)

Location
Watford, England, United Kingdom
improving. This is a hands‐on role. You’ll shape process, lead incidents, build the team, and move us from reactive firefighting to proactive reliability engineering. What You’ll Own Production Stability and availability of all live systems Operational readiness for new releases Safe production access and change coordination … Raise the bar on operational discipline You’re responsible for both system performance and team performance. What We’re Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under ...

Operations Team Lead (Production & Reliability)

Location
Cambridge, England, United Kingdom
improving. This is a hands‐on role. You’ll shape process, lead incidents, build the team, and move us from reactive firefighting to proactive reliability engineering. What You’ll Own Production Stability and availability of all live systems Operational readiness for new releases Safe production access and change coordination … Raise the bar on operational discipline You’re responsible for both system performance and team performance. What We’re Looking For Strong experience in SRE, DevOps, Infrastructure, or Production Engineering Prior experience leading technical teams Deep hands‐on incident management experience Strong observability and reliability mindset Calm under ...

Service Design Specialist

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
Overview: We are building a modern Service Management capability combining ITIL 4, SRE, automation and operational governance to enable fast, reliable delivery. This role translates business and technical requirements into practical, end-to-end service designs, creating the models, documentation and readiness evidence needed to ensure services are supportable, resilient … Skills and Experience: Experience with ServiceNow or a comparable ITSM platform, particularly service catalogue, CMDB, CSDM, service mapping, knowledge or workflow capabilities! Understanding of SRE concepts such as SLAs, SLOs, SLIs, error budgets, service health, reliability and observability. Experience in DevOps, CI/CD, cloud, SaaS, PaaS, platform engineering ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Birchanger, Hertfordshire, United Kingdom
Employment Type
Permanent
Salary
GBP 40,000 - 50,000 Annual
Aviva family, is growing our Technology capability and we're looking for a talented Performance and Monitoring Engineer to help us strengthen the stability, reliability and performance of our systems. If you're passionate about monitoring, observability and using data to proactively improve service health, this is a great … qualifications Microsoft certifications (AZ-900, AZ-104, AZ-305, AZ-500) or similar Experience with LogicMonitor admin, Grafana or other observability tools Familiarity with SRE concepts (SLIs, SLOs, error budgets) Understanding of ITIL processes Who are Solus? Solus, who are owned by Aviva, are one of the UK leaders ...

Performance and Monitoring Engineer

Hiring Organisation
Solus Accident Repair Centres
Location
Stansted, Essex, South East, United Kingdom
Employment Type
Permanent
Salary
£50,000
Aviva family, is growing our Technology capability and we're looking for a talented Performance and Monitoring Engineer to help us strengthen the stability, reliability and performance of our systems. If you're passionate about monitoring, observability and using data to proactively improve service health, this is a great … qualifications Microsoft certifications (AZ-900, AZ-104, AZ-305, AZ-500) or similar Experience with LogicMonitor admin, Grafana or other observability tools Familiarity with SRE concepts (SLIs, SLOs, error budgets) Understanding of ITIL processes Who are Solus? Solus, who are owned by Aviva, are one of the UK leaders ...