7 of 7 Remote/Hybrid Chaos Engineering Jobs in England

VodafoneThree - SRE III

Location
Greater London, England, United Kingdom
creating a connected future with technologies like Cloud, AI and big data. What you’ll do In this role, within VodafoneThree's Performance and Chaos Engineering (PaCE) team, you will play a key role in enabling engineering teams to deliver scalable, resilient, and high-performing digital services. … identify and address risks early and build confidence in their solutions before they reach production. Working closely with product, platform, development, and Site Reliability Engineering teams, you will champion a shift-left approach to performance and resilience engineering. Your focus will be on providing the frameworks, tooling, standards ...

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham Birmingham · United Kingdom · Vice President

Location
Birmingham, England, United Kingdom
Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham location_on Birmingham, West Midlands, England, United Kingdom WHAT WE DO Site Reliability Engineering at Goldman Sachs sits at the intersection of software engineering, systems design, and production excellence. In this VP role, you will help … engineer highly reliable, observable, and resilient platforms that support critical business services at scale. You will collaborate with multiple engineering teams to continually improve our production system architecture, facilitate fast delivery of new services, and reduce downtime. This role is for software engineers who enjoy solving complex distributed system ...

devops engineer for AI platforms

Location
Greater London, England, United Kingdom
Задачи: Own the DevOps and Platform roadmap, including Kubernetes platform evolution, application packaging, migration to EKS, and reliable production delivery; Lead by doing by engineering, reviewing, and enhancing Kubernetes and CNCF-aligned infrastructure while setting technical standards; Architect multi‐cluster, multi‐region environments using Istio/Linkerd, Cluster … meshes; Contribute to the Kong AI Gateway, including Dataplane deployments, ACM/SSL integration, and DataDog observability; Champion DevSecOps maturity through SAST/DAST, chaos engineering, and error budget monitoring; Collaborate with Security, Data, and AI teams to shape DevOps and AI platform architectures with regulatory compliance; Stay ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
hands‐on team Lead role where you will balance technical expertise with team leadership. Elliptic is building an AI fluent workforce. Across our Product, Engineering, Design and Intelligence we’re going beyond giving everyone the tools. We’re setting new standards. Using AI to transform how we build … one. What You Will Do Own the DevOps and Platform roadmap, including Kubernetes platform evolution, application packaging and migration to EKS, and enabling engineering teams to ship reliably to production. Lead by doing - engineer, review, and enhance Kubernetes and CNCF-aligned infrastructure, setting technical standards. Architect multi‐cluster, multi ...

Infrastructure / DevOps Engineer

Location
Birmingham, England, United Kingdom
lead resolution of infrastructure-level incidents; drive post-mortems Harden infrastructure for HIPAA compliance — encryption, access controls, audit logging, and network security Partner with engineering teams to optimize system performance, cost, and scalability Grow the SRE practice: runbooks, incident response playbooks, chaos engineering, and reliability reviews Location … area, this role is hybrid from our office (3 days in office/2 remote). 5+ years in infrastructure, DevOps, or cloud engineering roles Deep hands-on AWS experience (ECS/EKS, Lambda, RDS/Aurora, S3, VPC, IAM, CloudWatch, and related services) Proficiency with infrastructure-as-code ...

Staff Software Engineer, AI Reliability Engineering

Location
Greater London, England, United Kingdom
Staff Software Engineer, AI Reliability Engineering London, UK About Anthropic Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers … Role Claude has your back. AIRE has Claude's. Help us keep Claude reliable for everyone who depends on it. AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths -- every hop from the SDK through our network, API layers, serving ...

Senior Backend Engineer: Chaos & Reliability (Remote)

Location
Greater London, England, United Kingdom
Camunda seeks a Senior Software Engineer, Backend, to own automated reliability testing and chaos engineering for Camunda 8. You will break things in safe environments to strengthen the platform, guide product direction, and collaborate with friendly colleagues who live our FAITH values. You’ll design and run chaos ...