12 of 12 Chaos Engineering Jobs in London

DevOps Consultant | Harness London

Location
Greater London, England, United Kingdom
seeking a seasoned Harness Platform Specialist with deep expertise across the Harness Software Delivery Platform, including Continuous Integration, Continuous Delivery, Security Testing Orchestration (STO), Chaos Engineering, Cloud Cost Management (CCM), and Harness AI Agents. The ideal candidate will bring hands on experience in designing scalable CI/… platform, aligned with enterprise delivery, security, and compliance standards. Architect and maintain end‐to‐end delivery workflows across Harness modules, including CI, CD, STO, Chaos Engineering, CCM, and emerging AI‐driven capabilities. Implement and manage Harness Delegate installations, upgrades, image customizations, and integrations across cloud and on‐prem ...

Site Reliability Engineering Lead

Location
City Of London, England, United Kingdom
Ready to lead the reliability, scalability, and operational excellence of mission-critical platforms while shaping the future of Site Reliability Engineering? Would you like to mentor high-performing engineers, drive cloud modernization, and influence enterprise-wide engineering practices in a highly collaborative environment? About the Business LexisNexis Risk … establishing and driving reliability, observability, automation, and operational excellence standards across Insurance technology platforms. The team partners closely with application, infrastructure, database, and cloud engineering teams to improve platform availability, scalability, performance, and resilience. ICS leads strategic initiatives including SLO/SLI implementation, observability platform adoption, cloud modernization, operational ...

VodafoneThree - SRE III

Location
Greater London, England, United Kingdom
creating a connected future with technologies like Cloud, AI and big data. What you’ll do In this role, within VodafoneThree's Performance and Chaos Engineering (PaCE) team, you will play a key role in enabling engineering teams to deliver scalable, resilient, and high-performing digital services. … identify and address risks early and build confidence in their solutions before they reach production. Working closely with product, platform, development, and Site Reliability Engineering teams, you will champion a shift-left approach to performance and resilience engineering. Your focus will be on providing the frameworks, tooling, standards ...

devops engineer for AI platforms

Location
Greater London, England, United Kingdom
Задачи: Own the DevOps and Platform roadmap, including Kubernetes platform evolution, application packaging, migration to EKS, and reliable production delivery; Lead by doing by engineering, reviewing, and enhancing Kubernetes and CNCF-aligned infrastructure while setting technical standards; Architect multi‐cluster, multi‐region environments using Istio/Linkerd, Cluster … meshes; Contribute to the Kong AI Gateway, including Dataplane deployments, ACM/SSL integration, and DataDog observability; Champion DevSecOps maturity through SAST/DAST, chaos engineering, and error budget monitoring; Collaborate with Security, Data, and AI teams to shape DevOps and AI platform architectures with regulatory compliance; Stay ...

Sr. Observability Engineer – Kings Cross, London

Location
Greater London, England, United Kingdom
performance bottlenecks, optimize resource utilization, and guide capacity planning.* Lead & Mentor: Act as a technical leader and mentor for the observability team and wider engineering groups. Champion and enforce best practices, fostering a culture of proactive and data-informed decision-making.* Drive Incident & Problem Management: Working with Operations teams … part of this.**Job Requirements:**Essential Qualifications* Experience: 5-7+ years of hands-on experience in an Observability, Site Reliability Engineering (SRE), or DevOps role, with a proven track record of leading complex projects.* Technical Leadership: Demonstrated experience in architecting and designing large-scale monitoring and observability solutions. ...

Architect & Delivery Lead (68018)

Location
Greater London, England, United Kingdom
Make architectural decisions across IAM, Cloud, SRE, Network, Data, and Security — ensuring coherence, reusability, and alignment with business objectives Establish and chair the Architecture & Engineering Governance board, providing technical assurance across all workstreams Own the programme roadmap, resource plan, and financial model — tracking cost savings, team reduction trajectory … vendor and tool selection, ensuring standardisation across the programme and eliminating redundant tooling Build and lead high‐performing distributed teams, fostering a culture of engineering excellence, accountability, and continuous improvement Define the continuous improvement factory model, ensuring the transformation sustains beyond the initial programme Technical Skills & Expertise Broad ...

Lead DevOps Engineer

Location
Greater London, England, United Kingdom
hands‐on team Lead role where you will balance technical expertise with team leadership. Elliptic is building an AI fluent workforce. Across our Product, Engineering, Design and Intelligence we’re going beyond giving everyone the tools. We’re setting new standards. Using AI to transform how we build … one. What You Will Do Own the DevOps and Platform roadmap, including Kubernetes platform evolution, application packaging and migration to EKS, and enabling engineering teams to ship reliably to production. Lead by doing - engineer, review, and enhance Kubernetes and CNCF-aligned infrastructure, setting technical standards. Architect multi‐cluster, multi ...

Test Engineer - Cloud & Data Platform (Lakehouse)

Location
Greater London, England, United Kingdom
scanning tools (Checkov, tfsec) | Intermediate | Data platforms (Snowflake, Databricks, DBT, Apache Iceberg) | Intermediate | Python (test automation, pipeline utilities) | Intermediate | SRE practices (SLIs/SLOs, chaos engineering, performance testing) | Intermediate | Compliance testing (encryption, tagging, data residency) | Intermediate |Required Experience 3–5+ years\*\* in software testing, QA engineering, or platform engineering with a quality focus 2+ years\*\* testing cloud infrastructure and IaC (AWS + Terraform or CloudFormation) Hands-on experience with CI/CD pipeline testing (GitLab CI or equivalent) Familiarity with data platforms (Snowflake, Databricks, DBT, or AWS Glue) Python scripting for test automation ...

Observability Engineer - Assistant Vice President

Location
Greater London, England, United Kingdom
resiliency patterns to enhance application reliability. Recovery Testing Support: Support and participate in advanced recovery testing, including Production Swing Tests, Data Recovery Tests, and chaos engineering practices. Automation Drive: Drive the adoption and development of automation solutions (such as Ansible playbooks and Terraform) to minimize recovery time … application teams, SRE leads, and stakeholders. Qualifications Significant professional experience in software development, or an equivalent field, with a strong focus on Site Reliability Engineering and Observability. Expertise in analyzing complex application, database, network, and OS issues within large‐scale, customer‐facing systems. A service‐oriented attitude combined with ...

Staff Software Engineer, AI Reliability Engineering

Location
Greater London, England, United Kingdom
Staff Software Engineer, AI Reliability Engineering London, UK About Anthropic Anthropic's mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers … Role Claude has your back. AIRE has Claude's. Help us keep Claude reliable for everyone who depends on it. AIRE (AI Reliability Engineering) partners with teams across Anthropic to improve reliability across our most critical serving paths -- every hop from the SDK through our network, API layers, serving ...

Senior Backend Engineer: Chaos & Reliability (Remote)

Location
Greater London, England, United Kingdom
Camunda seeks a Senior Software Engineer, Backend, to own automated reliability testing and chaos engineering for Camunda 8. You will break things in safe environments to strengthen the platform, guide product direction, and collaborate with friendly colleagues who live our FAITH values. You’ll design and run chaos ...

SRE III: Performance & Chaos Engineer

Location
Greater London, England, United Kingdom
VodafoneThree PaCE team is hiring to enable engineering squads to deliver scalable, resilient digital services. You will embed performance, reliability and observability practices throughout the software lifecycle, helping teams identify risks before production. Join us to drive CI/CD integration, define meaningful SLOs, and develop self-service capabilities ...