10 of 10 Permanent Chaos Engineering Jobs in the UK excluding London

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham Birmingham · United Kingdom · Vice President

Location
Birmingham, England, United Kingdom
Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham location_on Birmingham, West Midlands, England, United Kingdom WHAT WE DO Site Reliability Engineering at Goldman Sachs sits at the intersection of software engineering, systems design, and production excellence. In this VP role, you will help … engineer highly reliable, observable, and resilient platforms that support critical business services at scale. You will collaborate with multiple engineering teams to continually improve our production system architecture, facilitate fast delivery of new services, and reduce downtime. This role is for software engineers who enjoy solving complex distributed system ...

Vice President - Site Reliability Engineering (SRE) - The Core Engineering - Birmingham

Location
West Midlands, England, United Kingdom
WHAT WE DO Site Reliability Engineering at Goldman Sachs sits at the intersection of software engineering, systems design, and production excellence. In this VP role, you will help engineer highly reliable, observable, and resilient platforms that support critical business services at scale. You will collaborate with multiple engineering … teams to continually improve our production system architecture, facilitate fast delivery of new services, and reduce downtime. Job Description WHAT WE DO Site Reliability Engineering at Goldman Sachs sits at the intersection of software engineering, systems design, and production excellence. In this VP role, you will help engineer ...

Sr. Manager, Site Reliability

Location
Manchester, England, United Kingdom
shifts from on-premise, hardware-centric products to a cloud-native, SaaS-delivered platform that hospitals depend on 24/7. The Site Reliability Engineering function is the reliability engine of that organization, and this role is the first senior SRE hire — the person who will design the practice … command structure: severity rubric, declaration criteria, war-room protocol, stakeholder communication cadence, and the postmortem template. Train the first cohort of incident commanders across Engineering and Support. Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry ...

Lead Site Reliability Engineer - Chief Technology Office

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
communicate effectively across different levels and stakeholder group Preferred qualifications, capabilities, and skills Experience mentoring or coaching engineers on site reliability practices and engineering standards Familiarity with cloud platforms and infrastructure-as-code practices in large-scale enterprise environments Experience contributing to or leading communities of practice, internal knowledge … sharing, or engineering guilds Exposure to chaos engineering or fault injection methodologies to proactively test system resilience Ability to evaluate and introduce emerging technologies that improve platform reliability and reduce operational toil ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice ...

Lead SRE - AWS,Python

Location
Glasgow, Scotland, United Kingdom
Drive reliability at scale — join a team where your engineering expertise shapes the resilience of critical systems. JPMorganChase is one of the world's leading financial services firms, and the technology that powers it demands the highest standards of reliability, performance, and scale. Here, you will work alongside talented … critical role in ensuring the availability, performance, and resilience of production systems that serve millions of customers and clients globally. You will partner with engineering and product teams to embed reliability practices into the software development lifecycle, driving a culture of operational excellence. Your work will directly impact ...

Infrastructure / DevOps Engineer

Location
Birmingham, England, United Kingdom
lead resolution of infrastructure-level incidents; drive post-mortems Harden infrastructure for HIPAA compliance — encryption, access controls, audit logging, and network security Partner with engineering teams to optimize system performance, cost, and scalability Grow the SRE practice: runbooks, incident response playbooks, chaos engineering, and reliability reviews Location … area, this role is hybrid from our office (3 days in office/2 remote). 5+ years in infrastructure, DevOps, or cloud engineering roles Deep hands-on AWS experience (ECS/EKS, Lambda, RDS/Aurora, S3, VPC, IAM, CloudWatch, and related services) Proficiency with infrastructure-as-code ...

Lead SRE - AWS Platform

Location
Glasgow, Scotland, United Kingdom
validating outputs and handling operational data according to sensitivity and security requirements Required qualifications, capabilities, and skills Formal training or certification on site reliability engineering concepts and advanced applied experience Demonstrated hands-on experience with Amazon Web Services (AWS), including deploying, operating, and maintaining resilient, highly available workloads … demonstrated ability to evaluate and recommend suitable new technologies Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve site reliability engineering workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity Ability to evaluate AI-assisted operational recommendations ...

Lead SRE - AWS Platform

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
validating outputs and handling operational data according to sensitivity and security requirements Required qualifications, capabilities, and skills Formal training or certification on site reliability engineering concepts and advanced applied experience Demonstrated hands-on experience with Amazon Web Services (AWS), including deploying, operating, and maintaining resilient, highly available workloads … demonstrated ability to evaluate and recommend suitable new technologies Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve site reliability engineering workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity Ability to evaluate AI-assisted operational recommendations ...

Principal Platform Engineer

Location
Caerphilly, Wales, United Kingdom
issues alongside long-term projects Focused on automation and reducing manual work Comfortable working within strict security and compliance rules Experience with service mesh, chaos engineering, or advanced observability tools VMware or Veeam certifications Public sector or government client experience What's on Offer A senior role with ...

Software Development Engineer in Test

Location
Swansea, Wales, United Kingdom
squad of civil servants and supplier staff. In this role, you will design, build, maintain, and continuously improve the automated test frameworks and quality‐engineering practices that underpin high‐volume, citizen‐facing digital services handling billions of interactions annually. You will own the full testing lifecycle from strategic planning … Level Test Automation Engineer certification. Advanced Testing Principles: Experience with contract testing (e.g., Pact) for microservices and API ecosystems, as well as familiarity with chaos engineering and site reliability testing principles. Tooling Proficiency: Advanced skills in test management tooling such as Jira, Xray, or Zephyr. Government Assessments: Prior ...