Director of Site Reliability Engineering
- Location
- Greater London, England, United Kingdom
Define and monitor KPIs for system reliability, performance, and operational efficiency Advance automation, Infrastructure as Code approaches, and promote self‐healing systems using AI / ML techniques Develop robust incident management frameworks and lead major incident response activities for critical systems Implement blameless postmortems and deliver systemic improvements across … error budgets across services Drive resilience strategies with highly available architectures and disaster recovery readiness Champion an automation‐first culture, leveraging CI / CD pipelines and operational tooling to reduce manual processes Requirements Strong background in Site Reliability Engineering, DevOps, or platform operations in complex ...