site reliability engineer
- Hiring Organisation
- Jobleads-UK
- Location
- Greater London, England, United Kingdom
monitor KPIs for system reliability, performance, and operational efficiency Advance automation, Infrastructure as Code approaches, and promote self-healing systems using AI/ML techniques Develop robust incident management frameworks and lead major incident response activities for critical systems Implement blameless postmortems and deliver systemic improvements across production environments Establish … incident management processes, ITSM standards, and ITIL principles Knowledge of resilience design patterns, high availability, and fault‐tolerant architectures Familiarity with AI/ML‐driven approaches for operational efficiency and system reliability Ability to lead transformation, influence across teams, and foster continuous improvement in culture Nice to have: Experience ...