Site Reliability Engineer / Production Support
- Hiring Organisation
- Hackajob Ltd
- Location
- South West London, London, United Kingdom
- Employment Type
- Permanent
system flows, the services, partners, and teams in play, and how to actively debug an incident. Lead reliability engineering: resilience patterns, performance tuning, capacity planning. Facilitate post-incident reviews and track actions to completion. Use AI tools for automated alert correlation, root cause analysis, and runbook generation. Hunt … alerting, logging, and distributed systems debugging. Experience managing and working with offshore support teams. Hands-on experience with reliability engineering: resilience patterns, performance tuning, capacity planning. Active use of AI tools for incident triage, automation, and runbook generation. Ability to understand complex system flows across multiple services and third ...