Software Engineer Lead - Site Reliability
- Location
- Telford, England, United Kingdom
failover, degradation handling and recovery testing. Embed reliability, security, performance and operational-readiness expectations into engineering delivery. Support major incidents, post-incident reviews and root-cause analysis, ensuring improvement actions are owned and delivered. Use DORA, availability, recovery and operational metrics to identify risks and evidence progress. … expertise in Terraform, GitHub, GitHub Actions, CI/CD pipeline design, deployment automation and release support. Practical knowledge of incident management, problem management, root-cause analysis and operational readiness. Good awareness of Site Reliability Engineering principles, with the ability to apply them pragmatically to DevOps and platform ...