Senior / Lead Site Reliability Engineer
- Location
- Watford, England, United Kingdom
post-incident review Ensure blameless post-mortems with clear remediation ownership Observability & service insight Define and evolve observability strategy using: Splunk (log analytics) CloudWatch (AWS telemetry) Grafana (metrics visualisation) Quantum Metric (user behaviour insight) Standardise: Alerting quality and signal-to-noise ratio Dashboards aligned to SLOs and customer impact Drive … Capacity & performance engineering Own capacity planning for: High-concurrency draw events Traffic spikes during jackpots Lead performance optimisation: Latency reduction Throughput scaling Cost efficiency (AWS utilisation and associated log costs, observability license consumption) Backlog ownership & reporting Own and prioritise the SRE backlog, balancing: Reliability improvements Technical debt Automation opportunities ...