Senior / Lead Site Reliability Engineer
- Hiring Organisation
- Jobleads-UK
- Location
- Watford, England, United Kingdom
resolution -> post‐incident review Ensure blameless post‐mortems with clear remediation ownership Observability & service insight Define and evolve observability strategy using: Splunk (log analytics) CloudWatch (AWS telemetry) Grafana (metrics visualisation) Quantum Metric (user behaviour insight) Standardise: Alerting quality and signal‐to‐noise ratio Dashboards aligned to SLOs and customer … cloud environments, ideally AWS (ECS, with exposure or experience in EKS/Kubernetes) Hands‐on experience with: Terraform (Infrastructure as Code) Observability stacks (Splunk, CloudWatch, Grafana) Strong programming skills (Python, Go, or similar) SRE practices Proven experience implementing: SLOs, SLIs, error budgets Incident management frameworks Observability strategies Strong experience ...