resolution post-incident review Ensure blameless post-mortems with clear remediation ownership Observability & service insight Define and evolve observability strategy using: Splunk (log analytics) CloudWatch (AWS telemetry) Grafana (metrics visualisation) Quantum Metric (user behaviour insight) Standardise: Alerting quality and signal-to-noise ratio Dashboards aligned to SLOs and customer … cloud environments, ideally AWS (ECS, with exposure or experience in EKS/Kubernetes) Hands-on experience with: Terraform (Infrastructure as Code) Observability stacks (Splunk, CloudWatch, Grafana) Strong programming skills (Python, Go, or similar) SRE practices Proven experience implementing: SLOs, SLIs, error budgets Incident management frameworks Observability strategies Strong experience ...