enhance observability and monitoring, ensuring meaningful alerting, clear operational dashboards, and rapid diagnosis of incidents. Raise the bar for operational maturity, including incident management, root‐cause analysis, and continuous improvement. Security, Compliance & Governance Ensure workloads are designed, deployed, and operated in line with cloud security, compliance, and governance requirements. … code (e.g. Terraform), CI/CD pipelines, and automation tooling. Hands‐on expertise in SRE and operational excellence, including monitoring, alerting, reliability improvement, and incident response. Proven ability to lead technical initiatives end‐to‐end, balancing robustness, efficiency, and developer usability. Experience operating and supporting Kubernetes‐based workloads ...