environmentsOversee Infrastructure as Code practices (e.g. Terraform) and ensure environments are consistent, auditable, and secureShape monitoring, observability and alerting across the engineering organisation (e.g. Datadog, CloudWatch) so issues are detected and resolved before customers are impacted4Lead incident management practices - on-call, triage, escalation, post-incident reviews, and runbook qualityChampion security … troubleshootingProficiency with Infrastructure as Code tools such as OpenTofu, Terraform, or Cloud formation (OpenTofu preferred)Expertise in observability - monitoring, logging, alerting, and APM (e.g. Datadog, Prometheus, Grafana, CloudWatch)Solid understanding of Docker and containerisation technologiesDeep CI/CD pipeline ownership and release governance in high-stakes environmentsKnowledge of AWS security ...