environment: high uptime, rapid incident resolution, and proactive capacity planning. Build and maintain CI/CD pipelines, deployment automation, and infrastructure-as-code using Terraform, Terragrunt, and Chef. Monitor system health using enterprise observability tools (Prometheus, Grafana, DataDog), defining metrics and response playbooks that get ahead of incidents before they … experience across IAM, VPC, EC2, ELB, RDS, S3, Lambda, API Gateway, Secrets Manager, KMS, CloudWatch, and CloudTrail — including multi-account management and SSO.Proficiency with Terraform for infrastructure-as-code, plus experience with configuration management tools like Chef or Ansible. Experience with CI/CD tooling including GitHub, GitHub Actions, JFrog ...