Site Reliability Engineer - Service Assurance Systems
- Location
- Greater London, England, United Kingdom
application deployments across development, staging and production environments, ensuring smooth and reliable release processes. Monitor application and infrastructure health using observability tools such as Prometheus and AWS CloudWatch; proactively identify and respond to anomalies and performance degradation. Own and maintain CI/CD pipelines, working to improve build, test … equivalent. Hands-on experience with AWS services (e.g. EC2, S3, ECS, Lambda, CloudWatch) and containerisation using Docker. Experience with monitoring and observability tooling, including Prometheus and AWS CloudWatch, for metrics collection, alerting and dashboarding. Proficiency in scripting with Python for automation, tooling and operational support tasks. Strong communication skills — able ...