Site Reliability Engineer - Service Assurance Systems
- Location
- Greater London, England, United Kingdom
across development, staging and production environments, ensuring smooth and reliable release processes. Monitor application and infrastructure health using observability tools such as Prometheus and AWS CloudWatch; proactively identify and respond to anomalies and performance degradation. Own and maintain CI/CD pipelines, working to improve build, test and deployment automation … towards greater automation, infrastructure-as-code, and operational maturity. Manage and maintain containerised workloads using Docker, and support the operation of services hosted on AWS cloud infrastructure. Write and maintain operational scripts (primarily Python) to automate routine tasks, support incident investigations and improve team efficiency. Collaborate with software developers ...