Site Reliability Engineer - Service Assurance Systems
- Location
- Greater London, England, United Kingdom
Prometheus and AWS CloudWatch; proactively identify and respond to anomalies and performance degradation. Own and maintain CI/CD pipelines, working to improve build, test and deployment automation to reduce manual effort and increase release confidence. Champion and drive the adoption of DevOps practices and culture within … group, working towards greater automation, infrastructure-as-code, and operational maturity. Manage and maintain containerised workloads using Docker, and support the operation of services hosted on AWS cloud infrastructure. Write and maintain operational scripts (primarily Python) to automate routine tasks, support incident investigations and improve team efficiency. Collaborate with ...