Site Reliability Engineer - Service Assurance Systems
- Location
- Greater London, England, United Kingdom
Manage and maintain containerised workloads using Docker, and support the operation of services hosted on AWS cloud infrastructure. Write and maintain operational scripts (primarily Python) to automate routine tasks, support incident investigations and improve team efficiency. Collaborate with software developers to identify recurring operational issues and feed findings back into … containerisation using Docker. Experience with monitoring and observability tooling, including Prometheus and AWS CloudWatch, for metrics collection, alerting and dashboarding. Proficiency in scripting with Python for automation, tooling and operational support tasks. Strong communication skills — able to clearly articulate technical issues and their impact to both technical and non-technical ...