Site Reliability Engineer - Service Assurance Systems
- Location
- Greater London, England, United Kingdom
Docker, and support the operation of services hosted on AWS cloud infrastructure. Write and maintain operational scripts (primarily Python) to automate routine tasks, support incident investigations and improve team efficiency. Collaborate with software developers to identify recurring operational issues and feed findings back into the development process to improve … application resilience. Maintain clear and up-to-date operational documentation including runbooks, deployment guides and incident post-mortems. Provide written and verbal progress updates on open incidents, deployments and operational improvements to the SAS group and wider stakeholders. Support on-call and out-of-hours incident response ...