4 of 4 Observability Jobs in Sutton

SRE & Reliability Manager — AI-Ops, Observability

Location
Carshalton, England, United Kingdom
Manager to lead the reliability function for production services. You will drive reliability improvements, automation and AI-Ops, and manage a team focused on observability, incident response and continuous improvement. You will translate priorities into operational plans, line-manage team leaders, and ensure RCAs, post-mortems and remediation actions ...

SRE & Reliability Lead — AI-Ops & Observability

Location
Carshalton, England, United Kingdom
services used by internal and external customers. You will drive reliability improvements, advance automation and AI-Ops capabilities, and lead a team focused on observability, incident response, operational excellence, and continuous improvement. Responsibilities include translating priorities into clear plans, line managing team leaders, ensuring RCAs and post-mortems are completed ...

Operations and SRE Manager

Location
Carshalton, England, United Kingdom
internal and external customers. You will be responsible for driving reliability improvements, advancing automation and AI-Ops capabilities, and leading a team focused on observability, incident response, operational excellence, and continuous improvement. Responsibilities Lead the implementation of the team’s strategic direction, translating priorities into clear operational plans, backlogs … improvement actions are owned, tracked and completed. Strengthen operational process adherence, ensuring responsibilities are clear and delegation is effective. Drive SRE practices across observability, automation, disaster recovery, design for reliability, on-call readiness and production support. Protect service levels by ensuring engineering effort is balanced across InfoSec commitments, operational tickets ...

Operations and SRE Manager

Location
Carshalton, England, United Kingdom
internal and external customers. You will be responsible for driving reliability improvements, advancing automation and AI-Ops capabilities, and leading a team focused on observability, incident response, operational excellence, and continuous improvement., Lead the implementation of the team's strategic direction, translating priorities into clear operational plans, backlogs and deliverables. … improvement actions are owned, tracked and completed. Strengthen operational process adherence, ensuring responsibilities are clear and delegation is effective. Drive SRE practices across observability, automation, disaster recovery, design for reliability, on-call readiness and production support. Protect service levels by ensuring engineering effort is balanced across InfoSec commitments, operational tickets ...