Sr. Manager, Site Reliability
- Location
- Manchester, England, United Kingdom
first cohort of incident commanders across Engineering and Support. Select and stand up the primary observability platform, preferring extension of existing Omnicell contracts (DataDog, IBM/Instana, Prometheus/Grafana, OpenTelemetry, or other tooling already in use) over net‐new procurement. Define the instrumentation standards all new services must meet. … model. Establish the on‐call rotation model, including fair distribution, compensation approach, paging discipline, and the handoff protocol with our existing managed services partners (IBM, HCL) who provide L1/L2 coverage. Develop and track operational KPIs — MTTR, SLO attainment, change‐failure rate, recurrence, cost per workload — and present reliability ...