Senior Cloud Engineer, AI Platform SRE
- Location
- Manchester, England, United Kingdom
prompt and configuration changes, with automated eval and regression checks before promotion. Reliability engineering: Define and track SLOs and SLAs for platform services, run incident response and root cause analysis, and write post-mortems people will read. On-call and runbooks: Take part in the programme … GitLab CI, Argo CD, Jenkins or similar. Observability: Hands‐on with Datadog, Prometheus, Grafana or OpenTelemetry, and opinionated about what's worth alerting on. Incident management: Calm, methodical instincts under pressure, and a habit of fixing the class of problem rather than the instance. AI and ML operations: Experience ...