Platform Engineer
- Location
- Greater London, England, United Kingdom
operate AI workload infrastructure, including model gateways, retrieval services, orchestration components, and supporting cloud or Kubernetes resources. Observability, Monitoring & Site Reliability (SRE) Instrument services and implement monitoring, logging, and alerting as code using standard tooling (Prometheus, Grafana, OpenTelemetry). Participate in the on‐call rotation, responding to incidents … least one major cloud platform (AWS or Azure) and Kubernetes/Docker. Familiarity with observability tooling (Grafana, Datadog, Splunk, ELK, OpenTelemetry) and basic SRE practices. Exposure to test automation, policy-as-code, and platform security practices. Familiarity with ITIL best practices (incident, change, and problem management) preferred. Experience with Lean ...