Senior Site Reliability Engineer
- Location
- Reading, England, United Kingdom
/staff engineers. Who you are (must-haves) 5+ years in infrastructure engineering, DevOps, or SRE, operating large-scale, high-availability production systems using Kubernetes Production Operational experience - a live cluster under real load, not a lab. Fluent with Helm, and Terraform or Cloudformation, on at least one major cloud … strength. Track record working with product, backend/frontend teams to pull through collective initiative. ML & AI platform (strongly preferred) Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management Model serving and inference at production scale (eg KServe, RayServe, Triton, vLLM, or similar) with real latency and cost ...