Staff SRE, AI Infrastructure
- Hiring Organisation
- wayve
- Location
- London, United Kingdom
- Salary
- £ 80 K
environments.Define and operationalise SLOs, SLIs, and error budgets across platform services.Improve capacity planning, scaling strategies, and resource efficiency across large GPU-backed clusters.Partner with ML, platform, and software teams to establish clear production readiness standards.Incident Response & On-CallParticipate in a 24/7 on-call rotation as first-line response … experience.Essential skillsProven experience in an SRE, Production Engineer, or Cloud Reliability role supporting large-scale cloud systems.Experience operating GPU-backed environments or large-scale ML infrastructure.Experience running model training or inference pipelines in production (MLOps).Strong Kubernetes experience, including operating production clusters.Hands-on experience running production workloads ...