Senior Site Reliability Engineer
- Location
- Reading, England, United Kingdom
strongly preferred) Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management Model serving and inference at production scale (eg KServe, RayServe, Triton, vLLM, or similar) with real latency and cost constraints(preferred RayServe) MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights ...