2 of 2 vLLM Jobs in the South East

Senior Site Reliability Engineer

Location
Reading, England, United Kingdom
strongly preferred) Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management Model serving and inference at production scale (eg KServe, RayServe, Triton, vLLM, or similar) with real latency and cost constraints(preferred RayServe) MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights ...

Software Inference Deployment Engineer

Location
Oxford, England, United Kingdom
PyTorch in particular) Practical experience with model deployment workflows - loading, format conversion, quantisation, or framework integration Comfortable working with inference serving stacks (for example vLLM, TensorRT‐LLM, or similar) Familiarity with Linux, containerisation (Docker), and cluster environments Comfortable in a customer‐facing role, able to communicate clearly with ...