1 of 1 vLLM Jobs in Yorkshire

MLOps Engineer (LLM/GenAI)

Hiring Organisation
17918
Location
Sheffield, Yorkshire, United Kingdom
heterogeneous hardware Optimise inference for latency, throughput, and cost (e.g., quantisation, KV-cache optimisation, dynamic/continuous batching) Evaluate and integrate inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang) to maximise performance on target hardware Own inference health/performance monitoring (latency, throughput, TTFT, memory, availability) and troubleshoot bottlenecks/deployment … architecture and HPC fundamentals Deep inference optimisation expertise: KV-cache, batching, quantisation (INT4/FP8/GPTQ/AWQ), operator optimisation, framework integration (vLLM/TensorRT-LLM/SGLang) Production hosting experience with Docker/Kubernetes and cloud platforms (AWS/GCP/Azure) End-to-end fine-tuning expertise ...