MLOps Engineer (LLM/GenAI)
- Location
- Sheffield, England, United Kingdom
/GenAI) In this fantastic role, you’ll engineer production-grade infrastructure for modern AI: hosting LLMs and speech/embedding models, pushing inference performance on real hardware, and building repeatable fine-tuning pipelines that ship domain-adapted models into production. If you like hard performance problems, platform … throughput, and cost (e.g., quantisation, KV-cache optimisation, dynamic/continuous batching) Evaluate and integrate inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang) to maximise performance on target hardware Own inference health/performance monitoring (latency, throughput, TTFT, memory, availability) and troubleshoot bottlenecks/deployment issues Build ...