businesses to thrive and economies to prosper, and, ultimately, helping people fulfil their hopes and realise their ambitions. We are seeking a MLOps Engineer (LLM/GenAI) In this fantastic role, you’ll engineer production-grade infrastructure for modern AI: hosting LLMs and speech/embedding models, pushing inference performance … Optimise inference for latency, throughput, and cost (e.g., quantisation, KV-cache optimisation, dynamic/continuous batching) Evaluate and integrate inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang) to maximise performance on target hardware Own inference health/performance monitoring (latency, throughput, TTFT, memory, availability) and troubleshoot bottlenecks/deployment issues Build ...