AI Inference Engineer
- Location
- Greater London, England, United Kingdom
above it: how models actually get served, scaled, and delivered against committed performance targets. The Opportunity Fuse is seeing significant demand for data centre capacity across the markets we operate in, primarily for inference. Few companies in the world can pair real power delivery with real compute … orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents). Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans. Act as a direct technical owner of inference performance and reliability. Work closely with the CUDA and GPU engineering teams to ensure custom ...