AI Inference Engineer
- Location
- Greater London, England, United Kingdom
cost per token, partnering with the CUDA/GPU engineers Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents) Translate throughput, latency and uptime commitments into concrete technical specifications and serving capacity plans Act as direct technical owner ...