AI Inference Engineer
- Hiring Organisation
- Fuse Energy Supply
- Location
- London, United Kingdom
- Salary
- £ 80 K
improve throughput and cost per token, partnering with the CUDA/GPU engineersMake the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)Translate throughput, latency and uptime commitments into concrete technical specifications and serving capacity plansAct as direct technical ...