Multimodal Inference & Serving Engineer
- Location
- Greater London, England, United Kingdom
multimodal agent inference stack in production, spanning from engine layers to serving architecture. You will help design and operate systems for low latency, high throughput, and cost efficiency while collaborating with research and engineering teams. The role focuses on research-driven production ML and scalable infrastructure, with exposure ...