Senior Software Engineer, Inference Platform
- Location
- Greater London, England, United Kingdom
looking for exceptional people to help us scale. Who You Are You’re a seasoned engineer who has built and scaled high-performance inference systems for AI/ML workloads. You understand the complexities of serving models at scale latency optimization, resource orchestration, autoscaling dynamics, and production reliability. … designed distributed systems that handle thousands of requests per second while maintaining sub‐second response times and cost efficiency. Experience with Golang is strongly preferred, and exposure to inference engines (vLLM, TGI, TensorRT), containerization, and distributed systems is an added bonus. You take ownership of platform ...