Staff Software Engineer, Inference
- Location
- Greater London, England, United Kingdom
latency (P95/P99) and platform reliability through metrics‐driven engineering. Proven experience owning system‐wide SLIs/SLOs, capacity planning, autoscaling strategies, and mentoring senior and mid‐level engineers. Preferred Direct open‐source or production contributions to modern inference frameworks (e.g., vLLM, Triton, TensorRT‐LLM, Ray Serve, or TorchServe ...