Staff Software Engineer, Inference
- Location
- Greater London, England, United Kingdom
autoscaling strategies, and mentoring senior and mid‐level engineers. Preferred Direct open‐source or production contributions to modern inference frameworks (e.g., vLLM, Triton, TensorRT‐LLM, Ray Serve, or TorchServe). Deep experience with GPU systems engineering and hardware performance optimisation (e.g., CUDA, NCCL, RDMA, NUMA, or GPU interconnects). Direct ...