Member of Technical Staff (AI Inference Engineer)
- Location
- Greater London, England, United Kingdom
generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway.* **GPU kernels migration to CuTe DSL.** Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera ...