Software Inference Deployment Engineer
- Location
- Oxford, England, United Kingdom
Strong Preference For Experience integrating accelerator hardware (GPUs, FPGAs, ASICs, NPUs, or novel architectures) into customer inference workflows Familiarity with the NVIDIA inference stack - CUDA, TensorRT, Triton Exposure to disaggregated inference architectures, prefill/decode separation, or KV cache management Compensation & Benefits Highly Competitive Salary: We are not saying ...