Performance Engineer, Containers/Serverless
- Hiring Organisation
- Jobleads-UK
- Location
- Greater London, England, United Kingdom
load is a ninety-second outage from the user's perspective - but the same fundamentals shape steady-state inference throughput, training step time, and checkpoint behavior. We want someone who understands how all of those parts interact. What you'll do Profile and optimize the end-to-end path … runtime, and GPU as one system rather than separate silos. Nice to have Experience of systems level programming (Go, Rust or Python) Experience of checkpoint/restore of CPU & GPU workloads Experience with container image acceleration (eStargz, SOCI, Nydus) or building a caching layer (FUSE-based, in-kernel, or sidecar ...