Senior ML Infrastructure Engineer Enterprise Operations Oxford, England, United Kingdom
- Location
- Oxford, England, United Kingdom
sensitive research environments. Partner with Research, Data, and Applied teams to forecast capacity and cost for GPU and storage needs, setting quotas and streamlining ML experimentation pipelines. Essential Skills, Qualifications & Experience: Proven experience leading the design, build, and operation of high-performance ML compute clusters at scale A proactive, autonomous … approach to systems design and the proven ability and desire to ideate, co-create and implementoptimalsolutions Exposure to migrating or transforming ML infrastructure from traditional schedulers to modern, containerised systems Expertisewith high-throughput storage systems for ML/HPC workloads Expert-level understanding of GPU architecture, high-speed networking ...