HPC Infrastructure Site Reliability Engineer
- Location
- Gloucester, England, United Kingdom
experience operating large‐scale distributed systems and recent hands‐on expertise in high‐performance computing (HPC) and AI infrastructure. This is an operations‐first SRE role, working in a 24/7/365 on‐call environment, responsible for ensuring reliability, performance, and continuous improvement of mission‐critical infrastructure. … This role sits within a cross‐functional organisation spanning network engineering, infrastructure SRE, Platform SRE, infrastructure tooling engineers (software) and data centre operations. The ideal candidate has progressed through large‐scale, globally distributed or multi‐site infrastructure environments and has more recently specialised in GPU‐accelerated HPC systems. This ...