HPC Infrastructure Site Reliability Engineer
- Location
- Gloucester, England, United Kingdom
mission‐critical environments where uptime, throughput, and ultra‐low latency are non‐negotiable. Role Overview We are looking for a senior Infrastructure Site Reliability Engineer with deep experience operating large‐scale distributed systems and recent hands‐on expertise in high‐performance computing (HPC) and AI infrastructure. This … infrastructure SRE, Platform SRE, infrastructure tooling engineers (software) and data centre operations. The ideal candidate has progressed through large‐scale, globally distributed or multi‐site infrastructure environments and has more recently specialised in GPU‐accelerated HPC systems. This role provides exposure to the latest high‐density AI compute platforms ...