HPC Infrastructure Site Reliability Engineer
- Location
- Gloucester, England, United Kingdom
generation GPU infrastructure at significant scale. You will bring strong breadth across bare metal, networking, storage, virtualisation, and orchestration, alongside deep HPC experience including NVIDIA GPU ecosystems, RDMA networking (RoCE and InfiniBand), and performance validation and benchmarking. Strong Linux and distributed systems expertise is essential. Alongside operational ownership, this … high‐performance computing infrastructure. As a HPC Infrastructure SRE, you’ll work hands‐on with cutting‐edge GPU and CPU platforms ‐ including the latest NVIDIA architectures ‐ powering dense, large‐scale compute environments used for AI, machine learning, and next‐generation workloads. This is an opportunity to build expertise ...