24/7 HPC Infra SRE for AI & GPU Compute
- Location
- Gloucester, England, United Kingdom
Radiant is seeking a senior Infrastructure Site Reliability Engineer to own and improve large‐scale GPU‐accelerated HPC infrastructure in a 24/7 production environment. You will work across network, storage, virtualization and orchestration with hands‐on Linux expertise, NVIDIA GPU ecosystems, RoCE/InfiniBand, and performance … benchmarking. This role champions observability, automation and on‐call reliability, shaping next‐gen HPC platforms within a globally distributed team. #J-18808-Ljbffr ...