HPC Infrastructure Site Reliability Engineer
- Location
- Gloucester, England, United Kingdom
This is an operations‐first SRE role, working in a 24/7/365 on‐call environment, responsible for ensuring reliability, performance, and continuous improvement of mission‐critical infrastructure. This role sits within a cross‐functional organisation spanning network engineering, infrastructure SRE, Platform SRE, infrastructure tooling … layers to validate readiness of high‐density GPU infrastructure and support safe, predictable deployment at scale. You will also play a central role in continuous service improvement (CSI)—reducing operational toil, increasing automation, and improving reliability, consistency, and operational efficiency across the platform. This includes strengthening observability ...