scientific level, to the org level. What you'd do Run and evolve our GPU clusters: scheduling, utilisation, debugging, performance. Scale the Kubernetes/Linux/networking/cloud stack end‐to‐end. Establish operational excellence: incident response, postmortem culture, on‐call health. Build agent‐driven automation for cluster lifecycle … optimise the platform they rely on every day. What we're looking for Experience operating infra at scale, preferably for LLM workloads. Depth in Kubernetes internals, cluster provisioning, and orchestration systems. Comfortable across the stack: Kubernetes, Linux, networking, containers, cloud environments. Strong systems thinking: you care about reliability, performance ...