Senior HPC & Cloud Engineer - AI/ML Compute Infrastructure
Senior HPC & Cloud Engineer - AI/ML Compute Infrastructure
Oxford | Hybrid (3 days in office) | Competitive, DOE
Sentinel is recruiting for several senior/staff-level engineers to design, build and operate a hybrid GPU compute environment, combining on-prem HPC clusters with public cloud infrastructure for large-scale AI research workloads.
Responsibilities:
- Build and operate high-performance GPU training/inference clusters, including scheduling, isolation and automated life cycle management
- Design high-throughput data paths across compute and storage, including parallel filesystems (eg Lustre)
- Benchmark and resolve performance bottlenecks across compute, network and orchestration layers
- Implement observability, resilience and security controls for a compliance-conscious research environment
- Work with research and applied ML teams to forecast GPU/storage capacity and streamline experimentation pipelines
Requirements:
- Experience with HPC/GPU clusters, including a strong understanding of GPU architecture, high-speed networking and distributed training performance
- Cloud platform experience (Azure, GCP, AWS or other)
- Experience with containerisation/Kubernetes
- Working knowledge of IaC and CI/CD (eg Terraform, Argo CD)