HPC Architect/Lead

HPC Architect/Lead, Crawley/Hybrid, £120k - £140k per annum

Owns the technical strategy and design of the client's hybrid HPC platform - responsible for what we build and why. The Architect defines the platform roadmap, leads technology evaluation, designs the hybrid cloud architecture, and ensures the platform meets the performance, security, and cost requirements of seismic processing and AI workloads. This role operates at the intersection of deep technical expertise and enterprise strategy, engaging with geoscience stakeholders, IT leadership, vendors, and finance to align platform capability with business demand.

Key Responsibilities:

1. Design the hybrid HPC platform architecture: on-premises NVIDIA H200 GPU compute integrated with scalable cloud infrastructure (AWS/GCP, to be selected).

2. Define and own the multi-year HPC platform roadmap, aligned to corporate strategy and geoscience business requirements.

3. Own HPC capability management, including current skills assessment and future capability planning.

4. Forecast medium- and long-term resource demand for the discipline in collaboration with other functions as needed.

5. Ensure appropriate deployment of resources across projects to balance delivery, development, and well-being.

6. Identify skills gaps and develop actions to close them through hiring, development, or external support

7. Sponsor and govern performance, development, and well-being outcomes across the HPC discipline

8. Lead and own annual HR people processes for the discipline, including but not limited to: Succession planning, Talent and workforce planning, Annual salary and reward processes, Approval of new roles and hiring decisions, Oversight of HPC-related expenses and people costs etc.

9. Provide clear, consistent line management and people leadership for team members.

10. Lead performance management, goal setting, feedback, and development planning.

11. Coach and develop as needed, building leadership capability within the team.

12. Foster a culture of psychological safety, accountability, learning, and inclusion.

13. Manage succession planning and retention of critical skills.

14. Role model and uphold the leadership behaviours outlined in the Behavioural Framework.

15. Sponsor and support technical learning, mentoring, and knowledge sharing across the team.

16. Encourage innovation, continuous improvement, and curiosity within the team.

17. Create opportunities for cross-disciplinary collaboration and learning.

18. Lead technology evaluation and vendor selection for storage (VAST, WEKA, DDN), cloud services, networking, and GPU compute evolution.

19. Design hybrid workload orchestration: policies determining where jobs execute based on cost, data locality, capacity, and performance requirements.

20. Architect cloud networking (VPN, dedicated InterconnecT, peering) and hybrid data movement pipelines between on-prem and cloud.

21. Define the AI/ML infrastructure strategy: platform capabilities required to support distributed training, inference, and MLOps at scale.

22. Establish FinOps practice for HPC: cloud cost visibility, chargeback/showback models, reserved capacity planning, and cost governance.

23. Design security and compliance architecture for the hybrid platform aligned to ISO 27001:2022, including multi tenancy, data sovereignty, access control, and firmware supply chain integrity.

24. Lead GPU-aware scheduling strategy: partition design, resource allocation policies, and capacity planning for mixed seismic and AI workloads.

25. Engage geoscience, seismic operations, IT, and finance stakeholders to translate business demand into platform capabilities and investment cases.

26. Manage vendor relationships with NVIDIA, storage vendors, and cloud providers at a technical and commercial level.

27. Own architecture governance: design review, technology standards, technical debt management, and exception handling.

28. Guide and technically direct the Principal HPC Engineer, ensuring implementation aligns with architectural intent.

29. Contribute to budget planning, procurement strategy, and investment cases for HPC capital and operational expenditure. 30. Identify and manage technical risks to platform delivery, performance, and continuity; develop mitigation and contingency plans.

Skills and experience:

1. Education to degree level in Computer Science, Engineering, Physics, or related discipline - or equivalent depth of experience and demonstrable architectural judgement.

2. Typically 10+ years of experience in HPC, cloud infrastructure, or enterprise platform architecture, with at least 3 years in a design or architecture role.

3. Demonstrated experience in people leadership, particularly in technical teams: objectives, performance reviews, recruitment, and development.

4. Demonstrated experience designing hybrid HPC architectures spanning on-prem GPU compute and cloud services.

5. Deep understanding of at least one major cloud platform's HPC and GPU service portfolio (AWS or GCP preferred).

6. Experience with NVIDIA GPU infrastructure at scale, including scheduling strategy, InterconnecT design, and capacity planning.

7. Track record of leading technology evaluations and vendor selection processes for storage, compute, or cloud services.

8. Experience with FinOps or cloud cost governance in a production environment.

9. Working knowledge of ISO 27001 or equivalent security framework in the context of infrastructure architecture.

10. Proven ability to engage non-technical stakeholders (business leaders, finance) and translate technical strategy into business terms.

11. Excellent interpersonal and communication skills with ability to work with diverse cultures across a global organisation.

12. Good spoken and written English skills.

13. Experience in the oil and gas, geoscience, or scientific computing sector is a strong advantage but not essential.

Job Details

Company
Norton Blake
Location
Crawley, Sussex, United Kingdom RH100
Hybrid / Remote Options
Employment Type
Permanent
Salary
GBP 120,000 - 140,000 Annual
Posted