AI/HPC Consultant
The Role:
- The AI/HPC Consultant will support the definition, design and validation of advanced AI infrastructure solutions based on NVIDIA reference architectures. The role is heavily focused on high-performance networking and physical connectivity design across GPU-accelerated platforms, with regular involvement in bill-of-materials reviews, design assurance and customer-facing technical guidance.
Responsibilities:
- Develop high-level and low-level designs for NVIDIA-based AI/HPC infrastructure, particularly GB300, B300, DGX, HGX and NVL72-aligned environments.
- Design InfiniBand Back End fabrics, Spectrum-X converged Front End networks and where required, Spectrum-X Ethernet Back End fabrics using NVIDIA reference architectures as the baseline.
- Provide expert technical guidance across RoCEv2, InfiniBand, VXLAN EVPN, leaf-spine architectures, congestion management, rail-optimised design and multi-plane fabrics.
- Create, review and optimise Bills of Material, including Switches, optics, DAC/ACC/AEC cables, fibre assemblies, adapters, DPUs/SuperNICs and associated licensing or support items.
- Support customer discovery workshops, requirements capture sessions, technical option analysis and architecture governance reviews.
- Act as a technical lead for defined work packages, coordinating with consultants, vendors, pre-sales teams and delivery stakeholders to ensure designs are technically correct and commercially practical.
- Maintain awareness of NVIDIA product roadmaps, networking reference architectures, switch platforms, adapter families, optics guidance and deployment best practice.
Skills Required:
- Strong networking background with demonstrable experience in data centre, HPC, AI infrastructure or large-scale low-latency fabric design.
- Excellent knowledge of high-performance AI/HPC networking technologies, including InfiniBand, RoCEv2, Spectrum-X Ethernet, congestion control concepts and lossless Ethernet design principles.
- Strong understanding of VXLAN EVPN, BGP, leaf-spine fabric design, routing design, resiliency models and operational troubleshooting in modern data centre networks.
- Good understanding of physical infrastructure design, including switch placement, row-level design, structured cabling, patching strategy, fibre polarity, cable routing and data centre implementation constraints.
- Practical knowledge of optical and copper connectivity choices, including MMF, SMF, DAC, ACC/AEC, breakout cables, MPO/MTP, LC, QSFP, QSFP112, OSFP and QSFP-DD form factors.
- Strong written communication skills, including the ability to produce clear HLDs, LLDs, design notes, BoM justifications and customer-facing technical explanations.
Technical capabilities:
- NVIDIA AI/HPC platforms
- GB300, B300, DGX, HGX, NVL72, GPU node connectivity, ConnectX, BlueField/SuperNIC concepts and NVIDIA reference architecture interpretation.
- InfiniBand Back End fabrics
- NDR/XDR/HDR concepts, rail-optimised design, spine/leaf sizing, congestion and resiliency considerations, and Quantum-class switching such as QM3400.
- Spectrum-X Ethernet
- Data centre networking
- Leaf-spine design, VXLAN EVPN, BGP, resilient routing, underlay/overlay separation, management, storage and in-band/out-of-band network segmentation.