HPC Network Engineer

DescriptionFuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy - fast. We're combining first-principles thinking with cutting-edge technology to build a radically better energy system. We raised $210M from top-tier investors including Multicoin, Balderton, Lakestar, Accel, Creandum, Lowercarbon, Ribbit, Box Group and strategic angels like Nico Rosberg, the Co-Founder of Solana and GPs behind Meta, Revolut, Spotify, Uber and more.The OpportunityYou'll design, deploy, and operate the network fabric for our multi-tenant AI cluster. This covers the full stack: the high-performance compute and storage fabrics carrying RDMA traffic between GPUs, the tenant-facing and management networks, fire walling and tenant isolation, and the out-of-band infrastructure that keeps it all recoverable. Beyond the data centre, you'll own the office network and act as the networking authority for the company, raising the bar for everyone by sharing what you know. You'll own the fabric from architecture through day-2 operations.ResponsibilitiesDesign and operate lossless, RDMA-capable fabrics (e.g. RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS, congestion control, and buffer tuning at scaleBuild and manage leaf-spine data centre fabrics, with routed underlay and overlay design (e.g. BGP, EVPN/VXLAN)Implement and maintain per-tenant network isolation across compute, storage, and management planes.Automate network provisioning, configuration, and validation, treating switch config as code (e.g. Ansible, Python, NetBox as source of truth), deployed through CIBuild telemetry and observability for the fabric: flow-level and buffer-level visibility, dashboards, and alerting that catches congestion and link degradation before tenants do (e.g. Prometheus/Grafana/Datadog, streaming telemetry)Troubleshoot performance issues end to end, from optics and cabling through switch buffers to NIC/DPU configuration and collective-communication behaviour on the hostsOperate the out-of-band management network, console access, and remote recovery pathsSupport tenant onboarding: segmentation and addressing, bandwidth and isolation guarantees, and capacity planning as the cluster scalesWrite clear design documentation capturing decisions, rationale, and rejected alternativesOwn and maintain the office network: wired and wireless infrastructure, firewalling, VPN/remote access, and connectivity between the office and data centre environmentsUpskill colleagues on networking: share knowledge through documentation, run-throughs, and pairing so the wider team can operate and troubleshoot the fabric confidentlyRequirements5+ years as a network engineer operating production data centre networksStrong dynamic routing experience (BGP in particular), plus overlay/encapsulation design and troubleshooting (e.g. EVPN/VXLAN)Hands-on experience with leaf-spine / Clos fabric design and operationExperience with modern data centre network operating systems and comfortable in the Linux networking stack, not just a vendor CLIPractical RDMA fabric experience: lossless Ethernet (e.g. RoCEv2 with PFC/ECN/DCQCN tuning) or InfiniBand, with an understanding of why lossless behaviour matters for GPU workloadsNetwork automation as a working practice, not an aspiration: scripting (e.g. Python), configuration management (e.g. Ansible), config generation from a source of truth, version-controlled changesSolid Linux administration fundamentals: you can debug from the host side as well as the switch sideExperience with network telemetry and monitoring (e.g. Prometheus/Grafana, sFlow/IPFIX, streaming telemetry)Experience running corporate/campus networks: wired and wireless, switching, NAC/802.1X, VPN and remote access (e.g. Cisco Catalyst/Meraki or comparable)Clear communicator who enjoys teaching: able to document, pair, and run sessions that bring less network-savvy colleagues up to speedNice to haveExperience with GPU cluster networking specifically (e.g. NVIDIA Spectrum-X or Quantum InfiniBand, ConnectX/BlueField NICs and DPUs, UFM, SHARP, or equivalent Broadcom/Arista AI fabric platforms)Container networking experience: CNI plugins and BGP integration between clusters and the fabric.Understanding of collective-communication libraries and how fabric behaviour shows up as training/inference performanceMulti-tenant network design: VRF-based isolation, tenant bandwidth guarantees, secure shared infrastructureExperience with enterprise firewall platforms (e.g. FortiGate, Palo Alto), including HA deployment and virtualised/segmented instancesStorage networking experience (e.g. NVMe-oF, lossless storage fabrics, per-tenant storage isolation)Bare-metal provisioning environments (e.g. MAAS, PXE, Redfish) and how network bootstrap fits into node lifecycleOptical layer knowledge at 200/400/800G: transceivers, MPO cabling, link qualificationExperience standing up a data centre network from greenfieldRelevant certifications (e.g. CCNP/CCIE or equivalent), valued as evidence of depth, not a gateBenefitsCompetitive salary and an equity sign-on bonusBiannual bonus schemeFully expensed tech to match your needsBreakfast and dinner allowance for office-based employeesJob SummaryID: 6F6AFE8C31Department:Type: full time

Job Details

Company
Appcast
Location
London, UK
Posted