26 to 42 of 42 InfiniBand Jobs in London

Network SRE - DC – Network WAN

Location
Greater London, England, United Kingdom
Operations: Participatein a shift-based model to ensurecontinuous availability of critical network services. Multi-Vendor Expertise: Operateacrossdiverse environments including Arista, Cisco, Cumulus, Spectrum Ethernet,InfiniBand, Palo Alto, Check Point, Mist, Aruba, A10,Netscaler, andF5. Security & Segmentation: Support networksegmentation, policy enforcement, and VPN solutions (GlobalProtect,AnyConnect). Automation & Observability: Utilizetoolslike Grafana … ServiceNow, ITMP, syslog, Splunk,Salt,Ansible, andPrometheusto enhance monitoring andautomation. Innovation Projects: Collaborate on wireless design and AI clusterdeployments to supportcutting-edgeinitiatives. PreferredSkills Experiencewith InfiniBand and AI cluster deployments . Familiaritywith network asset management systems (e.g.,Nautobot). Wirelessdesign experience with Cisco, Mist, Aruba . #J-18808-Ljbffr ...

GPU Infrastructure Lead - Systems Integrator

Location
Greater London, England, United Kingdom
ultimately build and lead the team responsible for the function. Responsibilities: Design and run cluster validation and certification: performance benchmarks, interconnect testing (NCCL, InfiniBand/RoCE), thermal and power verification, availability monitoring against SLAs. Build automation for cluster deployment, health checks, and continuous testing so certification scales without headcount scaling ...

Founding GPU Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
tooling/orchestrationExperience with performance profiling tools (Nsight Systems/Compute)Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand)Strong grasp of memory optimisation, kernel fusion and parallel algorithm designComfortable working across the stack, from low-level kernels to system-level infrastructureBonus: Triton, cuDNN, cuBLAS ...

Team Lead, Platform Engineering

Location
Greater London, England, United Kingdom
Intel TDX, or Confidential Containers (CoCo). Experience building SaaS or PaaS layers on top of an IaaS platform. Familiarity with RDMA, InfiniBand, or RoCE networking in GPU or HPC clusters. Experience working distributed across time zones with counterparts in other regions. Exposure to serverless or inference serving infrastructure. ...

Principal Network Engineer

Location
Greater London, England, United Kingdom
ongoing operation of the networking services underpinning both our internal management platform and customer-facing cloud infrastructure. This includes high-performance Ethernet fabrics, InfiniBand, RoCE, WAN connectivity, and large-scale data centre networking. In this role, you'll set technical direction across the low-latency, high-bandwidth networks supporting large … Platform Engineering, Systems, Storage, and technology vendors. What you'll be doing AI & High-Performance Network Architecture Define, design, validate, and evolve large-scale InfiniBand, RoCE, and Ethernet fabric architectures at rack, row, and data centre scale. Design networks that integrate closely with bare-metal provisioning and cluster management systems. ...

Founding GPU Engineer

Location
Greater London, England, United Kingdom
/orchestration. Experience with performance profiling tools (Nsight Systems/Compute). Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand). Strong grasp of memory optimisation, kernel fusion, and parallel algorithm design. Comfortable working across the stack from low-level kernels to system-level infrastructure. ...

HPC Operations Lead

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 70 K
supporting both internal workloads and external collaboration. The environment includes large scale HPC clusters, Linux based systems, workload schedulers such as Slurm, networking with Infiniband and parallel file systems such as GPFS. Experience with high performance storage at petabyte scale is particularly relevant, alongside a broader understanding of automation, data ...

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
cuBLAS. Intuition about the latency and throughput characteristics of CUDA graph launch, tensor core arithmetic, warp-level synchronization and asynchronous memory loads. Background in Infiniband, RoCE, GPUDirect, PXN, rail optimisation and NVLink, and how to use these networking technologies to link up GPU clusters. An understanding of the collective algorithms ...

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
cuBLAS Intuition about the latency and throughput characteristics of CUDA graph launch, tensor core arithmetic, warp-level synchronization and asynchronous memory loads Background in Infiniband, RoCE, GPUDirect, PXN, rail optimisation and NVLink, and how to use these networking technologies to link up GPU clusters An understanding of the collective algorithms ...

HPC Network Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
networking authority for the company. You own the fabric from architecture through day-2 operations.ResponsibilitiesDesign and operate lossless, RDMA-capable fabrics (e.g. RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS, congestion control and buffer tuning at scaleBuild and manage leaf-spine data centre fabrics, with routed underlay … Linux networking stack, not just a vendor CLIPractical RDMA fabric experience: lossless Ethernet (e.g. RoCEv2 with PFC/ECN/DCQCN tuning) or InfiniBand, understanding why lossless behaviour matters for GPU workloadsNetwork automation as a working practice: scripting (e.g. Python), configuration management (e.g. Ansible), config generation from a source ...

HPC Network Engineer

Location
Greater London, England, United Kingdom
authority for the company. You own the fabric from architecture through day-2 operations. Responsibilities Design and operate lossless, RDMA-capable fabrics (e.g. RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS, congestion control and buffer tuning at scale Build and manage leaf-spine data centre fabrics, with routed … Linux networking stack, not just a vendor CLI Practical RDMA fabric experience: lossless Ethernet (e.g. RoCEv2 with PFC/ECN/DCQCN tuning) or InfiniBand, understanding why lossless behaviour matters for GPU workloads Network automation as a working practice: scripting (e.g. Python), configuration management (e.g. Ansible), config generation from ...

C++ Software Engineer

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 80 K
hierarchies of modern processors.Experience with fast-packet processing in user space and familiarity with kernel-bypass techniques such as Solarflare OpenOnload, TPCDirect, ef_vi, InfiniBand verbs, DPDK or similarExperience optimising performance in managed runtime languages such as C# is desirableWhy join us Highly competitive compensation plus annual discretionary bonusLunch provided ...

C++ Software Engineer

Location
Greater London, England, United Kingdom
modern processors. Experience with fast‐packet processing in user space and familiarity with kernel‐bypass techniques such as Solarflare OpenOnload, TPCDirect, ef_vi, InfiniBand verbs, DPDK or similar. Experience optimising performance in managed runtime languages such as C# is desirable. Financial market knowledge is not required. Benefits Highly competitive compensation ...

Senior Solutions Engineer

Hiring Organisation
LJB & Co
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £1,000 to £1,200 per day
NVIDIA GPU clusters for AI and HPC workloads. Develop HLDs, LLDs, rack designs and detailed BOMs. Design NVLink/NVSwitch architectures. Design high-speed InfiniBand and RoCE/RoCEv2 networking. Develop infrastructure solutions using NVIDIA Blackwell, B300, GB300 and GB200 platforms. Design both bare-metal and Kubernetes-based GPU environments. … infrastructure. Experience with HGX/DGX/Blackwell/B300/GB300/GB200. Strong knowledge of NVLink/NVSwitch. Expert knowledge of InfiniBand and/or RoCE/RoCEv2. Experience with GPU clusters, HPC or AI infrastructure. Knowledge of Kubernetes and/or Slurm. Experience producing HLDs, LLDs, architecture ...

Technical Solutions Architect – Investors

Location
Greater London, England, United Kingdom
centre, power and ongoing operational expenditure Evaluate and recommend data centre/colocation partners capable of supporting high-density GPU racks Define networking topology (InfiniBand/Ethernet fabric) to optimise for distributed training and inference workloads Liaise with hardware vendors, OEMs and data centre providers to validate specifications and negotiate … NVIDIA H100/H200/B200/B300 class systems), server/rack design, and data centre infrastructure Methodology/Protocols: Strong grasp of InfiniBand/RDMA networking, NVLink/NVSwitch topologies, and cluster interconnect design for distributed compute workloads Soft Skills: Confident operating at boardroom level, able to translate ...

Staff Software Engineer, AI Reliability Engineering

Hiring Organisation
Humanloop
Location
London, United Kingdom
Salary
> £ 150 K
About AnthropicAnthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly ...

Senior Power Integrity Engineer

Location
Greater London, England, United Kingdom
About DriveNets DriveNets is a leader in high-scale networking software for AI infrastructure and service providers. The company pioneered a disaggregated networking architecture that transforms the economics of large-scale networks while maximizing performance ...