26 to 50 of 52 InfiniBand Jobs in England

Sr. Network Engineer - Data Center and Cloud (Hybrid, London)

Location
Greater London, England, United Kingdom
MSFT, OCI or similar) OpenStack/SONiC/Whitebox, FRR, open source network infrastructure experience AI/GPU cluster networking fabrics (RDMA, RoCE, InfiniBand) UK Government, G-Cloud, or Sovereign cloud experience Benefits of Working at CrowdStrike Market leader in compensation and equity awards Comprehensive physical and mental wellness programs ...

HPC System Administrator — On-Site/Remote Linux & HPC Ops

Location
Farnborough, England, United Kingdom
familiarity within a Managed Services context. The team seeks experience with HPC tooling, scripting (Python, Bash, Ansible), storage systems (GPFS, NFS), and networking technologies (InfiniBand, RDMA, ROCE). #J-18808-Ljbffr ...

Senior Network Engineer – GPU Infrastructure / Quantitative Trading

Hiring Organisation
Quant Capital
Location
London, United Kingdom
Salary
£ 100 K
similar• Good troubleshooting skills across network and host-level issues• Clear communication and strong ownershipUseful experience• GPU cluster, HPC or research infrastructure networking• InfiniBand, RoCE, RDMA or high-performance Ethernet• Network monitoring, telemetry or performance analysis• Low-latency, trading or colocation environmentsThis role suits a network engineer who wants ...

Field CTO

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
technology leader with startup energy in the AI Infrastructure practice Work on the most advanced AI infrastructure deployments in EMEA — GPU superclusters, liquid cooling, InfiniBand fabrics Direct access to NVIDIA, HPE, Dell, Cisco, and leading AI infrastructure vendors as strategic partners Enterprise-scale projects with Fortune 500 and regulated industry ...

Principal - AI & HPC Data Centre Compute

Location
Greater London, England, United Kingdom
solutioning and shaping complex technology engagements Nice to have Exposure to NVIDIA ecosystems (DGX, HGX, SuperPOD) or alternative accelerators (AMD or similar) Familiarity with InfiniBand/RoCE networking, PyTorch, high-density rack design, liquid cooling or GPU-as-a-Service deployments We offer EPAM Employee Stock Purchase Plan (ESPP) Protection ...

Member of Technical Staff (AI Inference Engineer)

Location
Greater London, England, United Kingdom
laid out for you.### **Nice-to-have:*** ML compilers and framework internals: PyTorch internals, torch.compile, custom operators.* Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism.* Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving.* Profiling and debugging tools: Nsight Compute/Systems, CUDA ...

Network SRE - DC – Network WAN

Location
Greater London, England, United Kingdom
Operations: Participatein a shift-based model to ensurecontinuous availability of critical network services. Multi-Vendor Expertise: Operateacrossdiverse environments including Arista, Cisco, Cumulus, Spectrum Ethernet,InfiniBand, Palo Alto, Check Point, Mist, Aruba, A10,Netscaler, andF5. Security & Segmentation: Support networksegmentation, policy enforcement, and VPN solutions (GlobalProtect,AnyConnect). Automation & Observability: Utilizetoolslike Grafana … ServiceNow, ITMP, syslog, Splunk,Salt,Ansible, andPrometheusto enhance monitoring andautomation. Innovation Projects: Collaborate on wireless design and AI clusterdeployments to supportcutting-edgeinitiatives. PreferredSkills Experiencewith InfiniBand and AI cluster deployments . Familiaritywith network asset management systems (e.g.,Nautobot). Wirelessdesign experience with Cisco, Mist, Aruba . #J-18808-Ljbffr ...

GPU Infrastructure Lead - Systems Integrator

Location
Greater London, England, United Kingdom
ultimately build and lead the team responsible for the function. Responsibilities: Design and run cluster validation and certification: performance benchmarks, interconnect testing (NCCL, InfiniBand/RoCE), thermal and power verification, availability monitoring against SLAs. Build automation for cluster deployment, health checks, and continuous testing so certification scales without headcount scaling ...

Founding GPU Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
tooling/orchestrationExperience with performance profiling tools (Nsight Systems/Compute)Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand)Strong grasp of memory optimisation, kernel fusion and parallel algorithm designComfortable working across the stack, from low-level kernels to system-level infrastructureBonus: Triton, cuDNN, cuBLAS ...

Team Lead, Platform Engineering

Location
Greater London, England, United Kingdom
Intel TDX, or Confidential Containers (CoCo). Experience building SaaS or PaaS layers on top of an IaaS platform. Familiarity with RDMA, InfiniBand, or RoCE networking in GPU or HPC clusters. Experience working distributed across time zones with counterparts in other regions. Exposure to serverless or inference serving infrastructure. ...

Principal Network Engineer

Location
Greater London, England, United Kingdom
ongoing operation of the networking services underpinning both our internal management platform and customer-facing cloud infrastructure. This includes high-performance Ethernet fabrics, InfiniBand, RoCE, WAN connectivity, and large-scale data centre networking. In this role, you'll set technical direction across the low-latency, high-bandwidth networks supporting large … Platform Engineering, Systems, Storage, and technology vendors. What you'll be doing AI & High-Performance Network Architecture Define, design, validate, and evolve large-scale InfiniBand, RoCE, and Ethernet fabric architectures at rack, row, and data centre scale. Design networks that integrate closely with bare-metal provisioning and cluster management systems. ...

Senior HPC Engineer

Location
Stevenage, England, United Kingdom
workloads Hardware, OS and scheduler troubleshooting ServiceNow or equivalent ITSM experience GPU computing and CUDA knowledge Ansible or similar automation tools HPC networking, including InfiniBand/high-speed Ethernet The role is hybrid, with 3 days per week onsite in Stevenage. #J-18808-Ljbffr ...

Senior HPC Engineer - Hybrid - Inside IR35

Location
Stevenage, England, United Kingdom
Docker or container technologies. Ansible or similar automation tools. GPU computing, CUDA and GPU-accelerated workloads . OpenMPI, MPICH or other MPI libraries . InfiniBand and high-speed networking. Web server and SSL certificate management. RHCSA, RHCE or equivalent Red Hat certification. #J-18808-Ljbffr ...

Founding GPU Engineer

Location
Greater London, England, United Kingdom
/orchestration. Experience with performance profiling tools (Nsight Systems/Compute). Familiarity with multi-GPU/multi-node scaling (NCCL, MPI, RDMA/InfiniBand). Strong grasp of memory optimisation, kernel fusion, and parallel algorithm design. Comfortable working across the stack from low-level kernels to system-level infrastructure. ...

Senior Solutions Architect – Large Scale AI Inference

Location
Layer-de-la-Haye, England, United Kingdom
crowd Hands-on experience with NVIDIA Dynamo, NIXL, Grove, or emerging disaggregated inference tooling. Understanding of GPU memory hierarchies and high-speed interconnects (NVLink, InfiniBand, RDMA, UCX). You have contributed to advanced AI lab or large scale AI infrastructure providers performing inference on thousands of GPUs. Published work ...

Linux/RHEL Engineer - HPC

Location
Stevenage, England, United Kingdom
environments MPI, compilers and scientific libraries Hardware, OS, scheduler and application troubleshooting ServiceNow or equivalent ITSM experience GPU/CUDA, Docker, Ansible and InfiniBand are desirable The role is based in Stevenage with a minimum of 3 days onsite each week. #J-18808-Ljbffr ...

HPC Operations Lead

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 70 K
supporting both internal workloads and external collaboration. The environment includes large scale HPC clusters, Linux based systems, workload schedulers such as Slurm, networking with Infiniband and parallel file systems such as GPFS. Experience with high performance storage at petabyte scale is particularly relevant, alongside a broader understanding of automation, data ...

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
cuBLAS. Intuition about the latency and throughput characteristics of CUDA graph launch, tensor core arithmetic, warp-level synchronization and asynchronous memory loads. Background in Infiniband, RoCE, GPUDirect, PXN, rail optimisation and NVLink, and how to use these networking technologies to link up GPU clusters. An understanding of the collective algorithms ...

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
cuBLAS Intuition about the latency and throughput characteristics of CUDA graph launch, tensor core arithmetic, warp-level synchronization and asynchronous memory loads Background in Infiniband, RoCE, GPUDirect, PXN, rail optimisation and NVLink, and how to use these networking technologies to link up GPU clusters An understanding of the collective algorithms ...

HPC Network Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
networking authority for the company. You own the fabric from architecture through day-2 operations.ResponsibilitiesDesign and operate lossless, RDMA-capable fabrics (e.g. RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS, congestion control and buffer tuning at scaleBuild and manage leaf-spine data centre fabrics, with routed underlay … Linux networking stack, not just a vendor CLIPractical RDMA fabric experience: lossless Ethernet (e.g. RoCEv2 with PFC/ECN/DCQCN tuning) or InfiniBand, understanding why lossless behaviour matters for GPU workloadsNetwork automation as a working practice: scripting (e.g. Python), configuration management (e.g. Ansible), config generation from a source ...

HPC Network Engineer

Location
Greater London, England, United Kingdom
authority for the company. You own the fabric from architecture through day-2 operations. Responsibilities Design and operate lossless, RDMA-capable fabrics (e.g. RoCEv2, InfiniBand) for GPU compute and storage traffic, including QoS, congestion control and buffer tuning at scale Build and manage leaf-spine data centre fabrics, with routed … Linux networking stack, not just a vendor CLI Practical RDMA fabric experience: lossless Ethernet (e.g. RoCEv2 with PFC/ECN/DCQCN tuning) or InfiniBand, understanding why lossless behaviour matters for GPU workloads Network automation as a working practice: scripting (e.g. Python), configuration management (e.g. Ansible), config generation from ...

C++ Software Engineer

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 80 K
hierarchies of modern processors.Experience with fast-packet processing in user space and familiarity with kernel-bypass techniques such as Solarflare OpenOnload, TPCDirect, ef_vi, InfiniBand verbs, DPDK or similarExperience optimising performance in managed runtime languages such as C# is desirableWhy join us Highly competitive compensation plus annual discretionary bonusLunch provided ...

C++ Software Engineer

Location
Greater London, England, United Kingdom
modern processors. Experience with fast‐packet processing in user space and familiarity with kernel‐bypass techniques such as Solarflare OpenOnload, TPCDirect, ef_vi, InfiniBand verbs, DPDK or similar. Experience optimising performance in managed runtime languages such as C# is desirable. Financial market knowledge is not required. Benefits Highly competitive compensation ...

Senior Solutions Engineer

Hiring Organisation
LJB & Co
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £1,000 to £1,200 per day
NVIDIA GPU clusters for AI and HPC workloads. Develop HLDs, LLDs, rack designs and detailed BOMs. Design NVLink/NVSwitch architectures. Design high-speed InfiniBand and RoCE/RoCEv2 networking. Develop infrastructure solutions using NVIDIA Blackwell, B300, GB300 and GB200 platforms. Design both bare-metal and Kubernetes-based GPU environments. … infrastructure. Experience with HGX/DGX/Blackwell/B300/GB300/GB200. Strong knowledge of NVLink/NVSwitch. Expert knowledge of InfiniBand and/or RoCE/RoCEv2. Experience with GPU clusters, HPC or AI infrastructure. Knowledge of Kubernetes and/or Slurm. Experience producing HLDs, LLDs, architecture ...

Technical Solutions Architect – Investors

Location
Greater London, England, United Kingdom
centre, power and ongoing operational expenditure Evaluate and recommend data centre/colocation partners capable of supporting high-density GPU racks Define networking topology (InfiniBand/Ethernet fabric) to optimise for distributed training and inference workloads Liaise with hardware vendors, OEMs and data centre providers to validate specifications and negotiate … NVIDIA H100/H200/B200/B300 class systems), server/rack design, and data centre infrastructure Methodology/Protocols: Strong grasp of InfiniBand/RDMA networking, NVLink/NVSwitch topologies, and cluster interconnect design for distributed compute workloads Soft Skills: Confident operating at boardroom level, able to translate ...