1 to 25 of 42 InfiniBand Jobs in London

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
infrastructure programmes.AI Hardware & Data Center InfrastructureGPU/accelerator architectures: NVIDIA/AMD, including multi-node scale-out design.Accelerator interconnects: NVLink, NVSwitchHigh-performance networking: InfiniBand and RoCEv2 fabric design, 400G/800G Ethernet, rail-optimized topologies for AI clusters.Data center facilities: power density, liquid cooling, and rack-level design considerations specific ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
programmes. AI Hardware & Data Center InfrastructureGPU/accelerator architectures: NVIDIA/AMD, including multi-node scale-out design. Accelerator interconnects: NVLink, NVSwitchHigh-performance networking: InfiniBand and RoCEv2 fabric design, 400G/800G Ethernet, rail-optimized topologies for AI clusters. Data center facilities: power density, liquid cooling, and rack-level design ...

Solutions Engineer - Cloud & AI Infrastructure

Hiring Organisation
CISCO Systems
Location
London, United Kingdom
Salary
£ 80 K
e.g., Cilium/Isovalent), GitOps workflows (ArgoCD/Flux), Helm, and service mesh.Familiarity with AI-ready infrastructure: GPU compute, high-performance fabrics (RoCE/InfiniBand), AI/ML reference architectures (e.g., Cisco + NVIDIA), and storage for AI workloads.Relevant industry certifications (e.g., CCNP/CCIE Data Center, VMware ...

Principal Cloud Architect – HPC/GPU & AI Platform Solutions

Hiring Organisation
Oracle Corporation
Location
London, United Kingdom
Salary
£ 80 K
Models (LLMs) Agentic AI AI Platform Architecture Inference Serving Programming & AutomationPython Bash PowerShell Automation Frameworks Infrastructure Automation HPC TechnologiesSlurm PBS Bright Cluster Manager RDMA InfiniBand MPI Distributed File Systems Customer & ConsultingSolution Architecture Technical Consulting Pre-Sales Executive Presentations Customer Workshops Technical Enablement AI Transformation Strategy Cloud Adoption Soft SkillsExcellent communication ...

Enterprise Architect - Network Infrastructure

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
DevNet-style software engineering.Data center compute and virtualization exposure (VMware, OpenStack, OpenShift), given how tightly fabric design is now coupled to compute platforms.AI networking (InfiniBand/RoCE fabrics for GPU clusters), an increasingly relevant adjacent specialism.Leadership & Delivery ExpectationsLeads technical delivery independently and is the de facto final word on routing ...

Platform Architect - Nvidia AI/GB300

Hiring Organisation
Oscar Associates (UK) Limited
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£700 - £765 per day
NVIDIA RTX 6000 series GPU servers. Strong understanding of GPU workload scheduling, partitioning and sharing, including MIG, vGPU and time-slicing Strong understanding of InfiniBand, RoCE, Spectrum-X, GPUDirect RDMA/Storage and high-performance AI fabrics. Experience with Terraform, Ansible, Python/shell, Git and CI/CD Highly ...

Platform Architect - Nvidia AI/GB300

Hiring Organisation
Oscar Technology
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£700.00 - £765.00 per day
NVIDIA RTX 6000 series GPU servers. Strong understanding of GPU workload scheduling, partitioning and sharing, including MIG, vGPU and time-slicing Strong understanding of InfiniBand, RoCE, Spectrum-X, GPUDirect RDMA/Storage and high-performance AI fabrics. Experience with Terraform, Ansible, Python/shell, Git and CI/CD Highly ...

Operations Engineering Manager (m/f/d)

Location
Greater London, England, United Kingdom
stakeholder management skills, able to work closely with Platform, Network, and leadership. Nice to Have Experience in HPC or GPU‐accelerated environments (NVIDIA GPUs, InfiniBand/RDMA, parallel file systems). Scripting skills in Python and/or Bash for automation and tooling. Understanding of performance tuning for HPC/ ...

Senior Product Manager - Storage & Networking

Location
Greater London, England, United Kingdom
storage, WEKA, VAST, DDN, or similar. Good understanding of modern data centre and high-performance networking technologies, with familiarity with technologies such as Ethernet, InfiniBand, RDMA, BGP, EVPN/VXLAN, and modern switching platforms. Familiarity with how storage and networking integrate with Kubernetes, bare-metal infrastructure, and large-scale compute ...

Network Engineer Apprentice

Hiring Organisation
QA
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£26,000 per annum
/VXLAN, leaf-spine architectures, network overlays, segmentation and software-defined networking concepts Awareness of AI and high-performance computing networking concepts, including NVIDIA InfiniBand, RDMA, low-latency fabrics and GPU cluster connectivity Interest in network automation, telemetry, observability and infrastructure-as-code practices Entry requirements: an A-Level ...

Principal Machine Learning Infrastructure Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
optimized collective communication, and know when to use FSDP vs. DDP vs. pipeline parallelism Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python ...

Staff Software Engineer, Node Infra

Hiring Organisation
Humanloop
Location
London, United Kingdom
Salary
> £ 150 K
node provisioning pipelinesLow-level systems experience: kernel, virtualization, device drivers, firmware, or hardware health/diagnostics daemonsFamiliarity with high-performance networking (EFA, RDMA, InfiniBand) for distributed ML workloads.Demonstrated ownership of production reliability for high-throughput, latency-sensitive systemsContributions to relevant open-source projects (Kubernetes, Linux kernel, container runtimes, etc.)Skill ...

Senior Infrastructure Engineer, Research Singapore

Location
Greater London, England, United Kingdom
optimized collective communication, and know when to use FSDP vs. DDP vs. pipeline parallelism Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python ...

Product Manager - Data

Location
Greater London, England, United Kingdom
connected device technology and ecosystems Familiarity with networking technologies - ethernet, IPv4 and IPv6, routing, firewalling, overlays such as OVN/OVS, VPNs, SR-IOV, infiniband Familiarity with telco networking - RAN, Core, CPE Experience in leading distributed teams across different time zones Demonstrated ability to foster collaboration and innovation in team ...

Account Solution Architect

Location
Greater London, England, United Kingdom
Familiarity with NVIDIA GPU architectures (H100, A100, H200) and the software stack around them: CUDA, NCCL, cuDNN. Working knowledge of high-performance networking concepts: InfiniBand, RDMA, RoCE, TCP/IP. Background working directly with AI labs, research institutions, or enterprise ML teams. Exposure to physical AI use cases ...

Senior Network Engineer – Data Center, Cloud

Location
Greater London, England, United Kingdom
MSFT, OCI or similar) OpenStack/SONiC/Whitebox, FRR, open source network infrastructure experience AI/GPU cluster networking fabrics (RDMA, RoCE, InfiniBand) UK Government, G-Cloud, or Sovereign cloud experience Core Competencies Demonstrates expertise in designing and implementing complex multi-cloud and hybrid data center networks, utilizing automation ...

Solutions Engineer

Location
Greater London, England, United Kingdom
partner-facing environment Experience supporting AI start-ups, research organisations, model builders, or enterprise AI teams Understanding of high-performance networking concepts such as InfiniBand, RDMA, or RoCE Experience working with APIs, developer platforms, or cloud-native application environments Experience with open-weight model training and inference What Success Looks ...

Staff HPC Systems Software Engineer

Location
Greater London, England, United Kingdom
Strong understanding of Slurm internals, scheduler behaviour, cluster lifecycle concernsand operational trade offs. Strong practical understanding of GPU-backed infrastructure and HPC networking,including InfiniBand, RoCE, RDMA and performance sensitive workload characteristics. Experience integrating HPC systems with cloud-native platforms, APIs, or service deliverymodels. Experience creating engineering leverage through standards ...

Sr. Network Engineer - Data Center and Cloud (Hybrid, London)

Location
Greater London, England, United Kingdom
MSFT, OCI or similar) OpenStack/SONiC/Whitebox, FRR, open source network infrastructure experience AI/GPU cluster networking fabrics (RDMA, RoCE, InfiniBand) UK Government, G-Cloud, or Sovereign cloud experience Benefits Of Working At CrowdStrike: Market leader in compensation and equity awards Comprehensive physical and mental wellness programs ...

Sr. Network Engineer - Data Center and Cloud (Hybrid, London)

Hiring Organisation
CrowdStrike
Location
London, United Kingdom
Salary
£ 80 K
MSFT, OCI or similar) OpenStack/SONiC/Whitebox, FRR, open source network infrastructure experienceAI/GPU cluster networking fabrics (RDMA, RoCE, InfiniBand)UK Government, G-Cloud, or Sovereign cloud experience#LI-GO1Benefits of Working at CrowdStrike: Market leader in compensation and equity awardsComprehensive physical and mental wellness programsCompetitive vacation ...

Sr. Network Engineer - Data Center and Cloud (Hybrid, London)

Location
Greater London, England, United Kingdom
MSFT, OCI or similar) OpenStack/SONiC/Whitebox, FRR, open source network infrastructure experience AI/GPU cluster networking fabrics (RDMA, RoCE, InfiniBand) UK Government, G-Cloud, or Sovereign cloud experience Benefits of Working at CrowdStrike Market leader in compensation and equity awards Comprehensive physical and mental wellness programs ...

Senior Network Engineer – GPU Infrastructure / Quantitative Trading

Hiring Organisation
Quant Capital
Location
London, United Kingdom
Salary
£ 100 K
similar• Good troubleshooting skills across network and host-level issues• Clear communication and strong ownershipUseful experience• GPU cluster, HPC or research infrastructure networking• InfiniBand, RoCE, RDMA or high-performance Ethernet• Network monitoring, telemetry or performance analysis• Low-latency, trading or colocation environmentsThis role suits a network engineer who wants ...

Field CTO

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
technology leader with startup energy in the AI Infrastructure practice Work on the most advanced AI infrastructure deployments in EMEA — GPU superclusters, liquid cooling, InfiniBand fabrics Direct access to NVIDIA, HPE, Dell, Cisco, and leading AI infrastructure vendors as strategic partners Enterprise-scale projects with Fortune 500 and regulated industry ...

Principal - AI & HPC Data Centre Compute

Location
Greater London, England, United Kingdom
solutioning and shaping complex technology engagements Nice to have Exposure to NVIDIA ecosystems (DGX, HGX, SuperPOD) or alternative accelerators (AMD or similar) Familiarity with InfiniBand/RoCE networking, PyTorch, high-density rack design, liquid cooling or GPU-as-a-Service deployments We offer EPAM Employee Stock Purchase Plan (ESPP) Protection ...

Member of Technical Staff (AI Inference Engineer)

Location
Greater London, England, United Kingdom
laid out for you.### **Nice-to-have:*** ML compilers and framework internals: PyTorch internals, torch.compile, custom operators.* Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism.* Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving.* Profiling and debugging tools: Nsight Compute/Systems, CUDA ...