17 of 17 Remote/Hybrid Slurm Workload Manager Jobs in the UK

Machine Learning Engineer (Mid to Principal)

Location
West of England, England, United Kingdom
Jira. Edge deployments: Nvidia Jetson (e.g. AGX Orin), Raspberry Pi, or other embedded accelerators. Distributed model training & infra: Pytorch DDP, FDSP and TorchTitan, Megatron, Slurm, Run:ai, DeepSpeed, Kubernetes, cloud or on‐prem GPU clusters. About you You’ve built ML systems that persist—deployed in real settings, iterated ...

Research Engineer – Human Influence

Location
Greater London, England, United Kingdom
scalable and maintainable production code in (at least) Python. Comfortabl e with serving, scaling, and containerising ML code , e.g. using Docker, Kubernetes, Ray, FastAPI , SLURM, e specially on large compute clusters. Good understanding of model internals, e.g. for mechanistic interpretability research, or analysing model activations and weights Experience shipping ...

Senior Software Engineer - Scientific / HPC

Hiring Organisation
Technical Futures Ltd
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£70,000 - £90,000 per annum
code quality. Some/most of the following should support the skills above: Experience of Cloud computing or HPC job management (such as Slurm). Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux (such as Podman). Infrastructure as code (such as Ansible ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, UK
Employment Type
Full-time
training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker, Singularity ...

Hybrid HPC System Administrator – Slurm, Linux & Storage

Location
Farnborough, England, United Kingdom
Lenovo in the United Kingdom (Hampshire) is hiring an HPC System Administrator to manage data center infrastructure, Linux/Unix environments, and monitoring to ensure uptime and security. This role blends on-site and remote ...

Senior Solutions Engineer

Hiring Organisation
LJB & Co
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £1,000 to £1,200 per day
RoCEv2 networking. Develop infrastructure solutions using NVIDIA Blackwell, B300, GB300 and GB200 platforms. Design both bare-metal and Kubernetes-based GPU environments. Work with Slurm, Kubernetes, NVIDIA GPU Operator, NCCL and GPUDirect. Design high-performance storage solutions for AI workloads. Lead technical discussions with CTOs, AI leaders and infrastructure ...

Senior HPC Engineer - Hybrid GPU Linux Clusters

Location
Stevenage, England, United Kingdom
researchers and technical teams. You will manage HPC clusters, ensure reliability and performance, and support scientific applications and complex workloads, including GPU computing and Slurm workload management. Hybrid onsite presence is required. #J-18808-Ljbffr ...

Linux HPC Specialist

Location
Stevenage, England, United Kingdom
fast paced development environment. Salary: Up to £75,000 depending on experience Dynamic (hybrid) working: 2-3 days per week on-site due to workload classification Security Clearance This role will require DV Clearance. Restrictions and/or limitations relating to nationality and/or rights to work … engineering and operations teams Performance & Scalability Ensure systems are designed for optimal throughput, latency, and scalability Lead performance benchmarking, capacity planning, and workload modellingIdentify and eliminate architectural bottlenecks Workload & Software Ecosystem Define strategies for workload orchestration (SLURM, Kubernetes for HPC, etc.) Guide software stack design ...

Technical Lead - GPU Infrastructure

Location
Greater London, England, United Kingdom
JavaScript. The platform is expanding from orchestrating workloads on a managed cluster to owning the full stack on bare metal GPU infrastructure: a managed Slurm scheduling layer for internal research and model-training teams first, and our own Kubernetes control plane for inference tenancy after that. The Technical Lead … about twelve engineers across backend, frontend, DevOps, QA and documentation, distributed across Europe and India. The role reports to the Senior Technical Product Manager for Cosmic AC, who owns scope, sequencing and partner commitments; the Technical Lead owns architecture, implementation and delivery plans, line-manages the engineers ...

Senior HPC Engineer - Hybrid - Inside IR35

Location
Stevenage, England, United Kingdom
Role We are looking for an experienced Senior HPC Engineer to support and maintain a scientific computing environment, with a strong focus on RHEL, Slurm and HPC infrastructure . You will work closely with research scientists and technical teams to ensure HPC services are secure, reliable and high performing. … Responsibilities Administer, patch and maintain RHEL 7, 8 and 9 across HPC clusters and workstations. Deploy, configure and manage Slurm , including queues, partitions and scheduling. Monitor cluster health, performance, storage, networking and resource utilisation. Install and support scientific applications, compilers, libraries and MPI environments. Work with scientists to optimise ...

Build Engineer

Location
Greater London, England, United Kingdom
hybrid working model, required onsite 3 days a week. Experience for the Build Engineer includes: Software development with Python Modern build systems, e.g. Bazel Workload management, e.g. Slurm, LSF or SGE Infrastructure as code or IAC, e.g. Ansible or Terraform Containerisation, e.g. Docker Desired: Experience supporting silicon ...

R&D Solution Architect

Location
Hook, England, United Kingdom
implementing solutions that adhere to FAIR data principles (Findable, Accessible, Interoperable, Reusable)." - Experience architecting for High-Performance Computing (HPC) environments, including knowledge of workload schedulers (e.g., SLURM) and applying cloud-native patterns to scientific, batch-processing workloads. - Familiarity with scientific workflow management tools (e.g., Nextflow, Snakemake ...

Senior Solution Engineer – GPU & AI Infrastructure

Location
United Kingdom
role, you will bridge the gap between customer business objectives and ultra-high-performance hardware execution. You will lead technical engagements, translate complex AI workload requirements into production-ready High-Level Designs (HLD), Low-Level Designs (LLD), and detailed Bills of Materials (BOM). Your expertise will span bare …/Spectrum-4) with lossless Ethernet mechanisms (PFC, ECN, Adaptive Routing). Multi-Tenant & Deployment Models: Deliver tailored architectures for both Bare-Metal (Slurm, OpenMPI, bare-metal provisioning) and Cloud-Native/Kubernetes environments (NVIDIA GPU Operator, Network Operator, Run:ai, KubeFlow). Storage Integration: Architect high-bandwidth parallel ...

Senior Principal AI Infrastructure Architect

Hiring Organisation
The Nippon Telegraph And Telephone Corporation (NTT)
Location
United Kingdom, UK
Employment Type
Full-time
land service-led AI solutions. Lead integration of compute, storage, networking, the AI software stack (CUDA, ROCm, Triton, NIM, NVIDIA AI Enterprise, Run:ai, Slurm, Kubernetes/Kubeflow) and managed-service operating models across multiple domains, delivery units and geographies. Build business cases, TCO and unit-economics models (cost … tree topologies. Working knowledge of the AI software and orchestration stack: CUDA, cuDNN, NCCL, ROCm, Triton Inference Server, NIM, vLLM, TensorRT-LLM, Slurm, Kubernetes (with GPU Operator), Kubeflow, Run:ai, MLflow and NVIDIA AI Enterprise. Familiarity with datacenter facilities engineering for AI workloads: high-density power, liquid cooling ...

Senior HPC Engineer

Location
West of England, England, United Kingdom
Linux system administration. Familiarity with at least one scripting language (e.g., Bash, Python). Interest in high-performance computing and willingness to learn Slurm, xCAT, and Ansible. Strong problem-solving skills and attention to detail. Good communication and collaboration skills. Ability to work independently with mentorship and as part … week. No hybrid/remote working option. Internship or academic experience in a research computing or HPC environment. Exposure to job schedulers (e.g., Slurm, LSF). Familiarity with version control systems (e.g., Git). Coursework or projects involving distributed systems, networking, or parallel computing. Understanding of basic cybersecurity concepts. ...

Senior HPC Engineer

Location
Stevenage, England, United Kingdom
supporting scientific applications and complex computational workloads. Key Skills: Strong hands-on administration of RHEL 7, 8 and 9 HPC cluster administration and troubleshooting Slurm deployment, configuration and workload management Scientific applications within Linux HPC environments MPI libraries and computational workloads Hardware, OS and scheduler troubleshooting ServiceNow ...

Principal - AI & HPC Data Centre Compute

Location
Greater London, England, United Kingdom
inference workloads (LLMs, multimodal and scientific AI) across distributed clusters for cost, throughput and latency targets Lead HPC cluster design and orchestration using Slurm, Kubernetes and parallel processing models (MPI) Enhance GPU utilisation by addressing bottlenecks across compute, memory and data pipelines; collaborate with energy teams on power-aware … infrastructure, accelerated computing or distributed systems architecture Deep knowledge of GPU architectures, AI workloads, networking and large-scale cluster operations Hands-on expertise with Slurm, Kubernetes and performance optimisation across multi-node environments Proven ability to advise senior stakeholders and influence technical strategy at C-level Demonstrated experience ...