7 of 7 Remote/Hybrid Slurm Workload Manager Jobs in the UK

Member of Technical Staff (AI Infrastructure Engineer)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Full time Location Type Hybrid Department AI We are looking for an AI Infra engineer to join our growing team. We work with Kubernetes, Slurm, Python, C++, PyTorch, and primarily on AWS. As an AI Infrastructure Engineer, you will be partnering closely with our Inference and Research teams … scale AI training and inference clusters. Responsibilities Design, deploy, and maintain scalable Kubernetes clusters for AI model inference and training workloads Manage and optimize Slurm-based HPC environments for distributed training of large language models Develop robust APIs and orchestration systems for both training pipelines and inference services Implement ...

Infrastructure / DevOps Lead

Hiring Organisation
Jobleads-UK
Location
United Kingdom
details. Nice to Have Experience managing physical data centres, co‐location facilities, or hybrid infrastructure environments. Working knowledge of ML orchestration frameworks (e.g., Ray, Slurm, Kubeflow). Background in media pipelines, VFX tooling, or media compliance standards (MPA, ISO 27001). Prior experience working in a hybrid startup/ ...

Research Engineer, Machine Learning – Paris/London/Zurich/Warsaw

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
+ years working on large‐scale ML codebases. Hands‐on with PyTorch, JAX or TensorFlow; comfortable with distributed training (DeepSpeed/FSDP/SLURM/K8s). Experience in deep learning, NLP or LLMs; bonus for CUDA or data‐pipeline chops. Strong software‐design instincts: testing, code review ...

R&D Solution Architect

Hiring Organisation
Jobleads-UK
Location
Hook, England, United Kingdom
implementing solutions that adhere to FAIR data principles (Findable, Accessible, Interoperable, Reusable)." - Experience architecting for High-Performance Computing (HPC) environments, including knowledge of workload schedulers (e.g., SLURM) and applying cloud-native patterns to scientific, batch-processing workloads. - Familiarity with scientific workflow management tools (e.g., Nextflow, Snakemake ...

Staff Software Engineer, Kubernetes Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure. We make sure the control plane is fast, correct, and always … accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption Scale the Kubernetes control plane (apiserver, etcd, controller‐manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds us Design, build, and operate core cluster services such ...

AI Infrastructure Engineer — Scalable ML Clusters

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
technology firm in London is seeking an AI Infrastructure Engineer to join their team. The role involves designing and maintaining scalable Kubernetes clusters, optimizing Slurm-based HPC environments, and developing APIs for AI workloads. Candidates should have expertise in Kubernetes, experience with Slurm, and skills in Python ...

Backend Engineer

Hiring Organisation
Technical Futures Ltd
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Permanent
Salary
GBP 55,000 - 70,000 Annual
between them; with experience of modern Python tooling and packaging, some knowledge of Cloud computing or HPC job management (such as Slurm), identity and authorization flows (such as OAuth2/OIDC). This cutting-edge technology company, focused on optimizing complex engineering systems, seeks a top class Backend Engineer … code quality. Some/most of the following should support the skills above: Experience of Cloud computing or HPC job management (such as Slurm). Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux (such as Podman). Infrastructure as code (such as Ansible ...