22 of 22 Slurm Workload Manager Jobs in London

ML Ops Engineer

Location
Greater London, England, United Kingdom
cost-efficiency. Your Impact Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray). Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed ...

Platform Application Specialist

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
AlertManager).Experience with a variety of database platforms (e.g., PostgreSQL, ClickHouse, MSSQL, Redis, FoundationDB).Familiarity with specific middleware (e.g., Kafka, Consul), HPC schedulers (e.g., Slurm), Kubernetes, or workflow orchestrators (e.g., Airflow, Prefect).Experience with on-premise deployments of development tools like GitLab, Artifactory, CICD, or JupyterHub. Knowledge of other ...

Senior Infrastructure Engineer, Research Singapore

Location
Greater London, England, United Kingdom
Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
high-throughput file systems (e.g. Everpure, WEKA, VAST, NetApp) sized for training and checkpointing workloads. AI Software, MLOps & Generative AIOrchestration & containers: Kubernetes, Docker, Slurm, Run:ai or equivalent GPU scheduling platforms. ML frameworks: PyTorch and TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron … continuous, declarative platform delivery. Pipeline orchestration: Kubeflow Pipelines, Apache Airflow, or Argo Workflows to orchestrate multi-stage training, fine-tuning, and inference pipelines. Cluster & workload scheduling: Slurm, Run:ai, and NVIDIA Base Command Manager for GPU job scheduling; Kubernetes-native GPU scheduling including device plugins ...

HPC Systems Engineer

Location
Greater London, England, United Kingdom
Need: At least 5+ years of professional experience in high performance computing (HPC), including parallel filesystems (e.g., Lustre, GPFS), batch systems (e.g., Slurm, Grid Engine), and high-performance network interconnects experience is a plus, but not required At least 5+ years of experience with Linux systems administration High proficiency ...

Senior ML Systems Engineer, Frameworks & Tooling

Location
Greater London, England, United Kingdom
training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar). Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, UK
Employment Type
Full-time
training or HPC systems. Deep familiarity with JAX internals, distributed training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker, Singularity ...

HPC Operations Engineer - Banking & Finance

Location
Greater London, England, United Kingdom
Python, Bash, Go or similar scripting experience. Experience supporting production infrastructure. Familiarity with automation and configuration management tools. Exposure to HPC technologies such as Slurm, Lustre or GPFS is advantageous. Location City of London (on-site) Benefits Work with cutting‐edge Linux infrastructure and distributed systems. Solve complex technical ...

Senior Machine Learning Systems Engineer (Frameworks & Tooling)

Location
Greater London, England, United Kingdom
allowance A monthly quality time allowance A track record of building tools that increase developer velocity for ML teamsExperience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar)Bonus: paper at top-tier venues (such as NeurIPS, ICML, ICLR, AIStats, MLSys, JAX, AAAI, Nature, COLING, ACL, EMNLP)Experience ...

AI infrastructure engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
throughput, and ensuring training is as efficient and cost-effective as possible. You'll also play a critical role in managing cluster orchestration with Slurm and Kubernetes while helping evolve the platform to support next-generation GPU infrastructure and specialised compute providers. This is an opportunity to work across ...

AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, UK
Employment Type
Full-time
will eliminate bottlenecks in the data path to ensure training is fast and as capital efficient as possible alongside managing cluster orchestration using slurm and Kubernetes while preparing to expand into specialised GPU providers. And finally you will master the stack from pytorch based learning libraries to complex data ...

Senior Solutions Engineer

Hiring Organisation
LJB & Co
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £1,000 to £1,200 per day
RoCEv2 networking. Develop infrastructure solutions using NVIDIA Blackwell, B300, GB300 and GB200 platforms. Design both bare-metal and Kubernetes-based GPU environments. Work with Slurm, Kubernetes, NVIDIA GPU Operator, NCCL and GPUDirect. Design high-performance storage solutions for AI workloads. Lead technical discussions with CTOs, AI leaders and infrastructure ...

AI Inference Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, UK
Employment Type
Full-time
scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused computeBenefitsCompetitive salary and eligibility for equityBiannual bonus schemeFully expensed tech to match your needsPrivate health ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
multi-tenant serving or SLA-driven infrastructure. Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system. Familiarity with Kubernetes/Slurm for cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Benefits Competitive salary and an equity sign-on bonus. ...

Platform Architect - Nvidia AI/GB300

Hiring Organisation
Oscar Associates (UK) Limited
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£700 - £765 per day
data-centre teams. Key Requirements Strong Platform/Infrastructure Architecture experience across compute, storage, networking and Linux. Expert-level Kubernetes architecture experience. Strong Slurm and Run experience - essential. Proven experience with GPU/HPC environments and large-scale AI platforms. Hands-on experience with NVIDIA HGX GB300/NVL72 … including NVLink, NVSwitch and Grace Blackwell architecture. Experience with NVIDIA RTX 6000 series GPU servers. Strong understanding of GPU workload scheduling, partitioning and sharing, including MIG, vGPU and time-slicing Strong understanding of InfiniBand, RoCE, Spectrum-X, GPUDirect RDMA/Storage and high-performance AI fabrics. Experience with Terraform ...

Platform Architect - Nvidia AI/GB300

Hiring Organisation
Oscar Technology
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£700.00 - £765.00 per day
data-centre teams. Key Requirements Strong Platform/Infrastructure Architecture experience across compute, storage, networking and Linux. Expert-level Kubernetes architecture experience. Strong Slurm and Run experience - essential. Proven experience with GPU/HPC environments and large-scale AI platforms. Hands-on experience with NVIDIA HGX GB300/NVL72 … including NVLink, NVSwitch and Grace Blackwell architecture. Experience with NVIDIA RTX 6000 series GPU servers. Strong understanding of GPU workload scheduling, partitioning and sharing, including MIG, vGPU and time-slicing Strong understanding of InfiniBand, RoCE, Spectrum-X, GPUDirect RDMA/Storage and high-performance AI fabrics. Experience with Terraform ...

Solutions Engineer

Location
Greater London, England, United Kingdom
engineering, and operations teams to ensure proposed solutions are realistic, scalable, and aligned with platform standards Provide guidance on compute, networking, storage, orchestration, and workload optimisation for AI and machine learning use cases Help create repeatable demo environments, technical playbooks, reference architectures, and sales enablement materials … data centre, or platform environments Good understanding of cloud infrastructure, GPU compute, AI/ML workloads, or high-performance infrastructure Familiarity with containers, Kubernetes, Slurm, orchestration platforms, or workload deployment models Understanding of networking, storage, and distributed compute concepts in modern infrastructure environments Ability to quickly learn ...

Staff Software Engineer, Kubernetes Platform

Hiring Organisation
Humanloop
Location
London, UK
Employment Type
Full-time
controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure. We make sure the control plane is fast, correct, and always … Anthropic's accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemptionScale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds usDesign, build, and operate core cluster services such ...

Senior Staff+ Software Engineer, Kubernetes Platform

Location
Greater London, England, United Kingdom
controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure. We make sure the control plane is fast, correct, and always … accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds us Design, build, and operate core cluster services such ...

Founding GPU Engineer

Location
Greater London, England, United Kingdom
node scaling using NCCL, MPI, or similar communication libraries. Work with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensity. Collaborate with ML/systems engineers to integrate custom kernels into training/inference pipelines. … data center power/thermal management or demand-response systems. Background in HPC, quantitative finance, or large-scale distributed systems. Familiarity with Kubernetes/Slurm for GPU cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Competitive salary and an equity sign-on bonus. ...

Research Software Engineer

Location
Greater London, England, United Kingdom
surface live metrics. Write efficient, well-tested Python and systems code; enforce code review, CI, and observability. Design and optimise distributed services (Kubernetes/SLURM, thousands-of-GPU jobs). Prototype utilities (CLI, dashboards) and carry them through to stable, shared libraries. About the Research Engineering team Based … observability. Fluency in Python plus one systems language (C++, Rust, Go or Java). Hands-on with container orchestration and schedulers (Kubernetes/K8s, SLURM, or similar). Comfortable profiling performance, optimising I/O, and automating workflows. Self-starter, low-ego, collaborative, high-energy. Nice-to-haves Exposure ...

Senior Software Engineer - Research Technology

Location
Greater London, England, United Kingdom
fundamentals: data structures, algorithms, networking, OS, concurrency, and system design. Experience running compute at cluster scale: job scheduling, resource management, retries, and reliability. Slurm, Kubernetes, Ray, Spark, or custom internal schedulers all count. Proven data‐engineering experience: schema design, storage formats, compression, I/O trade-offs, and pipelines … ship production software safely and repeatedly, with an obsession for data driven quality. Desirable/nice-to-have Rust experience alongside C++ and Python. Slurm or other cluster scheduler expertise. Familiarity with ML/Deep Learning frameworks. Prior finance or market‐data experience, including low‐level market connectivity. ...