1 to 25 of 32 Slurm Workload Manager Jobs in London

ML Ops Engineer

Location
Greater London, England, United Kingdom
cost-efficiency. Your Impact Provision and manage cloud-native AI/ML infrastructure utilising Kubernetes, Docker, and GPU orchestration frameworks (e.g., NVIDIA GPU Operator, Slurm, or Ray). Automate core platform infrastructure using Infrastructure as Code (IaC) tools like Terraform, Helm, and Ansible. Optimise GPU compute workloads, high-speed ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
environments Enjoys solving complex technical problems collaboratively Nice-to-Haves Experience with large scale model training or distributed inference systems Familiarity with Ray, Kubeflow, Slurm, or similar distributed scheduling platforms Experience with InfiniBand, RDMA, or high-performance networking Experience operating bare metal infrastructure Familiarity with storage systems commonly used ...

Principal Machine Learning Infrastructure Engineer London, United Kingdom

Location
Greater London, England, United Kingdom
Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar ...

Research HPC Support Engineer

Hiring Organisation
Appcast
Location
London, UK
have:- Strong technical expertise in high-performance computing (HPC) and GPU systems. - Proven experience administering, configuring, and optimising HPC clusters and GPU systems (e.g. Slurm, OpenPBS, K8).- Hands-on experience with open-source software build and installation frameworks specifically designed for High-Performance Computing (HPC) environments, such ...

HPC Systems Engineer

Location
Greater London, England, United Kingdom
Need: At least 5+ years of professional experience in high performance computing (HPC), including parallel filesystems (e.g., Lustre, GPFS), batch systems (e.g., Slurm, Grid Engine), and high-performance network interconnects experience is a plus, but not required At least 5+ years of experience with Linux systems administration High proficiency ...

HPC Operations Engineer - Banking & Finance

Location
Greater London, England, United Kingdom
Python, Bash, Go or similar scripting experience. Experience supporting production infrastructure. Familiarity with automation and configuration management tools. Exposure to HPC technologies such as Slurm, Lustre or GPFS is advantageous. Location City of London (on-site) Benefits Work with cutting‐edge Linux infrastructure and distributed systems. Solve complex technical ...

Senior Data & MLOps Engineer

Location
Greater London, England, United Kingdom
least one systems language (Go, Rust, or C++). Experience working with distributed compute or training systems (e.g., NCCL, PyTorch Distributed, Spark, Ray, Slurm). Familiarity with GPU telemetry systems such as NVML or DCGM and hardware‐level monitoring concepts. Demonstrated experience scaling systems from Proof‐of‐Concept ...

Senior Field Engineer - AI/ML HPC & Kubernetes Solutions

Location
Greater London, England, United Kingdom
drive proofs of concept to accelerate client deployments. You will collaborate with engineering teams, contribute to product direction, and help optimize workloads using Slurm, NCCL, and Infiniband in high-performance environments. #J-18808-Ljbffr ...

Research Engineer, Forge

Location
Greater London, England, United Kingdom
comfort in fast‐moving, under‐specified environments. Nice to have Distributed training experience (FSDP, DeepSpeed, Megatron, etc.). Cluster/orchestration experience (SLURM, Ray, Kubernetes, Kueue, Karpenter, Skypilot, etc.). Experience building reliable ML infrastructure, evaluation systems, or large‐scale data processing pipelines. Research experience in LLMs, agents, multimodal ...

IT Lead Engineer London

Location
Greater London, England, United Kingdom
slow" needs a real root cause, not a restart. What you will do Own and evolve our HPC environment: cluster administration, job scheduling (e.g., Slurm/PBS/LSF), performance tuning, and capacity planning for compute-heavy engineering workloads. Experience with Entra ID Governance: Access Reviews, Identity Protection, Privileged ...

Senior Solutions Engineer

Hiring Organisation
LJB & Co
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £1,000 to £1,200 per day
RoCEv2 networking. Develop infrastructure solutions using NVIDIA Blackwell, B300, GB300 and GB200 platforms. Design both bare-metal and Kubernetes-based GPU environments. Work with Slurm, Kubernetes, NVIDIA GPU Operator, NCCL and GPUDirect. Design high-performance storage solutions for AI workloads. Lead technical discussions with CTOs, AI leaders and infrastructure ...

Machine Learning Engineer

Hiring Organisation
Vertex Search
Location
London Area, United Kingdom
cross-node communication Working knowledge of PyTorch and/or JAX training loops Experience running workloads on HPC or multi-node GPU clusters (e.g. Slurm, distributed training) Bonus points Prior work with health records, single-cell/cytometry data, or time-series Prior experience working with Tabular data/ ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused compute Competitive salary and eligibility for equity Biannual bonus scheme Fully expensed tech to match ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
multi-tenant serving or SLA-driven infrastructure. Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system. Familiarity with Kubernetes/Slurm for cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Benefits Competitive salary and an equity sign-on bonus. ...

Senior Research HPC Engineer

Hiring Organisation
MRC Laboratory of Medical Sciences
Location
London, United Kingdom
Employment Type
Permanent
Salary
£65,000
operates its own dedicated HPC environment, which is extensively used by multidisciplinary research groups and Institute facilities. The LMS IT department currently manages a SLURM-based cluster comprising CPU, high-memory (HMEM), and GPU nodes, supporting a diverse range of computational research. About the role As the Institutes computational ...

Senior Research HPC Engineer

Location
Greater London, England, United Kingdom
administering Linux‐based systems in an HPC, research, academic, or production environment Experience in configuration and maintenance of multi‐queue job scheduling systems (e.g. SLURM) Experience in the use of scientific software compilation and deployment systems (e.g. Spack, EasyBuild, Lmod, conda) Virtualisation and containerisation deployment and management (e.g. Docker … verbal and written communication skills Able to self‐motivate when working independently, on projects, and collaboratively within a team Effectively plans, multitask, and prioritises workload to achieve results, adapting quickly while maintaining control in challenging situations Ensures appropriate engagement of colleagues with relevant stakeholders and issues Ability to develop ...

Senior Cloud Engineer (K8S)

Location
Greater London, England, United Kingdom
platform s. Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Programming experience with Python3 utilising classes and inheritance. Benefits In addition to a competitive salary flexible working ...

Senior DevOps Engineer

Location
Greater London, England, United Kingdom
high-performance, high-IOPS cloud storage solutions, ensuring seamless data synchronisation, backup strategies, and optimal throughput. Automate the deployment, configuration, and auto-scaling of workload orchestrators, container platforms, and core service components to ensure high resource utilisation and cost efficiency. Enforce strict security protocols across all environments by implementing … benefit updates for leadership. Experience - Desirable Proven track record of architecting and deploying production HPC workloads on AWS using AWS ParallelCluster, SOCA or custom Slurm fleets. Ability to architect for maximum cost efficiency, implementing automated spot-instance utilization, auto-scaling strategies, and Savings Plans optimization. Proficiency in monitoring infrastructure ...

Senior AI Platform Engineer

Location
Greater London, England, United Kingdom
engineering solutions. Partner with centralised infrastructure teams to design and deliver high-performance compute environments across AWS and on-premises platforms, including GPU infrastructure, Slurm clusters, and migration from ad hoc research workflows. Optimise LLM training and inference workloads, supporting research and product teams in maximising performance, scalability … Nsight, DCGM, and related ecosystem technologies. Strong background in AWS cloud services, high-performance computing, distributed systems, containerised environments, and infrastructure automation. Experience with workload orchestration technologies such as Slurm, Kubernetes, Ray, or equivalent distributed compute frameworks. Demonstrated success bridging research and production environments, enabling rapid experimentation while ...

Platform Architect - Nvidia AI/GB300

Hiring Organisation
Oscar Associates (UK) Limited
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£700 - £765 per day
data-centre teams. Key Requirements Strong Platform/Infrastructure Architecture experience across compute, storage, networking and Linux. Expert-level Kubernetes architecture experience. Strong Slurm and Run experience - essential. Proven experience with GPU/HPC environments and large-scale AI platforms. Hands-on experience with NVIDIA HGX GB300/NVL72 … including NVLink, NVSwitch and Grace Blackwell architecture. Experience with NVIDIA RTX 6000 series GPU servers. Strong understanding of GPU workload scheduling, partitioning and sharing, including MIG, vGPU and time-slicing Strong understanding of InfiniBand, RoCE, Spectrum-X, GPUDirect RDMA/Storage and high-performance AI fabrics. Experience with Terraform ...

Build Engineer

Location
Greater London, England, United Kingdom
hybrid working model, required onsite 3 days a week. Experience for the Build Engineer includes: Software development with Python Modern build systems, e.g. Bazel Workload management, e.g. Slurm, LSF or SGE Infrastructure as code or IAC, e.g. Ansible or Terraform Containerisation, e.g. Docker Desired: Experience supporting silicon ...

Solutions Engineer

Location
Greater London, England, United Kingdom
engineering, and operations teams to ensure proposed solutions are realistic, scalable, and aligned with platform standards Provide guidance on compute, networking, storage, orchestration, and workload optimisation for AI and machine learning use cases Help create repeatable demo environments, technical playbooks, reference architectures, and sales enablement materials … data centre, or platform environments Good understanding of cloud infrastructure, GPU compute, AI/ML workloads, or high-performance infrastructure Familiarity with containers, Kubernetes, Slurm, orchestration platforms, or workload deployment models Understanding of networking, storage, and distributed compute concepts in modern infrastructure environments Ability to quickly learn ...

High Performance Computing Architect (Linux)

Location
London, United Kingdom
parallel file systems, object storage, and tiered storage Provide technical leadership and oversight to engineering and operations teams Lead performance benchmarking, capacity planning, and workload modelling activities Identify and resolve system bottlenecks to ensure optimal throughput and scalability Establish standards, reference architectures, and best practices across HPC environments Collaborate … with internal stakeholders and external vendors to align infrastructure with evolving business needs Support the development of workload orchestration strategies using tools such as SLURM and Kubernetes The ideal candidate would have: Strong background in high-performance computing environments and infrastructure design Experience working with open-source technologies ...

Senior Specialist Field Engineer - HPC/AI/ML

Location
Greater London, England, United Kingdom
functionality, and performance, contributing regularly to discussions about product strategy and architecture. Conduct periodic technical reviews and assessments of customer workloads, pinpointing opportunities for workload optimization and suggesting suitable solutions. Stay informed of the latest developments and trends in Kubernetes, cloud computing and infrastructure, sharing your thought leadership with … related technical discipline, or equivalent experience 7+ years of proven experience as a Solutions Architect, Field Engineer, Engineer, Researcher, or Technical Account Manager in Cloud Infrastructure, focusing on building distributed systems or HPC/cloud services, with an expertise focused on AI/ML inference Fluency in cloud computing ...

Staff HPC Systems Software Engineer

Location
Greater London, England, United Kingdom
core HPC platform domain at Nscale. In this role, you will operate beyond a single team, shaping how multiple teams build, automate,and run Slurm-based capabilities within Nscale’s wider cloud-native platform. You’ll work acrossengineering boundaries to bring coherence to architecture, interfaces, lifecycle models, andoperational approaches … Domain Architecture & Technical Direction Own and evolve the technical direction for a defined HPC systems domain, such as Slurmplatform architecture, scheduler integrations, cluster lifecycle, workload environments orservice automation. Make architectural decisions that balance software quality, operational realities, customerneeds, and long-term maintainability. Define how proven Slurm implementations should ...