26 to 50 of 54 Slurm Workload Manager Jobs in England

Senior AI Infrastructure Engineer - Scale Multi-GPU Training

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
will architect and optimize distributed training across multiple GPUs and machines in AWS, eliminate bottlenecks in the data path, and manage cluster orchestration with Slurm and Kubernetes. The role requires deep PyTorch expertise, familiarity with transformer models, and experience deploying production AI systems. #J-18808-Ljbffr ...

HPC Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 60 K
This company is on the hunt for HPC Engineers to power their 25 Petabyte system.... Sound good? Well there's more! Imagine working with Slurm clusters and GPFS storage, all while being an integral part of groundbreaking translational research. You will work in a dynamic team of five, where ...

AI infrastructure engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 100 K
throughput, and ensuring training is as efficient and cost-effective as possible. You'll also play a critical role in managing cluster orchestration with Slurm and Kubernetes while helping evolve the platform to support next-generation GPU infrastructure and specialised compute providers. This is an opportunity to work across ...

AI Infrastructure Engineer

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
Salary
£ 80 K
will eliminate bottlenecks in the data path to ensure training is fast and as capital efficient as possible alongside managing cluster orchestration using slurm and Kubernetes while preparing to expand into specialised GPU providers. And finally you will master the stack from pytorch based learning libraries to complex data ...

AI Inference Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
workloads.Exposure to multi-tenant serving or SLA-driven infrastructure.Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system.Familiarity with Kubernetes/Slurm for cluster orchestration.Interest or experience in energy markets, grid systems, or sustainability-focused compute.BenefitsCompetitive salary and an equity sign-on bonus.Biannual bonus scheme.Fully expensed ...

AI Inference Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
multi-tenant serving or SLA-driven infrastructure. Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system. Familiarity with Kubernetes/Slurm for cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Benefits Competitive salary and an equity sign-on bonus. ...

Senior Applied Research Engineer - Video Team

Hiring Organisation
Synthesia
Location
London, United Kingdom
Salary
£ 80 K
human-centric generationFamiliarity with world/interactive modelsExperience with GANs or VAEsExperience optimizing inference systems for productionOur stackPython, PyTorch, CUDADeepSpeed, distributed training & inferenceSequence parallelismAWS, SLURM, DockerGitHub, CI/CD pipelinesWho you areYou are research-driven but outcome-focusedYou care about shipping, not just publishingYou can explore multiple ideas quickly ...

Quant Developer (C++)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
global team Python, Q/kdb+ Testing methodologies (unit tests, regression tests) Dev workflow – SVN, GIT, JIRA, Code Reviews, etc Grid & cluster tools (especially SLURM) The minimum base salary for this role is $120,000 if located in New York. This expectation is based on available information ...

Quant Developer (C++)

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, United Kingdom
Salary
£ 100 K
working on a global teamPython, Q/kdb+Testing methodologies (unit tests, regression tests)Dev workflow – SVN, GIT, JIRA, Code Reviews, etcGrid & cluster tools (especially SLURM) The minimum base salary for this role is $120,000 if located in New York. This expectation is based on available information ...

Software Engineer, General

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
observability tools (Prometheus, Grafana) is beneficial Bonus/Good to Have HPC & Cluster Management: Experience handling large-scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large-scale data processing frameworks Systems‐Level ...

HPC Senior Technology Consultant

Hiring Organisation
Hewlett Packard Enterprise
Location
Wokingham, Berkshire, United Kingdom
Salary
£ 70 K
customer-facing environment.Desirable ExperienceExperience in one or more of the following areas would be advantageous:Cluster management platforms such as Bright Cluster Manager, HPE Performance Cluster Manager (HPCM) or Cray Systems Manager (CSM).SLURM or PBS Pro job schedulers.High-performance networking technologies such as InfiniBand or Slingshot.Parallel ...

Senior Cloud Platform Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
focus on end‐user availability. Desirable but not required Experience with Openstack cloud platform s. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Experience with hardware offloading on RDMA‐capable NICs and how that integrates with virtual networking on Open V‐switch ...

MLOps Engineer

Hiring Organisation
Jobleads-UK
Location
Oxford, England, United Kingdom
across on-premises accelerator clusters and cloud (GPU/CPU) for training and simulation workloads Drive infrastructure-as-code practices: containerisation, orchestration (Kubernetes/Slurm), and reproducible environment management Contribute to the internal developer platform: self-service tooling, documentation, and runbooks that raise engineering productivity across the company What … Experience with experiment tracking and model lifecycle management tools (MLflow, W&B, DVC, or similar) Solid understanding of containerisation (Docker) and orchestration (Kubernetes or Slurm) for distributed compute workloads Infrastructure-as-code mindset: Terraform, Ansible, or equivalent; CI/CD pipelines (GitHub Actions, Jenkins, or similar) Experience with hardware ...

Solution Architect - GPU & HPC

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
winning them requires more than a great sales team. The Solutions Architect sits at the intersection of sales, infrastructure, and the customer, translating complex workload requirements into technically sound, commercially viable solutions on the Hyperstack platform. You’ll be the primary technical authority through the sales cycle: engaging directly … proposal, and delivery handover — acting as the primary technical authority for GPU cloud solution design. Engage directly with prospective and existing customers to understand workload requirements, technical constraints, and commercial objectives, producing detailed solution designs including architecture diagrams, network topology, storage configurations, and GPU resource allocation models. Collaborate closely ...

Senior Cloud Engineer (K8S)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
cloud platforms. Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Programming experience with Python3 utilising classes and inheritance. Benefits In addition to a competitive salary, Graphcore offers flexible ...

Senior AI Platform Engineer

Hiring Organisation
IQVIA
Location
London, United Kingdom
Salary
£ 80 K
scalable engineering solutions.Partner with centralised infrastructure teams to design and deliver high-performance compute environments across AWS and on-premises platforms, including GPU infrastructure, Slurm clusters, and migration from ad hoc research workflows.Optimise LLM training and inference workloads, supporting research and product teams in maximising performance, scalability, and reliability … NVIDIA Nsight, DCGM, and related ecosystem technologies.Strong background in AWS cloud services, high-performance computing, distributed systems, containerised environments, and infrastructure automation.Experience with workload orchestration technologies such as Slurm, Kubernetes, Ray, or equivalent distributed compute frameworks.Demonstrated success bridging research and production environments, enabling rapid experimentation while maintaining operational ...

Senior Software Engineer, Inference Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
model lifecycle management is highly desirable Bonus/Good to Have HPC & Cluster Management: Experience handling large‐scale HPC clusters using Kubernetes and Slurm for job scheduling, resource allocation, and workload orchestration Data Engineering: Expertise with data pipelines, ETL systems, and large‐scale data processing frameworks Systems-Level ...

Build Engineer

Hiring Organisation
Intellectual Capital Resources
Location
London, United Kingdom
Salary
£ 80 K
hybrid working model, required onsite 3 days a week. Experience for the Build Engineer includes: Software development with Python Modern build systems, e.g. Bazel Workload management, e.g. Slurm, LSF or SGE Infrastructure as code or IAC, e.g. Ansible or Terraform Containerisation, e.g. Docker Desired: Experience supporting silicon ...

R&D Solution Architect

Hiring Organisation
Jobleads-UK
Location
Hook, England, United Kingdom
implementing solutions that adhere to FAIR data principles (Findable, Accessible, Interoperable, Reusable)." - Experience architecting for High-Performance Computing (HPC) environments, including knowledge of workload schedulers (e.g., SLURM) and applying cloud-native patterns to scientific, batch-processing workloads. - Familiarity with scientific workflow management tools (e.g., Nextflow, Snakemake ...

Technical Account Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
services, we are scaling rapidly across multiple dimensions at once. As our customer base continues to grow, we are looking for a Technical Account Manager to become the trusted technical partner for our customers, helping them successfully deploy, scale, and operate large GPU and HPC workloads on our platform. … chance to join the ride. Join Verda while it's still being built, not once it's finished! Your responsibilities As a Technical Account Manager, you own the technical relationship with a portfolio of customers running demanding GPU and HPC workloads on Verda. The role works in two directions ...

Python Software Engineer - Intraday Trading

Hiring Organisation
Millennium Management
Location
London, United Kingdom
Salary
£ 100 K
strong understanding of Linux operating systems• Experience with grid scheduling and compute orchestration for real-time compute management at scale, including technologies such as SLURM• Strong understanding of event-driven architecture and experience with messaging and caching technologies such as Kafka, Solace, Pulsar, Memcache, and Redis• Experience building … scale real-time portfolio analytics tools, cloud platforms, containerization technologies such as Docker and Kubernetes, or multi-threaded C++ is a plusRecruiter:Ruby KazmiHiring Manager:Shashank GiriDepartment:Information Technology ...

Senior Staff+ Software Engineer, Kubernetes Platform

Hiring Organisation
Humanloop
Location
London, United Kingdom
Salary
> £ 150 K
controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure.We make sure the control plane is fast, correct, and always available. … Anthropic's accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemptionScale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds usDesign, build, and operate core cluster services such ...

Staff Software Engineer, Kubernetes Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure. We make sure the control plane is fast, correct, and always … accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption Scale the Kubernetes control plane (apiserver, etcd, controller‐manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds us Design, build, and operate core cluster services such ...

Senior Staff+ Software Engineer, Kubernetes Platform

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure. We make sure the control plane is fast, correct, and always … accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds us Design, build, and operate core cluster services such ...

Founding GPU Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
center systems. You'll work on low-level performance engineering for large-scale compute clusters, helping Fuse build the software layer that ties GPU workload behaviour to energy availability and grid demand.The OpportunityDemand for high-performance compute capacity across the markets we operate in significantly outpaces what … multi-node scaling using NCCL, MPI, or similar communication libraries.Work with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensity.Collaborate with ML/systems engineers to integrate custom kernels into training/inference pipelines.Benchmark against ...