51 to 75 of 81 Slurm Workload Manager Jobs in the UK

Build Engineer

Location
Greater London, England, United Kingdom
hybrid working model, required onsite 3 days a week. Experience for the Build Engineer includes: Software development with Python Modern build systems, e.g. Bazel Workload management, e.g. Slurm, LSF or SGE Infrastructure as code or IAC, e.g. Ansible or Terraform Containerisation, e.g. Docker Desired: Experience supporting silicon ...

Platform Architect - Nvidia AI/GB300

Hiring Organisation
Oscar Associates (UK) Limited
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£700 - £765 per day
data-centre teams. Key Requirements Strong Platform/Infrastructure Architecture experience across compute, storage, networking and Linux. Expert-level Kubernetes architecture experience. Strong Slurm and Run experience - essential. Proven experience with GPU/HPC environments and large-scale AI platforms. Hands-on experience with NVIDIA HGX GB300/NVL72 … including NVLink, NVSwitch and Grace Blackwell architecture. Experience with NVIDIA RTX 6000 series GPU servers. Strong understanding of GPU workload scheduling, partitioning and sharing, including MIG, vGPU and time-slicing Strong understanding of InfiniBand, RoCE, Spectrum-X, GPUDirect RDMA/Storage and high-performance AI fabrics. Experience with Terraform ...

Senior Solution Engineer – GPU & AI Infrastructure

Location
United Kingdom
role, you will bridge the gap between customer business objectives and ultra-high-performance hardware execution. You will lead technical engagements, translate complex AI workload requirements into production-ready High-Level Designs (HLD), Low-Level Designs (LLD), and detailed Bills of Materials (BOM). Your expertise will span bare …/Spectrum-4) with lossless Ethernet mechanisms (PFC, ECN, Adaptive Routing). Multi-Tenant & Deployment Models: Deliver tailored architectures for both Bare-Metal (Slurm, OpenMPI, bare-metal provisioning) and Cloud-Native/Kubernetes environments (NVIDIA GPU Operator, Network Operator, Run:ai, KubeFlow). Storage Integration: Architect high-bandwidth parallel ...

R&D Solution Architect

Location
Hook, England, United Kingdom
implementing solutions that adhere to FAIR data principles (Findable, Accessible, Interoperable, Reusable)." - Experience architecting for High-Performance Computing (HPC) environments, including knowledge of workload schedulers (e.g., SLURM) and applying cloud-native patterns to scientific, batch-processing workloads. - Familiarity with scientific workflow management tools (e.g., Nextflow, Snakemake ...

Solutions Engineer

Location
Greater London, England, United Kingdom
engineering, and operations teams to ensure proposed solutions are realistic, scalable, and aligned with platform standards Provide guidance on compute, networking, storage, orchestration, and workload optimisation for AI and machine learning use cases Help create repeatable demo environments, technical playbooks, reference architectures, and sales enablement materials … data centre, or platform environments Good understanding of cloud infrastructure, GPU compute, AI/ML workloads, or high-performance infrastructure Familiarity with containers, Kubernetes, Slurm, orchestration platforms, or workload deployment models Understanding of networking, storage, and distributed compute concepts in modern infrastructure environments Ability to quickly learn ...

High Performance Computing Architect (Linux)

Hiring Organisation
Eclectic Recruitment Ltd
Location
United Kingdom
Employment Type
Permanent
parallel file systems, object storage, and tiered storage Provide technical leadership and oversight to engineering and operations teams Lead performance benchmarking, capacity planning, and workload modelling activities Identify and resolve system bottlenecks to ensure optimal throughput and scalability Establish standards, reference architectures, and best practices across HPC environments Collaborate … with internal stakeholders and external vendors to align infrastructure with evolving business needs Support the development of workload orchestration strategies using tools such as SLURM and Kubernetes The ideal candidate would have: Strong background in high-performance computing environments and infrastructure design Experience working with open-source technologies ...

High Performance Computing Architect (Linux)

Location
London, United Kingdom
parallel file systems, object storage, and tiered storage Provide technical leadership and oversight to engineering and operations teams Lead performance benchmarking, capacity planning, and workload modelling activities Identify and resolve system bottlenecks to ensure optimal throughput and scalability Establish standards, reference architectures, and best practices across HPC environments Collaborate … with internal stakeholders and external vendors to align infrastructure with evolving business needs Support the development of workload orchestration strategies using tools such as SLURM and Kubernetes The ideal candidate would have: Strong background in high-performance computing environments and infrastructure design Experience working with open-source technologies ...

High Performance Computing Architect (Linux)

Hiring Organisation
Eclectic Recruitment
Location
Stevenage, Hertfordshire, United Kingdom
Employment Type
Permanent
Salary
£65000 - £80000/annum
parallel file systems, object storage, and tiered storage Provide technical leadership and oversight to engineering and operations teams Lead performance benchmarking, capacity planning, and workload modelling activities Identify and resolve system bottlenecks to ensure optimal throughput and scalability Establish standards, reference architectures, and best practices across HPC environments Collaborate … with internal stakeholders and external vendors to align infrastructure with evolving business needs Support the development of workload orchestration strategies using tools such as SLURM and Kubernetes The ideal candidate would have: Strong background in high-performance computing environments and infrastructure design Experience working with open-source technologies ...

Python Software Engineer - Intraday Trading

Hiring Organisation
Millennium Management
Location
London, UK
Employment Type
Full-time
strong understanding of Linux operating systems Experience with grid scheduling and compute orchestration for real-time compute management at scale, including technologies such as SLURM Strong understanding of event-driven architecture and experience with messaging and caching technologies such as Kafka, Solace, Pulsar, Memcache, and Redis Experience building … scale real-time portfolio analytics tools, cloud platforms, containerization technologies such as Docker and Kubernetes, or multi-threaded C++ is a plusRecruiter: Ruby KazmiHiring Manager: Shashank GiriDepartment: Information Technology ...

Senior Principal AI Infrastructure Architect

Hiring Organisation
The Nippon Telegraph And Telephone Corporation NTT
Location
United Kingdom
land service-led AI solutions. Lead integration of compute, storage, networking, the AI software stack (CUDA, ROCm, Triton, NIM, NVIDIA AI Enterprise, Run:ai, Slurm, Kubernetes/Kubeflow) and managed-service operating models across multiple domains, delivery units and geographies. Build business cases, TCO and unit-economics models (cost … tree topologies. Working knowledge of the AI software and orchestration stack: CUDA, cuDNN, NCCL, ROCm, Triton Inference Server, NIM, vLLM, TensorRT-LLM, Slurm, Kubernetes (with GPU Operator), Kubeflow, Run:ai, MLflow and NVIDIA AI Enterprise. Familiarity with datacenter facilities engineering for AI workloads: high-density power, liquid cooling ...

Senior HPC Engineer

Location
West of England, England, United Kingdom
Linux system administration. Familiarity with at least one scripting language (e.g., Bash, Python). Interest in high-performance computing and willingness to learn Slurm, xCAT, and Ansible. Strong problem-solving skills and attention to detail. Good communication and collaboration skills. Ability to work independently with mentorship and as part … week. No hybrid/remote working option. Internship or academic experience in a research computing or HPC environment. Exposure to job schedulers (e.g., Slurm, LSF). Familiarity with version control systems (e.g., Git). Coursework or projects involving distributed systems, networking, or parallel computing. Understanding of basic cybersecurity concepts. ...

Staff Software Engineer, Kubernetes Platform

Hiring Organisation
Humanloop
Location
London, UK
Employment Type
Full-time
controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure. We make sure the control plane is fast, correct, and always … Anthropic's accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemptionScale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds usDesign, build, and operate core cluster services such ...

Senior Staff+ Software Engineer, Kubernetes Platform

Location
Greater London, England, United Kingdom
controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure. We make sure the control plane is fast, correct, and always … accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds us Design, build, and operate core cluster services such ...

Senior Specialist Field Engineer - HPC/AI/ML

Location
Greater London, England, United Kingdom
functionality, and performance, contributing regularly to discussions about product strategy and architecture. Conduct periodic technical reviews and assessments of customer workloads, pinpointing opportunities for workload optimization and suggesting suitable solutions. Stay informed of the latest developments and trends in Kubernetes, cloud computing and infrastructure, sharing your thought leadership with … related technical discipline, or equivalent experience 7+ years of proven experience as a Solutions Architect, Field Engineer, Engineer, Researcher, or Technical Account Manager in Cloud Infrastructure, focusing on building distributed systems or HPC/cloud services, with an expertise focused on AI/ML inference Fluency in cloud computing ...

Founding GPU Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, UK
Employment Type
Full-time
Engineer to develop and optimise GPU-accelerated software for data centre systems: low-level performance engineering for large-scale compute clusters, tying GPU workload behaviour to energy availability and grid demand. This puts CUDA/GPU performance engineering at the centre of how Fuse scales its compute infrastructure. ResponsibilitiesDesign … multi-node scaling using NCCL, MPI, or similar communication librariesWork with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensityCollaborate with ML/systems engineers to integrate custom kernels into training/inference pipelinesBenchmark against ...

Staff HPC Systems Software Engineer

Location
Greater London, England, United Kingdom
core HPC platform domain at Nscale. In this role, you will operate beyond a single team, shaping how multiple teams build, automate,and run Slurm-based capabilities within Nscale’s wider cloud-native platform. You’ll work acrossengineering boundaries to bring coherence to architecture, interfaces, lifecycle models, andoperational approaches … Domain Architecture & Technical Direction Own and evolve the technical direction for a defined HPC systems domain, such as Slurmplatform architecture, scheduler integrations, cluster lifecycle, workload environments orservice automation. Make architectural decisions that balance software quality, operational realities, customerneeds, and long-term maintainability. Define how proven Slurm implementations should ...

Founding GPU Engineer

Location
Greater London, England, United Kingdom
node scaling using NCCL, MPI, or similar communication libraries. Work with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensity. Collaborate with ML/systems engineers to integrate custom kernels into training/inference pipelines. … data center power/thermal management or demand-response systems. Background in HPC, quantitative finance, or large-scale distributed systems. Familiarity with Kubernetes/Slurm for GPU cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Competitive salary and an equity sign-on bonus. ...

Sr Lead Software Engineer - C++, Python,

Location
Greater London, England, United Kingdom
solutions using Python, C++; and Go. Deploy and manage applications using Kubernetes (K8s) and HPC/Grid Computing/Big Data technologies such as Slurm, LSF, Spark, Ray and Symphony in cloud environments. Required Qualifications, Capabilities, And Skills Formal training or certification on Software Engineering concepts and 5+ years … applied experience. Expertise in C++, Python, Kubernetes (K8s), and HPC/Grid Computing technologies (Slurm, LSF, Symphony), GenAI, agent framework and tools, observability tooling such as OTel Proven experience with cloud technologies and environments Strong problem-solving skills and the ability to work collaboratively in a team setting. Experience ...

Principal - AI & HPC Data Centre Compute

Location
Greater London, England, United Kingdom
inference workloads (LLMs, multimodal and scientific AI) across distributed clusters for cost, throughput and latency targets Lead HPC cluster design and orchestration using Slurm, Kubernetes and parallel processing models (MPI) Enhance GPU utilisation by addressing bottlenecks across compute, memory and data pipelines; collaborate with energy teams on power-aware … infrastructure, accelerated computing or distributed systems architecture Deep knowledge of GPU architectures, AI workloads, networking and large-scale cluster operations Hands-on expertise with Slurm, Kubernetes and performance optimisation across multi-node environments Proven ability to advise senior stakeholders and influence technical strategy at C-level Demonstrated experience ...

Research Software Engineer

Location
Greater London, England, United Kingdom
surface live metrics. Write efficient, well-tested Python and systems code; enforce code review, CI, and observability. Design and optimise distributed services (Kubernetes/SLURM, thousands-of-GPU jobs). Prototype utilities (CLI, dashboards) and carry them through to stable, shared libraries. About the Research Engineering team Based … observability. Fluency in Python plus one systems language (C++, Rust, Go or Java). Hands-on with container orchestration and schedulers (Kubernetes/K8s, SLURM, or similar). Comfortable profiling performance, optimising I/O, and automating workflows. Self-starter, low-ego, collaborative, high-energy. Nice-to-haves Exposure ...

HPC Operations Lead

Hiring Organisation
LinuxRecruit
Location
London, United Kingdom
operation of high performance storage services, supporting both internal workloads and external collaboration. The environment includes large scale HPC clusters, Linux based systems, workload schedulers such as Slurm, networking with Infiniband and parallel file systems such as GPFS. Experience with high performance storage at petabyte scale is particularly ...

Senior Linux & HPC Systems Engineer

Location
Leatherhead, England, United Kingdom
Inc. seeks a Senior IT Systems Administrator (Linux & HPC) to own and operate Linux-based enterprise and HPC platforms, including SLURM scheduling, Nvidia Base Command Manager, and CycleCloud. You will work with engineering and scientific users to diagnose issues spanning compute, storage, and networking. The role emphasizes hands ...

Senior Manager – Counterparty Credit Risk & XVA

Location
Greater London, England, United Kingdom
Your role As a Senior Manager in Counterparty Credit Risk (CCR) and XVA at Zanders, you will join our global Financial Institutions team in London. Your remit is to lead quantitative traded risk engagements across CCR, XVA and the high-performance computing (HPC) that underpins them, working in multidisciplinary … sales in the CCR, XVA and HPC space. You can build on your existing UK network and expand it over time. As a Senior Manager you act as a career coach, responsible for the development of up to three direct reports, supporting them on project work and across their ...

HPC Architect: Storage & Infrastructure Lead (Hybrid)

Location
Stevenage, England, United Kingdom
will lead architecture, standards, and collaboration with engineering teams across national and international sites to deliver scalable, high-throughput infrastructure. You will drive workload strategies with SLURM, Kubernetes for HPC, MPI/CUDA stacks, and ensure reproducibility and portability of workloads while partnering with MBDA vendors #J ...

Director Customer Experience and Engagement

Location
Trellech, Wales, United Kingdom
function. This role will be pivotal in working with our customers across the AI lifecycle—from initial requirements gathering to contractual acceptance, onboarding and workload migration through production adoption, optimization, and expansion. This role sits at the intersection of customers, engineering, product, sales, and infrastructure. You will be accountable … segments to determine best path for each customer and increase deal velocity and efficiency Establish account-health frameworks that combine product usage, platform reliability, workload performance, customer sentiment, support activity, and commercial risk. Proactively identify at-risk enterprise accounts based on utilization patterns, support tickets, or competitive pressures ...