26 to 50 of 63 Slurm Workload Manager Jobs in the UK

AI Inference Engineer

Location
Greater London, England, United Kingdom
scale inference; multi-tenant serving or SLA-driven infrastructure; background at a hyperscaler, frontier AI lab or large-scale distributed inference system; Kubernetes/Slurm; interest in energy markets, grid systems or sustainability-focused compute Competitive salary and eligibility for equity Biannual bonus scheme Fully expensed tech to match ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
multi-tenant serving or SLA-driven infrastructure. Background at a hyperscaler, frontier AI lab, or large-scale distributed inference system. Familiarity with Kubernetes/Slurm for cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Benefits Competitive salary and an equity sign-on bonus. ...

Senior Hybrid HPC Engineer — Linux, Slurm & GPU

Location
Stevenage, England, United Kingdom
Gazelle Global seeks a Senior HPC Engineer to join a major scientific computing environment on a 12-month contract in Stevenage. The role is hands-on, focusing on maintaining, securing and optimising Linux-based HPC ...

Senior HPC Linux Engineer - Onsite, Slurm/PBS Platform Lead

Location
England, United Kingdom
Unknown is seeking an experienced HPC Engineer to join an elite engineering organisation in the East Midlands. This onsite role requires taking ownership of the HPC platform, improving services, and delivering best-in-class HPC ...

Senior HPC Engineer - Hybrid GPU Linux Clusters

Location
Stevenage, England, United Kingdom
researchers and technical teams. You will manage HPC clusters, ensure reliability and performance, and support scientific applications and complex workloads, including GPU computing and Slurm workload management. Hybrid onsite presence is required. #J-18808-Ljbffr ...

HPC & AI Platform Engineer (GPU/Networking)

Location
United Kingdom
with vendor engineering teams to ensure seamless AI platform operations. You will design, deploy and manage large-scale GPU-accelerated clusters using NVIDIA GPUs, Slurm, InfiniBand and high-availability practices, while automating provisioning, monitoring, and security. #J-18808-Ljbffr ...

Platform Engineer

Location
United Kingdom
managing large‐scale HPC and GPU‐accelerated clusters, including NVIDIA based compute environments. Implementing and administering HPC scheduling and resource‐management systems (e.g., Slurm), including GPU partitioning, workload scheduling, and capacity planning. Architecting and optimising InfiniBand and Ethernet network topologies. Ensuring high availability and resilience through failover strategies ...

Senior Cloud Engineer

Location
West of England, England, United Kingdom
platform s. Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Programming experience with Python3 utilising classes and inheritance. Benefits In addition to a competitive salary, Graphcore offers flexible ...

Senior Cloud Engineer (K8S)

Location
Greater London, England, United Kingdom
platform s. Experience with solutions for monitoring and observability. e.g. Grafana, Prometheus, OpenSearch/ElasticSearch, Loki. Experience with High Performance Computing (HPC) environments using SLURM or similar batch workload solutions. Programming experience with Python3 utilising classes and inheritance. Benefits In addition to a competitive salary flexible working ...

Senior AI Platform Engineer

Location
Greater London, England, United Kingdom
engineering solutions. Partner with centralised infrastructure teams to design and deliver high-performance compute environments across AWS and on-premises platforms, including GPU infrastructure, Slurm clusters, and migration from ad hoc research workflows. Optimise LLM training and inference workloads, supporting research and product teams in maximising performance, scalability … Nsight, DCGM, and related ecosystem technologies. Strong background in AWS cloud services, high-performance computing, distributed systems, containerised environments, and infrastructure automation. Experience with workload orchestration technologies such as Slurm, Kubernetes, Ray, or equivalent distributed compute frameworks. Demonstrated success bridging research and production environments, enabling rapid experimentation while ...

Build Engineer

Location
Greater London, England, United Kingdom
hybrid working model, required onsite 3 days a week. Experience for the Build Engineer includes: Software development with Python Modern build systems, e.g. Bazel Workload management, e.g. Slurm, LSF or SGE Infrastructure as code or IAC, e.g. Ansible or Terraform Containerisation, e.g. Docker Desired: Experience supporting silicon ...

Platform Architect - Nvidia AI/GB300

Hiring Organisation
Oscar Associates (UK) Limited
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
£700 - £765 per day
data-centre teams. Key Requirements Strong Platform/Infrastructure Architecture experience across compute, storage, networking and Linux. Expert-level Kubernetes architecture experience. Strong Slurm and Run experience - essential. Proven experience with GPU/HPC environments and large-scale AI platforms. Hands-on experience with NVIDIA HGX GB300/NVL72 … including NVLink, NVSwitch and Grace Blackwell architecture. Experience with NVIDIA RTX 6000 series GPU servers. Strong understanding of GPU workload scheduling, partitioning and sharing, including MIG, vGPU and time-slicing Strong understanding of InfiniBand, RoCE, Spectrum-X, GPUDirect RDMA/Storage and high-performance AI fabrics. Experience with Terraform ...

Senior HPC Engineer - Hybrid - Inside IR35

Location
Stevenage, England, United Kingdom
Role We are looking for an experienced Senior HPC Engineer to support and maintain a scientific computing environment, with a strong focus on RHEL, Slurm and HPC infrastructure . You will work closely with research scientists and technical teams to ensure HPC services are secure, reliable and high performing. … Responsibilities Administer, patch and maintain RHEL 7, 8 and 9 across HPC clusters and workstations. Deploy, configure and manage Slurm , including queues, partitions and scheduling. Monitor cluster health, performance, storage, networking and resource utilisation. Install and support scientific applications, compilers, libraries and MPI environments. Work with scientists to optimise ...

R&D Solution Architect

Location
Hook, England, United Kingdom
implementing solutions that adhere to FAIR data principles (Findable, Accessible, Interoperable, Reusable)." - Experience architecting for High-Performance Computing (HPC) environments, including knowledge of workload schedulers (e.g., SLURM) and applying cloud-native patterns to scientific, batch-processing workloads. - Familiarity with scientific workflow management tools (e.g., Nextflow, Snakemake ...

Linux/RHEL Engineer - HPC

Location
Stevenage, England, United Kingdom
Senior Linux/HPC Engineer to support a scientific computing environment in Stevenage. This hands-on role covers Linux infrastructure, HPC clusters, workload scheduling and scientific application support, working closely with research and technical teams. You'll keep the environment secure, stable and performing effectively, while troubleshooting infrastructure … application issues and improving cluster utilisation and job throughput. Key Skills: Strong RHEL 7, 8 and 9 administration and troubleshooting HPC cluster management and Slurm configuration Scientific or research application support in Linux HPC environments MPI, compilers and scientific libraries Hardware, OS, scheduler and application troubleshooting ServiceNow or equivalent ...

Senior Solution Engineer – GPU & AI Infrastructure

Location
United Kingdom
role, you will bridge the gap between customer business objectives and ultra-high-performance hardware execution. You will lead technical engagements, translate complex AI workload requirements into production-ready High-Level Designs (HLD), Low-Level Designs (LLD), and detailed Bills of Materials (BOM). Your expertise will span bare …/Spectrum-4) with lossless Ethernet mechanisms (PFC, ECN, Adaptive Routing). Multi-Tenant & Deployment Models: Deliver tailored architectures for both Bare-Metal (Slurm, OpenMPI, bare-metal provisioning) and Cloud-Native/Kubernetes environments (NVIDIA GPU Operator, Network Operator, Run:ai, KubeFlow). Storage Integration: Architect high-bandwidth parallel ...

Senior Solution Engineer, GPU & AI Infrastructure

Location
United Kingdom
engineering leads Architect and oversee proof-of-concept deployments, benchmarking with tools such as NCCL tests, GPUDirect RDMA and MLPerf to validate real-world workload performance Experience Required (NOT ALL ESSENTIAL): Experience: 5+ years in Solution Architecture, Systems Engineering or Technical Pre-Sales, focused on high-performance cloud … X800, Adaptive Routing) and RoCE/RoCEv2 (Spectrum-X/Spectrum-4), plus Kubernetes orchestration (NVIDIA GPU Operator, MPI Operator) and bare-metal tooling (Slurm, Ansible, Terraform) Soft Skills: Strong technical leadership and presentation skills, able to translate complex hardware and network trade-offs for executive stakeholders; strong spoken ...

Solutions Engineer

Location
Greater London, England, United Kingdom
engineering, and operations teams to ensure proposed solutions are realistic, scalable, and aligned with platform standards Provide guidance on compute, networking, storage, orchestration, and workload optimisation for AI and machine learning use cases Help create repeatable demo environments, technical playbooks, reference architectures, and sales enablement materials … data centre, or platform environments Good understanding of cloud infrastructure, GPU compute, AI/ML workloads, or high-performance infrastructure Familiarity with containers, Kubernetes, Slurm, orchestration platforms, or workload deployment models Understanding of networking, storage, and distributed compute concepts in modern infrastructure environments Ability to quickly learn ...

Staff Software Engineer, Kubernetes Platform

Hiring Organisation
Humanloop
Location
London, UK
Employment Type
Full-time
controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure. We make sure the control plane is fast, correct, and always … Anthropic's accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemptionScale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds usDesign, build, and operate core cluster services such ...

Python Software Engineer - Intraday Trading

Hiring Organisation
Millennium Management
Location
London, UK
Employment Type
Full-time
strong understanding of Linux operating systems Experience with grid scheduling and compute orchestration for real-time compute management at scale, including technologies such as SLURM Strong understanding of event-driven architecture and experience with messaging and caching technologies such as Kafka, Solace, Pulsar, Memcache, and Redis Experience building … scale real-time portfolio analytics tools, cloud platforms, containerization technologies such as Docker and Kubernetes, or multi-threaded C++ is a plusRecruiter: Ruby KazmiHiring Manager: Shashank GiriDepartment: Information Technology ...

Senior HPC Engineer

Location
West of England, England, United Kingdom
Linux system administration. Familiarity with at least one scripting language (e.g., Bash, Python). Interest in high-performance computing and willingness to learn Slurm, xCAT, and Ansible. Strong problem-solving skills and attention to detail. Good communication and collaboration skills. Ability to work independently with mentorship and as part … week. No hybrid/remote working option. Internship or academic experience in a research computing or HPC environment. Exposure to job schedulers (e.g., Slurm, LSF). Familiarity with version control systems (e.g., Git). Coursework or projects involving distributed systems, networking, or parallel computing. Understanding of basic cybersecurity concepts. ...

Senior HPC Engineer

Location
Stevenage, England, United Kingdom
supporting scientific applications and complex computational workloads. Key Skills: Strong hands-on administration of RHEL 7, 8 and 9 HPC cluster administration and troubleshooting Slurm deployment, configuration and workload management Scientific applications within Linux HPC environments MPI libraries and computational workloads Hardware, OS and scheduler troubleshooting ServiceNow ...

Senior Specialist Field Engineer - HPC/AI/ML

Location
Greater London, England, United Kingdom
functionality, and performance, contributing regularly to discussions about product strategy and architecture. Conduct periodic technical reviews and assessments of customer workloads, pinpointing opportunities for workload optimization and suggesting suitable solutions. Stay informed of the latest developments and trends in Kubernetes, cloud computing and infrastructure, sharing your thought leadership with … related technical discipline, or equivalent experience 7+ years of proven experience as a Solutions Architect, Field Engineer, Engineer, Researcher, or Technical Account Manager in Cloud Infrastructure, focusing on building distributed systems or HPC/cloud services, with an expertise focused on AI/ML inference Fluency in cloud computing ...

HPC Engineer

Location
Brixworth, England, United Kingdom
reliability, availability and performance of HPC platforms supporting simulation, analysis and engineering workloads Improve compute, storage, networking and scheduling services to enable efficient, scalable workload delivery Provide technical escalation for HPC incidents, capacity issues, performance bottlenecks and complex user problems Administering Linux-based HPC clusters, including compute nodes, schedulers … Translating technical user requirements into practical service improvements Have experience of... Supporting Linux-based HPC, scientific computing, simulation or high-throughput compute environments Diagnosing workload, queue, licence, performance, data movement and application issues Operating at a senior technical level in an enterprise or engineering-led environment Delivering maintenance, upgrades ...

Founding GPU Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, UK
Employment Type
Full-time
Engineer to develop and optimise GPU-accelerated software for data centre systems: low-level performance engineering for large-scale compute clusters, tying GPU workload behaviour to energy availability and grid demand. This puts CUDA/GPU performance engineering at the centre of how Fuse scales its compute infrastructure. ResponsibilitiesDesign … multi-node scaling using NCCL, MPI, or similar communication librariesWork with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensityCollaborate with ML/systems engineers to integrate custom kernels into training/inference pipelinesBenchmark against ...