13 of 13 Remote Slurm Workload Manager Jobs

Senior Software Engineer - Scientific / HPC

Hiring Organisation
Technical Futures Ltd
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£70,000 - £90,000 per annum
code quality. Some/most of the following should support the skills above: Experience of Cloud computing or HPC job management (such as Slurm). Identity and authorization flows such as OIDC/OAuth2. Deploying containerized services on Linux (such as Podman). Infrastructure as code (such as Ansible ...

Senior Infrastructure Engineer, Research Singapore

Location
Greater London, England, United Kingdom
Strong systems fundamentals: Linux, networking (including domain specific NVLink and InfiniBand), storage I/O, profiling and performance optimization Production experience with Kubernetes and SLURM for job orchestration on GPU clusters Proficiency in Python and ML frameworks (PyTorch strongly preferred) Experience with cloud GPU infrastructure; ideally CoreWeave or similar ...

Senior Applied Research Engineer - Video Team Synthesia Europe, Germany, Switzerland, UK

Location
United Kingdom
models Experience with GANs or VAEs Experience optimizing inference systems for production Our stack Python, PyTorch, CUDA DeepSpeed, distributed training & inference Sequence parallelism AWS, SLURM, Docker GitHub, CI/CD pipelines Who you are You are research-driven but outcome-focused You care about shipping, not just publishing ...

Senior Machine Learning Systems Engineer (Frameworks & Tooling)

Location
Greater London, England, United Kingdom
allowance A monthly quality time allowance A track record of building tools that increase developer velocity for ML teamsExperience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar)Bonus: paper at top-tier venues (such as NeurIPS, ICML, ICLR, AIStats, MLSys, JAX, AAAI, Nature, COLING, ACL, EMNLP)Experience ...

Senior Solutions Engineer

Hiring Organisation
LJB & Co
Location
City of London, London, United Kingdom
Employment Type
Contract, Work From Home
Contract Rate
From £1,000 to £1,200 per day
RoCEv2 networking. Develop infrastructure solutions using NVIDIA Blackwell, B300, GB300 and GB200 platforms. Design both bare-metal and Kubernetes-based GPU environments. Work with Slurm, Kubernetes, NVIDIA GPU Operator, NCCL and GPUDirect. Design high-performance storage solutions for AI workloads. Lead technical discussions with CTOs, AI leaders and infrastructure ...

Developer Experience Engineer

Hiring Organisation
Etched
Location
San Jose, California, United States
Employment Type
Permanent
Salary
USD Annual
cloud and on-prem hybrid environment. Key responsibilities Develop and maintain automation tools to streamline development, testing, and deployment workflows. Optimize and manage Slurm-based job scheduling for AI workloads, simulation, and chip design workflows. Build observability solutions using Grafana, Prometheus, and OpenTelemetry for monitoring pipelines, infrastructure, and compute … developer tooling and workflows. You may be a good fit if you have Strong Python skills for automation, scripting, and infrastructure development. Experience with Slurm job scheduling in an HPC or hybrid environment. Hands-on experience with observability and monitoring tools like Prometheus, Grafana, and OpenTelemetry. Expertise with Docker ...

Senior HPC Engineer - Hybrid - Inside IR35

Location
Stevenage, England, United Kingdom
Role We are looking for an experienced Senior HPC Engineer to support and maintain a scientific computing environment, with a strong focus on RHEL, Slurm and HPC infrastructure . You will work closely with research scientists and technical teams to ensure HPC services are secure, reliable and high performing. … Responsibilities Administer, patch and maintain RHEL 7, 8 and 9 across HPC clusters and workstations. Deploy, configure and manage Slurm , including queues, partitions and scheduling. Monitor cluster health, performance, storage, networking and resource utilisation. Install and support scientific applications, compilers, libraries and MPI environments. Work with scientists to optimise ...

R&D Solution Architect

Location
Hook, England, United Kingdom
implementing solutions that adhere to FAIR data principles (Findable, Accessible, Interoperable, Reusable)." - Experience architecting for High-Performance Computing (HPC) environments, including knowledge of workload schedulers (e.g., SLURM) and applying cloud-native patterns to scientific, batch-processing workloads. - Familiarity with scientific workflow management tools (e.g., Nextflow, Snakemake ...

Senior Solution Engineer – GPU & AI Infrastructure

Location
United Kingdom
role, you will bridge the gap between customer business objectives and ultra-high-performance hardware execution. You will lead technical engagements, translate complex AI workload requirements into production-ready High-Level Designs (HLD), Low-Level Designs (LLD), and detailed Bills of Materials (BOM). Your expertise will span bare …/Spectrum-4) with lossless Ethernet mechanisms (PFC, ECN, Adaptive Routing). Multi-Tenant & Deployment Models: Deliver tailored architectures for both Bare-Metal (Slurm, OpenMPI, bare-metal provisioning) and Cloud-Native/Kubernetes environments (NVIDIA GPU Operator, Network Operator, Run:ai, KubeFlow). Storage Integration: Architect high-bandwidth parallel ...

Senior Staff+ Software Engineer, Kubernetes Platform

Location
Greater London, England, United Kingdom
controllers — so it stays responsive as object counts and node counts grow by orders of magnitude. And we build the core cluster services every workload depends on, like service discovery, so they hold up under the same pressure. We make sure the control plane is fast, correct, and always … accelerator fleets, including custom scheduling plugins and policies for gang scheduling, topology awareness, and preemption Scale the Kubernetes control plane (apiserver, etcd, controller-manager) to support clusters far beyond typical limits, and find the next bottleneck before it finds us Design, build, and operate core cluster services such ...

Senior HPC Engineer

Hiring Organisation
Gazelle Global Consulting Ltd
Location
Stevenage, Hertfordshire, South East, United Kingdom
Employment Type
Contract, Work From Home
supporting scientific applications and complex computational workloads. Key Skills: Strong hands-on administration of RHEL 7, 8 and 9 HPC cluster administration and troubleshooting Slurm deployment, configuration and workload management Scientific applications within Linux HPC environments MPI libraries and computational workloads Hardware, OS and scheduler troubleshooting ServiceNow ...

HPC Platform Engineer

Hiring Organisation
Summa
Location
Houston, Texas, United States
Employment Type
Permanent
Salary
USD Annual
performance computing platforms used for complex engineering and scientific workloads. Responsibilities Administer and support HPC clusters in production environments Configure and tune job schedulers (Slurm preferred) Support distributed multi-node workloads (MPI) Install, upgrade, and maintain HPC infrastructure Configure and optimize parallel file systems (Lustre preferred) Support … equivalent experience 5+ years professional HPC experience Hands-on HPC cluster administration experience Linux systems administration experience (production environments) Experience configuring job schedulers (Slurm preferred) Experience supporting multi-node distributed workloads (MPI) Experience with HPC infrastructure components: Parallel file systems (Lustre preferred) High-speed interconnects GPU/accelerated computing ...

HPC Platform Engineer

Hiring Organisation
Summa
Location
The Woodlands, Texas, United States
Employment Type
Permanent
Salary
USD Annual
performance computing platforms used for complex engineering and scientific workloads. Responsibilities Administer and support HPC clusters in production environments Configure and tune job schedulers (Slurm preferred) Support distributed multi-node workloads (MPI) Install, upgrade, and maintain HPC infrastructure Configure and optimize parallel file systems (Lustre preferred) Support … equivalent experience 5+ years professional HPC experience Hands-on HPC cluster administration experience Linux systems administration experience (production environments) Experience configuring job schedulers (Slurm preferred) Experience supporting multi-node distributed workloads (MPI) Experience with HPC infrastructure components: Parallel file systems (Lustre preferred) High-speed interconnects GPU/accelerated computing ...