26 to 32 of 32 Slurm Workload Manager Jobs in London

Founding GPU Engineer

Location
Greater London, England, United Kingdom
node scaling using NCCL, MPI, or similar communication libraries. Work with data center infrastructure teams on power capping, dynamic voltage/frequency scaling, and workload scheduling strategies that reduce energy cost and carbon intensity. Collaborate with ML/systems engineers to integrate custom kernels into training/inference pipelines. … data center power/thermal management or demand-response systems. Background in HPC, quantitative finance, or large-scale distributed systems. Familiarity with Kubernetes/Slurm for GPU cluster orchestration. Interest or experience in energy markets, grid systems, or sustainability-focused compute. Competitive salary and an equity sign-on bonus. ...

Sr Lead Software Engineer - C++, Python,

Location
Greater London, England, United Kingdom
solutions using Python, C++; and Go. Deploy and manage applications using Kubernetes (K8s) and HPC/Grid Computing/Big Data technologies such as Slurm, LSF, Spark, Ray and Symphony in cloud environments. Required Qualifications, Capabilities, And Skills Formal training or certification on Software Engineering concepts and 5+ years … applied experience. Expertise in C++, Python, Kubernetes (K8s), and HPC/Grid Computing technologies (Slurm, LSF, Symphony), GenAI, agent framework and tools, observability tooling such as OTel Proven experience with cloud technologies and environments Strong problem-solving skills and the ability to work collaboratively in a team setting. Experience ...

Principal - AI & HPC Data Centre Compute

Location
Greater London, England, United Kingdom
inference workloads (LLMs, multimodal and scientific AI) across distributed clusters for cost, throughput and latency targets Lead HPC cluster design and orchestration using Slurm, Kubernetes and parallel processing models (MPI) Enhance GPU utilisation by addressing bottlenecks across compute, memory and data pipelines; collaborate with energy teams on power-aware … infrastructure, accelerated computing or distributed systems architecture Deep knowledge of GPU architectures, AI workloads, networking and large-scale cluster operations Hands-on expertise with Slurm, Kubernetes and performance optimisation across multi-node environments Proven ability to advise senior stakeholders and influence technical strategy at C-level Demonstrated experience ...

Research Software Engineer

Location
Greater London, England, United Kingdom
surface live metrics. Write efficient, well-tested Python and systems code; enforce code review, CI, and observability. Design and optimise distributed services (Kubernetes/SLURM, thousands-of-GPU jobs). Prototype utilities (CLI, dashboards) and carry them through to stable, shared libraries. About the Research Engineering team Based … observability. Fluency in Python plus one systems language (C++, Rust, Go or Java). Hands-on with container orchestration and schedulers (Kubernetes/K8s, SLURM, or similar). Comfortable profiling performance, optimising I/O, and automating workflows. Self-starter, low-ego, collaborative, high-energy. Nice-to-haves Exposure ...

Senior Manager – Counterparty Credit Risk & XVA

Location
Greater London, England, United Kingdom
Your role As a Senior Manager in Counterparty Credit Risk (CCR) and XVA at Zanders, you will join our global Financial Institutions team in London. Your remit is to lead quantitative traded risk engagements across CCR, XVA and the high-performance computing (HPC) that underpins them, working in multidisciplinary … sales in the CCR, XVA and HPC space. You can build on your existing UK network and expand it over time. As a Senior Manager you act as a career coach, responsible for the development of up to three direct reports, supporting them on project work and across their ...

Senior Software Engineer - Research Technology

Location
Greater London, England, United Kingdom
fundamentals: data structures, algorithms, networking, OS, concurrency, and system design. Experience running compute at cluster scale: job scheduling, resource management, retries, and reliability. Slurm, Kubernetes, Ray, Spark, or custom internal schedulers all count. Proven data‐engineering experience: schema design, storage formats, compression, I/O trade-offs, and pipelines … ship production software safely and repeatedly, with an obsession for data driven quality. Desirable/nice-to-have Rust experience alongside C++ and Python. Slurm or other cluster scheduler expertise. Familiarity with ML/Deep Learning frameworks. Prior finance or market‐data experience, including low‐level market connectivity. ...

Staff HPC Systems Engineer - Slurm & Cloud Platform Lead

Location
Greater London, England, United Kingdom
London is hiring a Staff HPC Systems Software Engineer to define the technical direction of a core HPC platform. You will scope architecture for Slurm-based services, shaping how multiple teams build, automate, and run the platform in a cloud-native environment. You’ll work across engineering boundaries ...