76 to 100 of 104 CUDA Jobs in London

Founding GPU Engineer - Equity & Biannual Bonus

Location
Greater London, England, United Kingdom
Fuse Energy is seeking a seasoned CUDA performance engineer to design and optimise kernels for high-throughput workloads in a high-performance compute environment. You will profile GPU bottlenecks, build tooling to align power draw with energy pricing, and optimise multi-GPU scaling across NCCL/MPI. Collaboration with ...

Senior MLOps Engineer

Location
Greater London, England, United Kingdom
processing jobs, pipelines, model registry, and endpoints Practical MLOps experience with MLflow, hyperparameter optimisation, and reproducible, configuration-driven training runs Experience with Docker and CUDA-based GPU images Terraform experience sufficient to ship models through code review Strong software engineering fundamentals, including Git, pull-request workflow, automated testing … alerting mechanisms. Highest-signal resume keywords Expert Python AWS SageMaker MLOps Experience PyTorch Terraform ATS Optimization Keywords Hard Skills Python PyTorch TensorFlow MLflow Docker CUDA Git Automated Testing CI/CD Event-Driven Architectures Soft Skills Clear Communication Industry Keywords Healthcare Medical Devices Regulated Industry Tools & Technologies SageMaker Hugging ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
reasoning and clear communication Debug ML Infrastructure & Distributed Workloads Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems Troubleshoot PyTorch, CUDA, NCCL, and inference serving related issues Analyze logs, metrics, traces, and system behavior to isolate root causes Debug containerized workloads running across Kubernetes … Infrastructure Experience Hands on experience operating machine learning workloads in production or research environments Experience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL Familiarity with GPU infrastructure and orchestration Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure Understanding of the operational challenges involved ...

Research Software Engineer

Location
Greater London, England, United Kingdom
About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and ...

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
We tackle the most complex problems in quantitative finance, by bringing scientific clarity to financial complexity. From our London HQ, we unite world-class researchers and engineers in an environment that values deep exploration and ...

Member of Technical Staff (AI Inference Engineer)

Location
Greater London, England, United Kingdom
engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL.### ### **Responsibilities:*** **New models support.** Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading … request scheduling and KV-cache management to support in API Gateway.* **GPU kernels migration to CuTe DSL.** Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.* **Rust-native serving runtime.** Develop our internal ...

ML Compiler & System Engineering & Research London

Location
Greater London, England, United Kingdom
Quadrants is an open-source translation layer that turns ordinary Python into optimized machine code at runtime, across CPU (Arm64, x86) and GPU (AMDGPU, CUDA, Metal, Vulkan) backends. CUDA is the primary target, with x86 the strong runner-up, especially for real-time simulation. The others drive adoption ...

Software Engineering Manager

Location
Greater London, England, United Kingdom
This role is based in our London (Kings Cross) office. Responsibilities: Manage and mentor a team of software engineers working across embedded C++/CUDA, GUI, and embedded linux, setting priorities, individual development plans, conducting performance reviews, and maintaining a high standard of engineering practice Own the technical roadmap … product software stack, from the C++/CUDA host application and real-time processing pipeline, aligning with product and system requirements Lead design and architecture reviews across the software stack, ensuring quality, consistency, and sound technical decisions Drive software from prototype through to production release, managing dependencies with ...

Genesis-World: Core Simulation Engine Engineer

Location
Greater London, England, United Kingdom
hack or compromise. Everything is Python-first and runs anywhere. Kernels are written once, and Quadrants, our in-house JIT compiler, lowers them to CUDA, AMD ROCm, Apple Metal, Vulkan, x86, and ARM64. A single laptop or a datacenter. Massively batched GPU simulation for learning at scale, and complex … cost. You will work closely with Quadrants. At ease debugging across the full platform matrix: Windows, Linux, and macOS, on x86 and arm64, over CUDA, AMD, and Apple Metal (and possibly Intel XPU via Vulkan soon) Open-source reflexes: triaging issues, reviewing external pull requests, communicating with a community. ...

Senior Research Engineer

Hiring Organisation
European Tech Recruit
Location
City of London, London, United Kingdom
Senior Research Engineer Cambridge or London, UK (100% Onsite) Our client is a global leader in semiconductor innovation, developing advanced technologies that power billions of smart devices worldwide. With significant investment in artificial intelligence research ...

Staff Robotics Engineer, Localisation

Location
Greater London, England, United Kingdom
About us Founded in 2017, Wayve is the leading developer of Embodied AI technology. Our advanced AI software and foundation models enable vehicles to perceive, understand, and navigate any complex environment, enhancing the usability and ...

Senior Solutions Architect, Higher Education and Research - Open Models and LLM

Location
Greater London, England, United Kingdom
good understanding of scientific policy engagement, grant processes, or national/European research program structures. Experience with NVIDIA's AI software stack, powered by CUDA and CUDA-X libraries, e.g. NVIDIA AI Enterprise, NeMo Framework, Megatron Bridge, NIM, TensorRT-LLM, Dynamo, NeMo Agent Toolkit, and Triton Inference Server ...

AI Research Engineer, Pre-Training

Hiring Organisation
Hudson River Trading
Location
London, UK
Employment Type
Full-time
business, and it will be challenging: this is a field with no easy or obvious solutions. QualificationsStrong engineering skills, especially any of: CUDA/Triton/Pallas/CuTe DSL kernel development, lower-level PyTorch/JAX/XLA development, CUDA Graphs, FPGA/ASIC experienceMust have ...

AI Research Engineer, Inference

Hiring Organisation
Hudson River Trading
Location
London, UK
Employment Type
Full-time
business, and it will be challenging: this is a field with no easy or obvious solutions. QualificationsStrong engineering skills, especially any of: CUDA/Triton/Pallas/CuTe DSL kernel development, lower-level PyTorch/JAX/XLA development, CUDA Graphs, FPGA/ASIC experienceMust have ...

Sr Biostatistician (R a must- EMEA BASED)

Hiring Organisation
Syneos Health
Location
London, UK
Employment Type
Full-time
experience in clinical data structures and programming with dataexpert in functional and object-oriented programming. Knowledgeable in Javascript/Typescript, HTML, WebGL, experience in CUDA/GPU-programing, cloud-computing, Github, web-hosting/Machine Learning. Strong communication skills and ability to work both, independently and collaboratively, clear … clinical data structures and programming with data -Expert in functional and object-oriented programming knowledgeable in Javascript/Typescript, HTML, WebGL -Experience in CUDA/GPU-programing, cloud-computing, Github, web-hosting -Strong communication skills and ability to work both, independently and collaboratively, clear in the presentation of complex ...

Campus ML Engineer: Build Scalable AI for Finance

Location
Greater London, England, United Kingdom
Jump Trading Group is seeking world-class engineers to collaborate with our research, trading and engineering teams to build state-of-the-art ML systems for quantitative finance. You will work on training pipelines, low ...

Founding GPU Engineer

Location
Greater London, England, United Kingdom
operate in significantly outpaces what we can currently build, meaning speed to power and reliability are critical to how we scale. This puts CUDA/GPU performance engineering at the center of how Fuse scales its compute infrastructure. Responsibilities Design, implement, and optimise CUDA kernels for high-throughput … baselines and drive continuous performance improvements. Contribute to internal libraries, documentation, and best practices for GPU performance engineering. 4+ years of experience writing production CUDA code, or equivalent strong project/industry experience. Deep understanding of GPU architecture (SMs, warps, memory hierarchy, occupancy). Proficiency in C++ and CUDA ...

Training / AI Infrastructure Engineering & Research London

Location
Greater London, England, United Kingdom
Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high-level training frameworks Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory … Bring Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years) Production-grade expertise in Python Low-level performance mastery: CUDA/cuDNN/Triton, CPU-GPU interactions, data movement, and kernel optimization Scaling at the frontier: experience with PyTorch and training jobs using data, context ...

CUDA Engineer

Location
Greater London, England, United Kingdom
optimising how power‐dense GPU workloads are scheduled, cooled, and balanced against grid conditions in real time. We're looking for a CUDA Engineer to write and optimise the low‐level GPU code that powers our inference workloads. You'll design custom CUDA kernels, tune performance across memory … operate in significantly outpaces what we can currently build, meaning speed to power and reliability are critical to how we scale. This puts CUDA/GPU performance engineering at the center of how Fuse scales its compute infrastructure. Responsibilities Write and optimise custom CUDA kernels for core transformer ...

CUDA Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, UK
Employment Type
Full-time
sources of electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're looking for a CUDA Engineer to write and optimise the low-level GPU code that powers our inference workloads: designing custom CUDA kernels, tuning performance across memory … squeezing maximum throughput out of every GPU in our fleet, working at the level of SMs, warps and memory hierarchies. ResponsibilitiesWrite and optimise custom CUDA kernels for core transformer inference operationsProfile kernels to identify and eliminate bottlenecks in occupancy, memory throughput and warp divergenceApply kernel fusion to reduce memory ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch -- and we're looking for the founding engineer … Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered ...

AI Inference Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, UK
Employment Type
Full-time
electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch, and we're looking for the founding engineer to own the latter. … serving, deciding where and how to apply quantisation, distillation, speculative decoding and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineersMake the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)Translate ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch, and we're looking for the founding engineer to own the latter. … serving, deciding where and how to apply quantisation, distillation, speculative decoding and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents ...

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
about efficient large-scale training, low-latency inference in real-time systems and high-throughput inference in research. Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking and host- and GPU-level considerations. Zooming in, we also want … level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores and the memory hierarchy Debugging and optimisation experience using tools like CUDA GDB, NSight Systems, NSight Computesight-systems and nsight-compute Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN and cuBLAS Intuition about the latency and throughput characteristics ...