101 to 108 of 108 CUDA Jobs in London

AI Inference Engineer

Location
Greater London, England, United Kingdom
demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch -- and we're looking for the founding engineer … Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered ...

AI Inference Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch, and we're looking for the founding engineer to own the latter. … serving, deciding where and how to apply quantisation, distillation, speculative decoding and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineersMake the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)Translate ...

Founding GPU Engineer — CUDA Performance for HPC

Location
Greater London, England, United Kingdom
Fuse Energy, LLC seeks an experienced CUDA performance engineer to design, optimize, and deploy high-throughput CUDA kernels across multi-GPU systems. You will profile hardware bottlenecks, develop tooling to monitor energy usage, and collaborate with ML engineers to integrate kernels into scalable training and inference pipelines. … role requires deep GPU architecture knowledge, strong C++/CUDA skills, and experience with Nsight, NCCL, and MPI. #J-18808-Ljbffr ...

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
about efficient large-scale training, low-latency inference in real-time systems and high-throughput inference in research. Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking and host- and GPU-level considerations. Zooming in, we also want … level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores and the memory hierarchy Debugging and optimisation experience using tools like CUDA GDB, NSight Systems, NSight Computesight-systems and nsight-compute Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN and cuBLAS Intuition about the latency and throughput characteristics ...

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
about efficient large-scale training, low-latency inference in real-time systems and high-throughput inference in research. Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking and host- and GPU-level considerations. Zooming in, we also want … end. Low-level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores and the memory hierarchy. Debugging and optimisation experience using tools like CUDA GDB, NSight Systems, NSight Compute-sight-systems and nsight-compute. Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN and cuBLAS. Intuition about the latency ...

ML Systems Performance Engineer

Location
Greater London, England, United Kingdom
Quant Blueprint LLC is seeking an engineer proficient in low-level systems programming and optimization to enhance our machine learning team. This role centers on optimizing model performance, both for training and real-time inference. ...

Senior CUDA Engineer - High-Performance GPU Inference

Location
Greater London, England, United Kingdom
Fuse Energy is seeking a CUDA Engineer to design and optimize low-level GPU kernels powering our transformer inference workloads. You will write custom CUDA kernels, tune memory bandwidth, and push the throughput of our GPU fleet at the SM, warp, and memory hierarchy level. Collaborate with ...

Senior CUDA Engineer for High-Performance Inference

Location
Greater London, England, United Kingdom
Fuse Energy is seeking an experienced CUDA Engineer to design and optimise low‐level GPU code for high‐throughput inference workloads. You will work on custom CUDA kernels, memory hierarchies, and mixed‐precision arithmetic to maximise throughput and minimise latency across GPU fleets. The role focuses on profiling ...