101 to 104 of 104 CUDA Jobs in London

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
about efficient large-scale training, low-latency inference in real-time systems and high-throughput inference in research. Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking and host- and GPU-level considerations. Zooming in, we also want … end. Low-level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores and the memory hierarchy. Debugging and optimisation experience using tools like CUDA GDB, NSight Systems, NSight Compute-sight-systems and nsight-compute. Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN and cuBLAS. Intuition about the latency ...

ML Systems Performance Engineer

Location
Greater London, England, United Kingdom
Quant Blueprint LLC is seeking an engineer proficient in low-level systems programming and optimization to enhance our machine learning team. This role centers on optimizing model performance, both for training and real-time inference. ...

Senior CUDA Engineer - High-Performance GPU Inference

Location
Greater London, England, United Kingdom
Fuse Energy is seeking a CUDA Engineer to design and optimize low-level GPU kernels powering our transformer inference workloads. You will write custom CUDA kernels, tune memory bandwidth, and push the throughput of our GPU fleet at the SM, warp, and memory hierarchy level. Collaborate with ...

Senior CUDA Engineer for High-Performance Inference

Location
Greater London, England, United Kingdom
Fuse Energy is seeking an experienced CUDA Engineer to design and optimise low‐level GPU code for high‐throughput inference workloads. You will work on custom CUDA kernels, memory hierarchies, and mixed‐precision arithmetic to maximise throughput and minimise latency across GPU fleets. The role focuses on profiling ...