126 to 140 of 140 Permanent CUDA Jobs

Senior Solutions Architect, Higher Education and Research, Multimodal and Physical AI

Location
Reading, England, United Kingdom
scientific policy engagement, grant processes, or national/European research program structures. Experience with NVIDIA's stack for visual and multimodal AI, built on CUDA and CUDA-X, including Cosmos world foundation models, NeMo Framework and Megatron for multimodal model training, multimodal Nemotron methodology, Isaac robotics platform, NuRec ...

Senior Embedded Imaging Software Engineer

Hiring Organisation
Urban Sky
Location
Denver, Colorado, United States
Employment Type
Permanent
Salary
USD Annual
plane to the command-and-control link on the ground. You'll write the firmware that drives our cameras and edge compute, own the CUDA image processing pipeline that turns raw pixels into clean, calibrated, geo-tagged imagery, and architect the high-rate data paths and C2 infrastructure that … develop board-support packages for custom carrier boards, system-on-modules, and companion microcontrollers Own the L0 image processing pipeline running in CUDA at the edge: debayering, white balance, gamma, and camera/lens correction - applied to raw frames before any codec Own IR-specific processing: non-uniformity correction ...

Campus ML Engineer: Build Scalable AI for Finance

Location
Greater London, England, United Kingdom
Jump Trading Group is seeking world-class engineers to collaborate with our research, trading and engineering teams to build state-of-the-art ML systems for quantitative finance. You will work on training pipelines, low ...

Founding GPU Engineer

Location
Greater London, England, United Kingdom
operate in significantly outpaces what we can currently build, meaning speed to power and reliability are critical to how we scale. This puts CUDA/GPU performance engineering at the center of how Fuse scales its compute infrastructure. Responsibilities Design, implement, and optimise CUDA kernels for high-throughput … baselines and drive continuous performance improvements. Contribute to internal libraries, documentation, and best practices for GPU performance engineering. 4+ years of experience writing production CUDA code, or equivalent strong project/industry experience. Deep understanding of GPU architecture (SMs, warps, memory hierarchy, occupancy). Proficiency in C++ and CUDA ...

Training / AI Infrastructure Engineering & Research London

Location
Greater London, England, United Kingdom
Design, build, and optimize distributed training systems (PyTorch) for multi-node GPU clusters, ensuring scalability, robustness, and high utilization Implement efficient low-level code (CUDA, cuDNN, Triton, custom kernels) and integrate it seamlessly into high-level training frameworks Optimize workloads for hardware efficiency: CPU/GPU compute balance, memory … Bring Deep experience in distributed systems, ML infrastructure, or high-performance computing (8+ years) Production-grade expertise in Python Low-level performance mastery: CUDA/cuDNN/Triton, CPU-GPU interactions, data movement, and kernel optimization Scaling at the frontier: experience with PyTorch and training jobs using data, context ...

CUDA Engineer

Location
Greater London, England, United Kingdom
optimising how power‐dense GPU workloads are scheduled, cooled, and balanced against grid conditions in real time. We're looking for a CUDA Engineer to write and optimise the low‐level GPU code that powers our inference workloads. You'll design custom CUDA kernels, tune performance across memory … operate in significantly outpaces what we can currently build, meaning speed to power and reliability are critical to how we scale. This puts CUDA/GPU performance engineering at the center of how Fuse scales its compute infrastructure. Responsibilities Write and optimise custom CUDA kernels for core transformer ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch -- and we're looking for the founding engineer … Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch, and we're looking for the founding engineer to own the latter. … serving, deciding where and how to apply quantisation, distillation, speculative decoding and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineers Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents ...

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
about efficient large-scale training, low-latency inference in real-time systems and high-throughput inference in research. Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking and host- and GPU-level considerations. Zooming in, we also want … level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores and the memory hierarchy Debugging and optimisation experience using tools like CUDA GDB, NSight Systems, NSight Computesight-systems and nsight-compute Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN and cuBLAS Intuition about the latency and throughput characteristics ...

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
about efficient large-scale training, low-latency inference in real-time systems and high-throughput inference in research. Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking and host- and GPU-level considerations. Zooming in, we also want … end. Low-level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores and the memory hierarchy. Debugging and optimisation experience using tools like CUDA GDB, NSight Systems, NSight Compute-sight-systems and nsight-compute. Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN and cuBLAS. Intuition about the latency ...

ML Systems Performance Engineer

Location
Greater London, England, United Kingdom
Quant Blueprint LLC is seeking an engineer proficient in low-level systems programming and optimization to enhance our machine learning team. This role centers on optimizing model performance, both for training and real-time inference. ...

LLM/GenAI MLOps Engineer - Platform Architect

Location
Sheffield, England, United Kingdom
慨正橡扯 is seeking an MLOps Engineer (LLM/GenAI) based in Sheffield, UK, to engineer production-grade infrastructure for modern AI. The ideal candidate will design scalable model hosting platforms, optimise inference performance, and build ...

Senior CUDA Engineer - High-Performance GPU Inference

Location
Greater London, England, United Kingdom
Fuse Energy is seeking a CUDA Engineer to design and optimize low-level GPU kernels powering our transformer inference workloads. You will write custom CUDA kernels, tune memory bandwidth, and push the throughput of our GPU fleet at the SM, warp, and memory hierarchy level. Collaborate with ...

Senior CUDA Engineer for High-Performance Inference

Location
Greater London, England, United Kingdom
Fuse Energy is seeking an experienced CUDA Engineer to design and optimise low‐level GPU code for high‐throughput inference workloads. You will work on custom CUDA kernels, memory hierarchies, and mixed‐precision arithmetic to maximise throughput and minimise latency across GPU fleets. The role focuses on profiling ...