101 to 108 of 108 Permanent CUDA Jobs

Software Engineering Manager

Location
Greater London, England, United Kingdom
This role is based in our London (Kings Cross) office. Responsibilities: Manage and mentor a team of software engineers working across embedded C++/CUDA, GUI, and embedded linux, setting priorities, individual development plans, conducting performance reviews, and maintaining a high standard of engineering practice Own the technical roadmap … product software stack, from the C++/CUDA host application and real-time processing pipeline, aligning with product and system requirements Lead design and architecture reviews across the software stack, ensuring quality, consistency, and sound technical decisions Drive software from prototype through to production release, managing dependencies with ...

Sr Biostatistician (R a must- EMEA BASED)

Hiring Organisation
Syneos Health
Location
London, United Kingdom
Salary
£ 80 K
experience in clinical data structures and programming with dataexpert in functional and object-oriented programming. Knowledgeable in Javascript/Typescript, HTML, WebGL, experience in CUDA/GPU-programing, cloud-computing, Github, web-hosting/Machine Learning. Strong communication skills and ability to work both, independently and collaboratively, clear … clinical data structures and programming with data -Expert in functional and object-oriented programming knowledgeable in Javascript/Typescript, HTML, WebGL -Experience in CUDA/GPU-programing, cloud-computing, Github, web-hosting -Strong communication skills and ability to work both, independently and collaboratively, clear in the presentation of complex ...

Campus ML Engineer: Build Scalable AI for Finance

Location
Greater London, England, United Kingdom
Jump Trading Group is seeking world-class engineers to collaborate with our research, trading and engineering teams to build state-of-the-art ML systems for quantitative finance. You will work on training pipelines, low ...

CUDA Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
sources of electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're looking for a CUDA Engineer to write and optimise the low-level GPU code that powers our inference workloads: designing custom CUDA kernels, tuning performance across memory … squeezing maximum throughput out of every GPU in our fleet, working at the level of SMs, warps and memory hierarchies.ResponsibilitiesWrite and optimise custom CUDA kernels for core transformer inference operationsProfile kernels to identify and eliminate bottlenecks in occupancy, memory throughput and warp divergenceApply kernel fusion to reduce memory round ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch -- and we're looking for the founding engineer … Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered ...

AI Inference Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
electricity demand, Fuse is expanding into high-performance compute infrastructure at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch, and we're looking for the founding engineer to own the latter. … serving, deciding where and how to apply quantisation, distillation, speculative decoding and similar techniques to improve throughput and cost per token, partnering with the CUDA/GPU engineersMake the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)Translate ...

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
about efficient large-scale training, low-latency inference in real-time systems and high-throughput inference in research. Part of this is improving straightforward CUDA, but the interesting part needs a whole-systems approach, including storage systems, networking and host- and GPU-level considerations. Zooming in, we also want … level GPU knowledge of PTX, SASS, warps, cooperative groups, Tensor Cores and the memory hierarchy Debugging and optimisation experience using tools like CUDA GDB, NSight Systems, NSight Computesight-systems and nsight-compute Library knowledge of Triton, CUTLASS, CUB, Thrust, cuDNN and cuBLAS Intuition about the latency and throughput characteristics ...