26 to 47 of 47 CUDA Jobs in London

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar). Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker, Singularity/Apptainer). A track record of building tools that ...

ML Systems Engineer Intern — High-Perf Finance

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
apply advanced techniques, optimize training pipelines on HPC clusters, and integrate ML models into latency-sensitive production environments, using C/C++, Python, and CUDA across large-scale systems. #J-18808-Ljbffr ...

ML Systems Engineer for Quantitative Finance

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
systems for quantitative finance. You will optimize training pipelines, deploy models to low-latency production environments, and work across C/C++, Python, CUDA, and related GPU technologies. Join a fast-paced, collaborative team focused on impactful projects that push the boundaries of AI research and its applications ...

Member of Technical Staff (Infrastructure Engineer, Training and Inference Systems)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
front and centre. Strong candidates may also have Experience with a systems programming language like Rust or C++. Experience with writing and profiling CUDA kernels. A track record of building reliable research tools. Why this is interesting You’ll shape the core technical foundation of a frontier ...

Autonomous AI Research Engineer — Hybrid-Remote

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
practical product impact and close collaboration with product teams. The role requires a strong math background, expertise in PyTorch, and experience with Rust/CUDA for high-performance inference. This is a unique opportunity to shape the AI strategy for a fast-growing platform. #J-18808-Ljbffr ...

Physical Design Engineer

Hiring Organisation
Oho Group
Location
Greater London, England, United Kingdom
technology that will power the next generation of AI What We’re Looking For Strong C/C++ programming experience Experience with GPU programming (CUDA, OpenCL, HIP, or similar) Strong understanding of parallel computing and computer architecture Experience optimising high-performance workloads Passion for building cutting-edge technology ...

Senior ML Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
optimize the end-to-end ML stack: data pipelines, training loops, inference serving, and deployment. Design and implement GPU-accelerated components, including custom CUDA kernels where off-the-shelf libraries are not enough. Work closely with the founders to translate product requirements into concrete optimization goals and technical roadmaps. ...

Developer Experience Engineer New London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
SDKs, APIs, or platforms that other engineers depend on. Experience developing for inference hardware and/or GPUs (e.g. one year or more using CUDA or ROCm). A good understanding of modern ML workloads and the practical challenges of deploying large language models to production. An instinct ...

Senior Machine Learning Engineer, AI Performance

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
deep learning models in PyTorch (not just using high‐level tooling). Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL) and confidence learning adjacent frameworks quickly. Comfort operating at multiple levels of abstraction — from high-level model behaviour down ...

Founding AI Inference Engineer – Scale & Serving Expert

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Inference Engineer to define and build how we serve AI workloads at scale, reporting to the CTO. You’ll own the layer above CUDA/GPU work, shaping inference delivery, throughput, and reliability as we scale data-centre compute for energy applications. With 4+ years in large-scale inference ...

Campus ML Research Engineer (Intern)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
resources. Integrate ML models into production systems where latency matters. Work across a mix of programming languages: C/C++/Python/CUDA and other low-level GPU languages. Build large scale ML systems that are observable, performant, and flexible. Help improve productivity by reducing the iteration cycle … Proficiency in Pytorch, JAX, Tensorflow or other DL library. Ability to thrive in a collaborative, team-oriented environment Expertise in GPU or Accelerator programming (CUDA, Triton, SYCL, ROCm or equivalent) Experience building ML systems at large scale (hundreds of TBs of training data, low latency or high throughput inference ...

Campus ML Research Engineer (Full-Time)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
resources. Integrate ML models into production systems where latency matters. Work across a mix of programming languages: C/C++/Python/CUDA and other low-level GPU languages. Build large scale ML systems that are observable, performant, and flexible. Help improve productivity by reducing the iteration cycle … Proficiency in Pytorch, JAX, Tensorflow or other DL library. Ability to thrive in a collaborative, team-oriented environment Expertise in GPU or Accelerator programming (CUDA, Triton, SYCL, ROCm or equivalent) Experience building ML systems at large scale (hundreds of TBs of training data, low latency or high throughput inference ...

HPC Specialist Architect - Energy Industry (AWS)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
such as oil and gas, automotive/aerospace, financial services, or pharmaceuticals. Proficient in one or more of the following programming languages: C++, Python, Cuda, or Bash. Experience in architecting an HPC platform with scheduling middleware (e.g., Slurm, Torque, Symphony or GridServer) and in deployment, tuning and management … programming models for both loosely and tightly coupled applications, including proficiency with MPI (OpenMPI, MPICH), OpenMP, and hybrid parallelization strategies, and GPU programming using CUDA, OpenACC, or similar technologies, with demonstrated ability to optimize applications for GPU architectures. Significant experience in infrastructure architecture and networking. Experience with infrastructure ...

Energy HPC Architect: AWS Solutions & PoC Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
such as oil and gas, automotive/aerospace, financial services or pharmaceuticals. Proficient in one or more of the following programming languages: C++, Python, Cuda, or Bash. - Experience in architecting an HPC platform with scheduling middleware (e.g. Slurm, Torque, Symphony or GridServer) and in deployment, tuning and management … programming models for both loosely and tightly coupled applications, including proficiency with MPI (OpenMPI, MPICH), OpenMP, and hybrid parallelization strategies, and GPU programming using CUDA, OpenACC, or similar technologies, with demonstrated ability to optimize applications for GPU architectures. - Significant experience in infrastructure architecture and networking. Experience with infrastructure ...

Research Software Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and ...

Research Engineer

Hiring Organisation
European Tech Recruit
Location
City of London, London, United Kingdom
Research Engineer Cambridge or London, UK (100% Onsite) Our client is a global leader in semiconductor innovation, developing advanced technologies that power billions of smart devices worldwide. With significant investment in artificial intelligence research, the ...

AI Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
engineers who have: A track record of working on model training or model inference at scale , or on low‐level GPU coding (e.g. CUDA, Triton). Experience with one is great, multiple is even better. What will I be doing? As a Senior AI Infrastructure Engineer focused on model … Model training (especially transformers and LLMs). Model inference at scale (again, especially transformers and LLMs). Low‐level GPU work , such as writing CUDA or Triton kernels. Comfortable working in production environments at meaningful scale (traffic, data, or organizational). You communicate clearly, can explain complex technical topics ...

ML Infrastructure Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
frameworks such as PyTorch or JAX Are proficient in Python, including concurrency, asynchronous programming, multiprocessing, and performance optimization Can debug distributed GPU workloads across CUDA runtime, container runtime, driver versions, NCCL or equivalent communication layers, networking, storage, scheduling, and checkpointing Have experience with profiling tools across the stack … Experience with GPU clusters on Kubernetes, Slurm, Ray, custom schedulers, or cloud GPU orchestration NCCL, UCX, NVSHMEM, RDMA, InfiniBand, RoCE, or EFA Rust, C++, CUDA, Go, or systems‐level performance work Why White Circle You will be able to propose and run your own experiments and research ideas ...

Software Engineering Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
This role is based in our London (Kings Cross) office. Responsibilities: Manage and mentor a team of software engineers working across embedded C++/CUDA, GUI, and embedded linux, setting priorities, individual development plans, conducting performance reviews, and maintaining a high standard of engineering practice Own the technical roadmap … product software stack, from the C++/CUDA host application and real-time processing pipeline, aligning with product and system requirements Lead design and architecture reviews across the software stack, ensuring quality, consistency, and sound technical decisions Drive software from prototype through to production release, managing dependencies with ...

AI Research Engineer, Pre-Training

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
will collaborate closely with researchers to co‐design and improve our models and shape the research agenda. Qualifications Strong engineering skills, especially in CUDA/Triton/Pallas/CuTe DSL kernel development, lower‐level PyTorch/JAX/XLA development, CUDA Graphs, FPGA/ASIC experience. ...

Member of Technical Staff (Infrastructure Engineer, Compute Infrastructure)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Member of Technical Staff, Compute Infrastructure — Inherent (London) At Inherent, we are on a mission to build AI that recursively self‐improves to discover new knowledge. Scientific advances are the backbone of our economic, technological ...

AI Inference Engineer

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
demand, Fuse is expanding into high-performance compute infrastructure that sits at the intersection of energy and AI. We're building the GPU/CUDA performance layer and the inference serving layer at the same time, from scratch -- and we're looking for the founding engineer … Founding AI Inference Engineer to define and build how Fuse serves AI inference workloads at scale, reporting directly to the CTO. Where our CUDA and GPU engineering hires own kernel-level and hardware performance, this role owns the layer above it: how models actually get served, scaled, and delivered ...