51 to 75 of 104 CUDA Jobs in London

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, UK
Employment Type
Full-time
training libraries, or custom kernels/fused ops. Experience with multi-node cluster orchestration (Slurm, Ray, Kubernetes, or similar).Comfort debugging performance issues across CUDA/NCCL, networking, IO, and data pipelines. Experience working with containerized environments (Docker, Singularity/Apptainer).A track record of building tools that increase ...

Electronic Engineer

Location
Greater London, England, United Kingdom
into existing AI and HPC infrastructure. Its current PT-2 system, a rack-mounted quantum computer, connects directly to GPU clusters via NVIDIA’s CUDA-Q platform, enabling accelerated AI workloads with significantly lower energy consumption than traditional silicon-based systems. Backed by a strong portfolio of patent families ...

ML Systems Engineer Intern — High-Perf Finance

Location
Greater London, England, United Kingdom
apply advanced techniques, optimize training pipelines on HPC clusters, and integrate ML models into latency-sensitive production environments, using C/C++, Python, and CUDA across large-scale systems. #J-18808-Ljbffr ...

ML Systems Engineer for Quantitative Finance

Location
Greater London, England, United Kingdom
systems for quantitative finance. You will optimize training pipelines, deploy models to low-latency production environments, and work across C/C++, Python, CUDA, and related GPU technologies. Join a fast-paced, collaborative team focused on impactful projects that push the boundaries of AI research and its applications ...

Software Engineer - Systems

Location
Greater London, England, United Kingdom
engineering Familiarity with tools like Apache Arrow, Parquet, DataFusion, Clickhouse, or DuckDB is a plus Understanding of cutting-edge ML infrastructure stack (e.g. PyTorch, CUDA) is also a plus Experience with Rust is a bonus Willingness to work in-person at our NYC or London office #J-18808-Ljbffr ...

Research Scientist

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
years of experience working in frontier AI research labs on pre-training or post-training teams. Experience writing TPU/GPU kernels (e.g., JAX, CUDA) to optimize model performance and real-time inference. Experience designing, training, and scaling generative pixel or real-time video architectures. Track record of cross ...

Research Scientist

Location
Westminster, West End, United Kingdom
years of experience working in frontier AI research labs on pre-training or post-training teams. Experience writing TPU/GPU kernels (e.g., JAX, CUDA) to optimize model performance and real-time inference. Experience designing, training, and scaling generative pixel or real-time video architectures. Track record of cross ...

AI Research Engineer

Location
Greater London, England, United Kingdom
plus; a strong background in mathematics is required. Alternatively, experience training models with strong math skills. Strong coding ability in Rust (preferred), CUDA, or C. Solid understanding of transformer architectures. A product mindset; you’ll build production‐ready products. Hybrid Remote with the London Office. What ...

Research Scientist, World Models, DeepMind

Location
Greater London, England, United Kingdom
years of experience working in frontier AI research labs on pre-training or post-training teams. Experience writing TPU/GPU kernels (e.g., JAX, CUDA) to optimize model performance and real-time inference. Experience designing, training, and scaling generative pixel or real-time video architectures. Track record of cross ...

Product Engineer, Physical AI

Location
Greater London, England, United Kingdom
these — as long as you're open to learning, please apply. Backend: Python Frontend: TypeScript and React Deployment: Kubernetes Infrastructure: GCP Machine learning: PyTorch, CUDA, Ray Why Encord Competitive salary, commission, and meaningful equity in a high-growth startup Strong in-person culture — most of the team works from ...

ML Performance Engineer – Scale GPU/CPU Workloads

Location
Greater London, England, United Kingdom
evolve the compute stack. The role shapes platform evolution and enables researchers to push the boundaries of machine learning. You will work with Python, CUDA, Kubernetes, and deep learning frameworks like PyTorch, applying #J-18808-Ljbffr ...

High Performance Computing Architect (Linux)

Location
London, United Kingdom
performance computing environments and infrastructure design Experience working with open-source technologies and software-defined data centres Knowledge of HPC software stacks including MPI, CUDA, compilers, and scientific libraries Familiarity with containerisation and reproducible workload environments Experience working within large, complex, or multi-national organisations Ability to influence technical ...

Junior Growth Sales Representative

Location
Greater London, England, United Kingdom
compute helps customers train and run AI workloads. Familiarity with cloud platforms (AWS, GCP, Azure) and an interest in GPUs and AI workloads (CUDA, PyTorch, vLLM, ...). Proficiency in English Nice to have Some exposure to sales, customer‐facing, or support work (internships, part‐time roles, or study ...

Lead AI Training Infrastructure Engineer

Location
Greater London, England, United Kingdom
systems on multi-node GPU clusters. You will push wall-clock time to convergence by profiling, refining data pipelines, and low-level kernels in CUDA/CuDNN/Triton. You will build scalable PyTorch-based workflows, balance CPU/GPU workloads, and develop monitoring tools to diagnose performance regressions ...

Senior ML Engineer

Location
Greater London, England, United Kingdom
optimize the end-to-end ML stack: data pipelines, training loops, inference serving, and deployment. Design and implement GPU-accelerated components, including custom CUDA kernels where off-the-shelf libraries are not enough. Work closely with the founders to translate product requirements into concrete optimization goals and technical roadmaps. ...

Developer Experience Engineer New London

Location
Greater London, England, United Kingdom
SDKs, APIs, or platforms that other engineers depend on. Experience developing for inference hardware and/or GPUs (e.g. one year or more using CUDA or ROCm). A good understanding of modern ML workloads and the practical challenges of deploying large language models to production. An instinct ...

Research Engineer (LLM Performance), London

Hiring Organisation
Isomorphic Labs
Location
London, UK
Employment Type
Full-time
more important than writing kernels from scratchExcellent collaboration skills. Nice to have: Experience with general LLM serving stacks. Knowledge of XLA, Triton, Pallas, CUDA or similar accelerator DSLs/compilers. Experience with optimising ML accuracy using low-precision formats. Prior experience building, deploying and maintaining production systems on GCP.Interest ...

Senior GPU Software Engineer

Location
Greater London, England, United Kingdom
week. Experience for the Senior GPU Software Engineer includes: 5+ years of software engineering experience with significant GPU or high-performance computing experience Strong CUDA and GPU programming experience Excellent C++ programming skills Strong understanding of CPU/GPU interaction and data movement #J-18808-Ljbffr ...

Simulation Infrastructure Engineering & Research London

Location
Greater London, England, United Kingdom
expertise in rigid body simulation and FEM Insights and requests from a robotic researcher/engineer’s perspective for improving user experience Proficiency in CUDA and GPU programming Bonus Points Built closed-loop or log-replay evaluation at scale. Experience with robotics simulation tools such as genesis-world, Isaac ...

Campus ML Research Engineer (Intern)

Location
Greater London, England, United Kingdom
resources. Integrate ML models into production systems where latency matters. Work across a mix of programming languages: C/C++/Python/CUDA and other low-level GPU languages. Build large scale ML systems that are observable, performant, and flexible. Help improve productivity by reducing the iteration cycle … Proficiency in Pytorch, JAX, Tensorflow or other DL library. Ability to thrive in a collaborative, team-oriented environment Expertise in GPU or Accelerator programming (CUDA, Triton, SYCL, ROCm or equivalent) Experience building ML systems at large scale (hundreds of TBs of training data, low latency or high throughput inference ...

Campus ML Research Engineer (Full-Time)

Location
Greater London, England, United Kingdom
resources. Integrate ML models into production systems where latency matters. Work across a mix of programming languages: C/C++/Python/CUDA and other low-level GPU languages. Build large scale ML systems that are observable, performant, and flexible. Help improve productivity by reducing the iteration cycle … Proficiency in Pytorch, JAX, Tensorflow or other DL library. Ability to thrive in a collaborative, team-oriented environment Expertise in GPU or Accelerator programming (CUDA, Triton, SYCL, ROCm or equivalent) Experience building ML systems at large scale (hundreds of TBs of training data, low latency or high throughput inference ...

Senior HPC Engineer

Hiring Organisation
Hays Specialist Recruitment Limited
Location
London, South East England, United Kingdom
Employment Type
Full-Time
Salary
£60,000 - £70,000 per annum
Senior HPC Engineer Please read the advert below carefully, and if you are a good match I want to speak to you ASAP. Please call or email Lorenz Pasch at Hays Recruitment - my contact details ...

Founding AI Inference Engineer for Energy AI

Location
Greater London, England, United Kingdom
serving layer, guiding performance, reliability, and deployment across a growing cluster. You will define the serving strategy, design high-throughput workloads, and collaborate with CUDA/GPU teams to optimize throughput and cost per token while meeting ambitious uptime targets. #J-18808-Ljbffr ...

Founding AI Inference Engineer – Scale & Serving Expert

Location
Greater London, England, United Kingdom
Inference Engineer to define and build how we serve AI workloads at scale, reporting to the CTO. You’ll own the layer above CUDA/GPU work, shaping inference delivery, throughput, and reliability as we scale data-centre compute for energy applications. With 4+ years in large-scale inference ...