101 to 125 of 143 CUDA Jobs

Lead AI Training Infrastructure Engineer

Location
Greater London, England, United Kingdom
systems on multi-node GPU clusters. You will push wall-clock time to convergence by profiling, refining data pipelines, and low-level kernels in CUDA/CuDNN/Triton. You will build scalable PyTorch-based workflows, balance CPU/GPU workloads, and develop monitoring tools to diagnose performance regressions ...

Simulation Infrastructure Engineering & Research London

Location
Greater London, England, United Kingdom
expertise in rigid body simulation and FEM Insights and requests from a robotic researcher/engineer’s perspective for improving user experience Proficiency in CUDA and GPU programming Bonus Points Built closed-loop or log-replay evaluation at scale. Experience with robotics simulation tools such as genesis-world, Isaac ...

Senior ML Engineer

Location
Greater London, England, United Kingdom
optimize the end-to-end ML stack: data pipelines, training loops, inference serving, and deployment. Design and implement GPU-accelerated components, including custom CUDA kernels where off-the-shelf libraries are not enough. Work closely with the founders to translate product requirements into concrete optimization goals and technical roadmaps. ...

Developer Experience Engineer New London

Location
Greater London, England, United Kingdom
SDKs, APIs, or platforms that other engineers depend on. Experience developing for inference hardware and/or GPUs (e.g. one year or more using CUDA or ROCm). A good understanding of modern ML workloads and the practical challenges of deploying large language models to production. An instinct ...

Senior Machine Learning Engineer, AI Performance London, United Kingdom

Location
Greater London, England, United Kingdom
deep learning models in PyTorch (not just using high‐level tooling). Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL) and confidence learning adjacent frameworks quickly. Comfort operating at multiple levels of abstraction — from high‐level model behaviour down ...

Senior GPU Networking Architect: AI-Scale Kernel Innovator

Location
United Kingdom
frameworks. You will meld GPU computing with networking, optimizing kernels for latency, throughput, and scalability across large AI systems. The role emphasizes in‐depth CUDA work, GPU architecture mastery, and collaboration with cross‐functional teams to push the boundaries of GPU‐driven networking. #J-18808-Ljbffr ...

HPC Architect: Storage & Infrastructure Lead (Hybrid)

Location
Stevenage, England, United Kingdom
teams across national and international sites to deliver scalable, high-throughput infrastructure. You will drive workload strategies with SLURM, Kubernetes for HPC, MPI/CUDA stacks, and ensure reproducibility and portability of workloads while partnering with MBDA vendors #J-18808-Ljbffr ...

Senior GPU Software Engineer

Location
Greater London, England, United Kingdom
week. Experience for the Senior GPU Software Engineer includes: 5+ years of software engineering experience with significant GPU or high-performance computing experience Strong CUDA and GPU programming experience Excellent C++ programming skills Strong understanding of CPU/GPU interaction and data movement #J-18808-Ljbffr ...

Campus ML Research Engineer (Intern)

Location
Greater London, England, United Kingdom
resources. Integrate ML models into production systems where latency matters. Work across a mix of programming languages: C/C++/Python/CUDA and other low-level GPU languages. Build large scale ML systems that are observable, performant, and flexible. Help improve productivity by reducing the iteration cycle … Proficiency in Pytorch, JAX, Tensorflow or other DL library. Ability to thrive in a collaborative, team-oriented environment Expertise in GPU or Accelerator programming (CUDA, Triton, SYCL, ROCm or equivalent) Experience building ML systems at large scale (hundreds of TBs of training data, low latency or high throughput inference ...

AI Startup Partnerships Lead

Location
United Kingdom
building on NVIDIA technologies, evaluating technical maturity and platform alignment. You will work with venture capital and NVIDIA teams to accelerate production adoption of CUDA libraries and SDKs, deliver technical workshops, and support startups across Europe. Travel to engage with startups and investors is expected. #J-18808-Ljbffr ...

Founding AI Inference Engineer for Energy AI

Location
Greater London, England, United Kingdom
serving layer, guiding performance, reliability, and deployment across a growing cluster. You will define the serving strategy, design high-throughput workloads, and collaborate with CUDA/GPU teams to optimize throughput and cost per token while meeting ambitious uptime targets. #J-18808-Ljbffr ...

Founding AI Inference Engineer – Scale & Serving Expert

Location
Greater London, England, United Kingdom
Inference Engineer to define and build how we serve AI workloads at scale, reporting to the CTO. You’ll own the layer above CUDA/GPU work, shaping inference delivery, throughput, and reliability as we scale data-centre compute for energy applications. With 4+ years in large-scale inference ...

Founding GPU Engineer - Equity & Biannual Bonus

Location
Greater London, England, United Kingdom
Fuse Energy is seeking a seasoned CUDA performance engineer to design and optimise kernels for high-throughput workloads in a high-performance compute environment. You will profile GPU bottlenecks, build tooling to align power draw with energy pricing, and optimise multi-GPU scaling across NCCL/MPI. Collaboration with ...

Principal AI Engineer

Location
United Kingdom
advancing AI and engineering best practices. Day-to-day tasking can include: Working on high-throughput vision systems on NVIDIA Jetson hardware, involving CUDA/TensorRT acceleration Multi-modal (thermal + optical) sensors Robust real-time tracking Exploring AI Assurance and AI Safety Delivering technical expertise on a variety … Skills: We are interested in the following skills, but they are not essential for you to apply: Knowledge of C++, Python, PyTorch, OpenCV, CUDA, TensorRT, Git, WSL, Docker, Jira/Confluence Machine Learning and Deep Learning Computer Vision methodologies and algorithms MLOps Developing AI for use on Edge Computers ...

Principal AI Engineer

Hiring Organisation
Synoptix Limited
Location
Bristol, Avon, South West, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£85,000
advancing AI and engineering best practices. Day-to-day tasking can include: Working on high-throughput vision systems on NVIDIA Jetson hardware, involving; CUDA/TensorRT acceleration Multi-modal (thermal + optical) sensors Robust real-time tracking Exploring AI Assurance and AI Safety Delivering technical expertise on a variety … Skills: We are interested in the following skills, but they are not essential for you to apply: Knowledge of C++, Python, PyTorch, OpenCV, CUDA, TensorRT, Git, WSL, Docker, Jira/Confluence Machine Learning and Deep Learning Computer Vision methodologies and algorithms MLOps Developing AI for use on Edge Computers ...

AI Platform Support Engineer (EMEA)

Location
Greater London, England, United Kingdom
reasoning and clear communication Debug ML Infrastructure & Distributed Workloads Investigate failures involving distributed training, Kubernetes orchestration, GPU allocation, networking, and storage systems Troubleshoot PyTorch, CUDA, NCCL, and inference serving related issues Analyze logs, metrics, traces, and system behavior to isolate root causes Debug containerized workloads running across Kubernetes … Infrastructure Experience Hands on experience operating machine learning workloads in production or research environments Experience with distributed ML systems and tooling such as PyTorch, CUDA, or NCCL Familiarity with GPU infrastructure and orchestration Experience troubleshooting performance, reliability, or scaling issues in ML infrastructure Understanding of the operational challenges involved ...

AI Researcher – 4D Reconstruction (Contractor)

Location
Cambridge, England, United Kingdom
About Huawei Research and Development UK Limited Founded in 1987, Huawei is a leading global provider of information and communications technology (ICT) infrastructure and smart devices. We have 207,000 employees and operate in over ...

Research Software Engineer

Location
Greater London, England, United Kingdom
About Mistral Mistral provides full-stack AI solutions: from frontier models to developer tools, applications, and compute. We partner with enterprises tackling the hardest problems—across high-stakes industries like finance, manufacturing, defense, healthcare, and ...

Machine Learning Performance Engineer

Location
Greater London, England, United Kingdom
We tackle the most complex problems in quantitative finance, by bringing scientific clarity to financial complexity. From our London HQ, we unite world-class researchers and engineers in an environment that values deep exploration and ...

Member of Technical Staff (AI Inference Engineer)

Location
Greater London, England, United Kingdom
engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL.### ### **Responsibilities:*** **New models support.** Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading … request scheduling and KV-cache management to support in API Gateway.* **GPU kernels migration to CuTe DSL.** Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow.* **Rust-native serving runtime.** Develop our internal ...

Campus ML Research Engineer (Full-Time)

Location
Greater London, England, United Kingdom
Full-time ML engineering role building financial-ML frameworks, HPC training pipelines, low-latency inference, and large-scale GPU systems with Python, C++, CUDA, and modern ML libraries. Build reusable frameworks for financial machine learning. Optimise training pipelines for high-performance computing. Integrate ML models into latency-sensitive production … systems. Build large-scale, observable ML systems. Work with C, C++, Python, CUDA, and other low-level GPU technologies. Python and/or C++. PyTorch, JAX, TensorFlow, or another deep-learning library. GPU programming with CUDA, Triton, SYCL, ROCm, or similar tools. Large-scale ML systems, including very ...

ML/AI Engineer

Location
Manchester, England, United Kingdom
training (CT) and continuous monitoring (CM) to keep models fresh and safe in production. Deploy and tune GPU‐backed inference services (e.g., A100), optimise CUDA environments, and leverage TensorRT where appropriate. Operate scalable serving frameworks (NVIDIA Triton, TorchServe) with attention to latency, efficiency, resilience, and cost. Implement … expertise having hands‐on experience with Harness (or similar) building multi‐stage pipelines; experience with GitOps, artefact repositories, and environment promotion. Practical experience with CUDA, TensorRT, Triton, TorchServe, and GPU scheduling/optimisation. Proficiency in Prometheus, Grafana, Dynatrace defining SLIs/SLOs and alert thresholds for ML systems. Experience ...

AI Infrastructure Engineer

Location
Greater London, England, United Kingdom
engineers who have: A track record of working on model training or model inference at scale , or on low‐level GPU coding (e.g. CUDA, Triton). Experience with one is great, multiple is even better. What will I be doing? As a Senior AI Infrastructure Engineer focused on model … Model training (especially transformers and LLMs). Model inference at scale (again, especially transformers and LLMs). Low‐level GPU work , such as writing CUDA or Triton kernels. Comfortable working in production environments at meaningful scale (traffic, data, or organizational). You communicate clearly, can explain complex technical topics ...

ML Compiler & System Engineering & Research London

Location
Greater London, England, United Kingdom
Quadrants is an open-source translation layer that turns ordinary Python into optimized machine code at runtime, across CPU (Arm64, x86) and GPU (AMDGPU, CUDA, Metal, Vulkan) backends. CUDA is the primary target, with x86 the strong runner-up, especially for real-time simulation. The others drive adoption ...

Software Engineering Manager

Location
Greater London, England, United Kingdom
This role is based in our London (Kings Cross) office. Responsibilities: Manage and mentor a team of software engineers working across embedded C++/CUDA, GUI, and embedded linux, setting priorities, individual development plans, conducting performance reviews, and maintaining a high standard of engineering practice Own the technical roadmap … product software stack, from the C++/CUDA host application and real-time processing pipeline, aligning with product and system requirements Lead design and architecture reviews across the software stack, ensuring quality, consistency, and sound technical decisions Drive software from prototype through to production release, managing dependencies with ...