26 to 50 of 143 CUDA Jobs

AI Infrastructure Software Engineer

Location
Greater London, England, United Kingdom
data structures and algorithms, and a rigorous coding style. Proficient in GPU/NPU architecture, with hands-on development and optimisation experience in CUDA C/C++ or Triton. Strong self-drive and technical curiosity; ability to proactively track the most cutting-edge AI Infra trends in the industry. ...

Senior Software Engineer, Machine Learning Services

Location
Greater London, England, United Kingdom
lifecycle of models in a multi‐tenant, high‐availability system. Familiarity with building ML inference services, model serialization (e.g., ONNX), and GPU programming (CUDA). You've built or worked on custom storage or job‐queueing systems before and have the scars to prove it. Maybe ...

Senior Research Engineer

Location
Greater London, England, United Kingdom
training frameworks, GPU/accelerator optimisation, data pipelines, experiment tracking, and making research reproducible and reliable. You have experience down to the level of CUDA kernels or XLA. You have contributed to open source projects like PyTorch, JAX, MLX or similar. You care about research outcomes as much ...

Sr. Software Engineer, Inference

Location
Greater London, England, United Kingdom
managing capacity planning, autoscaling policies, and driving post-incident remediation. Preferred Experience with low-level systems and high-performance computing components (e.g., C++ development, CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies). Active open-source or production contributions to modern inference frameworks (e.g., vLLM ...

UK - Simulation Software Engineer

Location
West of England, England, United Kingdom
architectures, including cross-architecture deployability and strict QoS timing constraints. Familiarity with High-Performance Computing (HPC), NVIDIA GPU acceleration, and hardware optimization (e.g., CUDA, TensorRT). Native fluency in Linux, Docker, and CI/CD pipelines—you know how to deploy and scale headless simulations unattended. Experience building ...

Software Inference Deployment Engineer

Location
Oxford, England, United Kingdom
Strong Preference For Experience integrating accelerator hardware (GPUs, FPGAs, ASICs, NPUs, or novel architectures) into customer inference workflows Familiarity with the NVIDIA inference stack - CUDA, TensorRT, Triton Exposure to disaggregated inference architectures, prefill/decode separation, or KV cache management Compensation & Benefits Highly Competitive Salary: We are not saying ...

Senior Software Engineer (vLLM)

Location
Cambridge, England, United Kingdom
experience working as a Senior Software Engineer with deep expertise in both Python and low-level programming (e.g. C/C++, Rust, assembly or CUDA). A history of direct contributions to vLLM or similar high-performance open-source ML/AI projects (e.g. PyTorch, Hugging Face TGI, TensorRT ...

Machine Learning Performance Engineer London, England, United Kingdom

Location
Greater London, England, United Kingdom
performance of XTX's training and inference platforms. The remit is wide, and you should expect to be challenged. We are not just writing CUDA code. Amongst other things, you will be working on a sophisticated optimising compiler for a variety of accelerated computing platforms. You should be proficient ...

AI Engineer (Fluent in Mandarin & English)

Location
Greater London, England, United Kingdom
systems, including vector database management and semantic search optimization. Preferred Qualifications Experience in the insurance or financial services sector. Deep knowledge of GPU architecture , CUDA, and hardware-level performance optimization. Familiarity with Document Intelligence frameworks (OCR, layout analysis, and multimodal extraction). MUST be fluent in Mandarin OR Cantonese ...

AI Integration Engineer

Hiring Organisation
Microtech Global Ltd
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
complex APIs, third-party libraries, and scalable data pipelines. Desirables: Direct experience with FFmpeg or GStreamer. Familiarity with GPU-accelerated video/AI tech (CUDA, PyTorch, TensorRT, DeepStream). Experience with multimodal AI (VLMs, LLM APIs, Retrieval-Augmented Generation). Cloud and containerization skills (AWS, Docker). ...

Senior AI Platform Engineer

Location
Greater London, England, United Kingdom
large-scale AI, machine learning, or distributed computing platforms in enterprise environments. Deep understanding of LLM architectures and their interaction with GPU infrastructure, including CUDA, cuDNN, NCCL, kernel-level acceleration libraries, and distributed training frameworks such as PyTorch. Strong knowledge of distributed training and inference strategies, including tensor, pipeline ...

Senior Performance Engineer

Location
Cambridge, England, United Kingdom
Excel/Google Sheets or Python modelling experience) to forecast system behaviour Optimisation of code running on GPUs and/or other accelerators (e.g. CUDA) Solid understanding of computer architecture fundamentals and how LLMs and Deep Learning models execute on that hardware (inference vs. training, matrix multiplication, KV-caching ...

Optical Systems Product Engineer (Photonic Quantum Computing)

Location
Greater London, England, United Kingdom
into existing AI and HPC infrastructure. Its current PT-2 system, a rack-mounted quantum computer, connects directly to GPU clusters via NVIDIA’s CUDA-Q platform, enabling accelerated AI workloads with significantly lower energy consumption than traditional silicon-based systems. Backed by a strong portfolio of patent families ...

Account Solution Architect

Location
Greater London, England, United Kingdom
Dutch, Swedish, Norwegian, Danish, or Finnish is a plus. Preferred Familiarity with NVIDIA GPU architectures (H100, A100, H200) and the software stack around them: CUDA, NCCL, cuDNN. Working knowledge of high-performance networking concepts: InfiniBand, RDMA, RoCE, TCP/IP. Background working directly with AI labs, research institutions ...

Senior Embedded Software Engineer

Location
Southampton, England, United Kingdom
knowledge of SOLID principles, OOD and software design patterns Security clearance will be needed Desirable Experience - Embedded Software Engineer -Southampton MATLAB OpenGL or Vulkan CUDA/OpenCL FPGA development and HDL #J-18808-Ljbffr ...

Senior Machine Learning Engineer

Hiring Organisation
Hexwired Recruitment Limited
Location
London, United Kingdom
Employment Type
Contract
they are looking for a Senior Machine Learning Engineer to join their expanding team in London. Required Experience Strong C++ development skills. Experience with CUDA for GPU acceleration. Python and PyTorch experience. Knowledge of Transformers, computer vision or image analysis. Postgraduate or commercial experience within AI/Machine Learning ...

NLP Performance Engineer

Location
Greater London, England, United Kingdom
PyTorch ecosystem Experience with inference optimisation techniques, including quantisation, speculative decoding and model parallelism across modern GPU architectures Strong software engineering skills, including Python, CUDA and building reliable systems for machine learning workloads Strong communication skills, with the ability to collaborate across research, infrastructure and engineering teams Why join ...

Senior Field Application Engineer

Hiring Organisation
Advanced Micro Devices
Location
Central London, London, United Kingdom
Employment Type
Permanent, Work From Home
Some Linux administration; understanding setup for HPC middleware. Nice to Haves: 5+ years HPC application experience Experience building and running HPC applications on GPU. CUDA or OpenACC or OpenMP paradigms. Experience running AI models on CPU or GPU. Any experience understanding/inspecting/writing assembly Understanding of memory ...

Staff Software Engineer, Inference

Location
Greater London, England, United Kingdom
modern inference frameworks (e.g., vLLM, Triton, TensorRT‐LLM, Ray Serve, or TorchServe). Deep experience with GPU systems engineering and hardware performance optimisation (e.g., CUDA, NCCL, RDMA, NUMA, or GPU interconnects). Direct exposure to large‐scale AI/ML infrastructure or hyperscale cloud environments. Wondering ...

Staff Software Engineer & Line Manager – Machine Learning (C++ & Vulkan)

Location
Cambridge, England, United Kingdom
recruitment. Building engineering teams through hiring, onboarding, mentoring or capability development. Open-source development or large software projects with complex dependencies. Vulkan, OpenCL, CUDA, DirectX, Metal or similar graphics and compute APIs. Machine learning frameworks, inference runtimes or model tooling. CI/CD, CMake, cross-compilation, packaging or developer ...

Staff Software Engineer & Line Manager – Machine Learning (C++ & Vulkan)

Location
United Kingdom
recruitment. Building engineering teams through hiring, onboarding, mentoring or capability development. Open‐source development or large software projects with complex dependencies. Vulkan, OpenCL, CUDA, DirectX, Metal or similar graphics and compute APIs. Machine learning frameworks, inference runtimes or model tooling. CI/CD, CMake, cross‐compilation, packaging or developer ...

Lead Software Engineer

Location
Nottingham, England, United Kingdom
/20), delivering robust, scalable, and maintainable solutions. Develop and optimize compute-intensive algorithms using multithreading, SIMD techniques, and GPU acceleration technologies such as CUDA and OpenCL. Analyze software performance, identify computational bottlenecks, and implement optimizations across CPU and GPU architectures. Leverage AI-assisted engineering tools such as GitHub ...

Machine Learning Engineer

Location
Greater London, England, United Kingdom
automation and orchestration systems for these platforms Proficiency in a low-level language such as C, C++, or Rust and in GPU frameworks like CUDA Competence in front-end web design to allow easy interfacing with large datasets Our Culture Follow the science. We prioritise rigorous scientific inquiry, relying ...

Tech Engagement Lead, AI Labs - EMEA

Location
United Kingdom
fixed.* Drive platform integration: Help partners adopt and optimize NVIDIA GPUs, systems, networking, and software libraries across their development pipelines. This may include CUDA, CUDA-X, NCCL, TensorRT-LLM, NeMo, Transformer Engine, CUTLASS, vLLM, SGLang, and related technologies.* Influence NVIDIA’s platform roadmap: Work with NVIDIA hardware … frontier AI research, model development, training, post-training, inference, and emerging AI workloads.* Practical understanding of GPU platforms and AI software libraries, including CUDA, CUDA-X, NCCL, PyTorch or JAX, and relevant training or inference systems.* Experience with large-scale GPU clusters, high-speed networking, distributed storage, workload ...