26 to 50 of 166 Permanent CUDA Jobs

Staff System Software Engineer

Location
West of England, England, United Kingdom
well within a multinational team and with multinational customers. Excellent cultural awareness is essential. Desirable Experience developing firmware or drivers for GPUs. Knowledge of CUDA or OpenCL Experience working on upstreaming of kernel code/contributing to Linux kernel. Exposure to integration with data centre/cloud service operations ...

Machine Learning Scientist — Large Multimodal Models (Post-Training)

Location
West of England, England, United Kingdom
Tune, or similar frameworks Strong engineering practices: reproducible experimentation, clean code, testing, and performance-aware debugging Comfort with modern ML infrastructure (e.g., Docker, CUDA, Kubernetes, experiment tracking tools such as Weights & Biases) PREFERRED Experience with multimodal or multi-task model architectures Training and inference optimization (e.g., mixed precision, kernel ...

Senior Research Engineer

Location
Greater London, England, United Kingdom
training frameworks, GPU/accelerator optimisation, data pipelines, experiment tracking, and making research reproducible and reliable. You have experience down to the level of CUDA kernels or XLA. You have contributed to open source projects like PyTorch, JAX, MLX or similar. You care about research outcomes as much ...

ML Research Engineer, London

Hiring Organisation
Isomorphic Labs
Location
London, UK
Employment Type
Full-time
models. Scale & Performance: Experience training models across distributed systems (multi-GPU/multi-node) and optimising training and inference performance (e.g., XLA, Triton, CUDA, Pallas).Domain Knowledge: A strong interest in, or knowledge of, biochemistry, computational biology, or drug discovery fundamentals. Industry Experience: Proven track record working in reputable ...

Software Inference Deployment Engineer

Location
Oxford, England, United Kingdom
Strong Preference For Experience integrating accelerator hardware (GPUs, FPGAs, ASICs, NPUs, or novel architectures) into customer inference workflows Familiarity with the NVIDIA inference stack - CUDA, TensorRT, Triton Exposure to disaggregated inference architectures, prefill/decode separation, or KV cache management Compensation & Benefits Highly Competitive Salary: We are not saying ...

Senior Embedded Software Engineer

Location
Southampton, England, United Kingdom
knowledge of SOLID principles, OOD and software design patterns Security clearance will be needed Desirable Experience - Embedded Software Engineer -Southampton MATLAB OpenGL or Vulkan CUDA/OpenCL FPGA development and HDL #J-18808-Ljbffr ...

Junior Algorithmic Developer, Analyst, London

Hiring Organisation
Jefferies Financial Group
Location
London, UK
Employment Type
Full-time
Computer Science, Computer Engineering, Mathematics, or a highly quantitative field. Preferred Qualifications: Experience with networking protocols (TCP/IP, UDP) or hardware acceleration (CUDA/GPUs).Familiarity with CVS, git and Linux/Unix environments. Basic understanding of quantitative finance and algorithmic trading. Jefferies is a leading global, full ...

Computer Vision Manager

Location
Middleton Stoney, England, United Kingdom
camera data processing. Experience with IMU/GNSS‐aided sensor fusion. Experience with C++, C# or Python in engineering software. Familiarity with Linux, Docker, CUDA or GPU acceleration. Understanding of navigation, localisation or mapping products. Familiarity with React or web‐based visualisation tools. Experience using AI coding assistants ...

Machine Learning Performance Engineer London, England, United Kingdom

Location
Greater London, England, United Kingdom
performance of XTX's training and inference platforms. The remit is wide, and you should expect to be challenged. We are not just writing CUDA code. Amongst other things, you will be working on a sophisticated optimising compiler for a variety of accelerated computing platforms. You should be proficient ...

Senior Software Engineer, Full-Stack

Location
Greater London, England, United Kingdom
these — as long as you're open to learning, Backend: Python Frontend: TypeScript and React Deployment: Kubernetes Infrastructure: GCP Machine learning: PyTorch, CUDA, Ray Why Encord Competitive salary, commission, and meaningful equity in a high-growth startup Strong in-person culture — most of the team works from our London ...

Scene Generation Software Engineer (RF/EO Simulation)

Location
West of England, England, United Kingdom
work in both domains. What we’re looking to see from you: Practical experience with C++ and GPU programming utilising one or more of: CUDA, OpenGL or Vulkan. An appreciation of the physics behind radio‐frequency (RF) or electro‐optical/infrared (EO/IR) propagation and sensing ...

AI Engineer (Fluent Portuguese & English)

Hiring Organisation
Chubb
Location
London, UK
Employment Type
Full-time
Generation systems, including vector database management and semantic search optimization. Preferred QualificationsExperience in the insurance or financial services sector. Deep knowledge of GPU architecture, CUDA, and hardware-level performance optimization. Familiarity with Document Intelligence frameworks (OCR, layout analysis, and multimodal extraction).MUST be fluent in Portuguese and EnglishWe offer ...

AI Engineer (Fluent in Mandarin & English)

Location
Greater London, England, United Kingdom
systems, including vector database management and semantic search optimization. Preferred Qualifications Experience in the insurance or financial services sector. Deep knowledge of GPU architecture , CUDA, and hardware-level performance optimization. Familiarity with Document Intelligence frameworks (OCR, layout analysis, and multimodal extraction). MUST be fluent in Mandarin OR Cantonese ...

Staff Software Engineer, Inference

Location
Greater London, England, United Kingdom
modern inference frameworks (e.g., vLLM, Triton, TensorRT‐LLM, Ray Serve, or TorchServe). Deep experience with GPU systems engineering and hardware performance optimisation (e.g., CUDA, NCCL, RDMA, NUMA, or GPU interconnects). Direct exposure to large‐scale AI/ML infrastructure or hyperscale cloud environments. Wondering ...

Lead Software Engineer

Location
Nottingham, England, United Kingdom
/20), delivering robust, scalable, and maintainable solutions. Develop and optimize compute-intensive algorithms using multithreading, SIMD techniques, and GPU acceleration technologies such as CUDA and OpenCL. Analyze software performance, identify computational bottlenecks, and implement optimizations across CPU and GPU architectures. Leverage AI-assisted engineering tools such as GitHub ...

Senior AI Platform Engineer

Location
Greater London, England, United Kingdom
large-scale AI, machine learning, or distributed computing platforms in enterprise environments. Deep understanding of LLM architectures and their interaction with GPU infrastructure, including CUDA, cuDNN, NCCL, kernel-level acceleration libraries, and distributed training frameworks such as PyTorch. Strong knowledge of distributed training and inference strategies, including tensor, pipeline ...

Senior Performance Engineer

Location
Cambridge, England, United Kingdom
Excel/Google Sheets or Python modelling experience) to forecast system behaviour Optimisation of code running on GPUs and/or other accelerators (e.g. CUDA) Solid understanding of computer architecture fundamentals and how LLMs and Deep Learning models execute on that hardware (inference vs. training, matrix multiplication, KV-caching ...

Account Solution Architect

Location
Greater London, England, United Kingdom
Dutch, Swedish, Norwegian, Danish, or Finnish is a plus. Preferred Familiarity with NVIDIA GPU architectures (H100, A100, H200) and the software stack around them: CUDA, NCCL, cuDNN. Working knowledge of high-performance networking concepts: InfiniBand, RDMA, RoCE, TCP/IP. Background working directly with AI labs, research institutions ...

Junior SRE

Hiring Organisation
SQUAREPOINT CAPITAL
Location
London, UK
Employment Type
Full-time
systems, HPC, or cloud platforms for CPU and GPU workloads. Experience with profiling, benchmarking, or performance analysis. Additional Qualifications (Nice to Have)Familiarity with CUDA or other GPU development frameworks. The minimum base salary for this role is $150,000 if located in New York. This expectation is based ...

MLOps Engineer (LLM/GenAI)

Location
Sheffield, England, United Kingdom
skills: Extensive experience in building AI platforms covering model hosting/inference optimisation and fine-tuning pipelines (LLM experience strongly preferred) Strong Python and CUDA engineering; solid understanding of GPU/CPU architecture and HPC fundamentals Deep inference optimisation expertise: KV-cache, batching, quantisation (INT4/FP8/GPTQ ...

Software Engineer, Model Inference, DeepMind

Hiring Organisation
Google
Location
London, UK
Employment Type
Full-time
Experience with developing serving infrastructure. Experience programming hardware accelerators (GPUs, TPUs) via ML frameworks (e.g., JAX, PyTorch) or low-level programming models (e.g., Pallas, CUDA, OpenCL).Experience profiling software to identify performance bottlenecks. Experience with distributed ML systems optimization and parallelism (e.g., data, model, or pipeline parallelism).Familiarity with ...

Software Engineer

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Experience with developing serving infrastructure. Experience programming hardware accelerators (GPUs, TPUs) via ML frameworks (e.g., JAX, PyTorch) or low-level programming models (e.g., Pallas, CUDA, OpenCL). Experience profiling software to identify performance bottlenecks. Experience with distributed ML systems optimization and parallelism (e.g., data, model, or pipeline parallelism). ...

Software Engineer - Model Inference

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Experience with developing serving infrastructure. Experience programming hardware accelerators (GPUs, TPUs) via ML frameworks (e.g., JAX, PyTorch) or low-level programming models (e.g., Pallas, CUDA, OpenCL). Experience profiling software to identify performance bottlenecks. Experience with distributed ML systems optimization and parallelism (e.g., data, model, or pipeline parallelism). ...

NLP Performance Engineer

Location
Greater London, England, United Kingdom
PyTorch ecosystem Experience with inference optimisation techniques, including quantisation, speculative decoding and model parallelism across modern GPU architectures Strong software engineering skills, including Python, CUDA and building reliable systems for machine learning workloads Strong communication skills, with the ability to collaborate across research, infrastructure and engineering teams Why join ...