51 to 75 of 126 CUDA Jobs in England

NLP Performance Engineer

Hiring Organisation
G Research
Location
London, UK
Employment Type
Full-time
PyTorch ecosystemExperience with inference optimisation techniques, including quantisation, speculative decoding and model parallelism across modern GPU architecturesStrong software engineering skills, including Python, CUDA and building reliable systems for machine learning workloadsStrong communication skills, with the ability to collaborate across research, infrastructure and engineering teamsWhy join us? Highly competitive compensation ...

Software Engineer, Model Inference, DeepMind

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
Experience with developing serving infrastructure. Experience programming hardware accelerators (GPUs, TPUs) via ML frameworks (e.g., JAX, PyTorch) or low-level programming models (e.g., Pallas, CUDA, OpenCL). Experience profiling software to identify performance bottlenecks. Experience with distributed ML systems optimization and parallelism (e.g., data, model, or pipeline parallelism). ...

NLP Performance Engineer

Location
Greater London, England, United Kingdom
PyTorch ecosystem Experience with inference optimisation techniques, including quantisation, speculative decoding and model parallelism across modern GPU architectures Strong software engineering skills, including Python, CUDA and building reliable systems for machine learning workloads Strong communication skills, with the ability to collaborate across research, infrastructure and engineering teams Why join ...

Product Engineer

Location
Greater London, England, United Kingdom
these — as long as you\'re open to learning, please apply. Backend: Python Frontend: TypeScript and React Deployment: Kubernetes Infrastructure: GCP Machine learning: PyTorch, CUDA, Ray Why Encord Competitive salary, commission, and meaningful equity in a high-growth startup Strong in-person culture — most of the team works from ...

Senior Software Engineer - Backend

Location
Greater London, England, United Kingdom
these. As long as you're open to learning, please apply. Backend: Python Frontend: TypeScript and React Deployment: Kubernetes Infrastructure: GCP Machine learning: PyTorch, CUDA, Ray Why Encord Competitive salary, commission, and meaningful equity in a high-growth startup Strong in-person culture: the team works from our London ...

AI Engineer (Fluent Portuguese & English)

Hiring Organisation
Chubb
Location
London, UK
Employment Type
Full-time
Generation systems, including vector database management and semantic search optimization. Preferred QualificationsExperience in the insurance or financial services sector. Deep knowledge of GPU architecture, CUDA, and hardware-level performance optimization. Familiarity with Document Intelligence frameworks (OCR, layout analysis, and multimodal extraction).MUST be fluent in Portuguese and EnglishWe offer ...

AI Engineer (Fluent in Mandarin & English)

Location
Greater London, England, United Kingdom
systems, including vector database management and semantic search optimization. Preferred Qualifications Experience in the insurance or financial services sector. Deep knowledge of GPU architecture , CUDA, and hardware-level performance optimization. Familiarity with Document Intelligence frameworks (OCR, layout analysis, and multimodal extraction). MUST be fluent in Mandarin OR Cantonese ...

Principal Biostatistician/ Sr Biostatistician (R/Rshiny - EMEA BASED)

Hiring Organisation
Syneos Health
Location
London, UK
Employment Type
Full-time
clinical data structures and programming with data expert in functional and object-oriented programming. Knowledgeable in Javascript/Typescript, HTML, WebGL, experience in CUDA/GPU-programing, cloud-computing, Github, web-hosting. Strong communication skills and ability to work both, independently and collaboratively, clear in the presentation of complex ...

Staff ML Performance Engineer (Compiler)

Location
Greater London, England, United Kingdom
with tight constraints (latency, memory, bandwidth, power/thermal, or cost). Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL, MLIR, ONNX) and confidence learning adjacent frameworks quickly. Comfort operating at multiple levels of abstraction — from high-level model behaviour ...

Staff ML Performance Engineer (Inference Optimisation)

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
systems with tight constraints (latency, memory, bandwidth, power/thermal, or cost).Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL) and confidence learning adjacent frameworks quickly. Comfort operating at multiple levels of abstraction — from high-level model behaviour down ...

Senior Embedded Software Engineer

Location
Southampton, England, United Kingdom
equivalent Due to security access you must have been in the U.K. for at least five years MATLAB/Simulink OpenGL or Vulkan CUDA or OpenCL FPGA or HDL exposure Video processing or machine learning Experience within defence, aerospace, automotive, or other high‐reliability industries #EmbeddedSoftware #EmbeddedEngineer #Cpp #Linux ...

Research Engineer (Inference & Serving)

Location
Greater London, England, United Kingdom
architectures Optional Bonus Research engagement: advanced degree with research output, top-tier publications (NeurIPS, ICML, MLSys, OSDI), or open-source contributions GPU kernel work - CUDA, Triton, or similar Experience with quantisation, speculative decoding, disaggregated inference, or KV-cache compression Shortlisted candidates will be contacted within 48 hours. #J ...

Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid) |

Hiring Organisation
Pure Resourcing Solutions
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£90,000 - £120,000 per annum
building performance models or calculators (Python or spreadsheet-based) that actually forecast how a system will behave Hands-on GPU/accelerator code optimisation, CUDA or similar Genuine understanding of how LLMs and deep learning models run on real hardware, training versus inference, matrix multiplication, KV-caching, that level ...

MLOps Engineer (LLM/GenAI)

Location
Sheffield, England, United Kingdom
skills: Extensive experience in building AI platforms covering model hosting/inference optimisation and fine-tuning pipelines (LLM experience strongly preferred) Strong Python and CUDA engineering; solid understanding of GPU/CPU architecture and HPC fundamentals Deep inference optimisation expertise: KV-cache, batching, quantisation (INT4/FP8/GPTQ ...

Machine Learning Researcher

Location
City Of London, England, United Kingdom
finance, trading, or quantitative research (not required). Publications, competition results (e.g., Kaggle, academic ML contests), or open-source contributions. Familiarity with C++, CUDA, or low-latency systems. Here is why you should join our dynamic team: Opportunity to work at one of the world's leading algorithmic trading ...

Product Engineer, Physical AI

Location
Greater London, England, United Kingdom
these — as long as you're open to learning, please apply. Backend: Python Frontend: TypeScript and React Deployment: Kubernetes Infrastructure: GCP Machine learning: PyTorch, CUDA, Ray Why Encord Competitive salary, commission, and meaningful equity in a high-growth startup Strong in-person culture — most of the team works from ...

GPU Infrastructure Lead - Systems Integrator

Location
Greater London, England, United Kingdom
deployments, monitoring and alerting stacks, and hardware acceptance and regression testing. Deep familiarity with the NVIDIA technology stack, including HGX platforms, NVLink/NVSwitch, CUDA-level debugging, and NCCL performance tuning. Strong leadership skills with the ability to set technical direction, solve complex problems, and guide engineering teams. Excellent ...

Electronic Engineer

Location
Greater London, England, United Kingdom
into existing AI and HPC infrastructure. Its current PT-2 system, a rack-mounted quantum computer, connects directly to GPU clusters via NVIDIA’s CUDA-Q platform, enabling accelerated AI workloads with significantly lower energy consumption than traditional silicon-based systems. Backed by a strong portfolio of patent families ...

ML Engineering Manager, Industrial Vision & Robotics

Location
England, United Kingdom
systems, collaborating with software, product management and research teams. You will require BS+10y or MS+7y in CS/ML/Robotics, expert Python and CUDA/PyTorch skills, and a track record in computer vision and model fine-tuning. #J-18808-Ljbffr ...

Software Engineer - Systems

Location
Greater London, England, United Kingdom
engineering Familiarity with tools like Apache Arrow, Parquet, DataFusion, Clickhouse, or DuckDB is a plus Understanding of cutting-edge ML infrastructure stack (e.g. PyTorch, CUDA) is also a plus Experience with Rust is a bonus Willingness to work in-person at our NYC or London office #J-18808-Ljbffr ...

AI Research Engineer

Location
Greater London, England, United Kingdom
plus; a strong background in mathematics is required. Alternatively, experience training models with strong math skills. Strong coding ability in Rust (preferred), CUDA, or C. Solid understanding of transformer architectures. A product mindset; you’ll build production‐ready products. Hybrid Remote with the London Office. What ...

Research Scientist, World Models, DeepMind

Location
Greater London, England, United Kingdom
years of experience working in frontier AI research labs on pre-training or post-training teams. Experience writing TPU/GPU kernels (e.g., JAX, CUDA) to optimize model performance and real-time inference. Experience designing, training, and scaling generative pixel or real-time video architectures. Track record of cross ...

ML Performance Engineer – Scale GPU/CPU Workloads

Location
Greater London, England, United Kingdom
evolve the compute stack. The role shapes platform evolution and enables researchers to push the boundaries of machine learning. You will work with Python, CUDA, Kubernetes, and deep learning frameworks like PyTorch, applying #J-18808-Ljbffr ...

ML Systems Engineer Intern — High-Perf Finance

Location
Greater London, England, United Kingdom
apply advanced techniques, optimize training pipelines on HPC clusters, and integrate ML models into latency-sensitive production environments, using C/C++, Python, and CUDA across large-scale systems. #J-18808-Ljbffr ...

ML Systems Engineer for Quantitative Finance

Location
Greater London, England, United Kingdom
systems for quantitative finance. You will optimize training pipelines, deploy models to low-latency production environments, and work across C/C++, Python, CUDA, and related GPU technologies. Join a fast-paced, collaborative team focused on impactful projects that push the boundaries of AI research and its applications ...