12 of 12 ONNX Jobs in London

Senior Machine Learning Engineer

Location
Greater London, England, United Kingdom
including quantization, batching, caching, compilation, and serving runtime tuning. Experience with large‐scale inference serving frameworks or runtimes such as SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, llama.cpp, or similar. Familiarity with LLMs, speech models, vision models, embeddings, multimodal models, retrieval-augmented generation, or other modern deep learning architectures. Experience ...

Machine Learning Engineer

Hiring Organisation
RedTech Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£75,000
following would be beneficial: LangChain and/or LangGraph Python memory management and performance optimisation GPU optimisation and acceleration PyTorch ONNX/ONNX Runtime Model inference and acceleration Responsibilities for Machine Learning Engineer: Develop and maintain high-quality Python software within advanced commercial AI products Deploy and integrate new machine …/Machine Learning Software Engineer/AI Integration Engineer/Python/Large Language Models/LLM/LangChain/LangGraph/PyTorch/ONNX/ONNX Runtime/NumPy/pandas/Machine Learning/Artificial Intelligence/GPU Optimisation/Model Inference/Model Acceleration/Deep Learning ...

Technical Architect

Hiring Organisation
Venturi
Location
City of London, London, United Kingdom
registry, versioning, CI/CD, drift detection, and retraining logic Working out how to run models in constrained or edge settings: quantisation, pruning, distillation, ONNX/TensorRT, and picking the right accelerators for the job Setting the evaluation approach - where precision and recall trade off, where thresholds ...

Senior Machine Learning Engineer

Location
Greater London, England, United Kingdom
quality is non-negotiable. ML Stack Highlights Python and Rust. We keep things simple but use the right tool for the job Rust with ONNX in-process model execution where throughput is critical Chalk.ai as our Feature Store GCP Agent Platform Endpoints for model serving Cloud-native on GCP. Services ...

Staff ML Performance Engineer (Compiler)

Location
Greater London, England, United Kingdom
power/thermal, or cost). Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL, MLIR, ONNX) and confidence learning adjacent frameworks quickly. Comfort operating at multiple levels of abstraction — from high-level model behaviour down to low-level kernel/runtime ...

AI Systems Researcher/Engineer New

Location
Greater London, England, United Kingdom
evaluating ML models and/or privacy & security technologies for on-device performance. Experience with modern Machine Learning (ML) frameworks (e.g., Pytorch, TensorFlow, MLX, ONNX Runtime, Core ML, TensorRT, etc.). Skills That Will Set You Apart from the Competition Open source projects on mobile/desktop systems optimization Experience ...

Founding Research Engineer

Location
Greater London, England, United Kingdom
exploration tooling to understand correlations, sparsity, and structure across sources Deploy models to production via the API and platform — including the gritty details when ONNX export or torch.compile breaks Iterate on model capabilities based on direct customer feedback Help shape the engineering and research culture as the team scales What ...

Senior ML Systems Engineer — Rust & Distributed Infra

Location
Greater London, England, United Kingdom
across Rust, Python, and cloud environments. Join a quick-moving, growth-focused team focused on distributed systems, Kubernetes orchestration, and performance-tuned GPU inference (ONNX) in production, while shaping robust, scalable architecture and contributing to a culture of #J-18808-Ljbffr ...

Senior Machine Learning Engineer

Hiring Organisation
Hackajob Ltd
Location
London, United Kingdom
Employment Type
Permanent, Work From Home
optimize open-source SLMs (e.g., Gemma 3, Llama 3) and vision-language models for execution on low-power edge runtimes (LiteRT/TensorFlow Lite, ONNX Runtime, ExecuTorch). Knowledge Analytics & Graph Processing: Design, implement, and maintain lightweight on-device graph databases and relationship extraction pipelines (Python, Rust, or C++ … V1.1/Dependable AI ). ML & Edge Inference Mastery 3+ years of production experience deploying ML models to edge runtime environments (LiteRT/TFLite, ONNX, C++ bindings). Experience in model quantization techniques (INT8, INT4, AWQ) and execution acceleration across NPU/GPU hardware. Proficiency in Python and PyTorch/ ...

Senior Machine Learning Engineer

Location
Greater London, England, United Kingdom
robust, scalable and reproducible training and inference workflows. What you will do Develop robust, scalable and reproducible inference pipelines. Deploy models into production using ONNX, TensorRT or similar frameworks. Build and optimise scalable machine learning training workflows. Optimise data loading, logging, checkpointing and resource utilisation for large-scale model training. … optimising and maintaining ML training and inference pipelines. Experience with Docker and containerised ML workflows. Experience with model deployment and optimisation frameworks such as ONNX, TensorRT or similar tools. Experience with GPU-based training, model serving and compute optimisation. Experience building and maintaining CI pipelines. Experience with cloud environments. Strong ...

AI Engineer

Location
City Of London, England, United Kingdom
DRW is a diversified trading firm with over 3 decades of experience bringing sophisticated technology and exceptional people together to operate in markets around the world. We value autonomy and the ability to quickly pivot ...

Senior Software Engineer, Machine Learning Services

Location
Greater London, England, United Kingdom
specific needs. Dive deep into the entire stack, from Kubernetes and container orchestration, through gRPC‐based service communication, to the performance tuning of ONNX‐based inference on GPU‐accelerated hardware. Write clean, efficient, and rigorously tested code. We value simplicity, correctness, and peer review. What you'll bring … challenges of managing the lifecycle of models in a multi‐tenant, high‐availability system. Familiarity with building ML inference services, model serialization (e.g., ONNX), and GPU programming (CUDA). You've built or worked on custom storage or job‐queueing systems before and have the scars to prove it. Maybe ...