1 to 25 of 30 Permanent ONNX Jobs

Senior Machine Learning Engineer

Location
Greater London, England, United Kingdom
including quantization, batching, caching, compilation, and serving runtime tuning. Experience with large‐scale inference serving frameworks or runtimes such as SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, llama.cpp, or similar. Familiarity with LLMs, speech models, vision models, embeddings, multimodal models, retrieval-augmented generation, or other modern deep learning architectures. Experience ...

Senior Audio AI Research Engineer

Hiring Organisation
Logitech
Location
London, UK
Employment Type
Full-time
coding practices (e.g., open-source contributions).Expertise in performance analysis and optimization of ML systems. Hands-on deployment experience with systems like TensorFlow Lite, ONNX, TVM, and Glow. Strong familiarity with cloud compute environments, ideally AWS, and data ingestion (e.g., TF Data pipelines).Knowledge of source control and project tracking ...

Machine Learning Engineer

Hiring Organisation
RedTech Recruitment Ltd
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£75,000
following would be beneficial: LangChain and/or LangGraph Python memory management and performance optimisation GPU optimisation and acceleration PyTorch ONNX/ONNX Runtime Model inference and acceleration Responsibilities for Machine Learning Engineer: Develop and maintain high-quality Python software within advanced commercial AI products Deploy and integrate new machine …/Machine Learning Software Engineer/AI Integration Engineer/Python/Large Language Models/LLM/LangChain/LangGraph/PyTorch/ONNX/ONNX Runtime/NumPy/pandas/Machine Learning/Artificial Intelligence/GPU Optimisation/Model Inference/Model Acceleration/Deep Learning ...

Senior Applied Diffusion Engineer

Hiring Organisation
ASOS
Location
London, UK
Employment Type
Full-time
pose estimation, and embeddings. Experience building production ML systems and inference pipelines. Strong Python engineering skills. Experience with GPU optimisation, CUDA fundamentals, TensorRT, or ONNX is advantageous. Experience in fashion, e-commerce, creative tooling, gaming avatars, or virtual try-on is highly desirable. Additional InformationBeneFITS' Employee discount (hello ASOS discount ...

Embedded Machine Learning Engineer

Hiring Organisation
KO2 Embedded Recruitment Solutions LTD
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent
Salary
£70,000
world experience working with sensor data, time-series data, or IoT data streams Familiarity with embedded ML tools and approaches: TensorFlow Lite, Edge Impulse, ONNX Runtime, or equivalent Hands-on mindset; comfortable getting close to hardware, firmware code, and the real-world constraints of device deployment Clear communication; ability ...

Machine Learning Framework/Runtime Software Engineer

Location
United Kingdom
building machine learning inference engines, runtime systems, or backend integration frameworks. Experience working with a machine learning inference framework such as LiteRT, TensorFlow Lite, ONNX Runtime, or a similar technology is also important. A good understanding of how AI models execute in practice is essential, including graph processing, operator execution ...

ML Integration Engineer – Cyber AI, Cambridge (Hybrid)

Location
Cambridge, England, United Kingdom
making independent decisions, while also being able to collaborate effectively within a team,* Familiar with common machine learning and model acceleration frameworks (e.g. PyTorch, ONNX, ONNX Runtime)* Experienced with Python data and matrix manipulation libraries (e.g. numpy and pandas),* Knowledgeable about Python memory management and optimising GPU usage (beneficial ...

Senior Machine Learning Engineer

Location
Greater London, England, United Kingdom
quality is non-negotiable. ML Stack Highlights Python and Rust. We keep things simple but use the right tool for the job Rust with ONNX in-process model execution where throughput is critical Chalk.ai as our Feature Store GCP Agent Platform Endpoints for model serving Cloud-native on GCP. Services ...

Staff ML Performance Engineer (Compiler)

Location
Greater London, England, United Kingdom
power/thermal, or cost). Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL, MLIR, ONNX) and confidence learning adjacent frameworks quickly. Comfort operating at multiple levels of abstraction — from high-level model behaviour down to low-level kernel/runtime ...

Staff ML Performance Engineer (Compiler)

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
bandwidth, power/thermal, or cost).Strong proficiency with at least one relevant stack/toolchain (e.g. TensorRT, CUDA, Qualcomm QNN, Triton, OpenCL, MLIR, ONNX) and confidence learning adjacent frameworks quickly. Comfort operating at multiple levels of abstraction — from high-level model behaviour down to low-level kernel/runtime ...

Technical Architect

Hiring Organisation
Venturi
Location
City of London, London, United Kingdom
registry, versioning, CI/CD, drift detection, and retraining logic Working out how to run models in constrained or edge settings: quantisation, pruning, distillation, ONNX/TensorRT, and picking the right accelerators for the job Setting the evaluation approach - where precision and recall trade off, where thresholds ...

Staff Engineering Manager

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
teams. Excellent communications skills both written and verbal"Nice To Have" Skills and Experience: Experience with AI/ML frameworks such as TensorFlow, PyTorch, ONNX, or inference runtimes. System bring-up and JTAG debugging expertiseGood background of system performance analysisExperience with RTL simulation tools and software development toolsIn Return ...

AI Systems Researcher/Engineer New

Location
Greater London, England, United Kingdom
evaluating ML models and/or privacy & security technologies for on-device performance. Experience with modern Machine Learning (ML) frameworks (e.g., Pytorch, TensorFlow, MLX, ONNX Runtime, Core ML, TensorRT, etc.). Skills That Will Set You Apart from the Competition Open source projects on mobile/desktop systems optimization Experience ...

Principal Product Manager - AI Tooling

Location
Cambridge, England, United Kingdom
demonstrate: Proven product management in development tools and software including CLIs, SDKs and APIs. Knowledge of machine learning frameworks, runtimes, infrastructure such as PyTorch, ONNX, ExecuTorch, Llama.cpp, vLLM and LiteRT. Understanding of model deployment to edge, embedded or heterogeneous computing, with an understanding of trade-offs between accuracy, performance ...

Senior ML Systems Engineer — Rust & Distributed Infra

Location
Greater London, England, United Kingdom
across Rust, Python, and cloud environments. Join a quick-moving, growth-focused team focused on distributed systems, Kubernetes orchestration, and performance-tuned GPU inference (ONNX) in production, while shaping robust, scalable architecture and contributing to a culture of #J-18808-Ljbffr ...

Machine Learning Framework/Runtime Software Engineer Arm

Location
United Kingdom
building machine learning inference engines, runtime systems, or backend integration frameworks. Experience working with a machine learning inference framework such as LiteRT, TensorFlow Lite, ONNX Runtime, or a similar technology is also important. A good understanding of how AI models execute in practice is essential, including graph processing, operator execution ...

Senior MLOps Engineer

Hiring Organisation
Searchability NS&D
Location
Brentford, England, United Kingdom
infrastructure experience Experience with distributed ML training using Ray, DeepSpeed or PyTorch Lightning Strong MLflow, model registry and ML lifecycle experience Experience with TensorRT, ONNX, quantisation and model deployment Strong Docker, CI/CD and GitHub Actions experience Experience with Prometheus/Grafana monitoring and production observability Cloud/edge …/AI Engineer/Machine Learning/ML/Python/PyTorch/Ray/DeepSpeed/PyTorch Lightning/MLflow/TensorRT/ONNX/Docker/Kubernetes/CI/CD/GitHub Actions/Prometheus/Grafana/Metaflow/Kafka/NVIDIA Jetson/DeepStream/ ...

AI Compiler Optimization Engineer

Location
City of Edinburgh, Scotland, United Kingdom
improvements Preferred: Experience with LLVM/MLIR development AI Model Profiling & Framework Optimization: Profile end-to-end inference workflows on frameworks like TensorFlow, PyTorch, ONNX, and llama.cpp to identify hotspots and bottlenecks Propose and implement optimization strategies (e.g., kernel tuning, graph-level optimizations) Preferred: Experience optimizing models on multiple ...

AI Compiler Optimization Engineer (Hybrid CPU/XPU)

Location
City of Edinburgh, Scotland, United Kingdom
model inference performance on CPU and CPU/XPU hybrid systems, using advanced compiler techniques. You will profile frameworks such as TensorFlow, PyTorch and ONNX, optimize graph execution, and contribute to open research with practical insights and publications. #J-18808-Ljbffr ...

AI Compiler Optimization Engineer - Edinburgh

Hiring Organisation
Microtech Global Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent
improvements Preferred: Experience with LLVM/MLIR development AI Model Profiling & Framework Optimization: Profile end-to-end inference workflows on frameworks like TensorFlow, PyTorch, ONNX, and llama.cpp to identify hotspots and bottlenecks Propose and implement optimization strategies (e.g., kernel tuning, graph-level optimizations) Preferred: Experience optimizing models on multiple ...

Senior Embedded Linux & Vision AI Architect

Location
United Kingdom
products, moving from prototypes to production devices deployed in customer fleets. The role combines kernel/driver work with edge AI deployment (TensorFlow Lite, ONNX Runtime) and requires mentoring peers, ensuring reliability, security, and field readiness. #J-18808-Ljbffr ...

Senior Software Engineer

Hiring Organisation
IBEX RECRUITMENT LTD
Location
London, United Kingdom
Employment Type
Permanent
Salary
£90,000
Zero Trust architectures and principles Data classification and secure data handling AI-enabled software platforms Supporting AI model deployment or model conversion ONNX (Open Neural Network Exchange) Secure-by-design software engineering practices The Ideal Candidate You'll likely be a Senior Product Engineer or Senior Software Engineer with experience ...

Senior Machine Learning Engineer

Location
Cambridge, England, United Kingdom
robust, scalable and reproducible training and inference workflows. What you will do Develop robust, scalable and reproducible inference pipelines. Deploy models into production using ONNX, TensorRT or similar frameworks. Build and optimise scalable machine learning training workflows. Optimise data loading, logging, checkpointing and resource utilisation for large-scale model training. … optimising and maintaining ML training and inference pipelines. Experience with Docker and containerised ML workflows. Experience with model deployment and optimisation frameworks such as ONNX, TensorRT or similar tools. Experience with GPU-based training, model serving and compute optimisation. Experience building and maintaining CI pipelines. Experience with cloud environments. Strong ...

Senior Machine Learning Engineer

Location
Greater London, England, United Kingdom
robust, scalable and reproducible training and inference workflows. What you will do Develop robust, scalable and reproducible inference pipelines. Deploy models into production using ONNX, TensorRT or similar frameworks. Build and optimise scalable machine learning training workflows. Optimise data loading, logging, checkpointing and resource utilisation for large-scale model training. … optimising and maintaining ML training and inference pipelines. Experience with Docker and containerised ML workflows. Experience with model deployment and optimisation frameworks such as ONNX, TensorRT or similar tools. Experience with GPU-based training, model serving and compute optimisation. Experience building and maintaining CI pipelines. Experience with cloud environments. Strong ...

Staff ML Engineer - Developer Tools

Location
Cambridge, England, United Kingdom
model optimisation techniques such as quantisation, graph optimisation, operator fusionandprecision reduction Experience across multiple ML frameworks, model formats and inference runtimes, such as PyTorch, ONNX/ONNX Runtime, ExecuTorch, TensorFlow/LiteRT and OpenVINO. Experience analysing, profiling and debugging ML workloads, including model compatibility, performance and the trade-offs between ...