26 to 50 of 85 vLLM Jobs in the UK

Software Engineer, Machine Learning Infrastructure

Hiring Organisation
Deliveroo
Location
London, United Kingdom
Salary
£ 80 K
full software development lifecycle, including designing, generating code, testing, monitoring and releasing softwareNice To HavesExperience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in productionExperience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluationGPU performance ...

Software Engineer, GenAI Platform

Hiring Organisation
Deliveroo
Location
London, UK
Employment Type
Full-time
full software development lifecycle, including designing, generating code, testing, monitoring and releasing softwareNice To HavesExperience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in productionExperience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluationGPU performance ...

AI Platform Engineer

Location
Greater London, England, United Kingdom
major cloud Hands‐on experience with LLM infrastructure: a model gateway pattern (such as LiteLLM or in‐house), model serving (such as vLLM or managed endpoints), and vector or retrieval systems Agent‐systems experience: you have built with an agent orchestration framework (such as LangGraph or first‐party agent SDKs ...

Senior / Principal Applied AI Engineer (UK / Europe, Remote)

Location
United Kingdom
haves Fintech/trading/market-data domain experience or experience as a trading-platform user. Open-source contributions to AI tooling (LangChain, LlamaIndex, vLLM, DSPy, or similar). Production RAG and vector-database experience; experience with code-generation, transpilation, or developer-experience tooling. Logistics Location: Europe (EU/ ...

AI Engineer (Fluent Portuguese & English)

Hiring Organisation
Chubb
Location
London, United Kingdom
Salary
£ 80 K
Hands-on experience with LLM training cycles, parameter-efficient fine-tuning (PEFT), and sophisticated prompt engineering.Inference Stack: Experience with high-performance inference servers (e.g., vLLM, TGI, or Triton) and an understanding of how to optimize models for GPU deployment.Infrastructure: Comfortable working in Linux-based environments and proficient in managing containerized ...

AI Engineer (Fluent in Mandarin & English)

Location
Greater London, England, United Kingdom
experience with LLM training cycles, parameter-efficient fine-tuning (PEFT), and sophisticated prompt engineering. Inference Stack: Experience with high-performance inference servers (e.g., vLLM, TGI, or Triton ) and an understanding of how to optimize models for GPU deployment. Infrastructure: Comfortable working in Linux-based environments and proficient in managing containerized ...

Machine Learning Research Engineer (Foundational Research)

Hiring Organisation
Thomson Reuters
Location
London, UK
Employment Type
Full-time
production-quality code and data pipelines for ML systemsProficiency in modern AI development frameworks including: PyTorch, Jax , HuggingFace Transformers, LLM APIs (litellm etc) and vLLM for building and deploying large-scale AI applicationsUnderstanding of LLM training methodologies including instruction fine-tuning, preference optimization, and reinforcement learning approachesStrong software engineering skills ...

Platform Lead - MLOps

Location
Greater London, England, United Kingdom
Infrastructure as Code (IaC), and deploy Large Language Models (LLMs) into production. By architecting resilient LLMOps pipelines and utilising serving frameworks like Triton, vLLM, or Hugging Face TGI, you will guarantee ultra-low latency, high availability, and proactive model monitoring. Crucially, you will instil financial accountability by establishing robust FinOps ...

AI Infrastructure Engineer, Serving Platform

Hiring Organisation
Scale AI
Location
London, United Kingdom
Salary
£ 100 K
Terraform).Proven ability to solve complex problems and work independently in fast-moving environments.Nice to haves:Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference.PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows ...

AI Infrastructure Engineer, Serving Platform London, UK Apply →

Location
Greater London, England, United Kingdom
Proven ability to solve complex problems and work independently in fast-moving environments. Nice to haves: Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference. PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This ...

Senior Machine Learning Engineer

Location
United Kingdom
systems thinking understanding how ML decisions impact infrastructure High ownership and comfort operating in a fast-paced startup environment Nice to have Experience with vLLM or custom inference servers Experience with Kubernetes or containerised ML workloads Experience working in high-throughput distributed systems Background in AI media generation (image, video ...

Forward Deployed Engineer, EMEA

Location
Greater London, England, United Kingdom
eligible for sponsorship Bonus Points AI Voice Assistants, conversational AI, STT/TTS, or real‐time AI experience Experience with open‐weight models, vLLM, SGLang, TGI, or Ollama LLM fine‐tuning, RAG, evaluation, or inference optimization Experience sizing and optimizing GPU infrastructure Background in telecom, CPaaS, cloud/AI infrastructure ...

Senior Principal AI Infrastructure Architect

Hiring Organisation
The Nippon Telegraph And Telephone Corporation (NTT)
Location
United Kingdom
Salary
£ 70 K
rail-optimised and fat-tree topologies. Working knowledge of the AI software and orchestration stack: CUDA, cuDNN, NCCL, ROCm, Triton Inference Server, NIM, vLLM, TensorRT-LLM, Slurm, Kubernetes (with GPU Operator), Kubeflow, Run:ai, MLflow and NVIDIA AI Enterprise. Familiarity with datacenter facilities engineering for AI workloads: high-density power ...

Sr. Software Engineer, Inference

Location
Greater London, England, United Kingdom
CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies). Active open-source or production contributions to modern inference frameworks (e.g., vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe). Experience leading multi-team technical initiatives or partnering directly with enterprise customers on mission-critical platform launches. ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
model training, evaluation, packaging, and production deployment. Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM). Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks. Monitor and optimise cloud spend across high-cost … Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio). Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins. Experience with vLLM, Ray, MLflow, LangChain/LangSmith, DeepSpeed, or Hugging Face TGI. Solid background in AWS/GCP/Azure, Kubecost, and GPU cost optimisation techniques. Strong ...

ML Ops Engineer

Hiring Organisation
Anaplan
Location
London, United Kingdom
Salary
£ 80 K
continuous model training, evaluation, packaging, and production deployment.Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM).Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks.Monitor and optimise cloud spend across high-cost GPU/… proficiency in Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio).Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.Experience with vLLM, Ray, MLflow, LangChain/LangSmith, DeepSpeed, or Hugging Face TGI.Solid background in AWS/GCP/Azure, Kubecost, and GPU cost optimisation techniques.Strong skills ...

Platform Lead, MLOps & LLM Infra

Location
Greater London, England, United Kingdom
define the AI/ML infrastructure roadmap, automate multi-cloud provisioning via IaC, and deploy LLMs into production with resilient LLMOps pipelines using Triton, vLLM, or Hugging Face TGI for ultra-low latency. You will establish robust FinOps frameworks to manage expensive GPU/CPU cloud budgets, enable auto-scaling ...

LLM & Generative AI Engineer

Location
Greater London, England, United Kingdom
Processing and Generative AI Hands-on experience with PyTorch, Hugging Face Transformers, and parameter-efficient fine-tuning (PEFT/LoRA) Experience deploying models with vLLM, Ollama, or Triton Inference Server Strong background in software engineering best practices and clean code #J-18808-Ljbffr ...

Lead Software Engineer - Python / Go & AI/ML

Location
Auchentibber, Scotland, United Kingdom
skills Formal training or certification on software engineering concepts and advanced applied experience – preferably Go/Python Hands-on experience with LLM inference systems — vLLM, TensorRT-LLM, SGLang, LLM-D, or equivalent production serving engines Strong understanding of GPU memory architecture, including KV cache sizing and dynamics, memory-bandwidth versus ...

Remote Senior Machine Learning Engineer

Location
Stirling, Scotland, United Kingdom
data, including classification, extraction, embeddings, re-rankers, clustering, and search. Experience in instruction fine-tuning and serving language models, familiarity with frameworks such as vLLM, DeepSpeed, or similar tools A solid grounding in classical ML and statistics, and the judgement to choose simpler methods when they’re the right solution. ...

Senior ML Systems Engineer, Frameworks & Tooling

Location
Greater London, England, United Kingdom
training LLMs or other large transformer architectures. Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.). Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches). Experience with data pipeline optimization, sharded datasets, or caching strategies. Background in performance engineering, profiling, or low-level systems. ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, United Kingdom
Salary
£ 80 K
Experience with training LLMs or other large transformer architectures.Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.).Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches).Experience with data pipeline optimization, sharded datasets, or caching strategies.Background in performance engineering, profiling, or low-level systems.Bonus: paper ...

Senior Research Scientist | Model Steering

Location
Greater London, England, United Kingdom
without degrading their reasoning capabilities. Experience with machine translation, multilingual NLP, or language quality estimation. Familiarity with inference and serving at scale (e.g. via vLLM, SGLang, TensorRT‐LLM, etc) and long‐context modelling. Publications at top‐tier venues. What we offer Diverse and internationally distributed team : joining our team means ...

LLM & Generative AI Engineer: RAG & Fine-Tuning

Location
Greater London, England, United Kingdom
semantic search solutions. The role requires hands-on experience with PyTorch, Hugging Face Transformers, and parameter-efficient fine-tuning, plus deploying models with vLLM, Ollama, or Triton. Strong software engineering practices are essential. #J-18808-Ljbffr ...

AI Research Scientist

Location
Greater London, England, United Kingdom
engineering & synthetic data pipelines Agentic frameworks, reasoning, tool use, memory systems Evaluation & benchmarking (LLM-as-judge, safety/alignment metrics) Inference optimisation (quantization, distillation, vLLM/TensorRT-LLM) Requirements: - PhD or Master's in CS/AI/ML/Physics/Maths or equivalent research experience - Strong hands ...