26 to 50 of 75 vLLM Jobs in the UK

AI Engineer (Fluent Portuguese & English)

Hiring Organisation
Chubb
Location
London, UK
Employment Type
Full-time
experience with LLM training cycles, parameter-efficient fine-tuning (PEFT), and sophisticated prompt engineering. Inference Stack: Experience with high-performance inference servers (e.g., vLLM, TGI, or Triton) and an understanding of how to optimize models for GPU deployment. Infrastructure: Comfortable working in Linux-based environments and proficient in managing containerized ...

AI Engineer (Fluent in Mandarin & English)

Location
Greater London, England, United Kingdom
experience with LLM training cycles, parameter-efficient fine-tuning (PEFT), and sophisticated prompt engineering. Inference Stack: Experience with high-performance inference servers (e.g., vLLM, TGI, or Triton ) and an understanding of how to optimize models for GPU deployment. Infrastructure: Comfortable working in Linux-based environments and proficient in managing containerized ...

Senior AI Platform Engineer

Hiring Organisation
IQVIA
Location
London, United Kingdom
Salary
£ 80 K
PyTorch.Strong knowledge of distributed training and inference strategies, including tensor, pipeline, data, and expert parallelism approaches.Experience optimising LLM inference workloads using technologies such as vLLM, TensorRT-LLM, NVIDIA NIM, SGLang, or similar high-performance serving frameworks.Expertise in model optimisation techniques including quantisation, mixed precision training and inference (FP8, GPTQ ...

Senior AI Platform Engineer

Hiring Organisation
IQVIA
Location
London, UK
Employment Type
Full-time
knowledge of distributed training and inference strategies, including tensor, pipeline, data, and expert parallelism approaches. Experience optimising LLM inference workloads using technologies such as vLLM, TensorRT-LLM, NVIDIA NIM, SGLang, or similar high-performance serving frameworks. Expertise in model optimisation techniques including quantisation, mixed precision training and inference (FP8, GPTQ ...

Principal Data Engineer

Location
Greater London, England, United Kingdom
Hands‐on experience with cloud‐native ML infrastructure platforms Knowledge of vector databases (Pinecone, Weaviate, Qdrant) and embedding models Experience with model serving frameworks (vLLM, TensorRT, Ray) Familiarity with Anaplan or similar enterprise planning platforms Experience with A/B testing and experimentation frameworks for AI features Experience with model ...

AI Infrastructure Engineer, Serving Platform

Hiring Organisation
Scale AI
Location
London, United Kingdom
Salary
£ 100 K
Terraform).Proven ability to solve complex problems and work independently in fast-moving environments.Nice to haves:Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference.PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows ...

AI Infrastructure Engineer, Serving Platform London, UK Apply →

Location
Greater London, England, United Kingdom
Proven ability to solve complex problems and work independently in fast-moving environments. Nice to haves: Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference. PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This ...

Software Engineer (LLM Engineering), London

Hiring Organisation
Isomorphic Labs
Location
London, United Kingdom
Salary
£ 80 K
Deep understanding of how models are trained, how they work internally, and their inherent limitations.LLM Serving stack: Experience with the LLM serving stack (e.g. vLLM) for open weight models for bringing the latest models to internal users.ML Literacy: Experience evaluating probabilistic ML systems and managing model "tool use", contexts ...

Machine Learning Research Engineer (Foundational Research)

Hiring Organisation
Thomson Reuters
Location
London, UK
Employment Type
Full-time
production-quality code and data pipelines for ML systemsProficiency in modern AI development frameworks including: PyTorch, Jax , HuggingFace Transformers, LLM APIs (litellm etc) and vLLM for building and deploying large-scale AI applicationsUnderstanding of LLM training methodologies including instruction fine-tuning, preference optimization, and reinforcement learning approachesStrong software engineering skills ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
model training, evaluation, packaging, and production deployment. Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM). Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks. Monitor and optimise cloud spend across high-cost … Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio). Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins. Experience with vLLM, Ray, MLflow, LangChain/LangSmith, DeepSpeed, or Hugging Face TGI. Solid background in AWS/GCP/Azure, Kubecost, and GPU cost optimisation techniques. Strong ...

ML Ops Engineer

Hiring Organisation
Anaplan
Location
London, United Kingdom
Salary
£ 80 K
continuous model training, evaluation, packaging, and production deployment.Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM).Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks.Monitor and optimise cloud spend across high-cost GPU/… proficiency in Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio).Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.Experience with vLLM, Ray, MLflow, LangChain/LangSmith, DeepSpeed, or Hugging Face TGI.Solid background in AWS/GCP/Azure, Kubecost, and GPU cost optimisation techniques.Strong skills ...

Senior Principal AI Infrastructure Architect

Hiring Organisation
The Nippon Telegraph And Telephone Corporation (NTT)
Location
United Kingdom
Salary
£ 70 K
rail-optimised and fat-tree topologies. Working knowledge of the AI software and orchestration stack: CUDA, cuDNN, NCCL, ROCm, Triton Inference Server, NIM, vLLM, TensorRT-LLM, Slurm, Kubernetes (with GPU Operator), Kubeflow, Run:ai, MLflow and NVIDIA AI Enterprise. Familiarity with datacenter facilities engineering for AI workloads: high-density power ...

Senior Machine Learning Engineer

Location
Greater London, England, United Kingdom
large-scale models, including quantization, batching, caching, compilation, and serving runtime tuning. Experience with large‐scale inference serving frameworks or runtimes such as SGLang, vLLM, TensorRT-LLM, ONNX Runtime, Triton, llama.cpp, or similar. Familiarity with LLMs, speech models, vision models, embeddings, multimodal models, retrieval-augmented generation, or other modern deep ...

LLM & Generative AI Engineer

Location
Greater London, England, United Kingdom
Processing and Generative AI Hands-on experience with PyTorch, Hugging Face Transformers, and parameter-efficient fine-tuning (PEFT/LoRA) Experience deploying models with vLLM, Ollama, or Triton Inference Server Strong background in software engineering best practices and clean code #J-18808-Ljbffr ...

Lead Software Engineer - Python / Go & AI/ML

Location
Auchentibber, Scotland, United Kingdom
skills Formal training or certification on software engineering concepts and advanced applied experience – preferably Go/Python Hands-on experience with LLM inference systems — vLLM, TensorRT-LLM, SGLang, LLM-D, or equivalent production serving engines Strong understanding of GPU memory architecture, including KV cache sizing and dynamics, memory-bandwidth versus ...

Remote Senior Machine Learning Engineer

Location
Stirling, Scotland, United Kingdom
data, including classification, extraction, embeddings, re-rankers, clustering, and search. Experience in instruction fine-tuning and serving language models, familiarity with frameworks such as vLLM, DeepSpeed, or similar tools A solid grounding in classical ML and statistics, and the judgement to choose simpler methods when they’re the right solution. ...

Lead Software Engineer - Python / Go & AI/ML

Hiring Organisation
JP Morgan Chase
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 80 K
skillsFormal training or certification on software engineering concepts and advanced applied experience – preferably Go/Python Hands-on experience with LLM inference systems — vLLM, TensorRT-LLM, SGLang, LLM-D, or equivalent production serving enginesStrong understanding of GPU memory architecture, including KV cache sizing and dynamics, memory-bandwidth versus compute bottlenecks ...

Senior ML Systems Engineer, Frameworks & Tooling

Location
Greater London, England, United Kingdom
training LLMs or other large transformer architectures. Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.). Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches). Experience with data pipeline optimization, sharded datasets, or caching strategies. Background in performance engineering, profiling, or low-level systems. ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, United Kingdom
Salary
£ 80 K
Experience with training LLMs or other large transformer architectures.Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.).Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches).Experience with data pipeline optimization, sharded datasets, or caching strategies.Background in performance engineering, profiling, or low-level systems.Bonus: paper ...

Senior Research Scientist | Model Steering

Location
Greater London, England, United Kingdom
without degrading their reasoning capabilities. Experience with machine translation, multilingual NLP, or language quality estimation. Familiarity with inference and serving at scale (e.g. via vLLM, SGLang, TensorRT‐LLM, etc) and long‐context modelling. Publications at top‐tier venues. What we offer Diverse and internationally distributed team : joining our team means ...

LLM & Generative AI Engineer: RAG & Fine-Tuning

Location
Greater London, England, United Kingdom
semantic search solutions. The role requires hands-on experience with PyTorch, Hugging Face Transformers, and parameter-efficient fine-tuning, plus deploying models with vLLM, Ollama, or Triton. Strong software engineering practices are essential. #J-18808-Ljbffr ...

AI Research Scientist

Location
Greater London, England, United Kingdom
engineering & synthetic data pipelines Agentic frameworks, reasoning, tool use, memory systems Evaluation & benchmarking (LLM-as-judge, safety/alignment metrics) Inference optimisation (quantization, distillation, vLLM/TensorRT-LLM) Requirements: - PhD or Master's in CS/AI/ML/Physics/Maths or equivalent research experience - Strong hands ...

AI Research Scientist

Hiring Organisation
Intellectual Capital Resources
Location
London, United Kingdom
Salary
£ 60 K
engineering & synthetic data pipelines Agentic frameworks, reasoning, tool use, memory systems Evaluation & benchmarking (LLM-as-judge, safety/alignment metrics) Inference optimisation (quantization, distillation, vLLM/TensorRT-LLM) Requirements: PhD or Master's in CS/AI/ML/Physics/Maths or equivalent research experience Strong hands ...

Natural Language Processing Researcher

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 70 K
ability to work both independently and as part of a team.Strong programming skills in Python and experience with machine learning libraries such as PyTorch, vLLM or similar are a prerequisiteYou will have, or be working towards gaining, a Masters or PhD degree in NLP or a related quantitative subject, such ...

Performance Engineer, Containers/Serverless

Location
Greater London, England, United Kingdom
compatible object storage - including performance-killing cases (small-object overhead, range-request patterns, eventual consistency, multipart tuning). Comfort with model-serving runtimes (vLLM, SGLang etc) and the formats they consume (safetensors, GGUF, sharded checkpoints). An end-to-end view: comfortable reasoning about NIC, switch, filesystem, cache, container runtime ...