1 to 25 of 49 Permanent vLLM Jobs in London

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
frameworks: PyTorch and TensorFlow at a working, hands-on level.Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent multi-node training frameworks.Inference & serving: NVIDIA Triton, vLLM, TensorRT-LLM, or equivalent high-throughput serving platforms.MLOps/LLMOps: Kubeflow, MLflow, and at least one hyperscaler ML platform (SageMaker, Azure ML, or Vertex ...

Enterprise Architect - AI

Location
Greater London, England, United Kingdom
TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent multi-node training frameworks. Inference & serving: NVIDIA Triton, vLLM, TensorRT-LLM, or equivalent high-throughput serving platforms. MLOps/LLMOps: Kubeflow, MLflow, and at least one hyperscaler ML platform (SageMaker, Azure ML, or Vertex ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent multi-node training frameworks. Inference & serving: NVIDIA Triton, vLLM, TensorRT-LLM, or equivalent high-throughput serving platforms. MLOps/LLMOps: Kubeflow, MLflow, and at least one hyperscaler ML platform (SageMaker, Azure ML, or Vertex ...

ML Ops Lead

Hiring Organisation
Anaplan
Location
London, United Kingdom
Salary
£ 100 K
Kubernetes (K8s), Docker, and service meshes.Expert knowledge of Terraform, Ansible, Jenkins, or GitHub Actions.Proficient in Python, Bash, or Go.Familiarity with Triton Inference Server, vLLM, or Hugging Face TGI.Our Commitment to Diversity, Equity, Inclusion and Belonging (DEIB) We believe attracting and retaining the best talent and fostering an inclusive culture strengthens ...

Senior Forward Deployed ML Engineer, Agents

Location
Greater London, England, United Kingdom
plus Model fine-tuning practical experience with LoRA/QLoRA, supervised fine-tuning, or RLHF workflows is a plus Inference optimization experience with vLLM, TensorRT-LLM, Triton, or model quantization techniques is desirable Observability tooling practical experience with LLM monitoring, tracing, and evaluation frameworks is a strong plus Familiarity with ...

Associate Director Lead AI Architect

Hiring Organisation
Anson Mccade
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Foundry and other AWS, Azure or GCP services . Work with self-hosted inference and model-serving technologies where appropriate, including tools such as vLLM, SGLang and Ollama . Design evaluation, observability and monitoring frameworks to assess model performance, system health, reliability and business impact. Optimise AI solutions for performance ...

Associate Director

Hiring Organisation
Anson McCade
Location
London, United Kingdom
Salary
£ 100 K
Azure AI Foundry and other AWS, Azure or GCP services.Work with self-hosted inference and model-serving technologies where appropriate, including tools such as vLLM, SGLang and Ollama.Design evaluation, observability and monitoring frameworks to assess model performance, system health, reliability and business impact.Optimise AI solutions for performance, scalability, security ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
London, United Kingdom
Salary
£ 100 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Principal AI Engineer

Location
Greater London, England, United Kingdom
Hands‐on experience with cloud‐native ML infrastructure platforms Knowledge of vector databases (Pinecone, Weaviate, Qdrant) and embedding models Experience with model serving frameworks (vLLM, TensorRT, Ray) Experience with A/B testing and experimentation frameworks for AI features Contributions to open‐source ML projects or research publications Experience with ...

Senior Software Engineer, GenAI Platform

Location
Greater London, England, United Kingdom
software development lifecycle, including designing, generating code, testing, monitoring and releasing software. Nice To Haves Experience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in production Experience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation ...

Senior Software Engineer, GenAI Platform

Hiring Organisation
Deliveroo
Location
London, United Kingdom
Salary
£ 80 K
full software development lifecycle, including designing, generating code, testing, monitoring and releasing softwareNice To HavesExperience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in productionExperience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluationGPU performance ...

Principal Machine Learning Engineer

Location
Greater London, England, United Kingdom
Hands-on experience with cloud-native ML infrastructure platforms Knowledge of vector databases (Pinecone, Weaviate, Qdrant) and embedding models Experience with model serving frameworks (vLLM, TensorRT, Ray) Experience with A/B testing and experimentation frameworks for AI features Contributions to open-source ML projects or research publications Experience with ...

Principal AI Engineer

Location
Greater London, England, United Kingdom
Hands‐on experience with cloud‐native ML infrastructure platforms Knowledge of vector databases (Pinecone, Weaviate, Qdrant) and embedding models Experience with model serving frameworks (vLLM, TensorRT, Ray) Experience with A/B testing and experimentation frameworks for AI features Contributions to open‐source ML projects or research publications Experience with ...

Principal AI Engineer

Hiring Organisation
Anaplan
Location
London, United Kingdom
Salary
£ 80 K
related quantitative fieldHands-on experience with cloud-native ML infrastructure platformsKnowledge of vector databases (Pinecone, Weaviate, Qdrant) and embedding modelsExperience with model serving frameworks (vLLM, TensorRT, Ray)Experience with A/B testing and experimentation frameworks for AI featuresContributions to open-source ML projects or research publicationsExperience with model observability ...

Senior Machine Learning Engineer

Hiring Organisation
Wellcome Trust
Location
London, United Kingdom
Salary
£ 70 K
PyTorch, TensorFlow or JAX. Strong knowledge of modern natural language processing techniques, large language and transformer models, and libraries such as Hugging Face, VLLM and LangChain. Proven experience owning the development of large scale, high impact machine learning solutions from initial experimentation through to deployment. Strong knowledge of data science ...

Software Engineer, GenAI Platform

Hiring Organisation
Deliveroo
Location
London, United Kingdom
Salary
£ 80 K
full software development lifecycle, including designing, generating code, testing, monitoring and releasing softwareNice To HavesExperience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in productionExperience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluationGPU performance ...

Software Engineer, Machine Learning Infrastructure

Hiring Organisation
Deliveroo
Location
London, UK
Employment Type
Full-time
full software development lifecycle, including designing, generating code, testing, monitoring and releasing softwareNice To HavesExperience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in productionExperience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluationGPU performance ...

Manager, Research Engineering (Foundational Research)

Hiring Organisation
Thomson Reuters
Location
London, United Kingdom
Salary
£ 80 K
purely technical preference.Technical Expertise:Deep proficiency in Python and modern software development practices.Hands-on experience with Distributed Training infrastructure (Multi-node GPU training, Kubernetes, vLLM).Familiarity with Deep Learning frameworks (PyTorch).Experience with MLOps tools and experiment tracking (e.g., ClearML, MLFlow, Weights & Biases).Research Fluency: Ability to read technical research ...

AI Engineer (Fluent Portuguese & English)

Hiring Organisation
Chubb
Location
London, United Kingdom
Salary
£ 80 K
Hands-on experience with LLM training cycles, parameter-efficient fine-tuning (PEFT), and sophisticated prompt engineering.Inference Stack: Experience with high-performance inference servers (e.g., vLLM, TGI, or Triton) and an understanding of how to optimize models for GPU deployment.Infrastructure: Comfortable working in Linux-based environments and proficient in managing containerized ...

AI Engineer (Fluent in Mandarin & English)

Location
Greater London, England, United Kingdom
experience with LLM training cycles, parameter-efficient fine-tuning (PEFT), and sophisticated prompt engineering. Inference Stack: Experience with high-performance inference servers (e.g., vLLM, TGI, or Triton ) and an understanding of how to optimize models for GPU deployment. Infrastructure: Comfortable working in Linux-based environments and proficient in managing containerized ...

Senior AI Platform Engineer

Hiring Organisation
IQVIA
Location
London, United Kingdom
Salary
£ 80 K
PyTorch.Strong knowledge of distributed training and inference strategies, including tensor, pipeline, data, and expert parallelism approaches.Experience optimising LLM inference workloads using technologies such as vLLM, TensorRT-LLM, NVIDIA NIM, SGLang, or similar high-performance serving frameworks.Expertise in model optimisation techniques including quantisation, mixed precision training and inference (FP8, GPTQ ...

Senior AI Platform Engineer

Hiring Organisation
IQVIA
Location
London, UK
Employment Type
Full-time
knowledge of distributed training and inference strategies, including tensor, pipeline, data, and expert parallelism approaches. Experience optimising LLM inference workloads using technologies such as vLLM, TensorRT-LLM, NVIDIA NIM, SGLang, or similar high-performance serving frameworks. Expertise in model optimisation techniques including quantisation, mixed precision training and inference (FP8, GPTQ ...

Principal Data Engineer

Location
Greater London, England, United Kingdom
Hands‐on experience with cloud‐native ML infrastructure platforms Knowledge of vector databases (Pinecone, Weaviate, Qdrant) and embedding models Experience with model serving frameworks (vLLM, TensorRT, Ray) Familiarity with Anaplan or similar enterprise planning platforms Experience with A/B testing and experimentation frameworks for AI features Experience with model ...

AI Infrastructure Engineer, Serving Platform

Hiring Organisation
Scale AI
Location
London, United Kingdom
Salary
£ 100 K
Terraform).Proven ability to solve complex problems and work independently in fast-moving environments.Nice to haves:Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference.PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows ...

AI Infrastructure Engineer, Serving Platform London, UK Apply →

Location
Greater London, England, United Kingdom
Proven ability to solve complex problems and work independently in fast-moving environments. Nice to haves: Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference. PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This ...