1 to 25 of 88 vLLM Jobs in the UK

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, United Kingdom
Salary
£ 70 K
frameworks: PyTorch and TensorFlow at a working, hands-on level.Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent multi-node training frameworks.Inference & serving: NVIDIA Triton, vLLM, TensorRT-LLM, or equivalent high-throughput serving platforms.MLOps/LLMOps: Kubeflow, MLflow, and at least one hyperscaler ML platform (SageMaker, Azure ML, or Vertex ...

Enterprise Architect - AI

Hiring Organisation
World Wide Technology
Location
London, UK
Employment Type
Full-time
TensorFlow at a working, hands-on level. Distributed training: Horovod, DeepSpeed, Megatron-LM, or equivalent multi-node training frameworks. Inference & serving: NVIDIA Triton, vLLM, TensorRT-LLM, or equivalent high-throughput serving platforms. MLOps/LLMOps: Kubeflow, MLflow, and at least one hyperscaler ML platform (SageMaker, Azure ML, or Vertex ...

Platform Lead - ML Ops

Location
Greater London, England, United Kingdom
Docker, and service meshes. Expert knowledge of Terraform, Ansible, Jenkins, or GitHub Actions. Proficient in Python, Bash, or Go. Familiarity with Triton Inference Server, vLLM, or Hugging Face TGI. Our Commitment to Diversity, Equity, Inclusionand Belonging (DEIB) We believe attracting and retaining the best talent and fostering an inclusive culture ...

ML Ops Lead

Hiring Organisation
Anaplan
Location
London, United Kingdom
Salary
£ 100 K
Kubernetes (K8s), Docker, and service meshes.Expert knowledge of Terraform, Ansible, Jenkins, or GitHub Actions.Proficient in Python, Bash, or Go.Familiarity with Triton Inference Server, vLLM, or Hugging Face TGI.Our Commitment to Diversity, Equity, Inclusion and Belonging (DEIB) We believe attracting and retaining the best talent and fostering an inclusive culture strengthens ...

Mid/Senior Solution Architect - UK

Hiring Organisation
Multiverse Computing
Location
London, United Kingdom
Employment Type
Permanent
Salary
GBP Annual
/ML services (e.g., SageMaker, Vertex AI, AzureML), including sovereign, on-premise, and hybrid deployment models. Understanding of LLM inference stacks (vLLM, llama.cpp, OpenVINO) and model delivery formats (ONNX, safetensors, HuggingFace model hub). Understanding of benchmarking and LLM performance evaluation (accuracy, latency, throughput). Hands-on coding skills ...

Senior Platform Engineer

Hiring Organisation
Lorien Resourcing
Location
London, United Kingdom
Salary
£ 80 K
workloads behave in production.Experience or exposure to areas such as:MLOps platforms (e.g. Kubeflow or similar frameworks)Model serving and inference platforms (e.g. KServe, vLLM, or equivalent)Supporting LLM‐based workloads, including performance and scaling considerationsNotebook environments such as JupyterHubAwareness of emerging tooling around Responsible/Trustworthy AI or comparable ...

Lead Platform Engineer

Hiring Organisation
Lorien Resourcing
Location
London, United Kingdom
Salary
£ 80 K
.Experience in areas such as:Building or operating MLOps platforms using tools like Kubeflow or similar frameworksRunning model serving and inference platforms (e.g. KServe, vLLM, or equivalent)Supporting LLM‐based workloads, including optimisation and serving considerationsProviding notebook‐based environments such as JupyterHub in secure platformsExposure to emerging tooling such ...

Lead Platform Engineer

Hiring Organisation
Lorien
Location
London, United Kingdom
Employment Type
Contract
Contract Rate
GBP Annual
areas such as: Building or operating MLOps platforms using tools like Kubeflow or similar frameworks Running model serving and inference platforms (eg KServe, vLLM, or equivalent) Supporting LLM-based workloads , including optimisation and serving considerations Providing notebook-based environments such as JupyterHub in secure platforms Exposure to emerging tooling such ...

Associate Director Lead AI Architect

Hiring Organisation
Anson Mccade
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Foundry and other AWS, Azure or GCP services . Work with self-hosted inference and model-serving technologies where appropriate, including tools such as vLLM, SGLang and Ollama . Design evaluation, observability and monitoring frameworks to assess model performance, system health, reliability and business impact. Optimise AI solutions for performance ...

Senior Forward Deployed ML Engineer, Agents

Location
Greater London, England, United Kingdom
plus Model fine-tuning practical experience with LoRA/QLoRA, supervised fine-tuning, or RLHF workflows is a plus Inference optimization experience with vLLM, TensorRT-LLM, Triton, or model quantization techniques is desirable Observability tooling practical experience with LLM monitoring, tracing, and evaluation frameworks is a strong plus Familiarity with ...

Sr. AI Architect Engineer

Location
City of Edinburgh, Scotland, United Kingdom
Google Professional Cloud Architect, TOGAF, or equivalent). Experience with large language model training and inference architecture at scale (e.g., PyTorch, DeepSpeed, Megatron-LM, vLLM). Experience with edge and on-device AI inference and the device-to-cloud continuum. Advanced Kubernetes experience: GPU scheduling plugins, multi-cluster and fleet ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
London, United Kingdom
Salary
£ 100 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Holywood, Down, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Edinburgh, Midlothian, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Associate Director

Hiring Organisation
Anson McCade
Location
London, United Kingdom
Salary
£ 100 K
Azure AI Foundry and other AWS, Azure or GCP services.Work with self-hosted inference and model-serving technologies where appropriate, including tools such as vLLM, SGLang and Ollama.Design evaluation, observability and monitoring frameworks to assess model performance, system health, reliability and business impact.Optimise AI solutions for performance, scalability, security ...

Senior Site Reliability Engineer

Location
Reading, England, United Kingdom
strongly preferred) Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management Model serving and inference at production scale (eg KServe, RayServe, Triton, vLLM, or similar) with real latency and cost constraints(preferred RayServe) MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights ...

Principal Engineer - AI

Location
Greater London, England, United Kingdom
Hands‐on experience with cloud‐native ML infrastructure platforms Knowledge of vector databases (Pinecone, Weaviate, Qdrant) and embedding models Experience with model serving frameworks (vLLM, TensorRT, Ray) Experience with A/B testing and experimentation frameworks for AI features Contributions to open‐source ML projects or research publications Experience with ...

Senior Software Engineer, GenAI Platform

Location
Greater London, England, United Kingdom
software development lifecycle, including designing, generating code, testing, monitoring and releasing software. Nice To Haves Experience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in production Experience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation ...

Senior Software Engineer, GenAI Platform

Hiring Organisation
Deliveroo
Location
London, United Kingdom
Salary
£ 80 K
full software development lifecycle, including designing, generating code, testing, monitoring and releasing softwareNice To HavesExperience with LLM inference engines and serving frameworks (e.g., vLLM, SGLang, TensorRT-LLM) in productionExperience with distributed/multi-node fine-tuning and training pipelines (SFT, DPO/RLHF, LoRA), including data preparation and evaluationGPU performance ...

Principal AI Engineer

Location
Greater London, England, United Kingdom
Hands‐on experience with cloud‐native ML infrastructure platforms Knowledge of vector databases (Pinecone, Weaviate, Qdrant) and embedding models Experience with model serving frameworks (vLLM, TensorRT, Ray) Experience with A/B testing and experimentation frameworks for AI features Contributions to open‐source ML projects or research publications Experience with ...

Principal Engineer - AI

Hiring Organisation
Anaplan
Location
London, United Kingdom
Salary
£ 80 K
related quantitative fieldHands-on experience with cloud-native ML infrastructure platformsKnowledge of vector databases (Pinecone, Weaviate, Qdrant) and embedding modelsExperience with model serving frameworks (vLLM, TensorRT, Ray)Experience with A/B testing and experimentation frameworks for AI featuresContributions to open-source ML projects or research publicationsExperience with model observability ...

Manager, Research Engineering (Foundational Research)

Hiring Organisation
Thomson Reuters
Location
London, United Kingdom
Salary
£ 80 K
purely technical preference.Technical Expertise:Deep proficiency in Python and modern software development practices.Hands-on experience with Distributed Training infrastructure (Multi-node GPU training, Kubernetes, vLLM).Familiarity with Deep Learning frameworks (PyTorch).Experience with MLOps tools and experiment tracking (e.g., ClearML, MLFlow, Weights & Biases).Research Fluency: Ability to read technical research ...

Senior AI Platform Engineer

Hiring Organisation
IQVIA
Location
London, United Kingdom
Salary
£ 80 K
PyTorch.Strong knowledge of distributed training and inference strategies, including tensor, pipeline, data, and expert parallelism approaches.Experience optimising LLM inference workloads using technologies such as vLLM, TensorRT-LLM, NVIDIA NIM, SGLang, or similar high-performance serving frameworks.Expertise in model optimisation techniques including quantisation, mixed precision training and inference (FP8, GPTQ ...