26 to 50 of 69 vLLM Jobs in England

AI Engineer (Fluent Portuguese & English)

Hiring Organisation
Chubb
Location
London, United Kingdom
Salary
£ 80 K
Hands-on experience with LLM training cycles, parameter-efficient fine-tuning (PEFT), and sophisticated prompt engineering.Inference Stack: Experience with high-performance inference servers (e.g., vLLM, TGI, or Triton) and an understanding of how to optimize models for GPU deployment.Infrastructure: Comfortable working in Linux-based environments and proficient in managing containerized ...

AI Engineer (Fluent in Mandarin & English)

Location
Greater London, England, United Kingdom
experience with LLM training cycles, parameter-efficient fine-tuning (PEFT), and sophisticated prompt engineering. Inference Stack: Experience with high-performance inference servers (e.g., vLLM, TGI, or Triton ) and an understanding of how to optimize models for GPU deployment. Infrastructure: Comfortable working in Linux-based environments and proficient in managing containerized ...

Machine Learning Research Engineer (Foundational Research)

Hiring Organisation
Thomson Reuters
Location
London, UK
Employment Type
Full-time
production-quality code and data pipelines for ML systemsProficiency in modern AI development frameworks including: PyTorch, Jax , HuggingFace Transformers, LLM APIs (litellm etc) and vLLM for building and deploying large-scale AI applicationsUnderstanding of LLM training methodologies including instruction fine-tuning, preference optimization, and reinforcement learning approachesStrong software engineering skills ...

Platform Lead - MLOps

Location
Greater London, England, United Kingdom
Infrastructure as Code (IaC), and deploy Large Language Models (LLMs) into production. By architecting resilient LLMOps pipelines and utilising serving frameworks like Triton, vLLM, or Hugging Face TGI, you will guarantee ultra-low latency, high availability, and proactive model monitoring. Crucially, you will instil financial accountability by establishing robust FinOps ...

AI Infrastructure Engineer, Serving Platform

Hiring Organisation
Scale AI
Location
London, United Kingdom
Salary
£ 100 K
Terraform).Proven ability to solve complex problems and work independently in fast-moving environments.Nice to haves:Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference.PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows ...

AI Infrastructure Engineer, Serving Platform London, UK Apply →

Location
Greater London, England, United Kingdom
Proven ability to solve complex problems and work independently in fast-moving environments. Nice to haves: Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference. PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This ...

Forward Deployed Engineer, EMEA

Location
Greater London, England, United Kingdom
eligible for sponsorship Bonus Points AI Voice Assistants, conversational AI, STT/TTS, or real‐time AI experience Experience with open‐weight models, vLLM, SGLang, TGI, or Ollama LLM fine‐tuning, RAG, evaluation, or inference optimization Experience sizing and optimizing GPU infrastructure Background in telecom, CPaaS, cloud/AI infrastructure ...

Sr. Software Engineer, Inference

Location
Greater London, England, United Kingdom
CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies). Active open-source or production contributions to modern inference frameworks (e.g., vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe). Experience leading multi-team technical initiatives or partnering directly with enterprise customers on mission-critical platform launches. ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
model training, evaluation, packaging, and production deployment. Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM). Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks. Monitor and optimise cloud spend across high-cost … Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio). Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins. Experience with vLLM, Ray, MLflow, LangChain/LangSmith, DeepSpeed, or Hugging Face TGI. Solid background in AWS/GCP/Azure, Kubecost, and GPU cost optimisation techniques. Strong ...

ML Ops Engineer

Hiring Organisation
Anaplan
Location
London, United Kingdom
Salary
£ 80 K
continuous model training, evaluation, packaging, and production deployment.Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM).Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks.Monitor and optimise cloud spend across high-cost GPU/… proficiency in Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio).Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.Experience with vLLM, Ray, MLflow, LangChain/LangSmith, DeepSpeed, or Hugging Face TGI.Solid background in AWS/GCP/Azure, Kubecost, and GPU cost optimisation techniques.Strong skills ...

Platform Lead, MLOps & LLM Infra

Location
Greater London, England, United Kingdom
define the AI/ML infrastructure roadmap, automate multi-cloud provisioning via IaC, and deploy LLMs into production with resilient LLMOps pipelines using Triton, vLLM, or Hugging Face TGI for ultra-low latency. You will establish robust FinOps frameworks to manage expensive GPU/CPU cloud budgets, enable auto-scaling ...

LLM & Generative AI Engineer

Location
Greater London, England, United Kingdom
Processing and Generative AI Hands-on experience with PyTorch, Hugging Face Transformers, and parameter-efficient fine-tuning (PEFT/LoRA) Experience deploying models with vLLM, Ollama, or Triton Inference Server Strong background in software engineering best practices and clean code #J-18808-Ljbffr ...

Senior ML Systems Engineer, Frameworks & Tooling

Location
Greater London, England, United Kingdom
training LLMs or other large transformer architectures. Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.). Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches). Experience with data pipeline optimization, sharded datasets, or caching strategies. Background in performance engineering, profiling, or low-level systems. ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, United Kingdom
Salary
£ 80 K
Experience with training LLMs or other large transformer architectures.Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.).Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches).Experience with data pipeline optimization, sharded datasets, or caching strategies.Background in performance engineering, profiling, or low-level systems.Bonus: paper ...

Senior Research Scientist | Model Steering

Location
Greater London, England, United Kingdom
without degrading their reasoning capabilities. Experience with machine translation, multilingual NLP, or language quality estimation. Familiarity with inference and serving at scale (e.g. via vLLM, SGLang, TensorRT‐LLM, etc) and long‐context modelling. Publications at top‐tier venues. What we offer Diverse and internationally distributed team : joining our team means ...

LLM & Generative AI Engineer: RAG & Fine-Tuning

Location
Greater London, England, United Kingdom
semantic search solutions. The role requires hands-on experience with PyTorch, Hugging Face Transformers, and parameter-efficient fine-tuning, plus deploying models with vLLM, Ollama, or Triton. Strong software engineering practices are essential. #J-18808-Ljbffr ...

AI Research Scientist

Location
Greater London, England, United Kingdom
engineering & synthetic data pipelines Agentic frameworks, reasoning, tool use, memory systems Evaluation & benchmarking (LLM-as-judge, safety/alignment metrics) Inference optimisation (quantization, distillation, vLLM/TensorRT-LLM) Requirements: - PhD or Master's in CS/AI/ML/Physics/Maths or equivalent research experience - Strong hands ...

AI Research Scientist

Hiring Organisation
Intellectual Capital Resources
Location
London, United Kingdom
Salary
£ 60 K
engineering & synthetic data pipelines Agentic frameworks, reasoning, tool use, memory systems Evaluation & benchmarking (LLM-as-judge, safety/alignment metrics) Inference optimisation (quantization, distillation, vLLM/TensorRT-LLM) Requirements: PhD or Master's in CS/AI/ML/Physics/Maths or equivalent research experience Strong hands ...

Senior Performance Engineer

Location
Cambridge, England, United Kingdom
also highly valued: Post-graduate degrees and research experience in relevant fields (please list your publications). Deep understanding of inference serving frameworks (e.g. vLLM) Background in statistical analysis Contributions to open source and/or research projects Benefits A collaborative and supportive work environment The opportunity to have ...

Natural Language Processing Researcher

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 70 K
ability to work both independently and as part of a team.Strong programming skills in Python and experience with machine learning libraries such as PyTorch, vLLM or similar are a prerequisiteYou will have, or be working towards gaining, a Masters or PhD degree in NLP or a related quantitative subject, such ...

Staff Software Engineer, Inference

Location
Greater London, England, United Kingdom
SLOs, capacity planning, autoscaling strategies, and mentoring senior and mid‐level engineers. Preferred Direct open‐source or production contributions to modern inference frameworks (e.g., vLLM, Triton, TensorRT‐LLM, Ray Serve, or TorchServe). Deep experience with GPU systems engineering and hardware performance optimisation (e.g., CUDA, NCCL, RDMA, NUMA ...

Principal Product Manager - AI Tooling

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 70 K
product management in development tools and software including CLIs, SDKs and APIs.Knowledge of machine learning frameworks, runtimes, infrastructure such as PyTorch, ONNX, ExecuTorch, Llama.cpp, vLLM and LiteRT.Understanding of model deployment to edge, embedded or heterogeneous computing, with an understanding of trade-offs between accuracy, performance and power consumption.You should also ...

Software Inference Deployment Engineer

Location
Oxford, England, United Kingdom
PyTorch in particular) Practical experience with model deployment workflows - loading, format conversion, quantisation, or framework integration Comfortable working with inference serving stacks (for example vLLM, TensorRT‐LLM, or similar) Familiarity with Linux, containerisation (Docker), and cluster environments Comfortable in a customer‐facing role, able to communicate clearly with ...

Principal Product Manager - AI Tooling

Location
Cambridge, England, United Kingdom
management in development tools and software including CLIs, SDKs and APIs. Knowledge of machine learning frameworks, runtimes, infrastructure such as PyTorch, ONNX, ExecuTorch, Llama.cpp, vLLM and LiteRT. Understanding of model deployment to edge, embedded or heterogeneous computing, with an understanding of trade-offs between accuracy, performance and power consumption. ...

Artificial Intelligence Engineer

Location
Greater London, England, United Kingdom
testing Optimise reliability, scalability, latency, and user experience Work closely with Product and Engineering teams to solve complex customer problems Python PyTorch JAX vLLM Vector Databases What We're Looking For Strong software engineering fundamentals Experience building AI agents, copilots, RAG systems, or workflow automation platforms Excellent Python and/ ...