26 to 50 of 58 vLLM Jobs in London

Machine Learning Research Engineer (Foundational Research)

Hiring Organisation
Thomson Reuters
Location
London, UK
Employment Type
Full-time
production-quality code and data pipelines for ML systemsProficiency in modern AI development frameworks including: PyTorch, Jax , HuggingFace Transformers, LLM APIs (litellm etc) and vLLM for building and deploying large-scale AI applicationsUnderstanding of LLM training methodologies including instruction fine-tuning, preference optimization, and reinforcement learning approachesStrong software engineering skills ...

Platform Lead - MLOps

Location
Greater London, England, United Kingdom
Infrastructure as Code (IaC), and deploy Large Language Models (LLMs) into production. By architecting resilient LLMOps pipelines and utilising serving frameworks like Triton, vLLM, or Hugging Face TGI, you will guarantee ultra-low latency, high availability, and proactive model monitoring. Crucially, you will instil financial accountability by establishing robust FinOps ...

AI Infrastructure Engineer, Serving Platform

Hiring Organisation
Scale AI
Location
London, United Kingdom
Salary
£ 100 K
Terraform).Proven ability to solve complex problems and work independently in fast-moving environments.Nice to haves:Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference.PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows ...

AI Infrastructure Engineer, Serving Platform London, UK Apply →

Location
Greater London, England, United Kingdom
Proven ability to solve complex problems and work independently in fast-moving environments. Nice to haves: Experience with modern LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM, or text-generation-inference. PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This ...

Forward Deployed Engineer, EMEA

Location
Greater London, England, United Kingdom
eligible for sponsorship Bonus Points AI Voice Assistants, conversational AI, STT/TTS, or real‐time AI experience Experience with open‐weight models, vLLM, SGLang, TGI, or Ollama LLM fine‐tuning, RAG, evaluation, or inference optimization Experience sizing and optimizing GPU infrastructure Background in telecom, CPaaS, cloud/AI infrastructure ...

Sr. Software Engineer, Inference

Location
Greater London, England, United Kingdom
CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies). Active open-source or production contributions to modern inference frameworks (e.g., vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe). Experience leading multi-team technical initiatives or partnering directly with enterprise customers on mission-critical platform launches. ...

ML Ops Engineer

Location
Greater London, England, United Kingdom
model training, evaluation, packaging, and production deployment. Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM). Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks. Monitor and optimise cloud spend across high-cost … Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio). Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins. Experience with vLLM, Ray, MLflow, LangChain/LangSmith, DeepSpeed, or Hugging Face TGI. Solid background in AWS/GCP/Azure, Kubecost, and GPU cost optimisation techniques. Strong ...

ML Ops Engineer

Hiring Organisation
Anaplan
Location
London, United Kingdom
Salary
£ 80 K
continuous model training, evaluation, packaging, and production deployment.Deploy Large Language Models (LLMs) and generative AI workloads using advanced inference engines (e.g., Triton Inference Server, vLLM, TensorRT-LLM).Enable automated model validation, monitoring for model drift, data drift, and latency bottlenecks.Monitor and optimise cloud spend across high-cost GPU/… proficiency in Kubernetes (K8s), Docker, Helm, KubeFlow, and service meshes (e.g., Istio).Hands-on experience with Terraform, Ansible, GitHub Actions, ArgoCD, or Jenkins.Experience with vLLM, Ray, MLflow, LangChain/LangSmith, DeepSpeed, or Hugging Face TGI.Solid background in AWS/GCP/Azure, Kubecost, and GPU cost optimisation techniques.Strong skills ...

Platform Lead, MLOps & LLM Infra

Location
Greater London, England, United Kingdom
define the AI/ML infrastructure roadmap, automate multi-cloud provisioning via IaC, and deploy LLMs into production with resilient LLMOps pipelines using Triton, vLLM, or Hugging Face TGI for ultra-low latency. You will establish robust FinOps frameworks to manage expensive GPU/CPU cloud budgets, enable auto-scaling ...

LLM & Generative AI Engineer

Location
Greater London, England, United Kingdom
Processing and Generative AI Hands-on experience with PyTorch, Hugging Face Transformers, and parameter-efficient fine-tuning (PEFT/LoRA) Experience deploying models with vLLM, Ollama, or Triton Inference Server Strong background in software engineering best practices and clean code #J-18808-Ljbffr ...

Senior ML Systems Engineer, Frameworks & Tooling

Location
Greater London, England, United Kingdom
training LLMs or other large transformer architectures. Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.). Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches). Experience with data pipeline optimization, sharded datasets, or caching strategies. Background in performance engineering, profiling, or low-level systems. ...

Senior ML Systems Engineer, Frameworks & Tooling

Hiring Organisation
Cohere
Location
London, United Kingdom
Salary
£ 80 K
Experience with training LLMs or other large transformer architectures.Contributions to ML frameworks (PyTorch, JAX, DeepSpeed, Megatron, xFormers, etc.).Familiarity with evaluation and serving frameworks (vLLM, TensorRT-LLM, custom KV caches).Experience with data pipeline optimization, sharded datasets, or caching strategies.Background in performance engineering, profiling, or low-level systems.Bonus: paper ...

Senior Research Scientist | Model Steering

Location
Greater London, England, United Kingdom
without degrading their reasoning capabilities. Experience with machine translation, multilingual NLP, or language quality estimation. Familiarity with inference and serving at scale (e.g. via vLLM, SGLang, TensorRT‐LLM, etc) and long‐context modelling. Publications at top‐tier venues. What we offer Diverse and internationally distributed team : joining our team means ...

LLM & Generative AI Engineer: RAG & Fine-Tuning

Location
Greater London, England, United Kingdom
semantic search solutions. The role requires hands-on experience with PyTorch, Hugging Face Transformers, and parameter-efficient fine-tuning, plus deploying models with vLLM, Ollama, or Triton. Strong software engineering practices are essential. #J-18808-Ljbffr ...

AI Research Scientist

Location
Greater London, England, United Kingdom
engineering & synthetic data pipelines Agentic frameworks, reasoning, tool use, memory systems Evaluation & benchmarking (LLM-as-judge, safety/alignment metrics) Inference optimisation (quantization, distillation, vLLM/TensorRT-LLM) Requirements: - PhD or Master's in CS/AI/ML/Physics/Maths or equivalent research experience - Strong hands ...

AI Research Scientist

Hiring Organisation
Intellectual Capital Resources
Location
London, United Kingdom
Salary
£ 60 K
engineering & synthetic data pipelines Agentic frameworks, reasoning, tool use, memory systems Evaluation & benchmarking (LLM-as-judge, safety/alignment metrics) Inference optimisation (quantization, distillation, vLLM/TensorRT-LLM) Requirements: PhD or Master's in CS/AI/ML/Physics/Maths or equivalent research experience Strong hands ...

Natural Language Processing Researcher

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 70 K
ability to work both independently and as part of a team.Strong programming skills in Python and experience with machine learning libraries such as PyTorch, vLLM or similar are a prerequisiteYou will have, or be working towards gaining, a Masters or PhD degree in NLP or a related quantitative subject, such ...

Staff Software Engineer, Inference

Location
Greater London, England, United Kingdom
SLOs, capacity planning, autoscaling strategies, and mentoring senior and mid‐level engineers. Preferred Direct open‐source or production contributions to modern inference frameworks (e.g., vLLM, Triton, TensorRT‐LLM, Ray Serve, or TorchServe). Deep experience with GPU systems engineering and hardware performance optimisation (e.g., CUDA, NCCL, RDMA, NUMA ...

Artificial Intelligence Engineer

Location
Greater London, England, United Kingdom
testing Optimise reliability, scalability, latency, and user experience Work closely with Product and Engineering teams to solve complex customer problems Python PyTorch JAX vLLM Vector Databases What We're Looking For Strong software engineering fundamentals Experience building AI agents, copilots, RAG systems, or workflow automation platforms Excellent Python and/ ...

Performance Engineer, Containers/Serverless

Location
Greater London, England, United Kingdom
compatible object storage - including performance-killing cases (small-object overhead, range-request patterns, eventual consistency, multipart tuning). Comfort with model-serving runtimes (vLLM, SGLang etc) and the formats they consume (safetensors, GGUF, sharded checkpoints). An end-to-end view: comfortable reasoning about NIC, switch, filesystem, cache, container runtime ...

Research Associate in Adaptive and Efficient LLM Architectures

Hiring Organisation
Imperial College London
Location
London, United Kingdom
Salary
£ 55 K
training, evaluation, RLVR, PEFT, quantisation, tensor/data parallelism.Ideally, the candidate should have familiarity with CUDA kernels and/or Triton, and inference engines (vLLM, SGLang, et cetera).Experience coding with deep learning libraries such as Pytorch/JAX is essential.Fluent written and spoken English skills as well as contributions ...

Backend Engineer - API

Hiring Organisation
X
Location
London, United Kingdom
Salary
> £ 150 K
operating commonly used databases such as PostgreSQL, Clickhouse, and MongoDBPREFERRED SKILLS AND EXPERIENCE:Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM)Experience designing or building with agent SDKs and agent orchestration frameworksExperience with Docker, Kubernetes, and containerized applicationsExpert knowledge of gRPC (unary, response streaming, bi-directional ...

NLP Performance Engineer

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 80 K
understanding of transformer inference, including prefill versus decode, KV-cache behaviour, attention variants and performance bottlenecksHands-on experience with LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM or TGI, and the PyTorch ecosystemExperience with inference optimisation techniques, including quantisation, speculative decoding and model parallelism across modern GPU architecturesStrong software ...

NLP Performance Engineer

Location
Greater London, England, United Kingdom
transformer inference, including prefill versus decode, KV‐cache behaviour, attention variants and performance bottlenecks Hands‐on experience with LLM serving frameworks such as vLLM, SGLang, TensorRT‐LLM or TGI, and the PyTorch ecosystem Experience with inference optimisation techniques, including quantisation, speculative decoding and model parallelism across modern GPU architectures Strong ...

Developer Experience Engineer New London

Location
Greater London, England, United Kingdom
user and a working results. Bonus Points These aren't requirements, but they'd make you stand out: Experience with inference serving stacks (vLLM, SGLang, TensorRT-LLM, Triton). A track record of contributing to or maintaining open source developer tools and ML ecosystem projects Experience writing technical documentation ...