51 to 75 of 85 vLLM Jobs in the UK

AI Research Scientist

Hiring Organisation
Intellectual Capital Resources
Location
London, United Kingdom
Salary
£ 60 K
engineering & synthetic data pipelines Agentic frameworks, reasoning, tool use, memory systems Evaluation & benchmarking (LLM-as-judge, safety/alignment metrics) Inference optimisation (quantization, distillation, vLLM/TensorRT-LLM) Requirements: PhD or Master's in CS/AI/ML/Physics/Maths or equivalent research experience Strong hands ...

Senior Performance Engineer

Location
Cambridge, England, United Kingdom
also highly valued: Post-graduate degrees and research experience in relevant fields (please list your publications). Deep understanding of inference serving frameworks (e.g. vLLM) Background in statistical analysis Contributions to open source and/or research projects Benefits A collaborative and supportive work environment The opportunity to have ...

Natural Language Processing Researcher

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 70 K
ability to work both independently and as part of a team.Strong programming skills in Python and experience with machine learning libraries such as PyTorch, vLLM or similar are a prerequisiteYou will have, or be working towards gaining, a Masters or PhD degree in NLP or a related quantitative subject, such ...

Staff Software Engineer, Inference

Location
Greater London, England, United Kingdom
SLOs, capacity planning, autoscaling strategies, and mentoring senior and mid‐level engineers. Preferred Direct open‐source or production contributions to modern inference frameworks (e.g., vLLM, Triton, TensorRT‐LLM, Ray Serve, or TorchServe). Deep experience with GPU systems engineering and hardware performance optimisation (e.g., CUDA, NCCL, RDMA, NUMA ...

Principal Product Manager - AI Tooling

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 70 K
product management in development tools and software including CLIs, SDKs and APIs.Knowledge of machine learning frameworks, runtimes, infrastructure such as PyTorch, ONNX, ExecuTorch, Llama.cpp, vLLM and LiteRT.Understanding of model deployment to edge, embedded or heterogeneous computing, with an understanding of trade-offs between accuracy, performance and power consumption.You should also ...

Software Inference Deployment Engineer

Location
Oxford, England, United Kingdom
PyTorch in particular) Practical experience with model deployment workflows - loading, format conversion, quantisation, or framework integration Comfortable working with inference serving stacks (for example vLLM, TensorRT‐LLM, or similar) Familiarity with Linux, containerisation (Docker), and cluster environments Comfortable in a customer‐facing role, able to communicate clearly with ...

Principal Product Manager - AI Tooling

Location
Cambridge, England, United Kingdom
management in development tools and software including CLIs, SDKs and APIs. Knowledge of machine learning frameworks, runtimes, infrastructure such as PyTorch, ONNX, ExecuTorch, Llama.cpp, vLLM and LiteRT. Understanding of model deployment to edge, embedded or heterogeneous computing, with an understanding of trade-offs between accuracy, performance and power consumption. ...

Artificial Intelligence Engineer

Location
Greater London, England, United Kingdom
testing Optimise reliability, scalability, latency, and user experience Work closely with Product and Engineering teams to solve complex customer problems Python PyTorch JAX vLLM Vector Databases What We're Looking For Strong software engineering fundamentals Experience building AI agents, copilots, RAG systems, or workflow automation platforms Excellent Python and/ ...

Performance Engineer, Containers/Serverless

Location
Greater London, England, United Kingdom
compatible object storage - including performance-killing cases (small-object overhead, range-request patterns, eventual consistency, multipart tuning). Comfort with model-serving runtimes (vLLM, SGLang etc) and the formats they consume (safetensors, GGUF, sharded checkpoints). An end-to-end view: comfortable reasoning about NIC, switch, filesystem, cache, container runtime ...

Research Associate in Adaptive and Efficient LLM Architectures

Hiring Organisation
Imperial College London
Location
London, United Kingdom
Salary
£ 55 K
training, evaluation, RLVR, PEFT, quantisation, tensor/data parallelism.Ideally, the candidate should have familiarity with CUDA kernels and/or Triton, and inference engines (vLLM, SGLang, et cetera).Experience coding with deep learning libraries such as Pytorch/JAX is essential.Fluent written and spoken English skills as well as contributions ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
reliability, performance, and cost-efficiency of the LLM inference platform end to end. You will operate large language model serving stacks (such as vLLM and llm-d) in production at scale, with deep instrumentation and strong operational rigor. You will partner across engineering to deliver secure software, improve stability … infrastructure Build backend services and APIs that enable reliable operation of AI infrastructure in production Operate and scale LLM serving infrastructure (such as vLLM and llm-d), including model hosting, request routing, continuous batching, and KV-cache optimization Deploy, host, and lifecycle-manage open-source and proprietary LLMs on Amazon ...

AI System Researcher

Hiring Organisation
Microtech Global Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent
Strong knowledge of distributed systems, operating systems, machine learning systems architecture, Inference serving, and AI Infrastructure. Hands-on experience with LLM serving frameworks (e.g., vLLM, Ray Serve, TensorRT-LLM, TGI) and distributed KV cache optimization. Proficiency in C/C++, with additional experience in Python for research prototyping. Solid grounding ...

Backend Engineer - API

Hiring Organisation
X
Location
London, United Kingdom
Salary
> £ 150 K
operating commonly used databases such as PostgreSQL, Clickhouse, and MongoDBPREFERRED SKILLS AND EXPERIENCE:Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM)Experience designing or building with agent SDKs and agent orchestration frameworksExperience with Docker, Kubernetes, and containerized applicationsExpert knowledge of gRPC (unary, response streaming, bi-directional ...

Project Technical Lead - AI Systems Simulation

Location
Cambridge, England, United Kingdom
infrastructure, ML systems, or computer architecture. Familiarity with Agile or other modern technical project management frameworks. Knowledge of modern inference‐serving frameworks (e.g., vLLM). Background in statistics, operations research, or large‐scale datacenter infrastructure. Contributions to open‐source AI or systems projects. Benefits High‐impact role in a rapidly ...

NLP Performance Engineer

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 80 K
understanding of transformer inference, including prefill versus decode, KV-cache behaviour, attention variants and performance bottlenecksHands-on experience with LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM or TGI, and the PyTorch ecosystemExperience with inference optimisation techniques, including quantisation, speculative decoding and model parallelism across modern GPU architecturesStrong software ...

NLP Performance Engineer

Location
Greater London, England, United Kingdom
transformer inference, including prefill versus decode, KV‐cache behaviour, attention variants and performance bottlenecks Hands‐on experience with LLM serving frameworks such as vLLM, SGLang, TensorRT‐LLM or TGI, and the PyTorch ecosystem Experience with inference optimisation techniques, including quantisation, speculative decoding and model parallelism across modern GPU architectures Strong ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Location
Auchentibber, Scotland, United Kingdom
guide peers on safe and effective usage within team practices Preferred qualifications, capabilities, and skills Experience operating large language model inference servers such as vLLM and llm-d or directly equivalent model serving stacks in production Experience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patterns … monitoring, including hallucination detection, toxicity filtering, and drift detection using open telemetry conventions Contributions to open-source large language model serving or inference projects, (vLLM, llm-d, Ray, KServe, Triton) #J-18808-Ljbffr ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Paisley, Renfrewshire, UK
Employment Type
Full-time
guide peers on safe and effective usage within team practices Preferred qualifications, capabilities, and skills Experience operating large language model inference servers such as vLLM and llm-d (or directly equivalent model serving stacks) in production Experience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patterns … monitoring, including hallucination detection, toxicity filtering, and drift detection using open telemetry conventions Contributions to open-source large language model serving or inference projects, (vLLM, llm-d, Ray, KServe, Triton) ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
guide peers on safe and effective usage within team practices Preferred qualifications, capabilities, and skills Experience operating large language model inference servers such as vLLM and llm-d (or directly equivalent model serving stacks) in production Experience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patterns … monitoring, including hallucination detection, toxicity filtering, and drift detection using open telemetry conventions Contributions to open-source large language model serving or inference projects, (vLLM, llm-d, Ray, KServe, Triton) ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
JP Morgan Chase
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 80 K
ability to guide peers on safe and effective usage within team practicesPreferred qualifications, capabilities, and skillsExperience operating large language model inference servers such as vLLM and llm-d (or directly equivalent model serving stacks) in productionExperience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patternsExperience building … quality monitoring, including hallucination detection, toxicity filtering, and drift detection using open telemetry conventionsContributions to open-source large language model serving or inference projects, (vLLM, llm-d, Ray, KServe, Triton)J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world’s most prominent ...

Developer Experience Engineer New London

Location
Greater London, England, United Kingdom
user and a working results. Bonus Points These aren't requirements, but they'd make you stand out: Experience with inference serving stacks (vLLM, SGLang, TensorRT-LLM, Triton). A track record of contributing to or maintaining open source developer tools and ML ecosystem projects Experience writing technical documentation ...

AI Inference Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
improve throughput and cost per token, partnering with the CUDA/GPU engineersMake the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)Translate throughput, latency and uptime commitments into concrete technical specifications and serving capacity plansAct as direct technical ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
throughput and cost per token, partnering with the CUDA/GPU engineers. Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents). Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans. ...

Senior Data Center, AI Software Product Strategy & Go-To-Market Lead — Qualcomm - Europe

Location
Greater London, England, United Kingdom
optimized in a Data Center AI software platforms Inference, model deployment, runtimes, SDKs, developer tools. AI orchestration frameworks, e.g., Kubernetes-based systems, Ray, vLLM, TGI. Cloud-native AI software stacks, deployment patterns, and data center operating models. Data center AI infrastructure CPUs, GPUs, NPUs, and heterogeneous accelerator environments. Kubernetes-based ...

Customer Solution Architect — Arango AI Product Suite

Location
United Kingdom
MLOps platforms and evaluation frameworks (MLflow, Weights & Biases, Ragas, promptfoo, DeepEval). Model adaptation and inference optimization awareness (LoRA/PEFT, DPO, distillation, quantization, vLLM/TGI/TensorRT-LLM), enough to advise on tradeoffs rather than to hand-build. Domain experience in finance, healthcare, public sector, manufacturing, or retail. … pgvector, Pinecone, Weaviate; rerankers (ColBERT, cross-encoders) Pipelines & Orchestration:LangChain, LlamaIndex, Ray, Airflow MLOps & Evals:MLflow, Weights & Biases, Ragas, promptfoo, Great Expectations Serving & Infra:vLLM, TGI, FastAPI/gRPC, Docker/K8s, Terraform, GitHub Actions Observability & Guardrails:OpenTelemetry, Prometheus/Grafana, Llama Guard/Content Safety, custom filters Data:Postgres ...