10 of 10 vLLM Jobs in the UK excluding London

Senior Platform Engineer

Hiring Organisation
Lorien
Location
London, South East, England, United Kingdom
Employment Type
Contractor
Contract Rate
Salary negotiable
behave in production. Experience or exposure to areas such as: MLOps platforms (e.g. Kubeflow or similar frameworks) Model serving and inference platforms (e.g. KServe, vLLM , or equivalent) Supporting LLM-based workloads , including performance and scaling considerations Notebook environments such as JupyterHub Awareness of emerging tooling around Responsible/Trustworthy ...

Lead Platform Engineer

Hiring Organisation
Lorien
Location
London, South East, England, United Kingdom
Employment Type
Contractor
Contract Rate
Salary negotiable
areas such as: Building or operating MLOps platforms using tools like Kubeflow or similar frameworks Running model serving and inference platforms (e.g. KServe, vLLM, or equivalent) Supporting LLM-based workloads , including optimisation and serving considerations Providing notebook-based environments such as JupyterHub in secure platforms Exposure to emerging tooling such ...

Senior Software Engineer

Hiring Organisation
Jobleads-UK
Location
Cambridge, England, United Kingdom
related field. Desirable: Exposure to machine learning frameworks such as PyTorch, JAX, Triton, TensorFlow Experience with distributed workload management systems such as Kubernetes, VLLM, Keras or MLOps pipelines Experience working with hardware simulators or emulators (e.g. QEMU). Experience developing for or working with FPGA-based systems. Experience with people ...

Remote Senior Machine Learning Engineer

Hiring Organisation
Jobleads-UK
Location
Stirling, Scotland, United Kingdom
data, including classification, extraction, embeddings, re-rankers, clustering, and search. Experience in instruction fine-tuning and serving language models, familiarity with frameworks such as vLLM, DeepSpeed, or similar tools A solid grounding in classical ML and statistics, and the judgement to choose simpler methods when they’re the right solution. ...

Software Inference Deployment Engineer

Hiring Organisation
Jobleads-UK
Location
Oxford, England, United Kingdom
PyTorch in particular) Practical experience with model deployment workflows - loading, format conversion, quantisation, or framework integration Comfortable working with inference serving stacks (for example vLLM, TensorRT‐LLM, or similar) Familiarity with Linux, containerisation (Docker), and cluster environments Comfortable in a customer‐facing role, able to communicate clearly with ...

AI System Researcher

Hiring Organisation
Microtech Global Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent
Strong knowledge of distributed systems, operating systems, machine learning systems architecture, Inference serving, and AI Infrastructure. Hands-on experience with LLM serving frameworks (e.g., vLLM, Ray Serve, TensorRT-LLM, TGI) and distributed KV cache optimization. Proficiency in C/C++, with additional experience in Python for research prototyping. Solid grounding ...

Lead Site Reliability Engineer – Operations Excellence

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
reliability, performance, and cost‐efficiency of the LLM inference platform end to end. You will operate large language model serving stacks (such as vLLM and llm‐d) in production at scale, with deep instrumentation and strong operational rigor. You will partner across engineering to deliver secure software, improve stability … infrastructure Build backend services and APIs that enable reliable operation of AI infrastructure in production Operate and scale LLM serving infrastructure (such as vLLM and llm‐d), including model hosting, request routing, continuous batching, and KV‐cache optimization Deploy, host, and lifecycle‐manage open‐source and proprietary LLMs on Amazon ...

Project Technical Lead - AI Systems Simulation

Hiring Organisation
Jobleads-UK
Location
Cambridge, England, United Kingdom
infrastructure, ML systems, or computer architecture. Familiarity with Agile or other modern technical project management frameworks. Knowledge of modern inference‐serving frameworks (e.g., vLLM). Background in statistics, operations research, or large‐scale datacenter infrastructure. Contributions to open‐source AI or systems projects. Benefits High‐impact role in a rapidly ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Jobleads-UK
Location
Glasgow, Scotland, United Kingdom
guide peers on safe and effective usage within team practices Preferred qualifications, capabilities, and skills Experience operating large language model inference servers such as vLLM and llm-d (or directly equivalent model serving stacks) in production Experience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patterns … monitoring, including hallucination detection, toxicity filtering, and drift detection using open telemetry conventions Contributions to open-source large language model serving or inference projects, (vLLM, llm-d, Ray, KServe, Triton) #J-18808-Ljbffr ...

Systems Research Engineer - Distributed Systems / C++

Hiring Organisation
European Tech Recruit
Location
Edinburgh, Scotland, United Kingdom
Conduct in-depth profiling and performance tuning of inference pipelines, focusing on KV cache management. Develop low-latency, fault-tolerant AI serving frameworks using vLLM, Ray Serve, and PyTorch Distributed. Research and prototype novel techniques for cache sharing, data locality, and resource orchestration. Translate innovative designs into publishable contributions … distributed systems, or related field. Strong knowledge of Distributed Systems, OS internals, and Machine Learning systems architecture. Hands-on experience with LLM serving frameworks (vLLM, Ray Serve, TensorRT-LLM, or TGI). Proficiency in C/C++ for systems development and Python for research prototyping. Solid grounding in distributed algorithms ...