25 of 25 vLLM Jobs in the UK excluding London

Sr. AI Architect Engineer

Location
City of Edinburgh, Scotland, United Kingdom
Google Professional Cloud Architect, TOGAF, or equivalent). Experience with large language model training and inference architecture at scale (e.g., PyTorch, DeepSpeed, Megatron-LM, vLLM). Experience with edge and on-device AI inference and the device-to-cloud continuum. Advanced Kubernetes experience: GPU scheduling plugins, multi-cluster and fleet ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Holywood, Down, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Edinburgh, Midlothian, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Senior Site Reliability Engineer

Location
Reading, England, United Kingdom
strongly preferred) Running ML workloads on Kubernetes - GPU scheduling, capacity, and cost management Model serving and inference at production scale (eg KServe, RayServe, Triton, vLLM, or similar) with real latency and cost constraints(preferred RayServe) MLOps pipeline tooling - training pipelines, model registries, feature stores, and lineage (Kubeflow, MLflow, Feast, Weights ...

Lead Software Engineer - Python / Go & AI/ML

Location
Auchentibber, Scotland, United Kingdom
skills Formal training or certification on software engineering concepts and advanced applied experience – preferably Go/Python Hands-on experience with LLM inference systems — vLLM, TensorRT-LLM, SGLang, LLM-D, or equivalent production serving engines Strong understanding of GPU memory architecture, including KV cache sizing and dynamics, memory-bandwidth versus ...

Remote Senior Machine Learning Engineer

Location
Stirling, Scotland, United Kingdom
data, including classification, extraction, embeddings, re-rankers, clustering, and search. Experience in instruction fine-tuning and serving language models, familiarity with frameworks such as vLLM, DeepSpeed, or similar tools A solid grounding in classical ML and statistics, and the judgement to choose simpler methods when they’re the right solution. ...

Senior Performance Engineer

Location
Cambridge, England, United Kingdom
also highly valued: Post-graduate degrees and research experience in relevant fields (please list your publications). Deep understanding of inference serving frameworks (e.g. vLLM) Background in statistical analysis Contributions to open source and/or research projects Benefits A collaborative and supportive work environment The opportunity to have ...

Principal Product Manager - AI Tooling

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 70 K
product management in development tools and software including CLIs, SDKs and APIs.Knowledge of machine learning frameworks, runtimes, infrastructure such as PyTorch, ONNX, ExecuTorch, Llama.cpp, vLLM and LiteRT.Understanding of model deployment to edge, embedded or heterogeneous computing, with an understanding of trade-offs between accuracy, performance and power consumption.You should also ...

Software Inference Deployment Engineer

Location
Oxford, England, United Kingdom
PyTorch in particular) Practical experience with model deployment workflows - loading, format conversion, quantisation, or framework integration Comfortable working with inference serving stacks (for example vLLM, TensorRT‐LLM, or similar) Familiarity with Linux, containerisation (Docker), and cluster environments Comfortable in a customer‐facing role, able to communicate clearly with ...

Principal Product Manager - AI Tooling

Location
Cambridge, England, United Kingdom
management in development tools and software including CLIs, SDKs and APIs. Knowledge of machine learning frameworks, runtimes, infrastructure such as PyTorch, ONNX, ExecuTorch, Llama.cpp, vLLM and LiteRT. Understanding of model deployment to edge, embedded or heterogeneous computing, with an understanding of trade-offs between accuracy, performance and power consumption. ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
reliability, performance, and cost-efficiency of the LLM inference platform end to end. You will operate large language model serving stacks (such as vLLM and llm-d) in production at scale, with deep instrumentation and strong operational rigor. You will partner across engineering to deliver secure software, improve stability … infrastructure Build backend services and APIs that enable reliable operation of AI infrastructure in production Operate and scale LLM serving infrastructure (such as vLLM and llm-d), including model hosting, request routing, continuous batching, and KV-cache optimization Deploy, host, and lifecycle-manage open-source and proprietary LLMs on Amazon ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Paisley, Scotland, United Kingdom
reliability, performance, and cost-efficiency of the LLM inference platform end to end. You will operate large language model serving stacks (such as vLLM and llm-d) in production at scale, with deep instrumentation and strong operational rigor. You will partner across engineering to deliver secure software, improve stability … infrastructure Build backend services and APIs that enable reliable operation of AI infrastructure in production Operate and scale LLM serving infrastructure (such as vLLM and llm-d), including model hosting, request routing, continuous batching, and KV-cache optimization Deploy, host, and lifecycle-manage open-source and proprietary LLMs on Amazon ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Milton Keynes, England, United Kingdom
reliability, performance, and cost-efficiency of the LLM inference platform end to end. You will operate large language model serving stacks (such as vLLM and llm-d) in production at scale, with deep instrumentation and strong operational rigor. You will partner across engineering to deliver secure software, improve stability … infrastructure Build backend services and APIs that enable reliable operation of AI infrastructure in production Operate and scale LLM serving infrastructure (such as vLLM and llm-d), including model hosting, request routing, continuous batching, and KV-cache optimization Deploy, host, and lifecycle-manage open-source and proprietary LLMs on Amazon ...

AI System Researcher

Hiring Organisation
Microtech Global Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent
Strong knowledge of distributed systems, operating systems, machine learning systems architecture, Inference serving, and AI Infrastructure. Hands-on experience with LLM serving frameworks (e.g., vLLM, Ray Serve, TensorRT-LLM, TGI) and distributed KV cache optimization. Proficiency in C/C++, with additional experience in Python for research prototyping. Solid grounding ...

Project Technical Lead - AI Systems Simulation

Location
Cambridge, England, United Kingdom
infrastructure, ML systems, or computer architecture. Familiarity with Agile or other modern technical project management frameworks. Knowledge of modern inference‐serving frameworks (e.g., vLLM). Background in statistics, operations research, or large‐scale datacenter infrastructure. Contributions to open‐source AI or systems projects. Benefits High‐impact role in a rapidly ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Location
Auchentibber, Scotland, United Kingdom
guide peers on safe and effective usage within team practices Preferred qualifications, capabilities, and skills Experience operating large language model inference servers such as vLLM and llm-d or directly equivalent model serving stacks in production Experience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patterns … monitoring, including hallucination detection, toxicity filtering, and drift detection using open telemetry conventions Contributions to open-source large language model serving or inference projects, (vLLM, llm-d, Ray, KServe, Triton) #J-18808-Ljbffr ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Paisley, Renfrewshire, UK
Employment Type
Full-time
guide peers on safe and effective usage within team practices Preferred qualifications, capabilities, and skills Experience operating large language model inference servers such as vLLM and llm-d (or directly equivalent model serving stacks) in production Experience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patterns … monitoring, including hallucination detection, toxicity filtering, and drift detection using open telemetry conventions Contributions to open-source large language model serving or inference projects, (vLLM, llm-d, Ray, KServe, Triton) ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
guide peers on safe and effective usage within team practices Preferred qualifications, capabilities, and skills Experience operating large language model inference servers such as vLLM and llm-d (or directly equivalent model serving stacks) in production Experience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patterns … monitoring, including hallucination detection, toxicity filtering, and drift detection using open telemetry conventions Contributions to open-source large language model serving or inference projects, (vLLM, llm-d, Ray, KServe, Triton) ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
JP Morgan Chase
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 80 K
ability to guide peers on safe and effective usage within team practicesPreferred qualifications, capabilities, and skillsExperience operating large language model inference servers such as vLLM and llm-d (or directly equivalent model serving stacks) in productionExperience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patternsExperience building … quality monitoring, including hallucination detection, toxicity filtering, and drift detection using open telemetry conventionsContributions to open-source large language model serving or inference projects, (vLLM, llm-d, Ray, KServe, Triton)J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world’s most prominent ...

Senior Software Engineer (vLLM)

Location
Cambridge, England, United Kingdom
Opportunity We are seeking a Senior Software Engineer with a passion for open-source AI infrastructure to work on deploying, extending and optimising vLLM (and potentially other inference serving engines) to support our projects. You will play a crucial, high-impact role across both our key programmes, the Scaling Inference … autonomous agents for software development. What You'll Do Deploy, instrument and monitor open weight models served using vLLM. Implement new features within vLLM to support novel hardware architectures as part of the Scaling Inference Lab. Work with the Panopticon team to identify opportunities to extend vLLM to enhance accuracy ...

Senior Software Engineer - Open-Source AI Inference & vLLM

Location
Cambridge, England, United Kingdom
CommonAI CIC is seeking a Senior Software Engineer to help deploy, extend and optimise vLLM and related inference engines for our AI infrastructure projects. You will work across the Scaling Inference Lab and the High Assurance programme to drive performance, reliability and safety in production‐grade systems. You will collaborate ...

Senior Researcher - AI Computer Architecture

Hiring Organisation
Microsoft
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 70 K
Job ID: 200037336Posted: 2026-09-07Location: United Kingdom, Cambridgeshire, CambridgeEmployment type: Full-TimeWork site: 3 days/week in-officeRole type: Individual ContributorTravel: Less than 25%Profession: Research, Applied, & Data SciencesDiscipline: Research SciencesCompany: MicrosoftOverviewThe ...

Systems Research Engineer

Hiring Organisation
European Tech Recruit
Location
Edinburgh, Scotland, United Kingdom
depth profiling of large-scale inference pipelines, specifically focusing on KV cache management and heterogeneous memory scheduling. AI Serving: Optimising high-throughput frameworks (vLLM, Ray Serve, PyTorch Distributed) to ensure low-latency, multi-tenant performance. Research Leadership: Contributing to top-tier venues (OSDI, NSDI, EuroSys, MLSys) and driving those innovations … Stack: Strong proficiency in C/C++ for systems work, with Python for rapid prototyping. Expertise: Hands-on experience with LLM serving frameworks ( vLLM, Ray Serve, TensorRT-LLM ) and distributed algorithms. Mindset: A solid grounding in systems research methodology and performance profiling tools. The "Value Add" (Desired): A PhD focused ...