11 of 11 vLLM Jobs in the East of England

Senior Performance Engineer

Location
Cambridge, England, United Kingdom
also highly valued: Post-graduate degrees and research experience in relevant fields (please list your publications). Deep understanding of inference serving frameworks (e.g. vLLM) Background in statistical analysis Contributions to open source and/or research projects Benefits A collaborative and supportive work environment The opportunity to have ...

Machine Learning Engineer / AI Research Engineer

Location
Cambridge, England, United Kingdom
with ambiguity and wants ownership. Desirable (or willing to learn) Experience with LLMs: fine-tuning (LoRA/PEFT/full), serving and hosting (e.g. vLLM, TGI, Ollama), and distributed or multi-GPU training. Familiarity with GPUs, cloud, and containers (Docker, etc.). Background in optimisation, Bayesian methods, reinforcement learning ...

Senior Performance Engineer | AI Infrastructure | Cambridge (Hybrid) |

Hiring Organisation
Pure Resourcing Solutions
Location
Cambridge, Cambridgeshire, United Kingdom
Employment Type
Full-Time
Salary
£90,000 - £120,000 per annum
plus general scripting Nice to have rather than essential: a postgraduate degree and research background (publications welcome), real depth on inference serving frameworks like vLLM, a stats background, and any open source or research contributions. Why look twice at this one It's a rare early seat at something with ...

Performance Engineer (Junior) | AI Infrastructure | Cambridge (Hybrid)

Hiring Organisation
Pure Resourcing Solutions Limited
Location
Linton, Dry Drayton, Cambridgeshire, United Kingdom
Employment Type
Permanent
Salary
£55000 - £70000/annum
monitoring stacks (Prometheus, Grafana) Strong Python for data work, Pandas and NumPy, genuine scripting ability Nice to have: exposure to inference serving frameworks like vLLM, published research, or open source contributions in this space. Why look twice at this one An early route into industry for someone whose academic record ...

Principal Product Manager - AI Tooling

Location
Cambridge, England, United Kingdom
management in development tools and software including CLIs, SDKs and APIs. Knowledge of machine learning frameworks, runtimes, infrastructure such as PyTorch, ONNX, ExecuTorch, Llama.cpp, vLLM and LiteRT. Understanding of model deployment to edge, embedded or heterogeneous computing, with an understanding of trade-offs between accuracy, performance and power consumption. ...

Senior Software Engineer - AI Compiler

Location
Cambridge, England, United Kingdom
Skills and Experience : Experience using software simulators and/or FPGAs Exposure to any of the following: driver development, inference engines like llama.cpp or vLLM, performance analysis of ML and GenAI workloads, knowledge of Neural Network Processing Units (NPU) or Graphics Processing Units (GPU) and how they are used ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Milton, Cambridgeshire, UK
reliability, performance, and cost-efficiency of the LLM inference platform end to end. You will operate large language model serving stacks (such as vLLM and llm-d) in production at scale, with deep instrumentation and strong operational rigor. You will partner across engineering to deliver secure software, improve stability … infrastructure Build backend services and APIs that enable reliable operation of AI infrastructure in production Operate and scale LLM serving infrastructure (such as vLLM and llm-d), including model hosting, request routing, continuous batching, and KV-cache optimization Deploy, host, and lifecycle-manage open-source and proprietary LLMs on Amazon ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Milton, Cambridgeshire, UK
guide peers on safe and effective usage within team practices Preferred qualifications, capabilities, and skills Experience operating large language model inference servers such as vLLM and llm-d (or directly equivalent model serving stacks) in production Experience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patterns … monitoring, including hallucination detection, toxicity filtering, and drift detection using open telemetry conventions Contributions to open-source large language model serving or inference projects, (vLLM, llm-d, Ray, KServe, Triton) ABOUT US Our client is a global leader in financial services, providing strategic advice and products to the world ...

Staff Software Engineer - AI Compiler

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, UK
Employment Type
Full-time
Have" Skills and Experience :Experience using software simulators and/or FPGAsExposure to any of the following: driver development, inference engines like llama.cpp or vLLM, performance analysis of ML and GenAI workloads, knowledge of Neural Network Processing Units (NPU) or Graphics Processing Units (GPU) and how they are used ...

Senior Software Engineer (vLLM)

Location
Cambridge, England, United Kingdom
Opportunity We are seeking a Senior Software Engineer with a passion for open-source AI infrastructure to work on deploying, extending and optimising vLLM (and potentially other inference serving engines) to support our projects. You will play a crucial, high-impact role across both our key programmes, the Scaling Inference … autonomous agents for software development. What You'll Do Deploy, instrument and monitor open weight models served using vLLM. Implement new features within vLLM to support novel hardware architectures as part of the Scaling Inference Lab. Work with the Panopticon team to identify opportunities to extend vLLM to enhance accuracy ...

Senior Software Engineer - Open-Source AI Inference & vLLM

Location
Cambridge, England, United Kingdom
CommonAI CIC is seeking a Senior Software Engineer to help deploy, extend and optimise vLLM and related inference engines for our AI infrastructure projects. You will work across the Scaling Inference Lab and the High Assurance programme to drive performance, reliability and safety in production‐grade systems. You will collaborate ...