51 to 69 of 69 vLLM Jobs in England

Performance Engineer, Containers/Serverless

Location
Greater London, England, United Kingdom
compatible object storage - including performance-killing cases (small-object overhead, range-request patterns, eventual consistency, multipart tuning). Comfort with model-serving runtimes (vLLM, SGLang etc) and the formats they consume (safetensors, GGUF, sharded checkpoints). An end-to-end view: comfortable reasoning about NIC, switch, filesystem, cache, container runtime ...

Research Associate in Adaptive and Efficient LLM Architectures

Hiring Organisation
Imperial College London
Location
London, United Kingdom
Salary
£ 55 K
training, evaluation, RLVR, PEFT, quantisation, tensor/data parallelism.Ideally, the candidate should have familiarity with CUDA kernels and/or Triton, and inference engines (vLLM, SGLang, et cetera).Experience coding with deep learning libraries such as Pytorch/JAX is essential.Fluent written and spoken English skills as well as contributions ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Milton Keynes, England, United Kingdom
reliability, performance, and cost-efficiency of the LLM inference platform end to end. You will operate large language model serving stacks (such as vLLM and llm-d) in production at scale, with deep instrumentation and strong operational rigor. You will partner across engineering to deliver secure software, improve stability … infrastructure Build backend services and APIs that enable reliable operation of AI infrastructure in production Operate and scale LLM serving infrastructure (such as vLLM and llm-d), including model hosting, request routing, continuous batching, and KV-cache optimization Deploy, host, and lifecycle-manage open-source and proprietary LLMs on Amazon ...

Backend Engineer - API

Hiring Organisation
X
Location
London, United Kingdom
Salary
> £ 150 K
operating commonly used databases such as PostgreSQL, Clickhouse, and MongoDBPREFERRED SKILLS AND EXPERIENCE:Experience with LLM inference engines and serving frameworks (e.g., SGLang, TensorRT, vLLM)Experience designing or building with agent SDKs and agent orchestration frameworksExperience with Docker, Kubernetes, and containerized applicationsExpert knowledge of gRPC (unary, response streaming, bi-directional ...

Project Technical Lead - AI Systems Simulation

Location
Cambridge, England, United Kingdom
infrastructure, ML systems, or computer architecture. Familiarity with Agile or other modern technical project management frameworks. Knowledge of modern inference‐serving frameworks (e.g., vLLM). Background in statistics, operations research, or large‐scale datacenter infrastructure. Contributions to open‐source AI or systems projects. Benefits High‐impact role in a rapidly ...

NLP Performance Engineer

Hiring Organisation
G Research
Location
London, United Kingdom
Salary
£ 80 K
understanding of transformer inference, including prefill versus decode, KV-cache behaviour, attention variants and performance bottlenecksHands-on experience with LLM serving frameworks such as vLLM, SGLang, TensorRT-LLM or TGI, and the PyTorch ecosystemExperience with inference optimisation techniques, including quantisation, speculative decoding and model parallelism across modern GPU architecturesStrong software ...

NLP Performance Engineer

Location
Greater London, England, United Kingdom
transformer inference, including prefill versus decode, KV‐cache behaviour, attention variants and performance bottlenecks Hands‐on experience with LLM serving frameworks such as vLLM, SGLang, TensorRT‐LLM or TGI, and the PyTorch ecosystem Experience with inference optimisation techniques, including quantisation, speculative decoding and model parallelism across modern GPU architectures Strong ...

Developer Experience Engineer New London

Location
Greater London, England, United Kingdom
user and a working results. Bonus Points These aren't requirements, but they'd make you stand out: Experience with inference serving stacks (vLLM, SGLang, TensorRT-LLM, Triton). A track record of contributing to or maintaining open source developer tools and ML ecosystem projects Experience writing technical documentation ...

AI Inference Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
improve throughput and cost per token, partnering with the CUDA/GPU engineersMake the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)Translate throughput, latency and uptime commitments into concrete technical specifications and serving capacity plansAct as direct technical ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
throughput and cost per token, partnering with the CUDA/GPU engineers. Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents). Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans. ...

Senior Data Center, AI Software Product Strategy & Go-To-Market Lead — Qualcomm - Europe

Location
Greater London, England, United Kingdom
optimized in a Data Center AI software platforms Inference, model deployment, runtimes, SDKs, developer tools. AI orchestration frameworks, e.g., Kubernetes-based systems, Ray, vLLM, TGI. Cloud-native AI software stacks, deployment patterns, and data center operating models. Data center AI infrastructure CPUs, GPUs, NPUs, and heterogeneous accelerator environments. Kubernetes-based ...

Senior Software Engineer (vLLM)

Location
Cambridge, England, United Kingdom
Opportunity We are seeking a Senior Software Engineer with a passion for open-source AI infrastructure to work on deploying, extending and optimising vLLM (and potentially other inference serving engines) to support our projects. You will play a crucial, high-impact role across both our key programmes, the Scaling Inference … autonomous agents for software development. What You'll Do Deploy, instrument and monitor open weight models served using vLLM. Implement new features within vLLM to support novel hardware architectures as part of the Scaling Inference Lab. Work with the Panopticon team to identify opportunities to extend vLLM to enhance accuracy ...

ML Research Engineer - Member of Technical Staff

Location
Greater London, England, United Kingdom
generation Depth in multi‐agent systems, planning, program synthesis, or retrieval over structured artefacts such as codebases Deep familiarity with the internals of SGLang, vLLM, or comparable inference serving frameworks - scheduler design, memory management, and execution pipelines A background in another field that studies systems of interacting heterogeneous components - neuroscience ...

MLOps Platform Developer / Full-Stack AI Engineer

Hiring Organisation
BluetownOnline Ltd
Location
London, United Kingdom
Employment Type
Permanent
results focus. What you will do: Operate the LLM estate: run LoRA fine-tuning cycles and evaluation gates on our GPU hardware, manage vLLM serving (including multi-adapter deployments) alongside production, promote or roll back model versions on the gate results, and keep the serving watchdogs healthy. The CEO retains … world system of record. Python for scripting, data processing or pipeline work. Preferable: Deeper LLM experience: fine-tuning methodology, corpus design, evaluation-harness construction, vLLM internals, or heavy daily use of AI coding agents. ERP, FSM or works-management domain experience job lifecycles, scheduling, SLAs, parts, timesheets, invoicing or heat ...

Senior Software Engineer - Open-Source AI Inference & vLLM

Location
Cambridge, England, United Kingdom
CommonAI CIC is seeking a Senior Software Engineer to help deploy, extend and optimise vLLM and related inference engines for our AI infrastructure projects. You will work across the Scaling Inference Lab and the High Assurance programme to drive performance, reliability and safety in production‐grade systems. You will collaborate ...

Research Engineer (Inference & Serving)

Location
Greater London, England, United Kingdom
backed challenger lab building state-of-the-art computer-use agents. The inference team owns the full stack from engine layer (vLLM, SGLang) through to serving architecture (disaggregated inference, intelligent routing). The team operates at the intersection of research and production - translating cutting-edge techniques directly into the systems … least one systems language - Rust, C++, or Go Hands-on experience with PyTorch or JAX in an industry setting Experience with inference frameworks: vLLM, SGLang, TensorRT-LLM Solid distributed systems fundamentals and experience operating production ML infrastructure Working knowledge of modern ML including transformers and multimodal architectures Optional Bonus Research ...

Senior Researcher - AI Computer Architecture

Hiring Organisation
Microsoft
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 70 K
Job ID: 200037336Posted: 2026-09-07Location: United Kingdom, Cambridgeshire, CambridgeEmployment type: Full-TimeWork site: 3 days/week in-officeRole type: Individual ContributorTravel: Less than 25%Profession: Research, Applied, & Data SciencesDiscipline: Research SciencesCompany: MicrosoftOverviewThe ...

Member of Technical Staff

Location
Greater London, England, United Kingdom
implementation level. Attention variants, KV cache strategies, quantisation schemes, and how they shape kernel design You've worked with production inference or training frameworks, vLLM, Megatron-LM, etc You've built performance-critical infrastructure before - compilers, profilers, auto-tuners, or search systems You have real intuition for evolutionary methods, fitness … work of François Chollet, Kenneth Stanley, Jeff Clune, Jurgen Schmidhuber, David Ha, and Christian Szegedy Bonus: Open-source kernel contributions (FlashAttention, FlashInfer, vLLM, Unsloth, Liger-Kernels, ThunderKittens) Publications in ML/AI, kernel optimisation or evolutionary methods (NeurIPS, ICLR, CVPR, GECCO or equivalent) Other HW experience (AMD, MLX, edge ...

Architect/Staff Systems Software Engineer

Location
Greater London, England, United Kingdom
direction you shape across the platform. Responsibilities Own the Runtime & Serving Stack: Design, build, and extend the distributed inference and serving stack (e.g. vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM) onto DX-1, rather than treating any layer as a black box. Scale Distributed Inference: Define how inference scales across many … runtime/network/accelerator boundary. Demonstrated ownership of a hard, end-to-end systems problem, ideally extending a distributed inference/serving stack (vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM) in production, with specifics on what you built or changed and why. Distributed inference at scale: parallelism strategies, collective communication ...