76 to 91 of 91 vLLM Jobs

Developer Experience Engineer New London

Location
Greater London, England, United Kingdom
user and a working results. Bonus Points These aren't requirements, but they'd make you stand out: Experience with inference serving stacks (vLLM, SGLang, TensorRT-LLM, Triton). A track record of contributing to or maintaining open source developer tools and ML ecosystem projects Experience writing technical documentation ...

AI Inference Engineer

Hiring Organisation
Fuse Energy Supply
Location
London, United Kingdom
Salary
£ 80 K
improve throughput and cost per token, partnering with the CUDA/GPU engineersMake the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server or equivalents)Translate throughput, latency and uptime commitments into concrete technical specifications and serving capacity plansAct as direct technical ...

AI Inference Engineer

Location
Greater London, England, United Kingdom
throughput and cost per token, partnering with the CUDA/GPU engineers. Make the core software architecture calls on serving frameworks and orchestration (e.g. vLLM, TensorRT-LLM, SGLang, Triton Inference Server, or equivalents). Translate throughput, latency, and uptime commitments into concrete technical specifications and serving capacity plans. ...

Customer Solution Architect — Arango AI Product Suite

Location
United Kingdom
MLOps platforms and evaluation frameworks (MLflow, Weights & Biases, Ragas, promptfoo, DeepEval). Model adaptation and inference optimization awareness (LoRA/PEFT, DPO, distillation, quantization, vLLM/TGI/TensorRT-LLM), enough to advise on tradeoffs rather than to hand-build. Domain experience in finance, healthcare, public sector, manufacturing, or retail. … pgvector, Pinecone, Weaviate; rerankers (ColBERT, cross-encoders) Pipelines & Orchestration:LangChain, LlamaIndex, Ray, Airflow MLOps & Evals:MLflow, Weights & Biases, Ragas, promptfoo, Great Expectations Serving & Infra:vLLM, TGI, FastAPI/gRPC, Docker/K8s, Terraform, GitHub Actions Observability & Guardrails:OpenTelemetry, Prometheus/Grafana, Llama Guard/Content Safety, custom filters Data:Postgres ...

Senior Software Engineer (vLLM)

Location
Cambridge, England, United Kingdom
Opportunity We are seeking a Senior Software Engineer with a passion for open-source AI infrastructure to work on deploying, extending and optimising vLLM (and potentially other inference serving engines) to support our projects. You will play a crucial, high-impact role across both our key programmes, the Scaling Inference … autonomous agents for software development. What You'll Do Deploy, instrument and monitor open weight models served using vLLM. Implement new features within vLLM to support novel hardware architectures as part of the Scaling Inference Lab. Work with the Panopticon team to identify opportunities to extend vLLM to enhance accuracy ...

ML Research Engineer - Member of Technical Staff

Location
Greater London, England, United Kingdom
generation Depth in multi‐agent systems, planning, program synthesis, or retrieval over structured artefacts such as codebases Deep familiarity with the internals of SGLang, vLLM, or comparable inference serving frameworks - scheduler design, memory management, and execution pipelines A background in another field that studies systems of interacting heterogeneous components - neuroscience ...

MLOps Engineer (LLM/GenAI)

Location
Sheffield, England, United Kingdom
heterogeneous hardware Optimise inference for latency, throughput, and cost (e.g., quantisation, KV-cache optimisation, dynamic/continuous batching) Evaluate and integrate inference frameworks (e.g., vLLM, TensorRT-LLM, SGLang) to maximise performance on target hardware Own inference health/performance monitoring (latency, throughput, TTFT, memory, availability) and troubleshoot bottlenecks/deployment … architecture and HPC fundamentals Deep inference optimisation expertise: KV-cache, batching, quantisation (INT4/FP8/GPTQ/AWQ), operator optimisation, framework integration (vLLM/TensorRT-LLM/SGLang) Production hosting experience with Docker/Kubernetes and cloud platforms (AWS/GCP/Azure) End-to-end fine-tuning expertise ...

MLOps Platform Developer / Full-Stack AI Engineer

Hiring Organisation
BluetownOnline Ltd
Location
London, United Kingdom
Employment Type
Permanent
results focus. What you will do: Operate the LLM estate: run LoRA fine-tuning cycles and evaluation gates on our GPU hardware, manage vLLM serving (including multi-adapter deployments) alongside production, promote or roll back model versions on the gate results, and keep the serving watchdogs healthy. The CEO retains … world system of record. Python for scripting, data processing or pipeline work. Preferable: Deeper LLM experience: fine-tuning methodology, corpus design, evaluation-harness construction, vLLM internals, or heavy daily use of AI coding agents. ERP, FSM or works-management domain experience job lifecycles, scheduling, SLAs, parts, timesheets, invoicing or heat ...

Senior AI Compute Infrastructure Engineer

Hiring Organisation
Kraken
Location
United Kingdom
Salary
£ 70 K
placement, quota management, and utilization systems across heterogeneous accelerator environments.Optimize inference pipelines for latency, throughput, reliability, memory efficiency, and cost using frameworks such as vLLM, Triton Inference Server, TensorRT, or equivalent serving stacks.Partner with ML engineers and researchers to remove bottlenecks in training, evaluation, batch inference, online inference, deployment … monitoring, and cost optimization.Strong systems engineering fundamentals across Linux, networking, storage, containers, Kubernetes, distributed runtimes, and production debugging.Experience with ML serving frameworks such as vLLM, Triton Inference Server, TensorRT, TorchServe, KServe, Ray Serve, or equivalent systems.Proficiency in Python for infrastructure automation, tooling, debugging, integration, and operational workflows.Practical understanding of performance ...

Senior Software Engineer - Open-Source AI Inference & vLLM

Location
Cambridge, England, United Kingdom
CommonAI CIC is seeking a Senior Software Engineer to help deploy, extend and optimise vLLM and related inference engines for our AI infrastructure projects. You will work across the Scaling Inference Lab and the High Assurance programme to drive performance, reliability and safety in production‐grade systems. You will collaborate ...

ML Software Engineer, Data Plane

Hiring Organisation
Annapurna Labs (U.S.) Inc
Location
Cupertino, California, United States
Employment Type
Permanent
Salary
USD Annual
architectures end-to-end - from PyTorch model definition through distributed execution on custom hardware. - Integrate custom accelerator backends into open-source ML serving frameworks (vLLM, PyTorch), including scheduler extensions, memory management, and model parallelism. - Build and maintain test infrastructure for model correctness validation across CPU, GPU, simulator, and hardware targets. … Knowledge of Machine Learning and LLM fundamentals, including transformer architecture, training/inference lifecycles, and optimization techniques - Knowledge of ML frameworks including JAX, PyTorch, vLLM, SGLang, Dynamo, TorchXLA, and TensorRT - Experience in developing and deploying LLMs in production on GPUs, Neuron, TPU or other AI acceleration hardware - Experience with distributed ...

Research Engineer (Inference & Serving)

Location
Greater London, England, United Kingdom
backed challenger lab building state-of-the-art computer-use agents. The inference team owns the full stack from engine layer (vLLM, SGLang) through to serving architecture (disaggregated inference, intelligent routing). The team operates at the intersection of research and production - translating cutting-edge techniques directly into the systems … least one systems language - Rust, C++, or Go Hands-on experience with PyTorch or JAX in an industry setting Experience with inference frameworks: vLLM, SGLang, TensorRT-LLM Solid distributed systems fundamentals and experience operating production ML infrastructure Working knowledge of modern ML including transformers and multimodal architectures Optional Bonus Research ...

Senior Researcher - AI Computer Architecture

Hiring Organisation
Microsoft
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 70 K
Job ID: 200037336Posted: 2026-09-07Location: United Kingdom, Cambridgeshire, CambridgeEmployment type: Full-TimeWork site: 3 days/week in-officeRole type: Individual ContributorTravel: Less than 25%Profession: Research, Applied, & Data SciencesDiscipline: Research SciencesCompany: MicrosoftOverviewThe ...

Systems Research Engineer

Hiring Organisation
European Tech Recruit
Location
Edinburgh, Scotland, United Kingdom
depth profiling of large-scale inference pipelines, specifically focusing on KV cache management and heterogeneous memory scheduling. AI Serving: Optimising high-throughput frameworks (vLLM, Ray Serve, PyTorch Distributed) to ensure low-latency, multi-tenant performance. Research Leadership: Contributing to top-tier venues (OSDI, NSDI, EuroSys, MLSys) and driving those innovations … Stack: Strong proficiency in C/C++ for systems work, with Python for rapid prototyping. Expertise: Hands-on experience with LLM serving frameworks ( vLLM, Ray Serve, TensorRT-LLM ) and distributed algorithms. Mindset: A solid grounding in systems research methodology and performance profiling tools. The "Value Add" (Desired): A PhD focused ...

Member of Technical Staff

Location
Greater London, England, United Kingdom
implementation level. Attention variants, KV cache strategies, quantisation schemes, and how they shape kernel design You've worked with production inference or training frameworks, vLLM, Megatron-LM, etc You've built performance-critical infrastructure before - compilers, profilers, auto-tuners, or search systems You have real intuition for evolutionary methods, fitness … work of François Chollet, Kenneth Stanley, Jeff Clune, Jurgen Schmidhuber, David Ha, and Christian Szegedy Bonus: Open-source kernel contributions (FlashAttention, FlashInfer, vLLM, Unsloth, Liger-Kernels, ThunderKittens) Publications in ML/AI, kernel optimisation or evolutionary methods (NeurIPS, ICLR, CVPR, GECCO or equivalent) Other HW experience (AMD, MLX, edge ...

Architect/Staff Systems Software Engineer

Location
Greater London, England, United Kingdom
direction you shape across the platform. Responsibilities Own the Runtime & Serving Stack: Design, build, and extend the distributed inference and serving stack (e.g. vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM) onto DX-1, rather than treating any layer as a black box. Scale Distributed Inference: Define how inference scales across many … runtime/network/accelerator boundary. Demonstrated ownership of a hard, end-to-end systems problem, ideally extending a distributed inference/serving stack (vLLM, SGLang, NVIDIA Dynamo, TensorRT-LLM) in production, with specifics on what you built or changed and why. Distributed inference at scale: parallelism strategies, collective communication ...