20 of 20 Remote/Hybrid vLLM Jobs

Associate Director Lead AI Architect

Hiring Organisation
Anson Mccade
Location
City of London, London, United Kingdom
Employment Type
Permanent, Work From Home
Foundry and other AWS, Azure or GCP services . Work with self-hosted inference and model-serving technologies where appropriate, including tools such as vLLM, SGLang and Ollama . Design evaluation, observability and monitoring frameworks to assess model performance, system health, reliability and business impact. Optimise AI solutions for performance ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
London, United Kingdom
Salary
£ 100 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Manchester, Greater Manchester, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Edinburgh, Midlothian, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Holywood, Down, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Senior / Principal Applied AI Engineer (UK / Europe, Remote)

Location
United Kingdom
haves Fintech/trading/market-data domain experience or experience as a trading-platform user. Open-source contributions to AI tooling (LangChain, LlamaIndex, vLLM, DSPy, or similar). Production RAG and vector-database experience; experience with code-generation, transpilation, or developer-experience tooling. Logistics Location: Europe (EU/ ...

Manager, Research Engineering (Foundational Research)

Hiring Organisation
Thomson Reuters
Location
London, United Kingdom
Salary
£ 80 K
purely technical preference.Technical Expertise:Deep proficiency in Python and modern software development practices.Hands-on experience with Distributed Training infrastructure (Multi-node GPU training, Kubernetes, vLLM).Familiarity with Deep Learning frameworks (PyTorch).Experience with MLOps tools and experiment tracking (e.g., ClearML, MLFlow, Weights & Biases).Research Fluency: Ability to read technical research ...

AI Engineer (Fluent Portuguese & English)

Hiring Organisation
Chubb
Location
London, United Kingdom
Salary
£ 80 K
Hands-on experience with LLM training cycles, parameter-efficient fine-tuning (PEFT), and sophisticated prompt engineering.Inference Stack: Experience with high-performance inference servers (e.g., vLLM, TGI, or Triton) and an understanding of how to optimize models for GPU deployment.Infrastructure: Comfortable working in Linux-based environments and proficient in managing containerized ...

AI Engineer (Fluent in Mandarin & English)

Location
Greater London, England, United Kingdom
experience with LLM training cycles, parameter-efficient fine-tuning (PEFT), and sophisticated prompt engineering. Inference Stack: Experience with high-performance inference servers (e.g., vLLM, TGI, or Triton ) and an understanding of how to optimize models for GPU deployment. Infrastructure: Comfortable working in Linux-based environments and proficient in managing containerized ...

Software Engineer (LLM Engineering), London

Hiring Organisation
Isomorphic Labs
Location
London, United Kingdom
Salary
£ 80 K
Deep understanding of how models are trained, how they work internally, and their inherent limitations.LLM Serving stack: Experience with the LLM serving stack (e.g. vLLM) for open weight models for bringing the latest models to internal users.ML Literacy: Experience evaluating probabilistic ML systems and managing model "tool use", contexts ...

Machine Learning Research Engineer (Foundational Research)

Hiring Organisation
Thomson Reuters
Location
London, United Kingdom
Salary
£ 80 K
production-quality code and data pipelines for ML systemsProficiency in modern AI development frameworks including: PyTorch, Jax , HuggingFace Transformers, LLM APIs (litellm etc) and vLLM for building and deploying large-scale AI applicationsUnderstanding of LLM training methodologies including instruction fine-tuning, preference optimization, and reinforcement learning approachesStrong software engineering skills ...

Senior Principal AI Infrastructure Architect

Hiring Organisation
NTT
Location
London, United Kingdom
Salary
£ 80 K
rail-optimised and fat-tree topologies. Working knowledge of the AI software and orchestration stack: CUDA, cuDNN, NCCL, ROCm, Triton Inference Server, NIM, vLLM, TensorRT-LLM, Slurm, Kubernetes (with GPU Operator), Kubeflow, Run:ai, MLflow and NVIDIA AI Enterprise. Familiarity with datacenter facilities engineering for AI workloads: high-density power ...

Senior Principal AI Infrastructure Architect

Hiring Organisation
The Nippon Telegraph And Telephone Corporation (NTT)
Location
United Kingdom
Salary
£ 70 K
rail-optimised and fat-tree topologies. Working knowledge of the AI software and orchestration stack: CUDA, cuDNN, NCCL, ROCm, Triton Inference Server, NIM, vLLM, TensorRT-LLM, Slurm, Kubernetes (with GPU Operator), Kubeflow, Run:ai, MLflow and NVIDIA AI Enterprise. Familiarity with datacenter facilities engineering for AI workloads: high-density power ...

Remote Senior Machine Learning Engineer

Location
Stirling, Scotland, United Kingdom
data, including classification, extraction, embeddings, re-rankers, clustering, and search. Experience in instruction fine-tuning and serving language models, familiarity with frameworks such as vLLM, DeepSpeed, or similar tools A solid grounding in classical ML and statistics, and the judgement to choose simpler methods when they’re the right solution. ...

Senior Research Scientist | Model Steering

Location
Greater London, England, United Kingdom
without degrading their reasoning capabilities. Experience with machine translation, multilingual NLP, or language quality estimation. Familiarity with inference and serving at scale (e.g. via vLLM, SGLang, TensorRT‐LLM, etc) and long‐context modelling. Publications at top‐tier venues. What we offer Diverse and internationally distributed team : joining our team means ...

Performance Engineer, Containers/Serverless

Location
Greater London, England, United Kingdom
compatible object storage - including performance-killing cases (small-object overhead, range-request patterns, eventual consistency, multipart tuning). Comfort with model-serving runtimes (vLLM, SGLang etc) and the formats they consume (safetensors, GGUF, sharded checkpoints). An end-to-end view: comfortable reasoning about NIC, switch, filesystem, cache, container runtime ...

Principal Product Manager - AI Tooling

Hiring Organisation
ARM
Location
Cambridge, Cambridgeshire, United Kingdom
Salary
£ 70 K
product management in development tools and software including CLIs, SDKs and APIs.Knowledge of machine learning frameworks, runtimes, infrastructure such as PyTorch, ONNX, ExecuTorch, Llama.cpp, vLLM and LiteRT.Understanding of model deployment to edge, embedded or heterogeneous computing, with an understanding of trade-offs between accuracy, performance and power consumption.You should also ...

Principal Product Manager - AI Tooling

Location
Cambridge, England, United Kingdom
management in development tools and software including CLIs, SDKs and APIs. Knowledge of machine learning frameworks, runtimes, infrastructure such as PyTorch, ONNX, ExecuTorch, Llama.cpp, vLLM and LiteRT. Understanding of model deployment to edge, embedded or heterogeneous computing, with an understanding of trade-offs between accuracy, performance and power consumption. ...

Developer Experience Engineer New London

Location
Greater London, England, United Kingdom
user and a working results. Bonus Points These aren't requirements, but they'd make you stand out: Experience with inference serving stacks (vLLM, SGLang, TensorRT-LLM, Triton). A track record of contributing to or maintaining open source developer tools and ML ecosystem projects Experience writing technical documentation ...