4 of 4 Permanent vLLM Jobs in Edinburgh

Sr. AI Architect Engineer

Location
City of Edinburgh, Scotland, United Kingdom
Google Professional Cloud Architect, TOGAF, or equivalent). Experience with large language model training and inference architecture at scale (e.g., PyTorch, DeepSpeed, Megatron-LM, vLLM). Experience with edge and on-device AI inference and the device-to-cloud continuum. Advanced Kubernetes experience: GPU scheduling plugins, multi-cluster and fleet ...

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Edinburgh, Midlothian, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

AI System Researcher

Hiring Organisation
Microtech Global Ltd
Location
Edinburgh, Midlothian, Scotland, United Kingdom
Employment Type
Permanent
Strong knowledge of distributed systems, operating systems, machine learning systems architecture, Inference serving, and AI Infrastructure. Hands-on experience with LLM serving frameworks (e.g., vLLM, Ray Serve, TensorRT-LLM, TGI) and distributed KV cache optimization. Proficiency in C/C++, with additional experience in Python for research prototyping. Solid grounding ...

Systems Research Engineer

Hiring Organisation
European Tech Recruit
Location
Edinburgh, Scotland, United Kingdom
depth profiling of large-scale inference pipelines, specifically focusing on KV cache management and heterogeneous memory scheduling. AI Serving: Optimising high-throughput frameworks (vLLM, Ray Serve, PyTorch Distributed) to ensure low-latency, multi-tenant performance. Research Leadership: Contributing to top-tier venues (OSDI, NSDI, EuroSys, MLSys) and driving those innovations … Stack: Strong proficiency in C/C++ for systems work, with Python for rapid prototyping. Expertise: Hands-on experience with LLM serving frameworks ( vLLM, Ray Serve, TensorRT-LLM ) and distributed algorithms. Mindset: A solid grounding in systems research methodology and performance profiling tools. The "Value Add" (Desired): A PhD focused ...