4 of 4 vLLM Jobs in Lanarkshire

Associate Director, Lead AI Architect, Security & Justice, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 80 K
optimise models using cloud services like AWS Bedrock and Azure AI Foundry, or self-host them on GPU/CPU hardware using tools like vLLM, SGLang, and Ollama.Solution Evaluation: Implement frameworks and approaches to evaluate model performance against business objectives, both pre-deployment and on an ongoing basis as part ...

Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
reliability, performance, and cost-efficiency of the LLM inference platform end to end. You will operate large language model serving stacks (such as vLLM and llm-d) in production at scale, with deep instrumentation and strong operational rigor. You will partner across engineering to deliver secure software, improve stability … infrastructure Build backend services and APIs that enable reliable operation of AI infrastructure in production Operate and scale LLM serving infrastructure (such as vLLM and llm-d), including model hosting, request routing, continuous batching, and KV-cache optimization Deploy, host, and lifecycle-manage open-source and proprietary LLMs on Amazon ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
Hackajob Ltd
Location
Glasgow, Lanarkshire, Scotland, United Kingdom
Employment Type
Permanent
guide peers on safe and effective usage within team practices Preferred qualifications, capabilities, and skills Experience operating large language model inference servers such as vLLM and llm-d (or directly equivalent model serving stacks) in production Experience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patterns … monitoring, including hallucination detection, toxicity filtering, and drift detection using open telemetry conventions Contributions to open-source large language model serving or inference projects, (vLLM, llm-d, Ray, KServe, Triton) ABOUT US J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world ...

Senior Lead Software Engineer - LLM Ops Platform Reliability

Hiring Organisation
JP Morgan Chase
Location
Glasgow, Lanarkshire, United Kingdom
Salary
£ 80 K
ability to guide peers on safe and effective usage within team practicesPreferred qualifications, capabilities, and skillsExperience operating large language model inference servers such as vLLM and llm-d (or directly equivalent model serving stacks) in productionExperience developing generative AI applications, AI agents, vector search, and retrieval-augmented generation patternsExperience building … quality monitoring, including hallucination detection, toxicity filtering, and drift detection using open telemetry conventionsContributions to open-source large language model serving or inference projects, (vLLM, llm-d, Ray, KServe, Triton)J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world’s most prominent ...