3 of 3 Permanent Reinforcement Learning Jobs in the Midlands

Lead Data Scientist

Hiring Organisation
Kainos
Location
Birmingham, West Midlands (County), United Kingdom
Salary
£ 80 K
others to deliver advanced AI and data solutions at citizen scale. Our 150-strong AI and Data Practice brings together deep expertise in machine learning, generative AI, agentic AI and data. We are pioneers in responsible AI, having authored the UK government’s AI Cyber Security Code of Practice … BUSINESS: As a Lead Data Scientist at Kainos, you will architect, design, and deliver advanced AI solutions leveraging state-of-the-art machine learning, generative and agentic AI technologies. You will drive the adoption of modern AI frameworks, AIOps best practices and scalable cloud-native architectures. Your role will ...

Data Scientist Manager

Hiring Organisation
Kainos
Location
Birmingham, West Midlands (County), United Kingdom
Salary
£ 80 K
others to deliver advanced AI and data solutions at citizen scale. Our 150-strong AI and Data Practice brings together deep expertise in machine learning, generative AI, agentic AI and data. We are pioneers in responsible AI, having authored the UK government’s AI Cyber Security Code of Practice … Data Scientist Manager at Kainos, you’ll be responsible for successful delivery of advanced AI solutions leveraging state-of-the-art machine learning, generative and agentic AI technologies. You will drive the adoption of modern AI development and scalable cloud-native architectures. Your role will involve technical leadership, engaging ...

Research Engineer

Location
Tipton, England, United Kingdom
product‐minded research engineer who enjoys building real systems people depend on. You’ll likely have: Strong technical background in software engineering, machine learning, or applied AI, demonstrated through an advanced degree and/or equivalent experience building production AI systems Strong software engineering fundamentals and good judgment … Practical understanding of modern model adaptation and post‐training methods, including LoRA/QLoRA, SFT, distillation, preference optimization, reward modeling, DPO/GRPO, and reinforcement learning from verifiable feedback Ownership mindset: you drive projects end‐to‐end, move quickly from real usage, and care about shipping measurable improvements ...