201 to 225 of 261 Reinforcement Learning Jobs in the UK

Research Scientist, Robotics Pre-Training and Data Quality, DeepMind

Hiring Organisation
Google
Location
London, United Kingdom
Salary
£ 70 K
tools and processes required to iterate on research ideas quickly and verify data quality through robot policy training.Contribute to our wider research agenda, including reinforcement learning, vision-language-action modeling, world-action models, and simulation, working in a fast-paced, collaborative team environment.Minimum qualifications:PhD in a technical … designing new architectures, our research scientists work on real-world problems that span the breadth of computer science, such as machine (and deep) learning, data mining, natural language processing, hardware and software performance analysis, improving compilers for mobile platforms, as well as core search and much more.As a Research ...

RLHF Manager

Location
Cheltenham, England, United Kingdom
Research & QualityHybridFull-timeDisability Confident Committed# RLHF ManagerLead Reinforcement Learning from Human Feedback (RLHF) programmes across Coaley Peak's internal AI engine development and external client AI projects, ensuring our models are safe, aligned, and genuinely useful.LocationRemote (UK)/Cheltenham, UKSalary£55,000 – £75,000 per annum (DOE)Positions1 … positionReferenceCP-RLF-2025-001Published21 March 2025Closing14 December 2026›AI Research & Quality›RLHF ManagerAbout the roleReinforcement Learning from Human Feedback is one of the most important mechanisms for making AI systems behave well in the real world. As RLHF Manager at Coaley Peak, you will design and manage ...

Machine Learning Engineer, Memory

Hiring Organisation
Epic Games
Location
London, United Kingdom
Salary
£ 100 K
that work across problems, modalities and productsDefine new ways to measure and evaluate memory performanceHelp create and deliver production-ready, scalable and high-quality learning models and associated algorithmsCritically assess the effectiveness of such models and make recommendations for the ongoing roadmapWhat we're looking forPhD in Computer Science … Mathematics or related field, or 3+ years of relevant industry experienceExperience creating machine learning algorithms for vision and/or language problems and deploying them as production-level systemsExpert, hands-on knowledge of:Foundational models and vision and/or language, incl. local deployment and model post-training ...

Weapon Systems Algorithms Engineer - Summer Placement 2027

Location
West of England, England, United Kingdom
close earlier depending on application volumes. What we can offer you: 10 week placement: Starting June 2027, that allows you to apply your university learning to real-world projects and technologies Perform well on your placement and you could secure a graduate role before your final year Pension: Maximum … future business teams to bring technical excellence to our future concepts. We leverage the use of leading technologies and tools such as deep reinforcement learning and test-driven algorithm development to provide optimised weapon system solutions across the whole of the MBDA product line. Over your time ...

Genesis-World: Core Simulation Engine Engineer

Location
Greater London, England, United Kingdom
compiler, lowers them to CUDA, AMD ROCm, Apple Metal, Vulkan, x86, and ARM64. A single laptop or a datacenter. Massively batched GPU simulation for learning at scale, and complex non-batched scenes where CPU wins outright. This is at the core of Genesis AI’s strategy. Evaluation … that improves at the speed of compute. The role Simulation is still a hard sell in robotics. Outside a few success stories, like reinforcement learning for locomotion, most teams skip it, and friction is a big part of why: painful to use, painful to debug. We live that ...

Senior AI Engineer — Design & Manufacturing (Generative AI)

Location
Cambridge, England, United Kingdom
agents for enterprise workflows, and create robust training pipelines that enable rapid experimentation and deployment at scale. We value hands-on expertise with LLMs, reinforcement learning, and agentic systems, along with strong Python and software engineering fundamentals. #J-18808-Ljbffr ...

Controllable Video Generation Researcher

Hiring Organisation
Tencent
Location
London, United Kingdom
Salary
£ 70 K
state retention.6. Collaborate with the Engine/Agent teams to define standardized interfaces for controllable signals.7. Drive the joint optimization of controllable generation and Reinforcement Learning post-training.Who We Look For1. Ph.D. in AI-related fields, focusing on Computer Vision/Deep Learning.2. Proficient in conditional generation models ...

Senior Software Engineer, Data Platform London, UK Apply →

Location
Greater London, England, United Kingdom
enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human evaluation and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT ...

Senior Software Engineer, Data Platform

Hiring Organisation
Scale AI
Location
London, United Kingdom
Salary
£ 80 K
enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human evaluation and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT ...

Senior RL Engineer for Humanoid Robots (Equity)

Location
Greater London, England, United Kingdom
Humanoid is seeking a Senior or Staff Reinforcement Learning Engineer to develop learning-based control policies for humanoid robots. You will design and train RL policies enabling dynamic locomotion and loco-manipulation on real robots, building scalable pipelines and improving sim-to-real transfer for reliable hardware ...

Research Engineer

Location
Tipton, England, United Kingdom
product‐minded research engineer who enjoys building real systems people depend on. You’ll likely have: Strong technical background in software engineering, machine learning, or applied AI, demonstrated through an advanced degree and/or equivalent experience building production AI systems Strong software engineering fundamentals and good judgment … Practical understanding of modern model adaptation and post‐training methods, including LoRA/QLoRA, SFT, distillation, preference optimization, reward modeling, DPO/GRPO, and reinforcement learning from verifiable feedback Ownership mindset: you drive projects end‐to‐end, move quickly from real usage, and care about shipping measurable improvements ...

AI Research Engineer

Hiring Organisation
Zealous Agency
Location
England, UK
well-funded AI research lab about to blast out of stealth mode, doing some novel work at the frontier of reinforcement learning. They’re moving into their next phase of growth after closing funding and are already working alongside major frontier AI labs. As a Founding RL Researcher … real influence over technical direction from day one and work directly with experienced second-time founders. Core skills and experience required Strong background in reinforcement learning Hands-on post-training experience: RLHF, GRPO etc Experience designing or building RL environments grounded in real-world workflows Distributed training ...

Advisory Director, EMEA

Location
United Kingdom
work at a startup. You are ambitious. You have fun solving problems that others think are impossible. You are curious. You find joy in learning about AI, technology, and finance. You are an owner. You are autonomous, self-directed, and comfortable working with ambiguity. You are collaborative, organized, thoughtful … which means you learn a lot and constantly take on more. Frontier technology: we're developing cutting‐edge AI systems, pushing the boundaries of reinforcement learning and published research, redefining what's possible, and inventing the future. Cutting Edge Product: Our platform is state ...

Director of Applied AI

Hiring Organisation
iCOMAT
Location
Gloucester, Gloucestershire, United Kingdom
Salary
£ 60 K
write code, design systems, and lead by example. You are not allergic to a whiteboard or a shop floor.You have shipped applied machine learning solutions in a real industrial, manufacturing, robotics, or hard tech setting. Not just dashboards. Not just demos.You speak fluent KPI. You can defend … production deployment experience at scale.Great plus:Background in composites, advanced manufacturing, aerospace, automotive, or defense.Experience with computer vision applied to automated inspection, generative design, reinforcement learning for process control.Prior experience standing up an AI function from zero inside a fast moving company.Our Values: Uncompromising ExcellenceDiscipline over excuses. ...

Director of Applied AI

Hiring Organisation
iCOMAT
Location
Gloucester, Gloucestershire, UK
Employment Type
Full-time
write code, design systems, and lead by example. You are not allergic to a whiteboard or a shop floor. You have shipped applied machine learning solutions in a real industrial, manufacturing, robotics, or hard tech setting. Not just dashboards. Not just demos. You speak fluent KPI. You can defend … experience at scale. Great plus: Background in composites, advanced manufacturing, aerospace, automotive, or defense. Experience with computer vision applied to automated inspection, generative design, reinforcement learning for process control. Prior experience standing up an AI function from zero inside a fast moving company. Our Values: Uncompromising ExcellenceDiscipline over ...

Platform Engineer

Location
Greater London, England, United Kingdom
superlearning capability - the ability to endlessly discover knowledge and skills, without relying on human data - will be driven by the world’s most powerful reinforcement learning algorithms. The superlearner is expected to rediscover and then transcend the greatest inventions in human history, such as language, science, mathematics … role where you inherit someone else's architecture, you'll be one of the people defining it. We're pushing these learning methods to a scale that hasn’t been tried before and we care deeply about the infrastructure that gets us there. You'll own the infrastructure that ...

Agent Product Manager

Location
Greater London, England, United Kingdom
usage and model performance, identify key issues, validate hypotheses, and continuously improve product experience and Agent capabilities. Requirements - Basic understanding of LLMs, AI Agents, reinforcement learning, and model training or post-training workflows. - Strong interest in general-purpose Agent products and enthusiasm for building product-data flywheels. - Experience … evaluation, RLHF/RLAIF, data annotation platforms, or model training data pipelines is a strong plus. Basic understanding of LLMs, AI Agents, and reinforcement learning.Strong product judgment, with the ability to identify user needs, define problems, and design practical product solutions.Strong structured thinking and analytical skills, with the ability ...

Simulation Engineer

Hiring Organisation
Helsing
Location
London, United Kingdom
Salary
£ 70 K
make a meaningful impact in the worldOur work frequently takes us right up to the state of the art in technical innovation, be it reinforcement learning, distributed systems, generative AI, or deployment infrastructure. The defence industry is entering the most exciting phase of the technological development curve. Advances ...

Video Generation Content Understanding and Feedback Researcher

Hiring Organisation
Tencent
Location
London, United Kingdom
Salary
£ 70 K
track industry-related work and formulate proprietary technical roadmaps.Who We Look For1.Ph.D. in AI-related fields, covering areas such as video understanding, video prediction, reinforcement learning, or multimodal generation and understanding.2.Deep understanding of the internal mechanisms of both video generation and VLM models; familiar with diffusion model principles.3.Research … generation.4.Ability to independently design complex model training pipelines with strong engineering skills.5.Proficient in Python/PyTorch, with publications in video understanding, controllable generation, or reinforcement learning.6.Experience with Game AI, simulation environments, or closed-loop systems is preferred.7.Relevant publications in top-tier conferences or journals are preferred.Equal Employment Opportunity ...

Senior Data Scientist in ADVANCED ANALYTICS

Location
Leeds, England, United Kingdom
opportunity to work at the heart of data science, research and public policy. Advanced Analytics is hiring a Senior Data Scientist to develop machine learning (ML) and Artificial Intelligence (AI) solutions, with a focus on robust prototyping. We have pioneered the application of new analytical approaches in the Bank … England, such as generative AI, machine learning, natural language processing, agent-based models, network analytics, and randomised control trials. We’re looking for someone with ML/AI expertise to help us continue pushing the organisation’s analytical frontier to inform policymaking. DAT is leading the transformation ...

Senior Data Scientist in ADVANCED ANALYTICS

Location
City Of London, England, United Kingdom
opportunity to work at the heart of data science, research and public policy. Advanced Analytics is hiring a Senior Data Scientist to develop machine learning (ML) and Artificial Intelligence (AI) solutions, with a focus on robust prototyping. We have pioneered the application of new analytical approaches in the Bank … England, such as generative AI, machine learning, natural language processing, agent-based models, network analytics, and randomised control trials. We’re looking for someone with ML/AI expertise to help us continue pushing the organisation’s analytical frontier to inform policymaking. DAT is leading the transformation ...

Vice President, Model Research & Development

Hiring Organisation
Thomson Reuters
Location
London, United Kingdom
Salary
£ 80 K
release decision) for models that ship to real users across CoCounsel, Westlaw, and Practical Law. Set the target for an online, agentic reinforcement-learning pipeline, training with subject-matter experts in the loop, and hold the work accountable to it. Ensure data strategy and evaluation approach are rigorous … weeks per year, empowering employees to achieve a better work-life balance.Career Development and Growth: By fostering a culture of continuous learning and skill development, we prepare our talent to tackle tomorrow’s challenges and deliver real-world solutions. Our Grow My Way programming and skills-first approach ensures ...

Onsite RL Engineer – Robotics & Embodied AI

Location
Greater London, England, United Kingdom
Randstad Technologies Recruitment is searching for a Reinforcement Learning (RL) Engineer to join a leading robotics company in London. This permanent position involves working collaboratively to define the company's 2026 technical roadmap while solving complex challenges in autonomous systems. The role requires expertise in AI, MLOps, Software ...

Senior RL Data Engineer: End-to-End Pipelines & QA

Location
Greater London, England, United Kingdom
Senior Engineer in Greater London to develop and manage data pipelines for AI systems. This role involves significant responsibilities, including ensuring the quality of reinforcement learning data, collaborating with various teams, and innovating operational frameworks. The ideal candidate will possess strong software engineering skills, have a background ...

(Senior) Product Marketing Manager Marketing & Communications Munich - Berlin

Location
Greater London, England, United Kingdom
safety, and ethical considerations are vital. Our work frequently takes us right up to the state of the art in technical innovation, be it reinforcement learning, distributed systems, generative AI, or deployment infrastructure. The defence industry is entering the most exciting phase of the technological development curve. Advances ...