1 to 25 of 261 Permanent Reinforcement Learning Jobs in the UK

Research Engineer, Machine Learning (Reinforcement Learning)

Hiring Organisation
Humanloop
Location
London, United Kingdom
Salary
> £ 150 K
team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.About the teamsOur Reinforcement Learning teams lead Anthropic's reinforcement learning research and development, playing a critical role in advancing our AI systems. … Claude Sonnet 4.5 and Opus 4.5. Our work spans several key areas:Developing systems that enable models to use computers effectivelyAdvancing code generation through reinforcement learningPioneering fundamental RL research for large language modelsBuilding scalable RL infrastructure and training methodologiesEnhancing model reasoning capabilitiesWe collaborate closely with Anthropic's alignment ...

Applied AI ML Engineer Director - NLP / LLM and Graphs

Location
Greater London, England, United Kingdom
making. The CDAO is also responsible for developing and implementing solutions that support the firm’s commercial goals by harnessing artificial intelligence and machine learning technologies to develop new products, improve productivity, and enhance risk management effectively and responsibly. As an Applied AI ML Director - NLP/… Graphs within the Chief Data & Analytics Office, Machine Learning Centre of Excellence, you will have the opportunity to apply sophisticated machine learning methods to complex tasks including natural language processing, graph analytics, speech analytics, time series, reinforcement learning and recommendation systems. You will collaborate with various ...

Applied AI ML Engineer Director – NLP / LLM and Graphs

Location
Greater London, England, United Kingdom
making. The CDAO is also responsible for developing and implementing solutions that support the firm's commercial goals by harnessing artificial intelligence and machine learning technologies to develop new products, improve productivity, and enhance risk management effectively and responsibly. As an Applied AI ML Director - NLP/… Graphs within the Chief Data & Analytics Office, Machine Learning Centre of Excellence, you will have the opportunity to apply sophisticated machine learning methods to complex tasks including natural language processing, graph analytics, speech analytics, time series, reinforcement learning and recommendation systems. You will collaborate with various ...

Research Scientist, Robotics, DeepMind

Location
Greater London, England, United Kingdom
designing new architectures, our research scientists work on real-world problems that span the breadth of computer science, such as machine (and deep) learning, data mining, natural language processing, hardware and software performance analysis, improving compilers for mobile platforms, as well as core search and much more. … applications and work with real robots inside and outside the lab to manage real-world use cases. A strong algorithmic background in scalable machine learning (e.g., reinforcement learning/imitation learning; multimodal foundation models) and experience with real robots/robot simulation and training setups ...

Applied AI ML Lead Engineer- (NLP/LLM/Graph)

Location
Greater London, England, United Kingdom
Description NLP/LLM Scientist – Applied AI ML Lead – Machine Learning Centre of Excellence The Machine Learning Center of Excellence invites the successful candidate to apply sophisticated machine learning methods to a wide variety of complex tasks including natural language processing, large language models, and recommendation systems. … environment together with the business, technologists and control partners to deploy solutions into production. The candidate must also have a strong passion for machine learning and invest independent time towards learning, researching and experimenting with new innovations in the field. The candidate must have solid expertise in Deep ...

Research Scientist, Robotics RL, DeepMind

Hiring Organisation
Google
Location
London, United Kingdom
Salary
£ 70 K
participate in a wide variety of other research themes in the context of robotics foundation models and physical agents (e.g. VLAs, WAMs, imitation learning, simulation-based learning, whole body control, dexterity, and more).Develop scalable research pipelines and write robust software to test hypotheses quickly and conduct research … presenting research findings clearly and efficiently both internally and externally.Minimum qualifications:PhD degree in a technical field or equivalent practical experience.2 years of experiencein reinforcement learning or other techniques for mid-/post-training of foundation models and self-improvement in robotics, and other areas such as multimodal ...

Applied AI ML Director - NLP / LLM and Graphs

Location
Greater London, England, United Kingdom
making. The CDAO is also responsible for developing and implementing solutions that support the firm’s commercial goals by harnessing artificial intelligence and machine learning technologies to develop new products, improve productivity, and enhance risk management effectively and responsibly. As an Applied AI ML Director - NLP/… Graphs within the Chief Data & Analytics Office, Machine Learning Centre of Excellence, you will have the opportunity to apply sophisticated machine learning methods to complex tasks including natural language processing, graph analytics, speech analytics, time series, reinforcement learning and recommendation systems. You will collaborate with various ...

Applied AI ML Engineer Director - NLP / LLM and Graphs

Hiring Organisation
JP Morgan Chase
Location
London, United Kingdom
Salary
£ 120 K
making. The CDAO is also responsible for developing and implementing solutions that support the firm’s commercial goals by harnessing artificial intelligence and machine learning technologies to develop new products, improve productivity, and enhance risk management effectively and responsibly.As an Applied AI ML Director - NLP/LLM and Graphs … within the Chief Data & Analytics Office, Machine Learning Centre of Excellence, you will have the opportunity to apply sophisticated machine learning methods to complex tasks including natural language processing, graph analytics, speech analytics, time series, reinforcement learning and recommendation systems. You will collaborate with various teams ...

Reinforcement Learning Engineer - Locomanipulation

Location
Greater London, England, United Kingdom
industrial pilots - and we’re growing the team to take it even further. About The Role We are looking for a Senior or Staff Reinforcement Learning Engineer to develop learning-based control policies for humanoid robots. You will design and train reinforcement learning policies that … will work on include dynamic locomotion, balance recovery, contact-rich manipulation, and multi-behavior policy learning. What You’ll Do Design and train reinforcement learning policies for humanoid robot control. Build scalable simulation and training pipelines (e.g., Isaac Lab, MuJoCo). Design reward functions, observation spaces, and curricula ...

Research Associate in Safe Reinforcement Learning

Hiring Organisation
Imperial College London
Location
London, United Kingdom
Salary
£ 55 K
Imperial College London, led by Dr. Francesco Belardinelli, in a fully funded postdoctoral research role to lead transformative research in formal methods for safe reinforcement learning. Overview. The FMAI lab at Imperial is seeking highly motivated and talented Postdoctoral Research Associates (PDRAs/PostDocs), who have demonstrated competence … primarily provide finite-horizon, statistical, or asymptotic guaranties, and fail to ensure strict safety compliance at runtime. This creates a fundamental gap between scalable learning and certifiable safety. To address this gap, this project aims at developing Certified Reinforcement Learning, a neuro-symbolic framework for learning ...

Principal Machine Learning Engineer, AI & Data Platforms (AiDP)

Location
Greater London, England, United Kingdom
Principal Machine Learning Engineer, AI & Data Platforms (AiDP) London, England, United Kingdom Corporate Functions At Apple, we build AI systems that define experiences for billions of people and we do it with an unwavering commitment to privacy, performance, and craft. The AI & Data Platforms (AiDP) team is seeking … Principlal Machine Learning Engineer to lead the design, fine‐tuning, evaluation, and productionisation of large language models and generative internal AI systems at global scale. This is a deeply hands‐on, high‐impact role: you will work across the full model lifecycle, from reinforcement learning and upstream ...

Reinforcement Learning Engineer - Manipulation

Location
Greater London, England, United Kingdom
running in real industrial pilots — and we’re growing the team to take it even further. About the Role We're hiring a Reinforcement Learning Engineer to join our Autonomy team based in London. In this role you will leverage reinforcement learning in both simulation … physical reality to build highly performant and robust manipulation policies. What You'll Do Train language‐vision conditioned manipulation policies via reinforcement learning (RL) in simulation and in the real world. Construct challenging and diverse suites of manipulation tasks in simulation. Partner with teleoperations to collect trajectories ...

Member of Technical Staff - Research Scientist

Location
Greater London, England, United Kingdom
Member of Technical Staff Research Scientist Push the frontier of reinforcement learning for long-horizon work London, UK Full-time Research Context Reinforcement learning is how language models learn to reason, use tools, and work autonomously over long horizons. But RL for long-horizon agents presents … capabilities of long-horizon agents. Before GR, our team previously worked on some of the leading open-source language model efforts and early reinforcement learning work on language models. About this Role As a Research Scientist, you'll design, implement, and evaluate agentic reinforcement learning methods ...

Machine Learning Engineer - Agentic AI & Reinforcement Learning

Hiring Organisation
CGG
Location
Oxford, Oxfordshire, United Kingdom
Salary
£ 70 K
interested in working on real world data like no other Viridien is expanding its capabilities in agentic AI, large language models and reinforcement learning for complex scientific and enterprise workflows.We are looking for a hands-on Machine Learning Engineer to design, develop and evaluate intelligent agent systems … real world data.About the TeamLearn more about our work and the team on our website, recent blog post featuring one of our senior machine learning engineers and our latest Voices of Viridien.Key responsibilitiesDesign and develop single-agent and multi-agent systems for complex workflows.Build reasoning, planning, task decomposition, memory ...

Machine Learning Engineer - Agentic AI & Reinforcement Learning

Hiring Organisation
CGG
Location
Asthall Leigh, Oxfordshire, UK
Employment Type
Full-time
interested in working on real world data like no other? Viridien is expanding its capabilities in agentic AI, large language models and reinforcement learning for complex scientific and enterprise workflows. We are looking for a hands-on Machine Learning Engineer to design, develop and evaluate intelligent agent … world data. About the TeamLearn more about our work and the team on our website, recent blog post featuring one of our senior machine learning engineers and our latest Voices of Viridien. Key responsibilitiesDesign and develop single-agent and multi-agent systems for complex workflows. Build reasoning, planning, task ...

Research Engineer / Scientist, Post-training - London

Location
Greater London, England, United Kingdom
seeking those who are dedicated as much to building safely and responsibly as to advancing disruptive agentic capabilities. We promote a mindset of openness, learning, and collaboration, where everyone has something to contribute. About the Research & Models Team The Models team builds the foundational models that power our cutting … perceive, understand, and act within complex environments. We own the entire pipeline including synthetic data generation, environment design, mid-training, supervised fine-tuning, offline reinforcement learning, online reinforcement learning, reward modelling, transition modelling, etc. Our team also has dedicated MLOps, Infra and Inference support at scale. ...

Research Scientist, Human Data, Robotics, DeepMind

Hiring Organisation
Google
Location
London, United Kingdom
Salary
£ 70 K
especially human data with or without capture devices into our robotics foundation models.Leverage broader expertise to participate in a wide variety of research, including learning from simulation, reinforcement learning, learning from demonstrations, vision-language-action models, transformers, video generation, robot control, humanoid robots and more.Report … findings clearly and efficiently both internally and externally.Minimum qualifications:PhD in Computer Science, a related field, or equivalent practical experience.2 years of experience with reinforcement learning and imitation learning.Experience with vision, vision-language, video, and other multimodal models.Experience with multimodal generative modeling, training and inference.Preferred qualifications:Experience with ...

Research Scientist

Location
Greater London, England, United Kingdom
Computer Science, Artificial Intelligence, or a related quantitative field, or equivalent practical experience. 4 years of experience in LLM fine-tuning/Reinforcement Learning, autonomous and human-on-loop agent development, and designing evaluation frameworks for real-world applications. Experience providing technical leadership to execute research from concept … NeurIPS, ICML, ICLR), or scientific/medical journals (e.g., Nature). About the job Our special interdisciplinary team combines the best techniques from deep learning and reinforcement learning to build general-purpose learning algorithms and apply these to medicine and healthcare. We have already made ...

Senior Research Scientist FMTA

Location
Greater London, England, United Kingdom
large language models. We focus on developing algorithms and systems that align pre-trained models with tasks and performance goals through techniques like reinforcement learning. As a research-driven team, we stay up to date with current literature to integrate cutting-edge ideas into our core stack. As part … controllability, and safer, more effective user experiences. Your responsibilities As a Senior Research Scientist, you'll design, implement, and deploy cutting-edge research in reinforcement learning and post-training at scale, driving innovations that make it into production. Build and deploy state-of-the-art reinforcement learning ...

Senior Robotic Controls & RL Simulation Engineer

Location
Greater London, England, United Kingdom
There is also the possibility of locating in Cambridge, depending on various factors including team member locations. Role Overview As the Senior Robotic Controls & Reinforcement Learning Simulation Engineer you will develop the "brains" that drive early simulations of robotic systems built with our breakthrough elastic synthetic muscle actuators. … will focus on designing, training, and exploring control concepts using both emerging methods for variable stiffness control and advanced reinforcement learning within the VSim physics environment. Given the highly elastic nature of our synthetic muscles, your deep understanding of compliance, stiffness control, and antagonistic/reciprocal actuation will ...

Machine Learning Engineer - Agentic AI & Reinforcement Learning

Location
Crawley, England, United Kingdom
solutions that efficiently and responsibly resolve complex natural resource, digital, energy transition and infrastructure challenges. We are looking for a hands‐on Machine Learning Engineer to design, develop and evaluate intelligent agent systems, turning emerging AI methods into reliable, scalable and reusable solutions for truly real world data. About … Team Learn more about our work and the team on our website, recent blog post featuring one of our senior machine learning engineers and our latest Voices of Viridien. Key responsibilities Design and develop single-agent and multi-agent systems for complex workflows. Build reasoning, planning, task decomposition, memory ...

Principal Data Scientist

Hiring Organisation
Aristocrat Technologies
Location
London, United Kingdom
Salary
£ 80 K
DoLead high-impact data science initiatives end-to-end, including problem framing, methodology selection, experiment development, implementation partnership, and impact measurement.Build and deliver machine learning and reinforcement learning solutions to improve player engagement, retention, monetization, and operational outcomes.Lead the modeling framework for complex systems, guaranteeing comprehensive evaluation … monitoring of causal inference, uplift modeling, sequential decisioning, bandits/reinforcement learning, and forecasting.Partner with game teams to define success metrics, guardrails, and decision frameworks, translating analytical results into actionable product and operational actions.Define and uphold engineering standards and guidelines for model development, including validation, uncertainty, reproducibility ...

Principal Data Scientist

Hiring Organisation
Aristocrat Technologies
Location
London, UK
Employment Type
Full-time
high-impact data science initiatives end-to-end, including problem framing, methodology selection, experiment development, implementation partnership, and impact measurement. Build and deliver machine learning and reinforcement learning solutions to improve player engagement, retention, monetization, and operational outcomes. Lead the modeling framework for complex systems, guaranteeing comprehensive … evaluation and monitoring of causal inference, uplift modeling, sequential decisioning, bandits/reinforcement learning, and forecasting. Partner with game teams to define success metrics, guardrails, and decision frameworks, translating analytical results into actionable product and operational actions. Define and uphold engineering standards and guidelines for model development, including ...

Data Scientist - Principal

Location
Greater London, England, United Kingdom
high-impact data science initiatives end-to-end, including problem framing, methodology selection, experiment development, implementation partnership, and impact measurement.* Build and deliver machine learning and reinforcement learning solutions to improve player engagement, retention, monetization, and operational outcomes.* Lead the modeling framework for complex systems, guaranteeing comprehensive … evaluation and monitoring of causal inference, uplift modeling, sequential decisioning, bandits/reinforcement learning, and forecasting.* Partner with game teams to define success metrics, guardrails, and decision frameworks, translating analytical results into actionable product and operational actions.* Define and uphold engineering standards and guidelines for model development, including ...

Research Scientist

Hiring Organisation
Google
Location
London, United Kingdom
Salary
£ 70 K
qualifications:PhD in Computer Science, Artificial Intelligence, or a related quantitative field, or equivalent practical experience.4 years of experience in LLM fine-tuning/Reinforcement Learning, autonomous and human-on-loop agent development, and designing evaluation frameworks for real-world applications.Experience providing technical leadership to execute research from … venues (e.g., NeurIPS, ICML, ICLR), or scientific/medical journals (e.g., Nature).Our special interdisciplinary team combines the best techniques from deep learning and reinforcement learning to build general-purpose learning algorithms and apply these to medicine and healthcare. We have already made a number ...