151 to 175 of 192 Reinforcement Learning Jobs in London

Senior Software Engineer, Data Platform London, UK Apply →

Location
Greater London, England, United Kingdom
enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human evaluation and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT ...

Senior Software Engineer, Data Platform

Hiring Organisation
Scale AI
Location
London, United Kingdom
Salary
£ 80 K
enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human evaluation and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT ...

Senior RL Engineer for Humanoid Robots (Equity)

Location
Greater London, England, United Kingdom
Humanoid is seeking a Senior or Staff Reinforcement Learning Engineer to develop learning-based control policies for humanoid robots. You will design and train RL policies enabling dynamic locomotion and loco-manipulation on real robots, building scalable pipelines and improving sim-to-real transfer for reliable hardware ...

Platform Engineer

Location
Greater London, England, United Kingdom
superlearning capability - the ability to endlessly discover knowledge and skills, without relying on human data - will be driven by the world’s most powerful reinforcement learning algorithms. The superlearner is expected to rediscover and then transcend the greatest inventions in human history, such as language, science, mathematics … role where you inherit someone else's architecture, you'll be one of the people defining it. We're pushing these learning methods to a scale that hasn’t been tried before and we care deeply about the infrastructure that gets us there. You'll own the infrastructure that ...

Agent Product Manager

Location
Greater London, England, United Kingdom
usage and model performance, identify key issues, validate hypotheses, and continuously improve product experience and Agent capabilities. Requirements - Basic understanding of LLMs, AI Agents, reinforcement learning, and model training or post-training workflows. - Strong interest in general-purpose Agent products and enthusiasm for building product-data flywheels. - Experience … evaluation, RLHF/RLAIF, data annotation platforms, or model training data pipelines is a strong plus. Basic understanding of LLMs, AI Agents, and reinforcement learning.Strong product judgment, with the ability to identify user needs, define problems, and design practical product solutions.Strong structured thinking and analytical skills, with the ability ...

Simulation Engineer

Hiring Organisation
Helsing
Location
London, United Kingdom
Salary
£ 70 K
make a meaningful impact in the worldOur work frequently takes us right up to the state of the art in technical innovation, be it reinforcement learning, distributed systems, generative AI, or deployment infrastructure. The defence industry is entering the most exciting phase of the technological development curve. Advances ...

Video Generation Content Understanding and Feedback Researcher

Hiring Organisation
Tencent
Location
London, United Kingdom
Salary
£ 70 K
track industry-related work and formulate proprietary technical roadmaps.Who We Look For1.Ph.D. in AI-related fields, covering areas such as video understanding, video prediction, reinforcement learning, or multimodal generation and understanding.2.Deep understanding of the internal mechanisms of both video generation and VLM models; familiar with diffusion model principles.3.Research … generation.4.Ability to independently design complex model training pipelines with strong engineering skills.5.Proficient in Python/PyTorch, with publications in video understanding, controllable generation, or reinforcement learning.6.Experience with Game AI, simulation environments, or closed-loop systems is preferred.7.Relevant publications in top-tier conferences or journals are preferred.Equal Employment Opportunity ...

Senior Data Scientist in ADVANCED ANALYTICS

Location
City Of London, England, United Kingdom
opportunity to work at the heart of data science, research and public policy. Advanced Analytics is hiring a Senior Data Scientist to develop machine learning (ML) and Artificial Intelligence (AI) solutions, with a focus on robust prototyping. We have pioneered the application of new analytical approaches in the Bank … England, such as generative AI, machine learning, natural language processing, agent-based models, network analytics, and randomised control trials. We’re looking for someone with ML/AI expertise to help us continue pushing the organisation’s analytical frontier to inform policymaking. DAT is leading the transformation ...

Vice President, Model Research & Development

Hiring Organisation
Thomson Reuters
Location
London, United Kingdom
Salary
£ 80 K
release decision) for models that ship to real users across CoCounsel, Westlaw, and Practical Law. Set the target for an online, agentic reinforcement-learning pipeline, training with subject-matter experts in the loop, and hold the work accountable to it. Ensure data strategy and evaluation approach are rigorous … weeks per year, empowering employees to achieve a better work-life balance.Career Development and Growth: By fostering a culture of continuous learning and skill development, we prepare our talent to tackle tomorrow’s challenges and deliver real-world solutions. Our Grow My Way programming and skills-first approach ensures ...

Onsite RL Engineer – Robotics & Embodied AI

Location
Greater London, England, United Kingdom
Randstad Technologies Recruitment is searching for a Reinforcement Learning (RL) Engineer to join a leading robotics company in London. This permanent position involves working collaboratively to define the company's 2026 technical roadmap while solving complex challenges in autonomous systems. The role requires expertise in AI, MLOps, Software ...

Senior RL Data Engineer: End-to-End Pipelines & QA

Location
Greater London, England, United Kingdom
Senior Engineer in Greater London to develop and manage data pipelines for AI systems. This role involves significant responsibilities, including ensuring the quality of reinforcement learning data, collaborating with various teams, and innovating operational frameworks. The ideal candidate will possess strong software engineering skills, have a background ...

(Senior) Product Marketing Manager Marketing & Communications Munich - Berlin

Location
Greater London, England, United Kingdom
safety, and ethical considerations are vital. Our work frequently takes us right up to the state of the art in technical innovation, be it reinforcement learning, distributed systems, generative AI, or deployment infrastructure. The defence industry is entering the most exciting phase of the technological development curve. Advances ...

Full Stack Engineer

Location
Greater London, England, United Kingdom
simulation and real hardware environments. You will be part of a focused team responsible for the application level software that connects control, navigation, perception, learning, and platform systems. Your work will ensure that these components operate as a coherent and reliable system that users can interact with seamlessly. This … will develop and maintain application level software for humanoid robots. You will integrate software components from controls, navigation, computer vision, reinforcement learning, and platform teams. You will contribute to the structure and evolution of the application architecture and its interfaces. You will work closely with the product ...

Research Scientist PhD Intern, 2027

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
your penultimate/final year of education. Experience in one or more research areas; natural language understanding/processing, computer vision, LLMs, machine learning, algorithmic foundations of optimization, data mining, machine intelligence (Artificial Intelligence), AI/ML algorithms, deep learning or image processing, multimodal, multilingual and reinforcement … extend and improve on Google's product offering. Contribute to a wide variety of projects utilising natural language processing, artificial intelligence, data compression, machine learning, and search technologies. Please complete your application before 23rd October 2026. We encourage you to apply as early as possible as we review applications ...

Research Scientist, RL for Long-Horizon Agents

Location
Greater London, England, United Kingdom
leading AI research firm in London is seeking a Research Scientist to push the frontier of reinforcement learning. The role involves designing, implementing, and evaluating novel RL methods for language models and requires experience in reinforcement learning and proficiency in Python. You will collaborate closely with engineering ...

Technical Recruiter

Hiring Organisation
Helsing
Location
London, United Kingdom
Salary
£ 70 K
methods beyond LinkedIn Recruiter across multiple countriesCommunicate clearly and openly with candidates and hiring stakeholders alikeStay curious about new technologies and enjoy learning enough to speak credibly with software engineers about their workManage multiple sourcing pipelines simultaneously while maintaining attention to detail, adapting strategies based on role, market … make a meaningful impact in the worldOur work frequently takes us right up to the state of the art in technical innovation, be it reinforcement learning, distributed systems, generative AI, or deployment infrastructure. The defence industry is entering the most exciting phase of the technological development curve. Advances ...

Hardware Engineer - Physical Products Hardware Engineering London; Oxford

Location
Greater London, England, United Kingdom
meaningful impact in the world. Our work frequently takes us right up to the state of the art in technical innovation, be it reinforcement learning, distributed systems, generative AI, or deployment infrastructure. The defence industry is entering the most exciting phase of the technological development curve. Advances … that you care deeply about. What we offer Competitive salary and VSOP options Relocation support: up to €2,500 and 4 weeks temporary accommodation Learning: €500/£450 yearly allowance Health & wellness: gym membership and mental health support (Nilo.health) Social: regular company events and monthly social allowances Enhanced parental ...

Research Engineer: Generative AI & ML Systems (Hybrid)

Location
Greater London, England, United Kingdom
seeking a Research Engineer to advance AI research and ML systems. The role blends hands-on engineering with rigorous experimentation on large models, reinforcement learning, synthetic data, and multi-agent setups. You will work within a world-leading technology organization, contributing to cutting-edge infrastructure and research initiatives ...

Senior Research Scientist - Multimodal Vision Translation

Location
Greater London, England, United Kingdom
DeepL is seeking a Senior Research Scientist to lead fine-tuning, post-training, and reinforcement learning for its document translation multimodal and vision models. You will develop models that reason about document layout, fuse data sources, and drive breakthroughs from prototype to deployment. The role requires hands ...

Senior PM: Commerce Personalization & ML-Powered Experience

Location
Greater London, England, United Kingdom
driven initiatives with cross-functional partners to improve customer outcomes at scale. You will partner with Applied Science and ML Engineering to develop reinforcement learning and optimization techniques, ensuring #J-18808-Ljbffr ...

Member of Technical Staff - Sandbox Service

Location
Greater London, England, United Kingdom
tuning filesystem performance on nodes. The Sandbox service enables Grok to safely run and test code in real‐time for user queries, and supports reinforcement learning in training, where models interactively explore tools ranging from compilers to productivity apps. About the role Expect deeply technical work ...

Technical Lead (Energy Systems)

Hiring Organisation
Gryd Energy
Location
London, England, United Kingdom
live deployment Strong documentation and code quality practices Helpful but not essential: Experience with constraint-based or multi-objective optimisation (linear programming, MPC, reinforcement learning) Familiarity with residential battery or inverter platforms and their cloud APIs Understanding of UK electricity markets, grid flexibility, or DNO constraints Startup ...

Field Engineer Deployed Engineering London

Location
Greater London, England, United Kingdom
meaningful impact in the world. Our work frequently takes us right up to the state of the art in technical innovation, be it reinforcement learning, distributed systems, generative AI, or deployment infrastructure. The defence industry is entering the most exciting phase of the technological development curve. Advances ...

Field Engineer

Hiring Organisation
Helsing
Location
London, United Kingdom
Salary
£ 70 K
meaningful impact in the world. Our work frequently takes us right up to the state of the art in technical innovation, be it reinforcement learning, distributed systems, generative AI, or deployment infrastructure. The defence industry is entering the most exciting phase of the technological development curve. Advances ...

RL Environment Architect & Data Research Engineer

Location
Greater London, England, United Kingdom
Eigent AI is seeking an RL Environment Data Engineer/Researcher to design and refine reinforcement learning training environments. This role emphasizes data collection, task definition, and the implementation of anti-reward-hacking mechanisms. The ideal candidate will have strong Python coding skills and a solid understanding … reinforcement learning. Responsibilities include collaborating with multiple teams to improve RL tasks and validation environments. #J-18808-Ljbffr ...