18 of 18 Remote/Hybrid Reinforcement Learning Jobs in London

Research Engineer / Scientist, Post-training - London

Location
Greater London, England, United Kingdom
seeking those who are dedicated as much to building safely and responsibly as to advancing disruptive agentic capabilities. We promote a mindset of openness, learning, and collaboration, where everyone has something to contribute. About the Research & Models Team The Models team builds the foundational models that power our cutting … perceive, understand, and act within complex environments. We own the entire pipeline including synthetic data generation, environment design, mid-training, supervised fine-tuning, offline reinforcement learning, online reinforcement learning, reward modelling, transition modelling, etc. Our team also has dedicated MLOps, Infra and Inference support at scale. ...

Senior Research Scientist FMTA

Location
Greater London, England, United Kingdom
large language models. We focus on developing algorithms and systems that align pre-trained models with tasks and performance goals through techniques like reinforcement learning. As a research-driven team, we stay up to date with current literature to integrate cutting-edge ideas into our core stack. As part … controllability, and safer, more effective user experiences. Your responsibilities As a Senior Research Scientist, you’ll design, implement, and deploy cutting-edge research in reinforcement learning and post-training at scale, driving innovations that make it into production. You will: Build and deploy state-of-the-art reinforcement ...

Senior Machine Learning Engineer - Messaging Platform

Location
Greater London, England, United Kingdom
moment. We’re evolving how messaging works at Spotify — moving from short-term optimization toward systems that understand long-term user journeys. By combining reinforcement learning approaches with deeper domain signals, we’re expanding how machine learning shapes the entire messaging funnel. What You’ll Do Design … build, and ship machine learning models that optimize messaging across push, email, and in-app channels Plan and run A/B experiments in a multi-objective environment, balancing conversion, engagement, retention, and reachability Contribute to reinforcement learning systems that optimize for long-term user outcomes rather ...

Senior AI Platform Engineer

Hiring Organisation
Capgemini
Location
Greater London, United Kingdom
Employment Type
Full Time
small senior team with heavy AI leverage. The ambition runs past serving frontier models: we close the loop from production feedback through reinforcement learning and fine-tuning, and train our own LLMs and SLMs where evaluations and economics justify it. What you will own One or more platform … guardrail engine, tool and agent registry, or shared product services The training and adaptation loop: pipelines that turn production traces and evaluation verdicts into reinforcement learning and fine-tuning datasets, and the infrastructure to train, evaluate, and serve NewCo-tuned LLMs and SLMs behind the same gates ...

Senior Research Scientist | Model Steering

Location
Greater London, England, United Kingdom
deliver perfect translations for the most demanding use cases. To that end, we take responsibility for the entire life cycle of the machine learning models that power our language AI products. This includes data, training, quality assurance, and operational aspects. In our highly collaborative teams, each person … drive impact across the company. Your responsibilities We are looking for a Senior Research Scientist to lead fine-tuning, post-training, model-steerability, and reinforcement learning for the next generation of DeepL's LLM-based translation models. This is a high-impact, hands‐on role for a researcher ...

Senior Operations Research Scientist

Location
Greater London, England, United Kingdom
ground operations. Designing and implementing high-performance optimisation algorithms and production-grade solutions using Python, C++ and commercial optimisation solvers. Integrating optimisation, simulation and reinforcement learning techniques to solve complex multi-stage decision problems. Partnering with senior stakeholders to translate business challenges into mathematical models with clear operational … teams. Significant relevant industry experience within operations research and optimisation, with evidence of technical leadership progression. Cloud-based optimisation platforms and distributed computing. Machine learning and hybrid optimisation approaches. Simulation, reinforcement learning and advanced heuristic methods. Commercial airline operations, crew planning systems or aircraft scheduling. Aviation regulations ...

Software Engineer, RL Data

Location
Greater London, England, United Kingdom
down to reading transcripts, supporting users, and wrangling vendors. The company's RL Data team builds the systems that produce high‐quality reinforcement learning data for Claude: data collection pipelines, human feedback tooling, the execution environments RL tasks run in, and the quality assurance that keeps training data … Effective use of AI tools in your own day‐to‐day work. Care about the societal impacts of your work. Preferred qualifications Experience with reinforcement learning on LLMs, particularly on the data side: creating evals, environments, rewards, graders, or training data. Experience helping organizations use AI more effectively ...

AI Harness Engineers

Hiring Organisation
Capgemini
Location
Greater London, United Kingdom
Employment Type
Full Time
deploying or operating AI coding agents (Claude Code, Cursor, Copilot, or in-house), beyond personal use You follow frontier agentic-systems research (harness design, reinforcement learning from execution feedback, evaluation methods) closely enough to put it into production within the quarter it lands A measurement habit … unattended agent pipelines with hard failure caps and human escalation You have turned agent execution traces into training or evaluation datasets, or built reinforcement learning pipelines from execution feedback You have written publicly about developer experience or agent harnesses Financial services engineering exposure (banks, insurers, or payment providers ...

Manager, Lead Research Scientist, LLM Agents (Foundational Research)

Location
Greater London, England, United Kingdom
curious and open-minded individual with an interest in conducting state-of-the-art foundational machine learning research? Thomson Reuters Labs is seeking Research Scientists with a passion for building complex agent-based AI systems in a data-rich, complex academic environment driven by real-world problems. Foundational Research … sleeves and participate in designing, coding, conducting experiments, and translating findings into concrete deliverables. Our focus areas are: LLM Training (Continued Pretraining, Instruction Tuning, Reinforcement Learning Alignment, Distributed Training, Efficient ML techniques) Post-training techniques for planning, reasoning & complex workflows (e.g., Reasoning Models, LLMs + Knowledge Graphs, Test ...

Manager, Research Engineering (Foundational Research)

Location
Greater London, England, United Kingdom
Manager, Research Engineering with the skills and drive to manage a top performing Engineering team focused on developing the latest in LLM and Machine Learning capabilities. Foundational Research Foundational Research is the dedicated core Machine Learning research division of Thomson Reuters. We are focused on research and development … with a particular focus on advanced algorithms and training techniques for Large Language Models (LLMs). LLM Training (Continued pretraining, instruction tuning, reinforcement learning, distributed training, efficient ML techniques) Post-training techniques for planning, reasoning & complex workflows (e.g., reasoning models, LLMs + knowledge graphs, test time compute ...

Senior Machine Learning Scientist (London)

Location
Greater London, England, United Kingdom
values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands. … move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting‐edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test and measure ...

Data Scientist Senior Manager

Location
Greater London, England, United Kingdom
ensure best practices in modelling, experimentation and deployment. Data Science & Advanced Analytics – Design and implement predictive and prescriptive models for airline operations, applying machine learning, optimisation and statistical analysis to large, complex data sets. Translation of Business Problems – Convert business challenges into analytical frameworks and actionable solutions to improve … driving high‐impact projects and aligning with business objectives. Advanced knowledge of the Data Science Toolbox – mathematics, statistics, programming, data ingestion, munging, visualisation, machine learning, optimisation, simulation, reinforcement learning, and big‐data technologies. Experience leading successful data‐science projects in an operational setting (preferred). Understanding ...

Senior Data Scientist & Operations Analytics Lead

Location
Greater London, England, United Kingdom
Ensure best practices in modelling, experimentation, and deployment. · Data Science & Advanced Analytics. Design and implement predictive and prescriptive models for airline operations. Apply machine learning, optimization techniques, and statistical analysis to large, complex datasets · Translate business problems into analytical frameworks and actionable solutions to improve decision making · Stakeholder engagement. … advanced knowledge level of the Data Science Toolbox (i.e. the fundamentals of Mathematics and Statistics, computer programming, Data Ingestion, Data Munging, Data visualisation, Machine Learning, Optimisation, Simulation, Reinforcement Learning and Big Data techniques and technologies) · Have experience of leading and landing successful Data Science projects ...

Software Engineer (Applied AI)

Location
Greater London, England, United Kingdom
iteration of our next-generation benefits platform features that leverage personalization, experimentation, and AI/ML methods (e.g. agents/LLMs, recommender systems, reinforcement learning) to enhance user experience in a meaningful business domain. Contribute across the tech stack : You’ll work in React (JavaScript/TypeScript … against important business goals that help the entire team win Pragmatic Best Practices : An overarching desire to build efficient, scalable, and maintainable code, while learning the tradeoffs between technical debt and delivery speed What we look for We’re a great bunch but we have some "Euph" cultural ...

Senior Robotics Software Engineer

Hiring Organisation
Your Tech Future
Location
South West London, London, United Kingdom
Employment Type
Permanent
Salary
£90,000
Mapping) Navigation and autonomy systems Sensor fusion and state estimation Gazebo, Isaac Sim, MATLAB/Simulink or similar simulation platforms Real-time systems development Reinforcement learning Pybind11 What We're Looking For Demonstrable experience delivering robotics projects Strong software engineering principles and coding standards Ability to take ownership ...

Applied Scientist (Integrity)

Location
Greater London, England, United Kingdom
What We Are Looking For Deep expertise in one applied ML, statistics or data science speciality. Examples could be risk modelling, LLM detection, or reinforcement learning – any area of genuine depth counts. Judgement about when to reach for simple statistics, classical ML, LLMs, or agentic approaches. 3+ years ...

Machine Learning Engineer (Closed Loop)

Location
Greater London, England, United Kingdom
embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future. At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster … that is faithful enough, and fast enough, to train and evaluate driving models in closed loop — not only to replay logs. As a Machine Learning Engineer in the Simulation team, you’ll play a key role in developing next-generation world models and planners that can simulate complex, diverse ...

Senior Research Scientist, RL & Post-Training (Hybrid)

Location
Greater London, England, United Kingdom
DeepL is seeking a Senior Research Scientist to design, implement, and deploy cutting-edge reinforcement learning and post-training research at scale. You will drive innovations that translate to production, collaborating with engineering, ML platforms, and HPC teams. The role emphasizes building scalable RL pipelines, aligning models with ...