1 to 25 of 46 Remote/Hybrid Reinforcement Learning Jobs in London

Machine Learning Research Engineer (Foundational Research)

Hiring Organisation
Thomson Reuters
Location
London, UK
Employment Type
Full-time
Join a cutting-edge research team working to deliver on the transformation promises of modern AI. We are seeking Machine Learning Research Engineers with the skills and drive to build and conduct experiments with advanced AI systems in an academic environment rich with high-quality data from real-world … problems. Foundational Research is the dedicated core Machine Learning research division of Thomson Reuters. We are focused on research and development, with a particular focus on advanced algorithms and training techniques for Large Language Models (LLMs). We are expanding our strong foundation of research capabilities across different areas ...

Research Engineer / Scientist, Post-training - London

Location
Greater London, England, United Kingdom
seeking those who are dedicated as much to building safely and responsibly as to advancing disruptive agentic capabilities. We promote a mindset of openness, learning, and collaboration, where everyone has something to contribute. About the Research & Models Team The Models team builds the foundational models that power our cutting … perceive, understand, and act within complex environments. We own the entire pipeline including synthetic data generation, environment design, mid-training, supervised fine-tuning, offline reinforcement learning, online reinforcement learning, reward modelling, transition modelling, etc. Our team also has dedicated MLOps, Infra and Inference support at scale. ...

Senior Research Scientist FMTA

Hiring Organisation
DeepL
Location
London, United Kingdom
Salary
£ 80 K
large language models. We focus on developing algorithms and systems that align pre-trained models with tasks and performance goals through techniques like reinforcement learning. As a research-driven team, we stay up to date with current literature to integrate cutting-edge ideas into our core stack. As part … capabilities, better controllability, and safer, more effective user experiences.Your responsibilitiesAs a Senior Research Scientist, you’ll design, implement, and deploy cutting-edge research in reinforcement learning and post-training at scale, driving innovations that make it into production.You will:Build and deploy state-of-the-art reinforcement ...

Senior Research Scientist FMTA

Location
Greater London, England, United Kingdom
large language models. We focus on developing algorithms and systems that align pre-trained models with tasks and performance goals through techniques like reinforcement learning. As a research-driven team, we stay up to date with current literature to integrate cutting-edge ideas into our core stack. As part … controllability, and safer, more effective user experiences. Your responsibilities As a Senior Research Scientist, you’ll design, implement, and deploy cutting-edge research in reinforcement learning and post-training at scale, driving innovations that make it into production. You will: Build and deploy state-of-the-art reinforcement ...

Senior Machine Learning Engineer - Messaging Platform

Location
Greater London, England, United Kingdom
moment. We’re evolving how messaging works at Spotify — moving from short-term optimization toward systems that understand long-term user journeys. By combining reinforcement learning approaches with deeper domain signals, we’re expanding how machine learning shapes the entire messaging funnel. What You’ll Do Design … build, and ship machine learning models that optimize messaging across push, email, and in-app channels Plan and run A/B experiments in a multi-objective environment, balancing conversion, engagement, retention, and reachability Contribute to reinforcement learning systems that optimize for long-term user outcomes rather ...

Senior Research Scientist | Model Steering

Location
Greater London, England, United Kingdom
deliver perfect translations for the most demanding use cases. To that end, we take responsibility for the entire life cycle of the machine learning models that power our language AI products. This includes data, training, quality assurance, and operational aspects. In our highly collaborative teams, each person … drive impact across the company. Your responsibilities We are looking for a Senior Research Scientist to lead fine-tuning, post-training, model-steerability, and reinforcement learning for the next generation of DeepL's LLM-based translation models. This is a high-impact, hands-on role for a researcher ...

Research Scientist, Wayve Labs

Hiring Organisation
wayve
Location
London, United Kingdom
Salary
£ 70 K
embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future.At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster an inclusive … Wayve Labs and help build the next generation of AI systems for autonomous driving and robotics. You’ll work at the intersection of machine learning, simulation, robotics, and real-world deployment, contributing to core innovations that push the boundaries of embodied AI.Situated within Wayve, we are a high-conviction ...

Senior Operations Research Scientist

Hiring Organisation
Easyjet
Location
London, United Kingdom
Salary
£ 80 K
ground operations.> Designing and implementing high-performance optimisation algorithms and production-grade solutions using Python, C++ and commercial optimisation solvers.> Integrating optimisation, simulation and reinforcement learning techniques to solve complex multi-stage decision problems.> Partnering with senior stakeholders to translate business challenges into mathematical models with clear operational … relevant industry experience within operations research and optimisation, with evidence of technical leadership progression.Desirable experience includes:> Cloud-based optimisation platforms and distributed computing.> Machine learning and hybrid optimisation approaches.> Simulation, reinforcement learning and advanced heuristic methods.> Commercial airline operations, crew planning systems or aircraft scheduling.> Aviation regulations ...

Senior Operations Research Scientist

Location
Greater London, England, United Kingdom
ground operations. Designing and implementing high-performance optimisation algorithms and production-grade solutions using Python, C++ and commercial optimisation solvers. Integrating optimisation, simulation and reinforcement learning techniques to solve complex multi-stage decision problems. Partnering with senior stakeholders to translate business challenges into mathematical models with clear operational … teams. Significant relevant industry experience within operations research and optimisation, with evidence of technical leadership progression. Cloud-based optimisation platforms and distributed computing. Machine learning and hybrid optimisation approaches. Simulation, reinforcement learning and advanced heuristic methods. Commercial airline operations, crew planning systems or aircraft scheduling. Aviation regulations ...

Machine Learning Engineer (Remote | $60–$120/hr)

Hiring Organisation
Synthires
Location
City of London, London, United Kingdom
codebase refactoring, and performance optimization to contribute to advanced AI research and evaluation projects . You'll apply your software engineering expertise to create Reinforcement Learning (RL) environments that test AI models on complex software engineering workflows. Tasks may involve debugging, feature implementation, code refactoring, performance optimization … working with Model Context Protocol (MCP) tools. No prior AI experience is required — your technical expertise is what matters. Responsibilities Design and develop Reinforcement Learning environments that evaluate AI models on realistic software engineering workflows. Create tasks involving bug fixing, feature implementation, codebase refactoring, and performance optimization . ...

Machine Learning Engineer (Remote | $60–$120/hr)

Hiring Organisation
Synthires
Location
East London, London, United Kingdom
codebase refactoring, and performance optimization to contribute to advanced AI research and evaluation projects . You'll apply your software engineering expertise to create Reinforcement Learning (RL) environments that test AI models on complex software engineering workflows. Tasks may involve debugging, feature implementation, code refactoring, performance optimization … working with Model Context Protocol (MCP) tools. No prior AI experience is required — your technical expertise is what matters. Responsibilities Design and develop Reinforcement Learning environments that evaluate AI models on realistic software engineering workflows. Create tasks involving bug fixing, feature implementation, codebase refactoring, and performance optimization . ...

AI Engineer (Remote | $60–$120/hr)

Hiring Organisation
Synthires
Location
London Area, United Kingdom
codebase refactoring, and performance optimization to contribute to advanced AI research and evaluation projects . You'll apply your software engineering expertise to create Reinforcement Learning (RL) environments that test AI models on complex software engineering workflows. Tasks may involve debugging, feature implementation, code refactoring, performance optimization … working with Model Context Protocol (MCP) tools. No prior AI experience is required — your technical expertise is what matters. Responsibilities Design and develop Reinforcement Learning environments that evaluate AI models on realistic software engineering workflows. Create tasks involving bug fixing, feature implementation, codebase refactoring, and performance optimization . ...

Software Engineer, RL Data

Location
Greater London, England, United Kingdom
down to reading transcripts, supporting users, and wrangling vendors. The company's RL Data team builds the systems that produce high‐quality reinforcement learning data for Claude: data collection pipelines, human feedback tooling, the execution environments RL tasks run in, and the quality assurance that keeps training data … Effective use of AI tools in your own day‐to‐day work. Care about the societal impacts of your work. Preferred qualifications Experience with reinforcement learning on LLMs, particularly on the data side: creating evals, environments, rewards, graders, or training data. Experience helping organizations use AI more effectively ...

Manager, Lead Research Scientist, Training Data (Foundational Research)

Hiring Organisation
Thomson Reuters
Location
London, United Kingdom
Salary
£ 120 K
curious and open-minded individual with an interest in conducting state-of-theart foundational machine learning research? Thomson Reuters Labs is seeking Research Scientists with a passion for building complex AI systems in a data-rich, complex academic environment driven by real-world problems.Foundational Research is the dedicated core … Machine Learning research division of Thomson Reuters. We are focused on research and development, with a particular focus on advanced algorithms and training techniques for Large Language Models (LLMs). We are building a strong foundation of research capabilities across different areas and are looking for managers ...

Manager, Lead Research Scientist, Training Data (Foundational Research)

Location
Greater London, England, United Kingdom
timeposted on: Posted Todayjob requisition id: JREQ193528Are you a curious and open-minded individual with an interest in conducting state-of-theart foundational machine learning research? Thomson Reuters Labs is seeking Research Scientists with a passion for building complex AI systems in a data-rich, complex academic environment driven … real-world problems.**Foundational Research** is the dedicated core Machine Learning research division of Thomson Reuters. We are focused on research and development, with a particular focus on advanced algorithms and training techniques for Large Language Models (LLMs). We are building a strong foundation of research capabilities across ...

Research Scientist, LLM Agents (Foundational Research)

Location
Greater London, England, United Kingdom
curious and open-minded individual with an interest in conducting state-of-the-art foundational machine learning research? Thomson Reuters Labs is seeking Research Scientists with a passion for building complex agent-based AI systems in a data-rich, complex academic environment driven by real-world problems. Foundational Research … dedicated core Machine Learning research division of Thomson Reuters. We are focused on research and development, with a particular focus on advanced algorithms and training techniques for Large Language Models (LLMs). We are building a strong foundation of research capabilities across different areas and are looking for scientists ...

Research Scientist, LLM Agents (Foundational Research)

Hiring Organisation
Thomson Reuters
Location
London, United Kingdom
Salary
£ 60 K
curious and open-minded individual with an interest in conducting state-of-theart foundational machine learning research? Thomson Reuters Labs is seeking Research Scientists with a passion for building complex agent-based AI systems in a data-rich, complex academic environment driven by real-world problems.Foundational Research … dedicated core Machine Learning research division of Thomson Reuters. We are focused on research and development, with a particular focus on advanced algorithms and training techniques for Large Language Models (LLMs). We are building a strong foundation of research capabilities across different areas and are looking for scientists ...

Manager, Lead Research Scientist, LLM Agents (Foundational Research)

Location
Greater London, England, United Kingdom
curious and open-minded individual with an interest in conducting state-of-the-art foundational machine learning research? Thomson Reuters Labs is seeking Research Scientists with a passion for building complex agent-based AI systems in a data-rich, complex academic environment driven by real-world problems. Foundational Research … sleeves and participate in designing, coding, conducting experiments, and translating findings into concrete deliverables. Our focus areas are: LLM Training (Continued Pretraining, Instruction Tuning, Reinforcement Learning Alignment, Distributed Training, Efficient ML techniques) Post-training techniques for planning, reasoning & complex workflows (e.g., Reasoning Models, LLMs + Knowledge Graphs, Test ...

Senior Consultant - Manager (ER&I), AI Engineer, AI Scaling and Transformation, Engineering, AI & Data, Technology & Transformation

Hiring Organisation
Deloitte
Location
London, United Kingdom
Salary
£ 80 K
design and delivery of data and AI solutions that address their needs.Take ownership of the development and implementation of AI and machine learning models, data pipelines, and analytics solutions, ensuring they are robust, scalable, and aligned with best practices.Support the deployment and operationalisation of AI systems, focusing on performance … analysts, to ensure successful project delivery and integration of AI solutions.Contribute to the mentoring and development of junior team members, promoting a culture of learning, innovation, and delivery excellence.Connect to your skills and professional experience We are looking for candidates who are able to demonstrate skills and experience ...

Manager, Research Engineering (Foundational Research)

Location
Greater London, England, United Kingdom
Manager, Research Engineering with the skills and drive to manage a top performing Engineering team focused on developing the latest in LLM and Machine Learning capabilities. Foundational Research Foundational Research is the dedicated core Machine Learning research division of Thomson Reuters. We are focused on research and development … with a particular focus on advanced algorithms and training techniques for Large Language Models (LLMs). LLM Training (Continued pretraining, instruction tuning, reinforcement learning, distributed training, efficient ML techniques) Post-training techniques for planning, reasoning & complex workflows (e.g., reasoning models, LLMs + knowledge graphs, test time compute ...

Research Scientist (Applied LLMs), London

Hiring Organisation
Isomorphic Labs
Location
London, United Kingdom
Salary
£ 60 K
advance human health by building on and beyond the Nobel-winning AlphaFold system. Since then, our interdisciplinary team of drug discovery experts and machine learning specialists has built powerful new predictive and generative AI models that accelerate scientific discovery at digital speed.Our name comes from the belief that there … capabilities and predictive power which will be critical to the organisation’s success. You will draw upon your existing deep understanding and experience whilst learning from those around you, to apply novel techniques and ideas to newly encountered computational biology and chemistry problems.Depending on your experience:You will both ...

Staff Machine Learning Scientist

Location
Greater London, England, United Kingdom
values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands. … move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting-edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test and measure ...

Senior Machine Learning Scientist

Hiring Organisation
Intercom
Location
London, United Kingdom
Salary
£ 80 K
core values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers.What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands.We … enable us to move to production fast, often shipping to beta in weeks after a successful offline test.We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting-edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test ...

Senior Machine Learning Scientist (London)

Location
Greater London, England, United Kingdom
values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands. … move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting‐edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test and measure ...

Staff Machine Learning Scientist

Hiring Organisation
Intercom
Location
London, UK
Employment Type
Full-time
values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers' hands. … move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting-edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test and measure ...