13 of 13 Remote/Hybrid Reinforcement Learning Jobs in London

Research Engineer / Scientist, Post-training - London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
seeking those who are dedicated as much to building safely and responsibly as to advancing disruptive agentic capabilities. We promote a mindset of openness, learning, and collaboration, where everyone has something to contribute. About the Research & Models Team The Models team builds the foundational models that power our cutting … perceive, understand, and act within complex environments. We own the entire pipeline including synthetic data generation, environment design, mid-training, supervised fine-tuning, offline reinforcement learning, online reinforcement learning, reward modelling, transition modelling, etc. Our team also has dedicated MLOps, Infra and Inference support at scale. ...

Senior Research Scientist | Model Steering

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
deliver perfect translations for the most demanding use cases. To that end, we take responsibility for the entire life cycle of the machine learning models that power our language AI products. This includes data, training, quality assurance, and operational aspects. In our highly collaborative teams, each person … drive impact across the company. Your responsibilities We are looking for a Senior Research Scientist to lead fine-tuning, post-training, model-steerability, and reinforcement learning for the next generation of DeepL's LLM-based translation models. This is a high-impact, hands‐on role for a researcher ...

Senior Machine Learning Engineer London

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
quantitative field (e.g., computer science, mathematics, statistics), another quantitative applied field or equivalent experience. Strong understanding of industry standard ML techniques (from classification to reinforcement learning and generative methods) Experience in applying machine learning algorithms to solve real-world problems. Experience with software development concepts like … strong desire to create engaging player experiences and drive business results. Bonus points if you also have Experience with more complex ML techniques (e.g. Reinforcement Learning). Experience with A/B testing methodologies in an online gaming context. Experience working in the games industry. Working at Tripledot ...

Software Engineer, RL Data

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
down to reading transcripts, supporting users, and wrangling vendors. The company's RL Data team builds the systems that produce high‐quality reinforcement learning data for Claude: data collection pipelines, human feedback tooling, the execution environments RL tasks run in, and the quality assurance that keeps training data … Effective use of AI tools in your own day‐to‐day work. Care about the societal impacts of your work. Preferred qualifications Experience with reinforcement learning on LLMs, particularly on the data side: creating evals, environments, rewards, graders, or training data. Experience helping organizations use AI more effectively ...

Software Engineer, RL Data

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
down to reading transcripts, supporting users, and wrangling vendors. Anthropic’s RL Data team builds the systems that produce high‐quality reinforcement learning data for Claude: data collection pipelines, human feedback tooling, the execution environments RL tasks run in, and the quality assurance that keeps training data trustworthy … Effective use of AI tools in your own day‐to‐day work. Care about the societal impacts of your work. Preferred qualifications Experience with reinforcement learning on LLMs, particularly on the data side: creating evals, environments, rewards, graders, or training data. Experience helping organizations use AI more effectively ...

Research Engineer, RL Scaling Science

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
policy experts, and business leaders working together to build beneficial AI systems. About the role The company's RL Scaling Science team studies how reinforcement learning behaves as we scale it (across model size, compute, and task horizon) and turns that understanding into the training recipes behind … scale Partner closely with adjacent RL teams across research and engineering and advance our overall RL stack Minimum qualifications Strong empirical research skills in Reinforcement Learning, large-scale ML training, or a closely adjacent area Demonstrated ability to own large experiments end-to-end, from design through interpretation ...

Manager, Lead Research Scientist, LLM Agents (Foundational Research)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
curious and open-minded individual with an interest in conducting state-of-the-art foundational machine learning research? Thomson Reuters Labs is seeking Research Scientists with a passion for building complex agent-based AI systems in a data-rich, complex academic environment driven by real-world problems. Foundational Research … sleeves and participate in designing, coding, conducting experiments, and translating findings into concrete deliverables. Our focus areas are: LLM Training (Continued Pretraining, Instruction Tuning, Reinforcement Learning Alignment, Distributed Training, Efficient ML techniques) Post-training techniques for planning, reasoning & complex workflows (e.g., Reasoning Models, LLMs + Knowledge Graphs, Test ...

Manager, Research Engineering (Foundational Research)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Manager, Research Engineering with the skills and drive to manage a top performing Engineering team focused on developing the latest in LLM and Machine Learning capabilities. Foundational Research Foundational Research is the dedicated core Machine Learning research division of Thomson Reuters. We are focused on research and development … with a particular focus on advanced algorithms and training techniques for Large Language Models (LLMs). LLM Training (Continued pretraining, instruction tuning, reinforcement learning, distributed training, efficient ML techniques) Post-training techniques for planning, reasoning & complex workflows (e.g., reasoning models, LLMs + knowledge graphs, test time compute ...

Data Scientist Senior Manager

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
ensure best practices in modelling, experimentation and deployment. Data Science & Advanced Analytics – Design and implement predictive and prescriptive models for airline operations, applying machine learning, optimisation and statistical analysis to large, complex data sets. Translation of Business Problems – Convert business challenges into analytical frameworks and actionable solutions to improve … driving high‐impact projects and aligning with business objectives. Advanced knowledge of the Data Science Toolbox – mathematics, statistics, programming, data ingestion, munging, visualisation, machine learning, optimisation, simulation, reinforcement learning, and big‐data technologies. Experience leading successful data‐science projects in an operational setting (preferred). Understanding ...

Senior Data Scientist & Operations Analytics Lead

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
Ensure best practices in modelling, experimentation, and deployment. · Data Science & Advanced Analytics. Design and implement predictive and prescriptive models for airline operations. Apply machine learning, optimization techniques, and statistical analysis to large, complex datasets · Translate business problems into analytical frameworks and actionable solutions to improve decision making · Stakeholder engagement. … advanced knowledge level of the Data Science Toolbox (i.e. the fundamentals of Mathematics and Statistics, computer programming, Data Ingestion, Data Munging, Data visualisation, Machine Learning, Optimisation, Simulation, Reinforcement Learning and Big Data techniques and technologies) · Have experience of leading and landing successful Data Science projects ...

Software Engineer (Applied AI)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
iteration of our next-generation benefits platform features that leverage personalization, experimentation, and AI/ML methods (e.g. agents/LLMs, recommender systems, reinforcement learning) to enhance user experience in a meaningful business domain. Contribute across the tech stack : You’ll work in React (JavaScript/TypeScript … against important business goals that help the entire team win Pragmatic Best Practices : An overarching desire to build efficient, scalable, and maintainable code, while learning the tradeoffs between technical debt and delivery speed What we look for We’re a great bunch but we have some "Euph" cultural ...

Applied Scientist (Integrity)

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
What We Are Looking For Deep expertise in one applied ML, statistics or data science speciality. Examples could be risk modelling, LLM detection, or reinforcement learning – any area of genuine depth counts. Judgement about when to reach for simple statistics, classical ML, LLMs, or agentic approaches. 3+ years ...

Applied AI Engineer - Senior to Principal

Hiring Organisation
Jobleads-UK
Location
Greater London, England, United Kingdom
into well‐defined AI tasks; work closely with Product Managers to define technical architectures and implement them. Select appropriate modelling approaches (classical ML, deep learning, foundation models, prompting, fine‐tuning) based on data, constraints and desired outcomes. Qualifications Proven experience rapidly developing and shipping AI products or models … . Hands‐on experience with a range of AI concepts such as LLMs, VLMs, RAG, agents, MCP, computer vision, GraphML, recommendation systems or reinforcement learning. Familiarity with ML frameworks (PyTorch, TensorFlow, etc.) and strong proficiency in Python and data structures and algorithms. Engineering excellence – writing high‐quality, maintainable code ...