1 to 25 of 32 Remote/Hybrid Reinforcement Learning Jobs in England

Research Engineer

Location
Oxford, England, United Kingdom
Role Research Engineer Salary Competitive Contract Perm Location Oxford The Role As our Research Engineer, you will sit at the boundary between frontier machine‐learning research and shipped product. You will take ideas from the research edge, reinforcement learning, RLHF, multi‐objective policy optimisation and LLM post … customers depend on, while connecting that research to OD’s product vision and to what our customers actually need. Responsibilities Apply and adapt machine‐learning and reinforcementlearning research to OD’s product roadmap, translating research direction into shippable capability. Design, train and evaluate models using reinforcement ...

Research Engineer / Scientist, Post-training - London

Location
Greater London, England, United Kingdom
seeking those who are dedicated as much to building safely and responsibly as to advancing disruptive agentic capabilities. We promote a mindset of openness, learning, and collaboration, where everyone has something to contribute. About the Research & Models Team The Models team builds the foundational models that power our cutting … perceive, understand, and act within complex environments. We own the entire pipeline including synthetic data generation, environment design, mid-training, supervised fine-tuning, offline reinforcement learning, online reinforcement learning, reward modelling, transition modelling, etc. Our team also has dedicated MLOps, Infra and Inference support at scale. ...

Senior Research Scientist FMTA

Location
Greater London, England, United Kingdom
large language models. We focus on developing algorithms and systems that align pre-trained models with tasks and performance goals through techniques like reinforcement learning. As a research-driven team, we stay up to date with current literature to integrate cutting-edge ideas into our core stack. As part … controllability, and safer, more effective user experiences. Your responsibilities As a Senior Research Scientist, you’ll design, implement, and deploy cutting-edge research in reinforcement learning and post-training at scale, driving innovations that make it into production. You will: Build and deploy state-of-the-art reinforcement ...

Machine Learning Engineer - Agentic AI & Reinforcement Learning

Location
Crawley, England, United Kingdom
solutions that efficiently and responsibly resolve complex natural resource, digital, energy transition and infrastructure challenges. We are looking for a hands‐on Machine Learning Engineer to design, develop and evaluate intelligent agent systems, turning emerging AI methods into reliable, scalable and reusable solutions for truly real world data. About … Team Learn more about our work and the team on our website, recent blog post featuring one of our senior machine learning engineers and our latest Voices of Viridien. Key responsibilities Design and develop single-agent and multi-agent systems for complex workflows. Build reasoning, planning, task decomposition, memory ...

Machine Learning Scientist — Large Multimodal Models (Post-Training)

Location
West of England, England, United Kingdom
SUMMARY We are seeking a Machine Learning Scientist to join the Enchant team at Iambic Therapeutics. Our mission is to deliver better medicines through innovation in AI-based discovery technologies. In this role, you will research and develop post-training methods for Enchant - our multimodal transformer model trained … discovery. The role centers on designing and evaluating post-training approaches for large multimodal language models including supervised fine-tuning, parameter-efficient fine-tuning, reinforcement learning, preference or reward-based optimization, and other emerging post-training methods. You will develop rigorous evaluations and training infrastructure that make ...

Senior Machine Learning Engineer - Messaging Platform

Location
Greater London, England, United Kingdom
moment. We’re evolving how messaging works at Spotify — moving from short-term optimization toward systems that understand long-term user journeys. By combining reinforcement learning approaches with deeper domain signals, we’re expanding how machine learning shapes the entire messaging funnel. What You’ll Do Design … build, and ship machine learning models that optimize messaging across push, email, and in-app channels Plan and run A/B experiments in a multi-objective environment, balancing conversion, engagement, retention, and reachability Contribute to reinforcement learning systems that optimize for long-term user outcomes rather ...

Senior AI Platform Engineer

Hiring Organisation
Capgemini
Location
Greater London, United Kingdom
Employment Type
Full Time
small senior team with heavy AI leverage. The ambition runs past serving frontier models: we close the loop from production feedback through reinforcement learning and fine-tuning, and train our own LLMs and SLMs where evaluations and economics justify it. What you will own One or more platform … guardrail engine, tool and agent registry, or shared product services The training and adaptation loop: pipelines that turn production traces and evaluation verdicts into reinforcement learning and fine-tuning datasets, and the infrastructure to train, evaluate, and serve NewCo-tuned LLMs and SLMs behind the same gates ...

Senior Research Scientist | Model Steering

Location
Greater London, England, United Kingdom
deliver perfect translations for the most demanding use cases. To that end, we take responsibility for the entire life cycle of the machine learning models that power our language AI products. This includes data, training, quality assurance, and operational aspects. In our highly collaborative teams, each person … drive impact across the company. Your responsibilities We are looking for a Senior Research Scientist to lead fine-tuning, post-training, model-steerability, and reinforcement learning for the next generation of DeepL's LLM-based translation models. This is a high-impact, hands‐on role for a researcher ...

Senior Software Engineer, Machine Learning

Location
Manchester, England, United Kingdom
relies on building customer relationships with Roku that delight and engage them. Within Advertising Engineering, the MarTech team builds the products, services, and machine learning systems that help Roku deliver the right communication, creative, and marketing experience to the right customer at the right time on the right marketing … business. The team owns server technologies, data platforms, and cloud services that power advertising and marketing use cases. We are recruiting a Senior Machine Learning Engineer to build and enhance intelligent systems that help the marketing team harness the power of data. About the Role The MarTech team ...

Senior Operations Research Scientist

Location
Luton, England, United Kingdom
ground operations. Designing and implementing high‐performance optimisation algorithms and production‐grade solutions using Python, C++ and commercial optimisation solvers. Integrating optimisation, simulation and reinforcement learning techniques to solve complex multi‐stage decision problems. Partnering with senior stakeholders to translate business challenges into mathematical models with clear operational … teams. Significant relevant industry experience within operations research and optimisation, with evidence of technical leadership progression. Cloud‐based optimisation platforms and distributed computing. Machine learning and hybrid optimisation approaches. Simulation, reinforcement learning and advanced heuristic methods. Commercial airline operations, crew planning systems or aircraft scheduling. Aviation regulations ...

Senior Operations Research Scientist

Location
Greater London, England, United Kingdom
ground operations. Designing and implementing high-performance optimisation algorithms and production-grade solutions using Python, C++ and commercial optimisation solvers. Integrating optimisation, simulation and reinforcement learning techniques to solve complex multi-stage decision problems. Partnering with senior stakeholders to translate business challenges into mathematical models with clear operational … teams. Significant relevant industry experience within operations research and optimisation, with evidence of technical leadership progression. Cloud-based optimisation platforms and distributed computing. Machine learning and hybrid optimisation approaches. Simulation, reinforcement learning and advanced heuristic methods. Commercial airline operations, crew planning systems or aircraft scheduling. Aviation regulations ...

AI/ML Engineer

Hiring Organisation
CBSbutler Holdings Limited trading as CBSbutler
Location
Romsey, Hampshire, United Kingdom
Employment Type
Permanent
Salary
£65000/annum plus 6% bonus
difference. What you'll be working on: You'll operate at the cutting edge of AI/ML, working across areas including: Computer Vision Reinforcement Learning Large Language Models (LLMs) Natural Language Processing (NLP) Deep Learning and advanced ML techniques Novel algorithm development and experimentation … Proven commercial experience in AI/ML R&D Strong Python development skills Experience in one or more of Computer Vision, NLP, LLMs or Reinforcement Learning A track record of developing novel AI/ML applications Experience adapting and applying academic research Strong software engineering fundamentals Experience working ...

AI/ML Engineer

Hiring Organisation
CBSbutler Holdings Limited trading as CBSbutler
Location
Romsey, Hampshire, United Kingdom
Employment Type
Permanent
Salary
£60000 - £65000/annum Benefits package
across Defence, National Security and other technically complex environments . This is an opportunity to work across areas including LLMs, NLP, Computer Vision and Reinforcement Learning , depending on your experience. What you'll be doing Researching and evaluating state-of-the-art AI/ML techniques. Adapting academic … with new technologies and solving problems where the solution isn't immediately obvious. Experience in one or more of LLMs, NLP, Computer Vision or Reinforcement Learning . Defence or National Security experience is advantageous but not essential. We're equally interested in candidates from other technically demanding environments ...

Software Engineer, RL Data

Location
Greater London, England, United Kingdom
down to reading transcripts, supporting users, and wrangling vendors. The company's RL Data team builds the systems that produce high‐quality reinforcement learning data for Claude: data collection pipelines, human feedback tooling, the execution environments RL tasks run in, and the quality assurance that keeps training data … Effective use of AI tools in your own day‐to‐day work. Care about the societal impacts of your work. Preferred qualifications Experience with reinforcement learning on LLMs, particularly on the data side: creating evals, environments, rewards, graders, or training data. Experience helping organizations use AI more effectively ...

AI Harness Engineers

Hiring Organisation
Capgemini
Location
Greater London, United Kingdom
Employment Type
Full Time
deploying or operating AI coding agents (Claude Code, Cursor, Copilot, or in-house), beyond personal use You follow frontier agentic-systems research (harness design, reinforcement learning from execution feedback, evaluation methods) closely enough to put it into production within the quarter it lands A measurement habit … unattended agent pipelines with hard failure caps and human escalation You have turned agent execution traces into training or evaluation datasets, or built reinforcement learning pipelines from execution feedback You have written publicly about developer experience or agent harnesses Financial services engineering exposure (banks, insurers, or payment providers ...

Manager, Lead Research Scientist, LLM Agents (Foundational Research)

Location
Greater London, England, United Kingdom
curious and open-minded individual with an interest in conducting state-of-the-art foundational machine learning research? Thomson Reuters Labs is seeking Research Scientists with a passion for building complex agent-based AI systems in a data-rich, complex academic environment driven by real-world problems. Foundational Research … sleeves and participate in designing, coding, conducting experiments, and translating findings into concrete deliverables. Our focus areas are: LLM Training (Continued Pretraining, Instruction Tuning, Reinforcement Learning Alignment, Distributed Training, Efficient ML techniques) Post-training techniques for planning, reasoning & complex workflows (e.g., Reasoning Models, LLMs + Knowledge Graphs, Test ...

Manager, Research Engineering (Foundational Research)

Location
Greater London, England, United Kingdom
Manager, Research Engineering with the skills and drive to manage a top performing Engineering team focused on developing the latest in LLM and Machine Learning capabilities. Foundational Research Foundational Research is the dedicated core Machine Learning research division of Thomson Reuters. We are focused on research and development … with a particular focus on advanced algorithms and training techniques for Large Language Models (LLMs). LLM Training (Continued pretraining, instruction tuning, reinforcement learning, distributed training, efficient ML techniques) Post-training techniques for planning, reasoning & complex workflows (e.g., reasoning models, LLMs + knowledge graphs, test time compute ...

Senior Machine Learning Scientist (London)

Location
Greater London, England, United Kingdom
values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands. … move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting‐edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test and measure ...

AI ML Researcher

Hiring Organisation
Hackajob Ltd
Location
Chelmsford, Essex, South East, United Kingdom
Employment Type
Permanent, Work From Home
Salary
£90,000
Systems could take you. Role Description BAE Systems Digital Intelligence Defence Innovation and Technology has a diverse range of AI teams working in: reinforcement learning, NLP, knowledge graphs, applications of LLMs, Agentic AI, Robotics, Autonomy, computer vision, AI for RF and EW, sonar and acoustics. We are looking … experience of working in projects on topics such as NLP, LLM applications/approaches (e.g. Agentic AI), knowledge graphs and/or graph machine learning and with a vision on how to develop solutions for practical applications of ML in these domains. You should have existing skills in Machine ...

AI/ML Researcher

Hiring Organisation
BAE Systems
Location
Essex, United Kingdom
Employment Type
Full Time
Systems could take you. Role Description BAE Systems Digital Intelligence Defence Innovation and Technology has a diverse range of AI teams working in: reinforcement learning, NLP, knowledge graphs, applications of LLMs, Agentic AI, Robotics, Autonomy, computer vision, AI for RF and EW, sonar and acoustics. We are looking … experience of working in projects on topics such as NLP, LLM applications/approaches (e.g. Agentic AI), knowledge graphs and/or graph machine learning and with a vision on how to develop solutions for practical applications of ML in these domains. You should have existing skills in Machine ...

Data Scientist Senior Manager

Location
Luton, England, United Kingdom
ensure best practices in modelling, experimentation and deployment. Data Science & Advanced Analytics – Design and implement predictive and prescriptive models for airline operations, applying machine learning, optimisation and statistical analysis to large, complex data sets. Translation of Business Problems – Convert business challenges into analytical frameworks and actionable solutions to improve … driving high‐impact projects and aligning with business objectives. Advanced knowledge of the Data Science Toolbox – mathematics, statistics, programming, data ingestion, munging, visualisation, machine learning, optimisation, simulation, reinforcement learning, and big‐data technologies. Experience leading successful data‐science projects in an operational setting (preferred). Understanding ...

Data Scientist Senior Manager

Location
Greater London, England, United Kingdom
ensure best practices in modelling, experimentation and deployment. Data Science & Advanced Analytics – Design and implement predictive and prescriptive models for airline operations, applying machine learning, optimisation and statistical analysis to large, complex data sets. Translation of Business Problems – Convert business challenges into analytical frameworks and actionable solutions to improve … driving high‐impact projects and aligning with business objectives. Advanced knowledge of the Data Science Toolbox – mathematics, statistics, programming, data ingestion, munging, visualisation, machine learning, optimisation, simulation, reinforcement learning, and big‐data technologies. Experience leading successful data‐science projects in an operational setting (preferred). Understanding ...

Data Scientist

Location
Kidlington, England, United Kingdom
e.g., NumPy, Pandas, scikit‐learn, TensorFlow/PyTorch). Demonstrable experience in creating and developing Python libraries. Demonstrable experience designing, implementing and training machine learning models from scratch. Strong foundations in applied mathematics and physics, particularly in statistical modelling, systems dynamics and differential equations. Familiarity with software engineering best … time series modelling techniques (e.g., ARIMA, VAR, Prophet, LSTM). Solid grasp of control theory concepts (e.g., PID controllers, Kalman Filters, Model Predictive Control, Reinforcement Learning). Familiarity with lower-level development of data pipelines in e.g. C++/Rust. Travel Up to 10% travel, including international travel ...

Senior Data Scientist & Operations Analytics Lead

Location
Luton, England, United Kingdom
Ensure best practices in modelling, experimentation, and deployment. · Data Science & Advanced Analytics. Design and implement predictive and prescriptive models for airline operations. Apply machine learning, optimization techniques, and statistical analysis to large, complex datasets · Translate business problems into analytical frameworks and actionable solutions to improve decision making · Stakeholder engagement. … advanced knowledge level of the Data Science Toolbox (i.e. the fundamentals of Mathematics and Statistics, computer programming, Data Ingestion, Data Munging, Data visualisation, Machine Learning, Optimisation, Simulation, Reinforcement Learning and Big Data techniques and technologies) · Have experience of leading and landing successful Data Science projects ...

Senior Data Scientist & Operations Analytics Lead

Location
Greater London, England, United Kingdom
Ensure best practices in modelling, experimentation, and deployment. · Data Science & Advanced Analytics. Design and implement predictive and prescriptive models for airline operations. Apply machine learning, optimization techniques, and statistical analysis to large, complex datasets · Translate business problems into analytical frameworks and actionable solutions to improve decision making · Stakeholder engagement. … advanced knowledge level of the Data Science Toolbox (i.e. the fundamentals of Mathematics and Statistics, computer programming, Data Ingestion, Data Munging, Data visualisation, Machine Learning, Optimisation, Simulation, Reinforcement Learning and Big Data techniques and technologies) · Have experience of leading and landing successful Data Science projects ...