1 to 25 of 40 Remote/Hybrid Reinforcement Learning Jobs in London

Research Engineer / Scientist, Post-training - London

Location
Greater London, England, United Kingdom
seeking those who are dedicated as much to building safely and responsibly as to advancing disruptive agentic capabilities. We promote a mindset of openness, learning, and collaboration, where everyone has something to contribute. About the Research & Models Team The Models team builds the foundational models that power our cutting … perceive, understand, and act within complex environments. We own the entire pipeline including synthetic data generation, environment design, mid-training, supervised fine-tuning, offline reinforcement learning, online reinforcement learning, reward modelling, transition modelling, etc. Our team also has dedicated MLOps, Infra and Inference support at scale. ...

Senior Machine Learning Engineer - Messaging Platform

Location
Greater London, England, United Kingdom
moment. We’re evolving how messaging works at Spotify — moving from short-term optimization toward systems that understand long-term user journeys. By combining reinforcement learning approaches with deeper domain signals, we’re expanding how machine learning shapes the entire messaging funnel. What You’ll Do Design … build, and ship machine learning models that optimize messaging across push, email, and in-app channels Plan and run A/B experiments in a multi-objective environment, balancing conversion, engagement, retention, and reachability Contribute to reinforcement learning systems that optimize for long-term user outcomes rather ...

Senior Research Scientist | Model Steering

Location
Greater London, England, United Kingdom
deliver perfect translations for the most demanding use cases. To that end, we take responsibility for the entire life cycle of the machine learning models that power our language AI products. This includes data, training, quality assurance, and operational aspects. In our highly collaborative teams, each person … drive impact across the company. Your responsibilities We are looking for a Senior Research Scientist to lead fine-tuning, post-training, model-steerability, and reinforcement learning for the next generation of DeepL's LLM-based translation models. This is a high-impact, hands‐on role for a researcher ...

Applied Scientist, Wayve Labs London, United Kingdom

Location
Greater London, England, United Kingdom
embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future. At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster … Wayve Labs and help build the next generation of AI systems for autonomous driving and robotics. You’ll work at the intersection of machine learning, simulation, robotics, and real-world deployment, contributing to core innovations that push the boundaries of embodied AI. Situated within Wayve, we are a high ...

Staff Research Scientist, Reinforce Learning

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future. At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster … join Wayve Labs and help build the next generation of AI systems for autonomous driving. You'll work at the intersection of machine learning, simulation, robotics, and real-world deployment, contributing to core innovations that push the boundaries of embodied AI.Situated within Wayve, we are a high-conviction research ...

Research Scientist, Wayve Labs

Hiring Organisation
wayve
Location
London, UK
Employment Type
Full-time
embrace uncertainty, leaning into complex challenges to unlock groundbreaking solutions. We aim high and stay humble in our pursuit of excellence, constantly learning and evolving as we pave the way for a smarter, safer future. At Wayve, your contributions matter. We value diversity, embrace new perspectives, and foster … Wayve Labs and help build the next generation of AI systems for autonomous driving and robotics. You'll work at the intersection of machine learning, simulation, robotics, and real-world deployment, contributing to core innovations that push the boundaries of embodied AI.Situated within Wayve, we are a high-conviction ...

Senior Operations Research Scientist

Location
Greater London, England, United Kingdom
ground operations. Designing and implementing high-performance optimisation algorithms and production-grade solutions using Python, C++ and commercial optimisation solvers. Integrating optimisation, simulation and reinforcement learning techniques to solve complex multi-stage decision problems. Partnering with senior stakeholders to translate business challenges into mathematical models with clear operational … teams. Significant relevant industry experience within operations research and optimisation, with evidence of technical leadership progression. Cloud-based optimisation platforms and distributed computing. Machine learning and hybrid optimisation approaches. Simulation, reinforcement learning and advanced heuristic methods. Commercial airline operations, crew planning systems or aircraft scheduling. Aviation regulations ...

[Expression of Interest] Research Engineer / Scientist, Alignment - London

Location
Greater London, England, United Kingdom
experts, and business leaders working together to build beneficial AI systems. About the role You want to build and run elegant and thorough machine learning experiments to help us understand and steer the behavior of powerful AI systems. You care about making AI helpful, honest, and harmless … safety techniques by training language models to subvert our safety techniques, and seeing how effective they are at subverting our interventions. Run multi‐agent reinforcement learning experiments to test out techniques like AI Debate. Build tooling to efficiently evaluate the effectiveness of novel LLM‐generated jailbreaks. Write scripts ...

Research Engineer, RL Scaling Science

Location
Greater London, England, United Kingdom
engineers, policy experts, and business leaders working together to build beneficial AI systems. About the role Anthropic's RL Scaling Science team studies how reinforcement learning behaves as we scale it (across model size, compute, and task horizon) and turns that understanding into the training recipes behind … appear at scale Partner closely with adjacent RL teams across research and engineering and advance our overall RL stack Strong empirical research skills in Reinforcement Learning, large-scale ML training, or a closely adjacent area Demonstrated ability to own large experiments end-to-end, from design through interpretation ...

Software Engineer, RL Data

Location
Greater London, England, United Kingdom
down to reading transcripts, supporting users, and wrangling vendors. The company's RL Data team builds the systems that produce high‐quality reinforcement learning data for Claude: data collection pipelines, human feedback tooling, the execution environments RL tasks run in, and the quality assurance that keeps training data … Effective use of AI tools in your own day‐to‐day work. Care about the societal impacts of your work. Preferred qualifications Experience with reinforcement learning on LLMs, particularly on the data side: creating evals, environments, rewards, graders, or training data. Experience helping organizations use AI more effectively ...

Head of Delivery (Public Sector)

Location
Greater London, England, United Kingdom
team. Experience of applying knowledge of cloud computing, including architecting and deploying solutions, and selecting appropriate tools and design patterns. Exposure to advanced machine learning techniques (e.g. reinforcement learning, deep learning). Experience owning the day-to-day resource allocation of a large team of technical … clients and partners within the public sector for successful technical delivery and to build strong long-term relationships Knowledge of data science and machine learning concepts and principles. Expertise in Python and SQL, both for development and to review code. Strong working knowledge of database systems (e.g., SQL, NoSQL ...

AI Application Developer

Location
Greater London, England, United Kingdom
applicable. Last date of application: 30th June 2024 Start Date: 1st August 2024 Key Responsibilities AI Model Development: Design, develop, and train machine learning models using frameworks like TensorFlow, PyTorch, or scikit-learn. Fine-tune and optimize models for performance and scalability. Software Development: Write clean, maintainable, and efficient … communication abilities. Capability to work collaboratively in a team environment. Preferred Skills Experience with natural language processing (NLP) and computer vision techniques. Knowledge of reinforcement learning and deep learning architectures. Familiarity with DevOps practices and continuous integration/continuous deployment (CI/CD) pipelines. Experience in agile ...

Global Banking & Markets - GSET - Quantitative Strategist - London - VP London · United Kingdom[...]

Location
Greater London, England, United Kingdom
responsible for the research, design, and continuous improvement of our execution algorithm platform. We combine deep expertise in market microstructure, statistical modelling, and machine learning with world-class engineering to build algorithms that optimise execution quality, minimise market impact, and adapt intelligently to real-time market conditions. Our work … stay at the frontier of quantitative research, whether that means reading the latest papers on optimal execution, experimenting with new ML techniques, or learning from post‐trade analytics. Culture carriers — You contribute to an inclusive, high‐performance team culture. You are willing to mentor others, share knowledge, and uphold ...

Senior Machine Learning Scientist

Location
Greater London, England, United Kingdom
right place. Agentforce is the future of AI, and you are the future of Salesforce. What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands. We are an extremely … move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting-edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test and measure ...

Staff Machine Learning Scientist

Location
Greater London, England, United Kingdom
values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands. … move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting-edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test and measure ...

Senior Machine Learning Scientist

Location
Greater London, England, United Kingdom
would like to apply, and the type of support you need to complete the application to adjustments@depop.com.RoleDepop is looking for a Senior Machine Learning Scientist to join our Pricing team in the UK. You will work alongside a cross-functional team of Product Managers, Engineers, Analysts and fellow … Machine Learning Scientists, playing a key role in building the machine learning models that power Depop’s pricing recommendations and our shipping pricing and optimisation strategies.As a senior member of the team, you will be expected to take ownership of high-impact projects, lead technical direction on core ...

Senior Machine Learning Scientist (London)

Location
Greater London, England, United Kingdom
values, we push boundaries, build with speed and intensity, and relentlessly deliver incredible value to our customers. What's the opportunity? Fin's Machine Learning team is responsible for defining new ML features, researching appropriate algorithms and technologies, and rapidly getting first prototypes in our customers’ hands. … move to production fast, often shipping to beta in weeks after a successful offline test. We are very passionate about applying machine learning technology, and have productized everything from classic supervised models, to cutting‐edge unsupervised clustering algorithms, to novel applications of transformer neural networks. We test and measure ...

Applied Value Engineer

Hiring Organisation
Celonis
Location
London, UK
Employment Type
Full-time
cloud reference architectures, identity and access, data governance, privacy, and compliance. Math & Operations Research: Applied knowledge of linear/nonlinear optimization, statistical analysis, reinforcement learning, and forecasting is a plus. Strong presentation skills to both internal and external stakeholders (including executives), whether whiteboarding sessions or formal readouts … flexible hybrid work model that balances remote focus with vibrant office collaboration. Continuous Growth: Elevate your skills through our 70-20-10 learning framework, mentorship programs, and access to a dedicated learning platform. Holistic Well-being: Prioritize your health with subsidized Wellhub memberships, mental health counseling, and dedicated ...

Principal Research Scientist - Frontier Industrial AI

Hiring Organisation
Aveva Group
Location
London, UK
Employment Type
Full-time
series and multimodal data, simulation and world models, trustworthy decision-support agents and staged autonomy, and AI-native approaches to industrial data foundations, ontology learning, and evaluation. This is a hands-on role for someone who thrives in uncertain problem spaces, works with high autonomy, and finds practical ways … AI.Develop new models, algorithms, evaluation methods, and system architectures. Build high-quality experimental systems and validated research prototypes. Publish at leading AI and machine learning venues and contribute to AVEVA's patent and IP portfolio. Create benchmarks, datasets, and evaluation frameworks for industrial AI.Partner with leading AI labs, universities ...

Data Scientist Senior Manager

Location
Greater London, England, United Kingdom
ensure best practices in modelling, experimentation and deployment. Data Science & Advanced Analytics – Design and implement predictive and prescriptive models for airline operations, applying machine learning, optimisation and statistical analysis to large, complex data sets. Translation of Business Problems – Convert business challenges into analytical frameworks and actionable solutions to improve … driving high‐impact projects and aligning with business objectives. Advanced knowledge of the Data Science Toolbox – mathematics, statistics, programming, data ingestion, munging, visualisation, machine learning, optimisation, simulation, reinforcement learning, and big‐data technologies. Experience leading successful data‐science projects in an operational setting (preferred). Understanding ...

Senior Data Scientist & Operations Analytics Lead

Location
Greater London, England, United Kingdom
Ensure best practices in modelling, experimentation, and deployment. · Data Science & Advanced Analytics. Design and implement predictive and prescriptive models for airline operations. Apply machine learning, optimization techniques, and statistical analysis to large, complex datasets · Translate business problems into analytical frameworks and actionable solutions to improve decision making · Stakeholder engagement. … advanced knowledge level of the Data Science Toolbox (i.e. the fundamentals of Mathematics and Statistics, computer programming, Data Ingestion, Data Munging, Data visualisation, Machine Learning, Optimisation, Simulation, Reinforcement Learning and Big Data techniques and technologies) · Have experience of leading and landing successful Data Science projects ...

Research Engineer, Pretraining

Location
Greater London, England, United Kingdom
Contribute to the entire stack, from low-level optimizations to high-level model design Qualifications: Advanced degree (MS or PhD) in Computer Science, Machine Learning, or a related field Strong software engineering skills with a proven track record of building complex systems Expertise in Python and experience with deep … learning frameworks (PyTorch preferred) Familiarity with large-scale machine learning, particularly in the context of language models Ability to balance research goals with practical engineering constraints Strong problem-solving skills and a results-oriented mindset Excellent communication skills and ability to work in a collaborative environment Care about ...

Machine Learning Researcher

Location
City Of London, England, United Kingdom
plans, 10% employer pension contributions and an extensive structured professional development programme. Keywords Defence, EM, Electromagnetic, RF, Radio Frequency, AI, Artificial Intelligence, ML, Machine Learning, GNNs, Transformers, Autoencoders, Reinforcement Learning, Multi-Modal AI, Data Fusion, EPM, Electronic Protection Measures, ESM, Electronics Support Measures, EA, Electronics Attack, Python ...

AI Engineer

Location
City Of London, England, United Kingdom
organization with multiple teams to improve the efficiency of company processes and provide insights and assistance to users. Responsibilities Design, implement, and deploy machine learning models and systems to solve complex problems and drive business outcomes; Research, develop, and implement machine learning algorithms and models for tasks such … implement AI tools such as NLP, LLM, and IA; Collect, preprocess, and curate large datasets required for training generative models; Experiment with different machine learning techniques and algorithms, including supervised, unsupervised, semi-supervised, reinforcement, and deep learning; Design and optimize machine learning pipelines and workflows, incorporating ...

Research Scientist, Personalization

Location
Greater London, England, United Kingdom
helps millions of listeners discover what they love. Within this space, our research team focuses on advancing the state of the art in machine learning and AI to shape the future of personalization. We explore new approaches, challenge existing assumptions, and contribute to the broader research community while influencing … long-term product direction. What You’ll Do Conduct original research in machine learning and AI, with a focus on large-scale foundation models, generative AI, and post-training methods, including reinforcement learning. Develop novel methodologies, models, and evaluation frameworks to advance personalization systems. Design and execute rigorous ...