126 to 150 of 165 Reinforcement Learning Jobs in England

VP of Engineering - Cloud, AI, and Microservices. London, United Kingdom

Location
City Of London, England, United Kingdom
integrations; oversee LLM (Large Language Model) integration, leveraging models like GPT, and implement advanced AI architectures such as Retrieval-Augmented Generation (RAG) and Reinforcement Learning with Human Feedback (RLHF); lead the development of applications using React, .NET, Python, or Node.js, ensuring scalable, maintainable, and performant solutions. Team Management … Manage and mentor engineering managers, leads, and developers, fostering a culture of excellence and continuous learning; ensure the team adheres to best practices in coding, architecture, testing, and deployment. Project Delivery & Client Satisfaction: Ensure client projects meet deadlines, scope, and quality standards while managing resources effectively; resolve technical challenges ...

RLHF Manager

Location
Cheltenham, England, United Kingdom
Research & QualityHybridFull-timeDisability Confident Committed# RLHF ManagerLead Reinforcement Learning from Human Feedback (RLHF) programmes across Coaley Peak's internal AI engine development and external client AI projects, ensuring our models are safe, aligned, and genuinely useful.LocationRemote (UK)/Cheltenham, UKSalary£55,000 – £75,000 per annum (DOE)Positions1 … positionReferenceCP-RLF-2025-001Published21 March 2025Closing14 December 2026›AI Research & Quality›RLHF ManagerAbout the roleReinforcement Learning from Human Feedback is one of the most important mechanisms for making AI systems behave well in the real world. As RLHF Manager at Coaley Peak, you will design and manage ...

Weapon Systems Algorithms Engineer - Summer Placement 2027

Location
West of England, England, United Kingdom
close earlier depending on application volumes. What we can offer you: 10 week placement: Starting June 2027, that allows you to apply your university learning to real-world projects and technologies Perform well on your placement and you could secure a graduate role before your final year Pension: Maximum … future business teams to bring technical excellence to our future concepts. We leverage the use of leading technologies and tools such as deep reinforcement learning and test-driven algorithm development to provide optimised weapon system solutions across the whole of the MBDA product line. Over your time ...

Genesis-World: Core Simulation Engine Engineer

Location
Greater London, England, United Kingdom
compiler, lowers them to CUDA, AMD ROCm, Apple Metal, Vulkan, x86, and ARM64. A single laptop or a datacenter. Massively batched GPU simulation for learning at scale, and complex non-batched scenes where CPU wins outright. This is at the core of Genesis AI’s strategy. Evaluation … that improves at the speed of compute. The role Simulation is still a hard sell in robotics. Outside a few success stories, like reinforcement learning for locomotion, most teams skip it, and friction is a big part of why: painful to use, painful to debug. We live that ...

Senior AI Engineer — Design & Manufacturing (Generative AI)

Location
Cambridge, England, United Kingdom
agents for enterprise workflows, and create robust training pipelines that enable rapid experimentation and deployment at scale. We value hands-on expertise with LLMs, reinforcement learning, and agentic systems, along with strong Python and software engineering fundamentals. #J-18808-Ljbffr ...

Senior Software Engineer, Data Platform London, UK Apply →

Location
Greater London, England, United Kingdom
enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human evaluation and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT ...

Senior RL Engineer for Humanoid Robots (Equity)

Location
Greater London, England, United Kingdom
Humanoid is seeking a Senior or Staff Reinforcement Learning Engineer to develop learning-based control policies for humanoid robots. You will design and train RL policies enabling dynamic locomotion and loco-manipulation on real robots, building scalable pipelines and improving sim-to-real transfer for reliable hardware ...

Research Engineer

Location
Tipton, England, United Kingdom
product‐minded research engineer who enjoys building real systems people depend on. You’ll likely have: Strong technical background in software engineering, machine learning, or applied AI, demonstrated through an advanced degree and/or equivalent experience building production AI systems Strong software engineering fundamentals and good judgment … Practical understanding of modern model adaptation and post‐training methods, including LoRA/QLoRA, SFT, distillation, preference optimization, reward modeling, DPO/GRPO, and reinforcement learning from verifiable feedback Ownership mindset: you drive projects end‐to‐end, move quickly from real usage, and care about shipping measurable improvements ...

Platform Engineer

Location
Greater London, England, United Kingdom
superlearning capability - the ability to endlessly discover knowledge and skills, without relying on human data - will be driven by the world’s most powerful reinforcement learning algorithms. The superlearner is expected to rediscover and then transcend the greatest inventions in human history, such as language, science, mathematics … role where you inherit someone else's architecture, you'll be one of the people defining it. We're pushing these learning methods to a scale that hasn’t been tried before and we care deeply about the infrastructure that gets us there. You'll own the infrastructure that ...

Agent Product Manager

Location
Greater London, England, United Kingdom
usage and model performance, identify key issues, validate hypotheses, and continuously improve product experience and Agent capabilities. Requirements - Basic understanding of LLMs, AI Agents, reinforcement learning, and model training or post-training workflows. - Strong interest in general-purpose Agent products and enthusiasm for building product-data flywheels. - Experience … evaluation, RLHF/RLAIF, data annotation platforms, or model training data pipelines is a strong plus. Basic understanding of LLMs, AI Agents, and reinforcement learning.Strong product judgment, with the ability to identify user needs, define problems, and design practical product solutions.Strong structured thinking and analytical skills, with the ability ...

Senior Data Scientist in ADVANCED ANALYTICS

Location
Leeds, England, United Kingdom
opportunity to work at the heart of data science, research and public policy. Advanced Analytics is hiring a Senior Data Scientist to develop machine learning (ML) and Artificial Intelligence (AI) solutions, with a focus on robust prototyping. We have pioneered the application of new analytical approaches in the Bank … England, such as generative AI, machine learning, natural language processing, agent-based models, network analytics, and randomised control trials. We’re looking for someone with ML/AI expertise to help us continue pushing the organisation’s analytical frontier to inform policymaking. DAT is leading the transformation ...

Senior Data Scientist in ADVANCED ANALYTICS

Location
City Of London, England, United Kingdom
opportunity to work at the heart of data science, research and public policy. Advanced Analytics is hiring a Senior Data Scientist to develop machine learning (ML) and Artificial Intelligence (AI) solutions, with a focus on robust prototyping. We have pioneered the application of new analytical approaches in the Bank … England, such as generative AI, machine learning, natural language processing, agent-based models, network analytics, and randomised control trials. We’re looking for someone with ML/AI expertise to help us continue pushing the organisation’s analytical frontier to inform policymaking. DAT is leading the transformation ...

Onsite RL Engineer – Robotics & Embodied AI

Location
Greater London, England, United Kingdom
Randstad Technologies Recruitment is searching for a Reinforcement Learning (RL) Engineer to join a leading robotics company in London. This permanent position involves working collaboratively to define the company's 2026 technical roadmap while solving complex challenges in autonomous systems. The role requires expertise in AI, MLOps, Software ...

Senior RL Data Engineer: End-to-End Pipelines & QA

Location
Greater London, England, United Kingdom
Senior Engineer in Greater London to develop and manage data pipelines for AI systems. This role involves significant responsibilities, including ensuring the quality of reinforcement learning data, collaborating with various teams, and innovating operational frameworks. The ideal candidate will possess strong software engineering skills, have a background ...

(Senior) Product Marketing Manager Marketing & Communications Munich - Berlin

Location
Greater London, England, United Kingdom
safety, and ethical considerations are vital. Our work frequently takes us right up to the state of the art in technical innovation, be it reinforcement learning, distributed systems, generative AI, or deployment infrastructure. The defence industry is entering the most exciting phase of the technological development curve. Advances ...

Full Stack Engineer

Location
Greater London, England, United Kingdom
simulation and real hardware environments. You will be part of a focused team responsible for the application level software that connects control, navigation, perception, learning, and platform systems. Your work will ensure that these components operate as a coherent and reliable system that users can interact with seamlessly. This … will develop and maintain application level software for humanoid robots. You will integrate software components from controls, navigation, computer vision, reinforcement learning, and platform teams. You will contribute to the structure and evolution of the application architecture and its interfaces. You will work closely with the product ...

Research Scientist PhD Intern, 2027

Hiring Organisation
Hackajob Ltd
Location
South West London, London, United Kingdom
Employment Type
Permanent
your penultimate/final year of education. Experience in one or more research areas; natural language understanding/processing, computer vision, LLMs, machine learning, algorithmic foundations of optimization, data mining, machine intelligence (Artificial Intelligence), AI/ML algorithms, deep learning or image processing, multimodal, multilingual and reinforcement … extend and improve on Google's product offering. Contribute to a wide variety of projects utilising natural language processing, artificial intelligence, data compression, machine learning, and search technologies. Please complete your application before 23rd October 2026. We encourage you to apply as early as possible as we review applications ...

Research Scientist, RL for Long-Horizon Agents

Location
Greater London, England, United Kingdom
leading AI research firm in London is seeking a Research Scientist to push the frontier of reinforcement learning. The role involves designing, implementing, and evaluating novel RL methods for language models and requires experience in reinforcement learning and proficiency in Python. You will collaborate closely with engineering ...

Hardware Engineer - Physical Products Hardware Engineering London; Oxford

Location
Oxford, England, United Kingdom
meaningful impact in the world. Our work frequently takes us right up to the state of the art in technical innovation, be it reinforcement learning, distributed systems, generative AI, or deployment infrastructure. The defence industry is entering the most exciting phase of the technological development curve. Advances … that you care deeply about. What we offer Competitive salary and VSOP options Relocation support: up to €2,500 and 4 weeks temporary accommodation Learning: €500/£450 yearly allowance Health & wellness: gym membership and mental health support (Nilo.health) Social: regular company events and monthly social allowances Enhanced parental ...

Hardware Engineer - Physical Products Hardware Engineering London; Oxford

Location
Greater London, England, United Kingdom
meaningful impact in the world. Our work frequently takes us right up to the state of the art in technical innovation, be it reinforcement learning, distributed systems, generative AI, or deployment infrastructure. The defence industry is entering the most exciting phase of the technological development curve. Advances … that you care deeply about. What we offer Competitive salary and VSOP options Relocation support: up to €2,500 and 4 weeks temporary accommodation Learning: €500/£450 yearly allowance Health & wellness: gym membership and mental health support (Nilo.health) Social: regular company events and monthly social allowances Enhanced parental ...

Research Engineer: Generative AI & ML Systems (Hybrid)

Location
Greater London, England, United Kingdom
seeking a Research Engineer to advance AI research and ML systems. The role blends hands-on engineering with rigorous experimentation on large models, reinforcement learning, synthetic data, and multi-agent setups. You will work within a world-leading technology organization, contributing to cutting-edge infrastructure and research initiatives ...

Senior Research Scientist - Multimodal Vision Translation

Location
Greater London, England, United Kingdom
DeepL is seeking a Senior Research Scientist to lead fine-tuning, post-training, and reinforcement learning for its document translation multimodal and vision models. You will develop models that reason about document layout, fuse data sources, and drive breakthroughs from prototype to deployment. The role requires hands ...

Senior PM: Commerce Personalization & ML-Powered Experience

Location
Greater London, England, United Kingdom
driven initiatives with cross-functional partners to improve customer outcomes at scale. You will partner with Applied Science and ML Engineering to develop reinforcement learning and optimization techniques, ensuring #J-18808-Ljbffr ...

Member of Technical Staff - Sandbox Service

Location
Greater London, England, United Kingdom
tuning filesystem performance on nodes. The Sandbox service enables Grok to safely run and test code in real‐time for user queries, and supports reinforcement learning in training, where models interactively explore tools ranging from compilers to productivity apps. About the role Expect deeply technical work ...

Technical Lead (Energy Systems)

Hiring Organisation
Gryd Energy
Location
London, England, United Kingdom
live deployment Strong documentation and code quality practices Helpful but not essential: Experience with constraint-based or multi-objective optimisation (linear programming, MPC, reinforcement learning) Familiarity with residential battery or inverter platforms and their cloud APIs Understanding of UK electricity markets, grid flexibility, or DNO constraints Startup ...