176 to 200 of 228 Reinforcement Learning Jobs in England

Quantitative Researcher

Location
Greater London, England, United Kingdom
Good Markets is building advanced algorithmic trading systems across FX and crypto, combining dynamic hedging, multi-agent grid logic, machine-learning techniques, and high-frequency data insights. We’re expanding the research team with exceptional talent who want to push the boundaries of systematic trading and simulation at scale. … data Experience designing or evaluating algos, simulations, or optimisation routines Understanding of statistics, stochastic processes, PDEs, or signal processing Curiosity about agent-based modelling, reinforcement learning, or market regimes Why Join Us Work directly with founders on real trading systems deployed in live markets Zero bureaucracy, we build ...

Senior AI Software Engineer

Location
Cambridge, England, United Kingdom
novel model architectures and iterate rapidly toward measurable outcomes Solid experience with automated training pipelines, training orchestration, and scalable model evaluation Practical experience applying reinforcement learning in real product or research settings Excellent Python skills and strong software engineering fundamentals for building reliable, maintainable AI systems Strong communication ...

Research Scientist, Robotics Pre-Training and Data Quality, DeepMind

Location
Greater London, England, United Kingdom
designing new architectures, our research scientists work on real-world problems that span the breadth of computer science, such as machine (and deep) learning, data mining, natural language processing, hardware and software performance analysis, improving compilers for mobile platforms, as well as core search and much more. … real world and eager to get direct with foundation model training. In this role, you will have experience with robot policy training using imitation learning and related techniques, alongside experience navigating the realities of pre-training. You should be enthusiastic about—and experience with—practical considerations such as data ...

Research Scientist, Robotics Pre-Training and Data Quality, DeepMind

Location
Greater London, England, United Kingdom
designing new architectures, our research scientists work on real-world problems that span the breadth of computer science, such as machine (and deep) learning, data mining, natural language processing, hardware and software performance analysis, improving compilers for mobile platforms, as well as core search and much more. … real world and eager to get direct with foundation model training. In this role, you will have experience with robot policy training using imitation learning and related techniques, alongside experience navigating the realities of pre-training. You should be enthusiastic about—and experience with—practical considerations such as data ...

VP of Engineering - Cloud, AI, and Microservices. London, United Kingdom

Location
City Of London, England, United Kingdom
integrations; oversee LLM (Large Language Model) integration, leveraging models like GPT, and implement advanced AI architectures such as Retrieval-Augmented Generation (RAG) and Reinforcement Learning with Human Feedback (RLHF); lead the development of applications using React, .NET, Python, or Node.js, ensuring scalable, maintainable, and performant solutions. Team Management … Manage and mentor engineering managers, leads, and developers, fostering a culture of excellence and continuous learning; ensure the team adheres to best practices in coding, architecture, testing, and deployment. Project Delivery & Client Satisfaction: Ensure client projects meet deadlines, scope, and quality standards while managing resources effectively; resolve technical challenges ...

Research Scientist, Robotics Pre-Training and Data Quality, DeepMind

Hiring Organisation
Google
Location
London, United Kingdom
Salary
£ 70 K
tools and processes required to iterate on research ideas quickly and verify data quality through robot policy training.Contribute to our wider research agenda, including reinforcement learning, vision-language-action modeling, world-action models, and simulation, working in a fast-paced, collaborative team environment.Minimum qualifications:PhD in a technical … designing new architectures, our research scientists work on real-world problems that span the breadth of computer science, such as machine (and deep) learning, data mining, natural language processing, hardware and software performance analysis, improving compilers for mobile platforms, as well as core search and much more.As a Research ...

RLHF Manager

Location
Cheltenham, England, United Kingdom
Research & QualityHybridFull-timeDisability Confident Committed# RLHF ManagerLead Reinforcement Learning from Human Feedback (RLHF) programmes across Coaley Peak's internal AI engine development and external client AI projects, ensuring our models are safe, aligned, and genuinely useful.LocationRemote (UK)/Cheltenham, UKSalary£55,000 – £75,000 per annum (DOE)Positions1 … positionReferenceCP-RLF-2025-001Published21 March 2025Closing14 December 2026›AI Research & Quality›RLHF ManagerAbout the roleReinforcement Learning from Human Feedback is one of the most important mechanisms for making AI systems behave well in the real world. As RLHF Manager at Coaley Peak, you will design and manage ...

Machine Learning Engineer, Memory

Hiring Organisation
Epic Games
Location
London, United Kingdom
Salary
£ 100 K
that work across problems, modalities and productsDefine new ways to measure and evaluate memory performanceHelp create and deliver production-ready, scalable and high-quality learning models and associated algorithmsCritically assess the effectiveness of such models and make recommendations for the ongoing roadmapWhat we're looking forPhD in Computer Science … Mathematics or related field, or 3+ years of relevant industry experienceExperience creating machine learning algorithms for vision and/or language problems and deploying them as production-level systemsExpert, hands-on knowledge of:Foundational models and vision and/or language, incl. local deployment and model post-training ...

Weapon Systems Algorithms Engineer - Summer Placement 2027

Location
West of England, England, United Kingdom
close earlier depending on application volumes. What we can offer you: 10 week placement: Starting June 2027, that allows you to apply your university learning to real-world projects and technologies Perform well on your placement and you could secure a graduate role before your final year Pension: Maximum … future business teams to bring technical excellence to our future concepts. We leverage the use of leading technologies and tools such as deep reinforcement learning and test-driven algorithm development to provide optimised weapon system solutions across the whole of the MBDA product line. Over your time ...

Genesis-World: Core Simulation Engine Engineer

Location
Greater London, England, United Kingdom
compiler, lowers them to CUDA, AMD ROCm, Apple Metal, Vulkan, x86, and ARM64. A single laptop or a datacenter. Massively batched GPU simulation for learning at scale, and complex non-batched scenes where CPU wins outright. This is at the core of Genesis AI’s strategy. Evaluation … that improves at the speed of compute. The role Simulation is still a hard sell in robotics. Outside a few success stories, like reinforcement learning for locomotion, most teams skip it, and friction is a big part of why: painful to use, painful to debug. We live that ...

Senior AI Engineer — Design & Manufacturing (Generative AI)

Location
Cambridge, England, United Kingdom
agents for enterprise workflows, and create robust training pipelines that enable rapid experimentation and deployment at scale. We value hands-on expertise with LLMs, reinforcement learning, and agentic systems, along with strong Python and software engineering fundamentals. #J-18808-Ljbffr ...

Controllable Video Generation Researcher

Hiring Organisation
Tencent
Location
London, United Kingdom
Salary
£ 70 K
state retention.6. Collaborate with the Engine/Agent teams to define standardized interfaces for controllable signals.7. Drive the joint optimization of controllable generation and Reinforcement Learning post-training.Who We Look For1. Ph.D. in AI-related fields, focusing on Computer Vision/Deep Learning.2. Proficient in conditional generation models ...

Senior Software Engineer, Data Platform London, UK Apply →

Location
Greater London, England, United Kingdom
enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human evaluation and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT ...

Senior Software Engineer, Data Platform

Hiring Organisation
Scale AI
Location
London, United Kingdom
Salary
£ 80 K
enterprises figure out how to add it to their products. To make them safe, aligned and actually useful, these models need human evaluation and reinforcement learning through human feedback (RLHF) during pre-training, fine-tuning, and production evaluations. This is the main innovation that’s enabled ChatGPT ...

Senior RL Engineer for Humanoid Robots (Equity)

Location
Greater London, England, United Kingdom
Humanoid is seeking a Senior or Staff Reinforcement Learning Engineer to develop learning-based control policies for humanoid robots. You will design and train RL policies enabling dynamic locomotion and loco-manipulation on real robots, building scalable pipelines and improving sim-to-real transfer for reliable hardware ...

Research Engineer

Location
Tipton, England, United Kingdom
product‐minded research engineer who enjoys building real systems people depend on. You’ll likely have: Strong technical background in software engineering, machine learning, or applied AI, demonstrated through an advanced degree and/or equivalent experience building production AI systems Strong software engineering fundamentals and good judgment … Practical understanding of modern model adaptation and post‐training methods, including LoRA/QLoRA, SFT, distillation, preference optimization, reward modeling, DPO/GRPO, and reinforcement learning from verifiable feedback Ownership mindset: you drive projects end‐to‐end, move quickly from real usage, and care about shipping measurable improvements ...

Director of Applied AI

Hiring Organisation
iCOMAT
Location
Gloucester, Gloucestershire, United Kingdom
Salary
£ 60 K
write code, design systems, and lead by example. You are not allergic to a whiteboard or a shop floor.You have shipped applied machine learning solutions in a real industrial, manufacturing, robotics, or hard tech setting. Not just dashboards. Not just demos.You speak fluent KPI. You can defend … production deployment experience at scale.Great plus:Background in composites, advanced manufacturing, aerospace, automotive, or defense.Experience with computer vision applied to automated inspection, generative design, reinforcement learning for process control.Prior experience standing up an AI function from zero inside a fast moving company.Our Values: Uncompromising ExcellenceDiscipline over excuses. ...

Platform Engineer

Location
Greater London, England, United Kingdom
superlearning capability - the ability to endlessly discover knowledge and skills, without relying on human data - will be driven by the world’s most powerful reinforcement learning algorithms. The superlearner is expected to rediscover and then transcend the greatest inventions in human history, such as language, science, mathematics … role where you inherit someone else's architecture, you'll be one of the people defining it. We're pushing these learning methods to a scale that hasn’t been tried before and we care deeply about the infrastructure that gets us there. You'll own the infrastructure that ...

Agent Product Manager

Location
Greater London, England, United Kingdom
usage and model performance, identify key issues, validate hypotheses, and continuously improve product experience and Agent capabilities. Requirements - Basic understanding of LLMs, AI Agents, reinforcement learning, and model training or post-training workflows. - Strong interest in general-purpose Agent products and enthusiasm for building product-data flywheels. - Experience … evaluation, RLHF/RLAIF, data annotation platforms, or model training data pipelines is a strong plus. Basic understanding of LLMs, AI Agents, and reinforcement learning.Strong product judgment, with the ability to identify user needs, define problems, and design practical product solutions.Strong structured thinking and analytical skills, with the ability ...

Simulation Engineer

Hiring Organisation
Helsing
Location
London, United Kingdom
Salary
£ 70 K
make a meaningful impact in the worldOur work frequently takes us right up to the state of the art in technical innovation, be it reinforcement learning, distributed systems, generative AI, or deployment infrastructure. The defence industry is entering the most exciting phase of the technological development curve. Advances ...

Video Generation Content Understanding and Feedback Researcher

Hiring Organisation
Tencent
Location
London, United Kingdom
Salary
£ 70 K
track industry-related work and formulate proprietary technical roadmaps.Who We Look For1.Ph.D. in AI-related fields, covering areas such as video understanding, video prediction, reinforcement learning, or multimodal generation and understanding.2.Deep understanding of the internal mechanisms of both video generation and VLM models; familiar with diffusion model principles.3.Research … generation.4.Ability to independently design complex model training pipelines with strong engineering skills.5.Proficient in Python/PyTorch, with publications in video understanding, controllable generation, or reinforcement learning.6.Experience with Game AI, simulation environments, or closed-loop systems is preferred.7.Relevant publications in top-tier conferences or journals are preferred.Equal Employment Opportunity ...

Senior Data Scientist in ADVANCED ANALYTICS

Location
Leeds, England, United Kingdom
opportunity to work at the heart of data science, research and public policy. Advanced Analytics is hiring a Senior Data Scientist to develop machine learning (ML) and Artificial Intelligence (AI) solutions, with a focus on robust prototyping. We have pioneered the application of new analytical approaches in the Bank … England, such as generative AI, machine learning, natural language processing, agent-based models, network analytics, and randomised control trials. We’re looking for someone with ML/AI expertise to help us continue pushing the organisation’s analytical frontier to inform policymaking. DAT is leading the transformation ...

Senior Data Scientist in ADVANCED ANALYTICS

Location
City Of London, England, United Kingdom
opportunity to work at the heart of data science, research and public policy. Advanced Analytics is hiring a Senior Data Scientist to develop machine learning (ML) and Artificial Intelligence (AI) solutions, with a focus on robust prototyping. We have pioneered the application of new analytical approaches in the Bank … England, such as generative AI, machine learning, natural language processing, agent-based models, network analytics, and randomised control trials. We’re looking for someone with ML/AI expertise to help us continue pushing the organisation’s analytical frontier to inform policymaking. DAT is leading the transformation ...

Vice President, Model Research & Development

Hiring Organisation
Thomson Reuters
Location
London, United Kingdom
Salary
£ 80 K
release decision) for models that ship to real users across CoCounsel, Westlaw, and Practical Law. Set the target for an online, agentic reinforcement-learning pipeline, training with subject-matter experts in the loop, and hold the work accountable to it. Ensure data strategy and evaluation approach are rigorous … weeks per year, empowering employees to achieve a better work-life balance.Career Development and Growth: By fostering a culture of continuous learning and skill development, we prepare our talent to tackle tomorrow’s challenges and deliver real-world solutions. Our Grow My Way programming and skills-first approach ensures ...

Onsite RL Engineer – Robotics & Embodied AI

Location
Greater London, England, United Kingdom
Randstad Technologies Recruitment is searching for a Reinforcement Learning (RL) Engineer to join a leading robotics company in London. This permanent position involves working collaboratively to define the company's 2026 technical roadmap while solving complex challenges in autonomous systems. The role requires expertise in AI, MLOps, Software ...