New · France 2027 presidential election: what the candidates propose on AI, quoted and sourced. Explore the tracker →

Last reviewed:

What is reinforcement learning? Definition, examples and role in generative AI

Reinforcement learning is a machine learning method in which a system learns which actions to choose by trying them and receiving a reward or a penalty, so as to maximise a cumulative gain over time. Long associated with games and robotics, it is also used to adjust the behaviour of conversational assistants (RLHF) and to train reasoning models.

The CNIL, the French data protection authority, defines it as a machine learning process in which an autonomous system learns the actions to take, from experience, so as to optimise a quantitative reward over time. There are no corrected examples, as in supervised learning, and no mere data exploration, as in unsupervised learning: an “agent” acts in an environment, observes the result, receives a reward signal and adjusts its strategy. Andrew Barto and Richard Sutton laid its conceptual and algorithmic foundations from the 1980s onwards, earning them the 2024 Turing Award, announced by the ACM on 5 March 2025. The general public discovered it with AlphaGo: DeepMind's program, first trained on expert games and then by playing against itself thousands of times, beat champion Lee Sedol 4 games to 1 in March 2016 in Seoul. Since then, reinforcement learning has entered generative AI through two doors. The first is RLHF, which adjusts a language model according to the preferences of human raters; the ACM notes that ChatGPT was trained in two phases, the second of which uses this technique. The second is the training of reasoning models: in January 2025, in a paper later published in Nature, DeepSeek showed that the reasoning abilities of a language model can be developed through pure reinforcement learning, without human-annotated reasoning examples, with the reward coming from verifiable answers. The ACM also cites industrial uses: supply chain optimisation, chip design, online advertising, network congestion control. Its main limit: the system optimises exactly the reward it is given, not the intention of whoever defined it.

Concrete example

Illustrative case (fictitious company). A French chain of 25 DIY stores wants to improve its replenishment. Rather than forecasting demand product by product, its provider trains a reinforcement learning agent in a simulator built from two years of sales: at each cycle, the agent decides the quantities to order and receives a reward combining revenue achieved, stock-outs avoided and inventory cost. After the simulation phase, the agent's proposals are compared for three months with those of the buyers in a few pilot stores, who keep control over validation. For most SMEs, however, reinforcement learning remains indirect: it is already present in the AI assistants they use, tuned with RLHF, and in reasoning models.

Comparison

Three uses of reinforcement learning, from AlphaGo to reasoning models
CriterionClassic reinforcement learningRLHFReinforcement with verifiable rewards
Source of the rewardThe environment: game score, cost, lead timeA model trained on human preferencesAn automatic check: correct answer, code that passes tests
ExampleAlphaGo (2016), robotics, logisticsInstructGPT (2022), ChatGPTDeepSeek-R1 (2025), reasoning models
What is optimisedAn action strategyThe style and perceived usefulness of answersAccuracy on problems with verifiable solutions
Main riskUnexpected shortcuts to maximise the scoreSycophancy towards the userPerformance limited to verifiable domains

FAQ

What is reinforcement learning in simple terms?

It is learning by trial and error, guided by a reward. The system tries an action, sees whether the result is good or bad according to a score defined in advance, and adjusts its strategy to get a better score next time. It resembles the way an animal is trained with treats.

What is an example of reinforcement learning?

The best-known example is DeepMind's AlphaGo, which improved by playing against itself thousands of times, then beat champion Lee Sedol in March 2016. Other examples: robots learning movements in simulation, logistics optimisation, online advertising, and the tuning of assistants such as ChatGPT with RLHF.

How does it differ from supervised and unsupervised learning?

Supervised learning learns from examples whose right answer is provided. Unsupervised learning looks for structure in data without answers. Reinforcement learning has no ready-made answer: it only receives a signal (reward or penalty) after each action, and must discover the best long-term strategy itself.

How is it related to ChatGPT and reasoning models?

There are two direct links. RLHF, a form of reinforcement learning guided by human preferences, was used to make language models useful and able to follow instructions. And reasoning models, such as DeepSeek-R1 (January 2025), are trained by reinforcement on problems whose answer can be verified, which pushes them to produce longer reasoning before answering.

Who invented reinforcement learning?

The idea of learning from reward is old; Alan Turing mentioned it as early as 1950. Andrew Barto and Richard Sutton formalised its framework and main algorithms from the 1980s, and published the field's reference textbook in 1998. They received the 2024 Turing Award for this work, announced on 5 March 2025.

Where can I find a course on reinforcement learning?

Sutton and Barto's textbook, Reinforcement Learning: An Introduction (2nd edition, 2018), is freely available online and remains the reference. It assumes a mathematical background. For an executive, understanding the agent, environment and reward triad is enough to assess a project.

See also

Further reading

Reinforcement Learning: An Introduction, Richard S. Sutton and Andrew G. Barto, 2nd edition, MIT Press, 2018 (reference textbook, available online) (external resource)

Sources

  1. Apprentissage par renforcement (reinforcement learning), definition, CNIL (French data protection authority). https://www.cnil.fr/fr/definition/apprentissage-par-renforcement (accessed 2026-09-30)
  2. ACM A.M. Turing Award Honors Two Researchers Who Led the Development of Cornerstone AI Technology, ACM press release, 5 March 2025. https://www.acm.org/media-center/2025/march/turing-award-2024 (accessed 2026-09-30)
  3. AlphaGo, overview page, Google DeepMind. https://deepmind.google/research/alphago/ (accessed 2026-09-30)
  4. DeepSeek-AI, DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, arXiv, 22 January 2025 (published in Nature, vol. 645, 2025). https://arxiv.org/abs/2501.12948 (accessed 2026-09-30)

← Back to glossary

Address copied