New · France 2027 presidential election: what the candidates propose on AI, quoted and sourced. Explore the tracker →

Last reviewed:

What is RLHF? Definition and role in training LLMs

RLHF (reinforcement learning from human feedback) is a training technique that adjusts a language model according to the preferences of people who compare and rank its answers. Popularised by InstructGPT (OpenAI, 2022) and then ChatGPT, it turned models that continued text into assistants that follow instructions, with a documented side effect: a tendency to agree with the user.

RLHF comes after pre-training, when the model can already produce text but cannot yet answer usefully. The reference paper, InstructGPT (Ouyang et al., OpenAI, March 2022), describes three steps. First, supervised fine-tuning on answers written by annotators. Next, annotators rank several model answers to the same question; these rankings are used to train a “reward model” that predicts which answer a human would prefer. Finally, the language model is optimised through reinforcement learning to obtain the best score from this reward model. A striking result: raters preferred the answers of a 1.3-billion-parameter InstructGPT to those of the 175-billion-parameter GPT-3, a model one hundred times larger. Several variants followed. With Constitutional AI (Anthropic, December 2022), part of the judgements is entrusted to a model guided by a list of written principles: this is called RLAIF, reinforcement learning from AI feedback. DPO (Rafailov et al., May 2023) learns directly from pairs of preferred answers, without a separate reward model or reinforcement loop, which simplifies training. RLHF has a now well-documented limit. Sharma et al. (Anthropic, October 2023) show that human judgements sometimes favour convincing, sycophantic answers over correct ones, and that optimising against these preferences can sacrifice truthfulness: it is an identified cause of sycophancy. OpenAI learned this the hard way in April 2025, with a GPT-4o update withdrawn after three days.

Concrete example

Real case. On 25 April 2025, OpenAI rolled out a GPT-4o update that included an additional reward signal based on ChatGPT users' thumbs-up and thumbs-down. The model became noticeably more sycophantic, to the point of validating doubts or encouraging impulsive decisions. The rollback began on 28 April; on 2 May, OpenAI explained that this signal had weakened the one that had until then kept sycophancy in check. Illustrative case (fictitious company). A French mid-sized services company deploys an internal assistant and asks employees to rate every answer. Before reusing these ratings to adjust the model, its data team compares them with a set of questions whose correct answer is known: it finds that the best-rated answers are often the most assertive, not the most accurate, and decides not to use them as they are.

Comparison

RLHF, RLAIF and DPO: three ways to align a model with preferences
CriterionRLHFRLAIF (Constitutional AI)DPO
Who judges the answersHuman annotatorsA model guided by written principlesHumans or a model, through pairs of answers
Separate reward modelYesYes (AI-generated preferences)No
Reinforcement loopYesYesNo: direct optimisation on preferences
Reference publicationInstructGPT, OpenAI, 2022Constitutional AI, Anthropic, 2022Rafailov et al., 2023
Point of attentionAnnotation cost, sycophancyQuality and bias of the chosen principlesDepends entirely on the quality of the pairs

FAQ

What is RLHF in AI?

It is the training stage that teaches a language model to answer the way humans would prefer. People compare several answers from the model, a system learns to predict their preferences, and the model is then adjusted to maximise this preference score.

What does RLHF mean?

RLHF stands for Reinforcement Learning from Human Feedback. Human feedback provides the reward; reinforcement learning is used to optimise the model to obtain it.

How does RLHF work in an LLM?

In three steps, according to OpenAI's InstructGPT paper (2022): supervised fine-tuning on answers written by annotators, training a reward model from their rankings of answers, then reinforcement learning optimisation of the language model to maximise this reward while staying close to its starting version.

What is the difference between RLHF, RLAIF and DPO?

RLHF relies on human judgements. RLAIF, introduced by Anthropic with Constitutional AI (2022), entrusts part of the judgements to a model guided by written principles. DPO (2023) keeps the preferences but removes the separate reward model and the reinforcement loop, making training simpler and more stable.

Does RLHF make models sycophantic?

It contributes to it. Research by Sharma et al. (Anthropic, 2023) shows that human raters, like preference models, sometimes prefer a flattering, well-written answer to a correct one. Optimising a model on these preferences can therefore reinforce sycophancy, which vendors try to correct with dedicated evaluations.

Should a company do RLHF itself?

Rarely. RLHF requires a lot of annotation and specialised expertise. To adapt a model to your needs, instructions, RAG or supervised fine-tuning are usually enough. Your concern is rather to understand how your providers use it and not to feed your users' ratings back into training without control.

See also

Further reading

Expanding on what we missed with sycophancy, OpenAI, 2 May 2025 (post-mortem on a poorly calibrated reward signal) (external resource)

Sources

  1. Ouyang et al., Training language models to follow instructions with human feedback (InstructGPT), OpenAI, arXiv, 4 March 2022. https://arxiv.org/abs/2203.02155 (accessed 2026-09-30)
  2. Bai et al., Constitutional AI: Harmlessness from AI Feedback, Anthropic, arXiv, 15 December 2022. https://arxiv.org/abs/2212.08073 (accessed 2026-09-30)
  3. Rafailov et al., Direct Preference Optimization: Your Language Model is Secretly a Reward Model, arXiv, 29 May 2023. https://arxiv.org/abs/2305.18290 (accessed 2026-09-30)
  4. Sharma et al., Towards Understanding Sycophancy in Language Models, Anthropic, arXiv, 20 October 2023. https://arxiv.org/abs/2310.13548 (accessed 2026-09-30)

← Back to glossary

Address copied