New · France 2027 presidential election: what the candidates propose on AI, quoted and sourced. Explore the tracker →

Last reviewed:

What is a reasoning model? Definition and difference from a standard LLM

A reasoning model is a large language model trained to produce internal, step-by-step thinking before giving its answer. It is more reliable on complex problems (analysis, calculation, code, agent planning), but slower and more expensive, because its thinking tokens are billed.

The first mainstream reasoning model was OpenAI o1, released as a preview on September 12, 2024: OpenAI described it as designed to spend more time thinking before responding. On a qualifying exam for the International Mathematics Olympiad, GPT-4o solved 13% of problems, the reasoning model 83%. In January 2025, DeepSeek published R1 and showed that reinforcement learning alone can make this behaviour emerge, without human-written reasoning examples. The principle is simple. Before the visible answer, the model generates a chain of thought: it breaks down the problem, tries several approaches, checks its calculations, corrects its mistakes. These thinking tokens are largely invisible, but they occupy the context window and are billed as output tokens (OpenAI and Anthropic documentation). In 2026, the line between standard LLMs and reasoning models is blurring: flagship models decide for themselves whether to think. At Anthropic, Claude's thinking is adaptive. The model assesses each request, answers a simple question directly and thinks longer on a multi-step problem. Developers set an “effort” level (low, medium, high, xhigh, max; medium by default on Claude Opus 5.5). OpenAI offers an equivalent setting, reasoning effort, from none to max. The higher the effort, the more careful the answer, and the slower and more expensive it is. “Reasoning” remains a metaphor: the model does not think like a human and can, after long deliberation, reach a wrong answer stated with confidence.

Concrete example

Illustrative case: a French retail company with 600 employees uses an AI assistant for two tasks. The first is sorting 3,000 supplier emails a day (invoice, dispute, reminder). The second is preparing a monthly analysis of margin variances by product family, from several accounting exports. For email sorting, thinking effort set to the minimum gives the same quality as high effort, with answers in one to two seconds. For the margin analysis, the same model at high effort spots a unit conversion error that the no-thinking version missed. The finance department therefore keeps two settings: low effort for volume, high effort for the monthly analysis, whose extra cost stays marginal since it runs once a month.

Comparison

Standard LLM vs reasoning model: what changes for the business
Standard LLM (direct answer)Reasoning model
How it worksGenerates the answer immediatelyGenerates internal thinking, then the answer
Cost per requestLowerHigher: thinking tokens are billed as output
LatencyOne to a few secondsFrom a few seconds to several minutes depending on effort
Suitable use casesSorting, extraction, rewording, short answersNumerical analysis, code, compliance, agent planning
Main riskErrors on multi-step problemsExtra cost and slowness when thinking is unnecessary
SettingChoice of modelEffort level (from low to maximum)

FAQ

What is a reasoning model in AI?

It is a large language model trained to think step by step before answering. It generates internal thinking (breaking down the problem, trying approaches, checking), then its answer. It is more reliable on multi-step problems, slower and more expensive per request.

What is the difference between a reasoning model and a standard LLM?

A standard LLM produces its answer directly, token by token. A reasoning model first produces thinking tokens, billed as output tokens. In 2026, most flagship models do both: they decide for themselves whether to think, according to an effort level set by the developer.

What are examples of reasoning models?

OpenAI o1 (preview in September 2024) led the way, followed by DeepSeek-R1 (January 2025), released with open weights. In 2026, Anthropic's Claude Opus 5.5 includes adaptive thinking controlled by effort level, and OpenAI recommends GPT-6 Astra for most reasoning tasks.

Do reasoning models make fewer mistakes?

They make fewer mistakes on multi-step problems (calculation, logic, code, planning), but they are not infallible. Long thinking can end in a wrong answer stated with confidence. Source checks and human review remain necessary on high-stakes topics.

Why does a reasoning model cost more?

Because internal thinking consumes tokens, billed as output tokens even when you do not see them. On a hard problem, thinking can account for most of the billed tokens. The effort level is the main lever to control this cost.

Should reasoning always be switched on?

No. For sorting, extraction or rewording, low effort is usually enough and cuts cost and latency. Keep high effort for tasks where an error is costly, and measure the quality gap on your own cases before deciding.

See also

Further reading

Steering thinking (effort, adaptive thinking), Claude documentation, Anthropic (external resource)

Sources

  1. Introducing OpenAI o1-preview, OpenAI, September 12, 2024. https://openai.com/index/introducing-openai-o1-preview/ (accessed 2026-09-30)
  2. DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning, DeepSeek-AI, arXiv, January 22, 2025. https://arxiv.org/abs/2501.12948 (accessed 2026-09-30)
  3. Reasoning models (reasoning effort, billing of reasoning tokens), OpenAI documentation. https://developers.openai.com/api/docs/guides/reasoning (accessed 2026-09-30)
  4. Steering thinking (effort levels, adaptive thinking, pricing), Claude documentation, Anthropic. https://platform.claude.com/docs/en/build-with-claude/thinking-steering-and-cost (accessed 2026-09-30)

← Back to glossary

Address copied