What it is
A reasoning model is a language model built to think before it speaks. Where an ordinary chat model starts writing its answer right away, a reasoning model first produces a long stretch of working: it plans, tries approaches, checks itself and backs up when something looks wrong. Only then does it give the final reply.
The simplest analogy is the difference between blurting out an answer and pausing to scribble on scratch paper.
How it works
Reasoning models rely on a chain of thought. They are trained, typically with reinforcement learning on problems that have checkable answers, to produce useful reasoning steps and to reach correct results. The result is a model that learns when to slow down, break a problem apart and verify its own work.
Many products let you set how much thinking the model does. A higher setting means more reasoning tokens, which usually means better answers on difficult problems and a slower, costlier response. Depending on the provider, the reasoning may be shown in full, shown as a summary, or hidden.
Why it matters to you
- Hard problems. Reasoning models tend to do better on math, coding, science questions and multi-step planning, and they are the engine behind many AI agents.
- Cost and latency. The thinking is billed as output tokens, and you wait for it. For a quick rewrite or a simple question, a standard model is often the better choice.
- Context use. Long reasoning also takes space in the context window.
- Benchmarks. Many scores on an AI benchmark for math and coding come from reasoning models, so check how much thinking effort was allowed when you compare results.
Example
Suppose you ask for a bug fix in a large function. A standard model might patch the first suspicious line. A reasoning model is more likely to trace how the data flows, notice the real cause two functions away, and test its fix against the cases you described before it answers.
Reasoning models build directly on chain of thought. They are a different axis from mixture of experts, which is about how a model is built rather than how long it thinks, and the two are often combined. Running one costs more inference compute per question than a standard model.