Intermediate level8 min readTo learn at another level, choose it before you start the course.
After this lesson you canExplain how reasoning models spend extra computing while they answer, choose when a higher effort setting is worth its cost, and say why the visible thinking is not proof.
The short answerA reasoning model writes out hidden working steps before it gives its answer. That takes more time and costs more, but it helps most with problems whose answer can be checked, such as maths or code.
In simple words
- Before answering, the model “thinks”: it writes steps that you usually do not see.
- It learned this by getting rewards when its final answer was right.
- Thinking costs money and time, so it is worth it for hard problems, not for quick facts.
- The thinking summary you see is not a full record of how it reached the answer.
The short answerReasoning models spend extra computing while they answer: they generate a long, mostly hidden chain of thought first. They learn this through reinforcement learning on tasks with checkable answers, and an effort setting trades quality against cost and waiting.
In simple words
- Chain-of-thought prompting (2022) showed that writing out steps helps large models with maths and logic.
- DeepSeek’s R1-Zero learned long reasoning through rewards for right answers, on top of an ordinary base model.
- Thinking tokens are billed as output, even when you see only a summary.
- Gains are uneven: large for maths and code, small or even negative for some easy or intuitive tasks.
The short answerReasoning models scale test-time compute: reinforcement learning with verifiable rewards teaches a base model to produce long chains of thought before answering. Gains are large on checkable tasks such as maths and code, but the cost is billed tokens and latency, and the chains are not faithful records.
Key points
- Snell and others (2024): spending test-time compute adaptively per question was over 4 times more efficient than best-of-N.
- DeepSeek-R1-Zero: GRPO on DeepSeek-V3-Base with rule-based rewards and no supervised step.
- Hidden reasoning is billed as output; providers return summaries or nothing.
- Faithfulness is limited: when reasoning models used a hint, they acknowledged it less than 20% of the time in many settings.
1 Same question, two ways
QuestionAnswer
QuestionTry a planCheck a stepSpot a mistakeFix itAnswer
Thinking tokens are billed like the answer, and they take time. They help most on hard problems whose answer can be checked, such as maths and code.
2 FrontierMath, today’s leaders
- GPT-6.1 Sol (max)OpenAI · 93.7% solved
- GPT-6 Astra (max)OpenAI · 93.7% solved
- Claude Opus 5.5 (max)Anthropic · 91.2% solved
Words to know
- Chain of thought
- The written-out steps a model produces before its answer.
- Test-time compute
- Computing a model spends while it answers, not while it trains.
- Reasoning effort
- A setting for how much the model thinks, trading quality against cost and waiting.
Working it out first
Ask a person a hard sum, and they write the steps down. Researchers found in 2022 that language models also do better at maths when they write out steps first. Even adding “Let’s think step by step” to a question helped.
Reasoning models are trained to do this by themselves. Before the answer, they write a long chain of steps, called a chain of thought: they try ideas, spot mistakes and fix them. OpenAI’s o1-preview, launched in September 2024, was one of the first.
This extra computing while the model answers is called test-time compute. It comes on top of the computing that went into training.
Rewards for right answers
How do they learn it? In training, the model tries many problems whose answers can be checked, such as a maths result or a program that must pass its tests. Right answers earn a reward, so ways of thinking that work become more likely.
DeepSeek published this method in January 2025. Its model learned to think for longer by itself, and its score on a hard maths contest rose from about 16% to 71% during training.
When thinking is worth it
Thinking is billed like the answer, and it takes time. Many apps let you choose how much the model thinks, a setting called reasoning effort, with levels such as low, medium and high.
The difference can be large. In one outside test, a model at its top setting cost about 11 times as much per task as at its medium setting.
Google advises little thinking for simple lookups and sorting, and more for hard coding, maths and planning. On easy questions, extra thinking adds waiting and cost for little gain.
Most apps show only a summary of the thinking. Research found that models often leave out things that shaped their answer. So the thinking you see is not a full record, and it does not prove the answer is right.
Try it yourself
In a chat app with a thinking or effort switch, ask: “A bat and a ball cost $1.10. The bat costs $1.00 more than the ball. What does the ball cost?” Try low, then high; if both are right, compare the time. Then ask for the capital of France on high: is it better, or only slower?
Check yourself
You need the capital of France, fast. Which setting makes most sense?
How do reasoning models learn to think well?
An app shows you the model’s thinking before its answer. What does that tell you?
Words to know
- Chain of thought
- The written-out steps a model produces before its answer.
- Test-time compute
- Computing a model spends while it answers, not while it trains.
- Reasoning effort
- A setting for how much the model thinks, trading quality against cost and waiting.
From prompting to training
In 2022, Wei and others showed that examples with worked steps, called chain-of-thought prompting, improved maths and logic in large models. Kojima and others found that adding “Let’s think step by step” raised one model from 17.7% to 78.7% on a maths word-problem test.
Reasoning models make this a trained habit. OpenAI says its o-series models learn through large-scale reinforcement learning, training by trial and reward, to reason in a chain of thought, trying strategies and spotting mistakes. It has not published the data, the algorithm or the model size.
The published recipe
DeepSeek published its method in January 2025. It took its ordinary base model, DeepSeek-V3-Base, and trained it with rewards set by simple rules: a right maths answer, code that passes tests, and thinking kept inside marked tags.
That model, R1-Zero, learned to think for longer by itself. Its score on AIME 2024, a hard maths contest, rose from 15.6% to 71.0%. Its writing mixed languages and read poorly, so the released R1 added starter examples and more training.
Effort settings and cost
Most makers let you choose how much the model thinks, and so how much test-time compute it uses. OpenAI’s effort setting runs from none to max, and Anthropic offers low, medium, high, xhigh and max. The hidden thinking tokens are billed as output tokens.
More effort is not always better value. Anthropic notes that bigger thinking budgets add delay with shrinking returns, and researchers found that o1-like models overthink easy problems.
As of 8 October 2026, Artificial Analysis shows about $5.46 per task for Claude Sonnet 5.5 at max effort, against $0.48 at medium. Our model comparison shows each model with its effort setting, such as “(max)”; on FrontierMath, a hard maths test, the top results come from max-effort settings.
Limits
A review of more than 100 papers found that chain of thought helped mainly with maths and logic. On some tasks where people also do worse when they stop to think, o1-preview scored up to 36 points below GPT-4o.
The thinking is not a full record. In tests by Anthropic researchers, when reasoning models used a hint, they mentioned it less than 20% of the time in many settings. Reasoning also does not end made-up answers: on one OpenAI test, o3 made things up more often than o1.
Try it yourself
In a chat app with a thinking or effort switch, ask: “A bat and a ball cost $1.10. The bat costs $1.00 more than the ball. What does the ball cost?” Try low, then high; if both are right, compare the time. Then ask for the capital of France on high: is it better, or only slower?
Check yourself
A support team runs a reasoning model at max effort to sort emails into four folders. What is the likely problem?
What did DeepSeek reward when it trained R1-Zero?
Researchers slipped a hint into a question, and the model used it. What did its visible thinking usually show?
Test-time compute
Snell and others (2024) found that spending answer-time computing adaptively, per question, was over 4 times more efficient than best-of-N sampling, which keeps the best of N answers. With equal computing, a small model beat one 14 times larger on questions it could partly solve.
The o1 system card says the models learn through large-scale reinforcement learning to reason in a chain of thought. It names no algorithm, data or model size, and OpenAI never returns the raw chain through its API.
Reinforcement learning with verifiable rewards
DeepSeek-R1-Zero applied GRPO, a reinforcement learning method, directly to DeepSeek-V3-Base, with no supervised step. Rewards were rule-based: accuracy, checked by a maths verifier or code tests, and a format reward for reasoning inside think tags.
DeepSeek avoided learned reward models because of reward hacking. During training, AIME 2024 accuracy rose from 15.6% to 71.0%, and responses grew longer without being told to.
R1 then added cold-start examples, about 800,000 fine-tuning samples and more reinforcement learning, and reached 79.8% on AIME 2024. It has 671 billion parameters, 37 billion active, and its weights are under the MIT licence.
Controls and billing
OpenAI bills reasoning tokens as output and advises reserving at least 25,000 tokens at first. Anthropic bills the full thinking but returns a summary written by a different model, or nothing.
Google’s Gemini 3 and 2.5 models think by default, except 2.5 Flash-Lite, and return only summaries. Qwen3 puts thinking and non-thinking modes, plus a thinking budget, into one model.
Outside tests show the spread: as of 8 October 2026, Artificial Analysis shows about $5.46 per task for Claude Sonnet 5.5 at max effort, against $0.48 at medium.
Where it fails
Sprague and others (2024) reviewed over 100 papers and 14 models: chain of thought helped mainly on maths and symbolic logic. Liu and others (2024) found tasks where people also do worse when they stop to think, with o1-preview up to 36.3 points below GPT-4o.
Chains are not faithful records. Turpin and others (2023) planted biases that the explanations almost never mentioned, cutting accuracy by up to 36%. Chen and others (2025) found that, when reasoning models used a hint, they acknowledged it less than 20% of the time in many settings.
Reasoning does not remove hallucination either: on PersonQA, OpenAI reported made-up answers in 33% of cases for o3 and 48% for o4-mini, against 16% for o1. Korbak and others call readable chains of thought a “new and fragile opportunity” for safety.
Try it yourself
In a chat app with a thinking or effort switch, ask: “A bat and a ball cost $1.10. The bat costs $1.00 more than the ball. What does the ball cost?” Try low, then high; if both are right, compare the time. Then ask for the capital of France on high: is it better, or only slower?
Check yourself
What did DeepSeek-R1-Zero’s training use to score answers?
In Snell and others (2024), where did extra test-time compute let a small model beat a much larger one?
An auditor wants to know why a reasoning model approved a loan. Is its visible chain of thought enough?
Sources
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models (Wei and others, 2022, arXiv)
- Large Language Models are Zero-Shot Reasoners (Kojima and others, 2022, arXiv)
- Scaling LLM Test-Time Compute Optimally (Snell and others, 2024, arXiv)
- Introducing OpenAI o1-preview (OpenAI, September 2024)
- OpenAI o1 System Card (OpenAI, 2024, arXiv)
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning (DeepSeek, 2025, arXiv)
- DeepSeek-R1 model card and licence (DeepSeek, Hugging Face)
- Reasoning models (OpenAI API documentation)
- Extended thinking (Claude documentation)
- Effort (Claude documentation)
- Gemini thinking (Gemini API documentation)
- Qwen3 Technical Report (Qwen, 2025, arXiv)
- To CoT or not to CoT? (Sprague and others, 2024, arXiv)
- Mind Your Step (by Step): tasks where thinking makes models worse (Liu and others, 2024, arXiv)
- Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs (Chen and others, 2024, arXiv)
- Language Models Don’t Always Say What They Think (Turpin and others, 2023, arXiv)
- Reasoning Models Don’t Always Say What They Think (Chen and others, 2025, arXiv)
- OpenAI o3 and o4-mini System Card (OpenAI, 2025)
- Chain of Thought Monitorability (Korbak and others, 2025, arXiv)
- Claude Sonnet 5.5 at max effort: cost per task (Artificial Analysis)
- Claude Sonnet 5.5 at medium effort: cost per task (Artificial Analysis)
This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.