Intermediate level8 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain how reasoning models spend extra computing while they answer, choose when a higher effort setting is worth its cost, and say why the visible thinking is not proof.

The short answerReasoning models spend extra computing while they answer: they generate a long, mostly hidden chain of thought first. They learn this through reinforcement learning on tasks with checkable answers, and an effort setting trades quality against cost and waiting.

In simple words

  • Chain-of-thought prompting (2022) showed that writing out steps helps large models with maths and logic.
  • DeepSeek’s R1-Zero learned long reasoning through rewards for right answers, on top of an ordinary base model.
  • Thinking tokens are billed as output, even when you see only a summary.
  • Gains are uneven: large for maths and code, small or even negative for some easy or intuitive tasks.
Thinking before answeringIllustration + live data

1 Same question, two ways

Answers at once

QuestionAnswer

Thinks first

QuestionTry a planCheck a stepSpot a mistakeFix itAnswer

Thinking tokens are billed like the answer, and they take time. They help most on hard problems whose answer can be checked, such as maths and code.

2 FrontierMath, today’s leaders

  1. GPT-6.1 Sol (max)OpenAI · 93.7% solved
  2. GPT-6 Astra (max)OpenAI · 93.7% solved
  3. Claude Opus 5.5 (max)Anthropic · 91.2% solved
See the full comparison
The two paths are a simplified picture. The leaders come from our model comparison of October 5, 2026: FrontierMath Tiers 1–3, run by Epoch AI. Epoch AI runs each model on hundreds of new, unpublished problems written by mathematicians. OpenAI funded the benchmark and has access to part of the problem set. “Max” is the highest thinking setting the maker offers.

Words to know

Chain of thought
The written-out steps a model produces before its answer.
Test-time compute
Computing a model spends while it answers, not while it trains.
Reasoning effort
A setting for how much the model thinks, trading quality against cost and waiting.

From prompting to training

In 2022, Wei and others showed that examples with worked steps, called chain-of-thought prompting, improved maths and logic in large models. Kojima and others found that adding “Let’s think step by step” raised one model from 17.7% to 78.7% on a maths word-problem test.

Reasoning models make this a trained habit. OpenAI says its o-series models learn through large-scale reinforcement learning, training by trial and reward, to reason in a chain of thought, trying strategies and spotting mistakes. It has not published the data, the algorithm or the model size.

The published recipe

DeepSeek published its method in January 2025. It took its ordinary base model, DeepSeek-V3-Base, and trained it with rewards set by simple rules: a right maths answer, code that passes tests, and thinking kept inside marked tags.

That model, R1-Zero, learned to think for longer by itself. Its score on AIME 2024, a hard maths contest, rose from 15.6% to 71.0%. Its writing mixed languages and read poorly, so the released R1 added starter examples and more training.

Effort settings and cost

Most makers let you choose how much the model thinks, and so how much test-time compute it uses. OpenAI’s effort setting runs from none to max, and Anthropic offers low, medium, high, xhigh and max. The hidden thinking tokens are billed as output tokens.

More effort is not always better value. Anthropic notes that bigger thinking budgets add delay with shrinking returns, and researchers found that o1-like models overthink easy problems.

As of 8 October 2026, Artificial Analysis shows about $5.46 per task for Claude Sonnet 5.5 at max effort, against $0.48 at medium. Our model comparison shows each model with its effort setting, such as “(max)”; on FrontierMath, a hard maths test, the top results come from max-effort settings.

Limits

A review of more than 100 papers found that chain of thought helped mainly with maths and logic. On some tasks where people also do worse when they stop to think, o1-preview scored up to 36 points below GPT-4o.

The thinking is not a full record. In tests by Anthropic researchers, when reasoning models used a hint, they mentioned it less than 20% of the time in many settings. Reasoning also does not end made-up answers: on one OpenAI test, o3 made things up more often than o1.

Try it yourself

In a chat app with a thinking or effort switch, ask: “A bat and a ball cost $1.10. The bat costs $1.00 more than the ball. What does the ball cost?” Try low, then high; if both are right, compare the time. Then ask for the capital of France on high: is it better, or only slower?

Check yourself

  1. A support team runs a reasoning model at max effort to sort emails into four folders. What is the likely problem?

  2. What did DeepSeek reward when it trained R1-Zero?

  3. Researchers slipped a hint into a question, and the model used it. What did its visible thinking usually show?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.