Beginner level7 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain how reasoning models spend extra computing while they answer, choose when a higher effort setting is worth its cost, and say why the visible thinking is not proof.

The short answerA reasoning model writes out hidden working steps before it gives its answer. That takes more time and costs more, but it helps most with problems whose answer can be checked, such as maths or code.

In simple words

  • Before answering, the model “thinks”: it writes steps that you usually do not see.
  • It learned this by getting rewards when its final answer was right.
  • Thinking costs money and time, so it is worth it for hard problems, not for quick facts.
  • The thinking summary you see is not a full record of how it reached the answer.
Thinking before answeringIllustration + live data

1 Same question, two ways

Answers at once

QuestionAnswer

Thinks first

QuestionTry a planCheck a stepSpot a mistakeFix itAnswer

Thinking tokens are billed like the answer, and they take time. They help most on hard problems whose answer can be checked, such as maths and code.

2 FrontierMath, today’s leaders

  1. GPT-6.1 Sol (max)OpenAI · 93.7% solved
  2. GPT-6 Astra (max)OpenAI · 93.7% solved
  3. Claude Opus 5.5 (max)Anthropic · 91.2% solved
See the full comparison
The two paths are a simplified picture. The leaders come from our model comparison of October 5, 2026: FrontierMath Tiers 1–3, run by Epoch AI. Epoch AI runs each model on hundreds of new, unpublished problems written by mathematicians. OpenAI funded the benchmark and has access to part of the problem set. “Max” is the highest thinking setting the maker offers.

Words to know

Chain of thought
The written-out steps a model produces before its answer.
Test-time compute
Computing a model spends while it answers, not while it trains.
Reasoning effort
A setting for how much the model thinks, trading quality against cost and waiting.

Working it out first

Ask a person a hard sum, and they write the steps down. Researchers found in 2022 that language models also do better at maths when they write out steps first. Even adding “Let’s think step by step” to a question helped.

Reasoning models are trained to do this by themselves. Before the answer, they write a long chain of steps, called a chain of thought: they try ideas, spot mistakes and fix them. OpenAI’s o1-preview, launched in September 2024, was one of the first.

This extra computing while the model answers is called test-time compute. It comes on top of the computing that went into training.

Rewards for right answers

How do they learn it? In training, the model tries many problems whose answers can be checked, such as a maths result or a program that must pass its tests. Right answers earn a reward, so ways of thinking that work become more likely.

DeepSeek published this method in January 2025. Its model learned to think for longer by itself, and its score on a hard maths contest rose from about 16% to 71% during training.

When thinking is worth it

Thinking is billed like the answer, and it takes time. Many apps let you choose how much the model thinks, a setting called reasoning effort, with levels such as low, medium and high.

The difference can be large. In one outside test, a model at its top setting cost about 11 times as much per task as at its medium setting.

Google advises little thinking for simple lookups and sorting, and more for hard coding, maths and planning. On easy questions, extra thinking adds waiting and cost for little gain.

Most apps show only a summary of the thinking. Research found that models often leave out things that shaped their answer. So the thinking you see is not a full record, and it does not prove the answer is right.

Try it yourself

In a chat app with a thinking or effort switch, ask: “A bat and a ball cost $1.10. The bat costs $1.00 more than the ball. What does the ball cost?” Try low, then high; if both are right, compare the time. Then ask for the capital of France on high: is it better, or only slower?

Check yourself

  1. You need the capital of France, fast. Which setting makes most sense?

  2. How do reasoning models learn to think well?

  3. An app shows you the model’s thinking before its answer. What does that tell you?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.