One request, start to finishIllustration + live data

Cost of 1,000 such answers

  1. GPT-6 Luna$0.350
  2. Gemini 3.7 Flash$2.63
  3. Grok 4.7$5.00
  4. Claude Sonnet 5.5$7.00
  5. Claude Opus 5.5$14.00
  6. Claude Opus 4.8$17.50
  7. GPT-6 Astra$35.00

List prices per million tokens from each company’s pricing page, as in our model comparison of September 30, 2026. Output tokens cost more than input, so long answers add up fast.

The journey is a simplified picture of serving. The prices are real list prices from our model comparison.

The short answerEvery request runs the trained model on chips in a data centre to produce an answer. This step is called inference: training happens once, inference happens every time someone uses the model.

In simple words

  • Inference means using a trained model to answer a request.
  • It runs on chips such as GPUs, in large data centres.
  • Services charge per token, so speed and price matter.
  • Data centres need a lot of electricity, which keeps them in the news.

Training once, inference every time

A model is trained once, at great cost. After that, every question, image or agent step needs the model to run again. That is inference.

Because it happens for every user and every step, the cost of inference decides how cheap, fast and widely used AI can be.

Tokens and prices

AI services measure work in tokens. You pay for the tokens you send in and the tokens the model writes back, and prices differ a lot between models.

Our model comparison lists prices next to independent test results, so you can see what you get for the money.

Chips and data centres

Models run on specialised chips: Nvidia’s GPUs, Google’s TPUs and a growing number of custom chips. Thousands of them sit in data centres that need a lot of power and cooling.

That is why chip export rules, power deals and new data centres appear in the news so often.

Labels on what comes out

Some laws cover the output itself. In the EU, generative AI must mark what it creates in a machine-readable way, and California requires hidden provenance data in AI-made images, video and audio.

Check yourself

  1. What is inference?

  2. How do most AI services charge?

  3. Why do power and chips appear so often in AI news?

0 of 3 answered

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.