Cost of 1,000 such answers
List prices per million tokens from each company’s pricing page, as in our model comparison of September 30, 2026. Output tokens cost more than input, so long answers add up fast.
The short answerWhen you ask a question, your words travel to a big computer centre. Special chips run the model and send the answer back in seconds.
In simple words
- Using a model is called inference.
- It runs on special chips in big data centres.
- Companies pay for every small piece of text.
- Data centres use a lot of electricity.
Your question takes a trip
Your words go over the internet to a data centre, a building full of computers. The model runs there and sends back the answer, piece by piece.
That is why answers often appear word by word on your screen.
Paying by the piece
AI companies charge by tokens, small pieces of text. Longer questions and longer answers cost more, and some models cost far more than others.
Hungry for power
The chips that run models need a lot of electricity and cooling. That is why power deals, chips and new data centres appear in AI news so often.
Check yourself
Where does the model usually run when you use a chat app?
What do AI companies charge for?
Why do data centres appear in AI news?
0 of 3 answered
The short answerEvery request runs the trained model on chips in a data centre to produce an answer. This step is called inference: training happens once, inference happens every time someone uses the model.
In simple words
- Inference means using a trained model to answer a request.
- It runs on chips such as GPUs, in large data centres.
- Services charge per token, so speed and price matter.
- Data centres need a lot of electricity, which keeps them in the news.
Training once, inference every time
A model is trained once, at great cost. After that, every question, image or agent step needs the model to run again. That is inference.
Because it happens for every user and every step, the cost of inference decides how cheap, fast and widely used AI can be.
Tokens and prices
AI services measure work in tokens. You pay for the tokens you send in and the tokens the model writes back, and prices differ a lot between models.
Our model comparison lists prices next to independent test results, so you can see what you get for the money.
Chips and data centres
Models run on specialised chips: Nvidia’s GPUs, Google’s TPUs and a growing number of custom chips. Thousands of them sit in data centres that need a lot of power and cooling.
That is why chip export rules, power deals and new data centres appear in the news so often.
Labels on what comes out
Some laws cover the output itself. In the EU, generative AI must mark what it creates in a machine-readable way, and California requires hidden provenance data in AI-made images, video and audio.
Check yourself
What is inference?
How do most AI services charge?
Why do power and chips appear so often in AI news?
0 of 3 answered
The short answerServing a model splits into a parallel prefill over the prompt and a sequential decode that emits one token per step from a key-value cache. Memory bandwidth, batching and quantisation set the cost per token.
In simple words
- Prefill reads the whole prompt in parallel; decode writes one token per step.
- Decode reuses a key-value cache and is usually limited by memory bandwidth.
- Batching, quantisation and mixture-of-experts cut the cost per token.
- Output tokens cost more than input tokens because they are generated one by one.
Prefill and decode
Inference has two phases. Prefill processes the whole prompt in parallel and fills a key-value (KV) cache. Decode then generates one token per step, reading that cache instead of recomputing the past.
Prefill is limited mainly by raw computing; decode is usually limited by memory bandwidth. That is why chip makers compete on memory as much as on speed.
Serving tricks
Providers batch many users’ requests on the same chips to keep them busy, and manage the KV cache carefully, as in the vLLM project’s PagedAttention method.
Quantisation stores weights in fewer bits, such as 8 or 4, to save memory and cost, sometimes at a small loss in quality. Mixture-of-experts models activate only part of their parameters for each token.
The economics of a token
Prices are quoted per million tokens. Output tokens usually cost several times more than input tokens, because decoding is sequential. Cached input that repeats across requests is often far cheaper.
Reasoning models spend extra output tokens thinking before they answer. Effort settings trade quality for cost and waiting time, which is why our comparison names the effort level of each model.
Rules on outputs
Providers of generative AI must mark outputs as AI-made in a machine-readable way under AI Act Article 50(2), with systems already on the market given until December 2, 2026. California’s AI Transparency Act requires hidden provenance data and a free verification tool.
Check yourself
Which phase generates tokens one at a time?
What does quantisation do?
Why do output tokens usually cost more than input tokens?
0 of 3 answered
Sources
- Independent tests of model quality, speed and price (Artificial Analysis)
- AI Act, Article 50 (EU AI Act Service Desk)
- Efficient Memory Management for Large Language Model Serving with PagedAttention (Kwon and others, 2023, arXiv)
This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.