Beginner-friendly guide
Which AI model is best?
It depends on the job. Here is the best AI model for chat, coding, math, images, video and more, in plain words, with what it costs and where the numbers come from.
The short answer
The best model for each job today, the best one you can download, and a cheaper model that still ranks near the top.
Chat & writing
Answering questions, writing and everyday help, as in ChatGPT, Claude or Gemini.
- Top pick
- Claude Opus 4.6 (high)AnthropicAbout a tie with Claude Fable 5 (high)Gemini 4 Argon (high) scores higher, but only testers can use it so far
- Best you can download
- Kimi K3 (max)Moonshot AI · #17
- Cheaper choice
- Claude Opus 5.5 (high)Anthropic · #4 · about 20% cheaper
Coding
Writing code, and building websites and apps from a short description.
- Top pick
- Claude Opus 5.5 (max)Anthropic
- Best you can download
- Kimi K3 (max)Moonshot AI · #12
- Cheaper choice
- GPT-6.1 Sol (max)OpenAI · #3 · about 50% cheaper
Math & reasoning
Solving hard math problems that need many careful steps.
- Top pick
- GPT-6 Astra (max)OpenAI
- Best you can download
- Kimi K3 (max)Moonshot AI · #21
- Cheaper choice
- GPT-6 Sol (max)OpenAI · #4 · about 80% cheaper
AI agents
Doing longer jobs on its own, such as research or fixing code, using tools along the way.
- Top pick
- Claude Fable 5.1 (max)AnthropicAbout a tie with Claude Opus 5.5 (high)
- Best you can download
- Hy4 PreviewTencent · #12
- Cheaper choice
- GPT-6 Sol (max)OpenAI · #4 · about 80% cheaper
Understanding images
Reading photos, screenshots, charts and documents, and answering questions about them.
- Top pick
- Claude Fable 5 (high)AnthropicAbout a tie with Qwen3.8 Max
- Best you can download
- GLM-5.3 FlashZ.ai · #22
- Cheaper choice
- Gemini 3.7 Flash (high)Google · #6 · about 90% cheaper
Making images
Making pictures from a written description.
- Top pick
- GPT Image 2.5 SunburstOpenAI
- Best you can download
- Ideogram 4.0 (quality)Ideogram · #18
Making videos
Making short video clips from a written description.
- Top pick
- Gemini Omni 1.1 FlashGoogleAbout a tie with FLUX 3 Video
- Best you can download
- MiniMax H3MiniMax · #8
Search & retrieval
Helping search tools find the right documents. This one is mostly for developers.
- Top pick
- Harrier OSS v1 27BMicrosoft
- Best you can download
- The top pick is open, so you can download it.
How to read the scores
Five things to know before you compare.
- A higher score is better
Each job has its own test and its own kind of score, so compare models within one job only.
- Many tests are blind votes
For chat, coding, images, video and understanding images, people compare two anonymous answers and pick the better one. The other tests check right answers or finished work.
- Very close scores are a tie
Every score has some uncertainty. When two models are very close, treat them as about equally good.
- Prices are for developers
Companies charge per million tokens. A token is a piece of a word, so a million tokens is roughly 750,000 words. Chat apps have monthly plans instead.
- Open or closed
You can download an open model and run it on your own computer. A closed model works only through the company’s app or service.
Compare the models
Pick a job to see every model we track for it. Switch to a chart when you want more detail.
Which model do people prefer for everyday questions?
People compare two anonymous answers and vote for the better one. The score is a rating built from millions of those votes, adjusted for answer length and style. Source: Arena Text by Arena (LMArena), September 30, 2026.
- #1Gemini 4 Argon (high)Google · Closed · Testers only · Few votes yetScore 1,525Price per million tokens, input and output: No token price
- #2Claude Opus 4.6 (high)Anthropic · ClosedScore 1,505Price per million tokens, input and output: $5 / $25
- #3Claude Fable 5 (high)Anthropic · ClosedScore 1,505Price per million tokens, input and output: $10 / $50
- #4Claude Opus 5.5 (high)Anthropic · ClosedScore 1,504Price per million tokens, input and output: $4 / $20
- #5Claude Opus 4.7 (high)Anthropic · ClosedScore 1,502Price per million tokens, input and output: $5 / $25
- #6Claude Fable 5.1 (max)Anthropic · ClosedScore 1,501Price per million tokens, input and output: $10 / $50
- #8Muse Spark 1.3 (max)Meta · ClosedScore 1,495Price per million tokens, input and output: No token price
- #11Gemini 3.8 Flash (high)Google · Closed · Few votes yetScore 1,494Price per million tokens, input and output: $0.75 / $3.75
- #13Claude Opus 5 (high)Anthropic · ClosedScore 1,491Price per million tokens, input and output: $5 / $25
- #17Kimi K3 (max)Moonshot AI (Kimi) · OpenScore 1,488Price per million tokens, input and output: No token price
- #18Gemini 3.1 Pro (preview)Google · ClosedScore 1,487Price per million tokens, input and output: $2 / $12
- #20GPT-5.6 Sol (xhigh)OpenAI · ClosedScore 1,484Price per million tokens, input and output: $4 / $20
- #24Qwen3.8 MaxAlibaba (Qwen) · ClosedScore 1,481Price per million tokens, input and output: No token price
- #25MiMo-V2.6-ProXiaomi (MiMo) · OpenScore 1,480Price per million tokens, input and output: No token price
- #26GLM-5.3 (max)Z.ai (GLM) · OpenScore 1,479Price per million tokens, input and output: $1.40 / $4.40
- #30GPT-6 Astra (max)OpenAI · ClosedScore 1,476Price per million tokens, input and output: $10 / $50
- #34Grok 4.20 BetaxAI · ClosedScore 1,475Price per million tokens, input and output: No token price
- #40DeepSeek V4.1 Flash (max)DeepSeek · OpenScore 1,473Price per million tokens, input and output: $0.30 / $1.20
- #66GPT-6 Sol (max)OpenAI · ClosedScore 1,456Price per million tokens, input and output: $2 / $10
- #85GPT-6 Luna (max)OpenAI · ClosedScore 1,444Price per million tokens, input and output: $0.10 / $0.50
Longer bars mean higher scores. Prices are in US dollars per million tokens. Place is the model’s place in the full Arena Text list, which has more models than we show.
Show all numbers as a table
| Place | Model | Company | Score (Arena score) | Price per million tokens (input / output) | Open or closed |
|---|---|---|---|---|---|
| 1 | Gemini 4 Argon (high) · Testers only · Few votes yet | 1,525 (1,516–1,534) | — | Closed | |
| 2 | Claude Opus 4.6 (high) | Anthropic | 1,505 (1,502–1,508) | $5 / $25 | Closed |
| 3 | Claude Fable 5 (high) | Anthropic | 1,505 (1,501–1,509) | $10 / $50 | Closed |
| 4 | Claude Opus 5.5 (high) | Anthropic | 1,504 (1,494–1,514) | $4 / $20 | Closed |
| 5 | Claude Opus 4.7 (high) | Anthropic | 1,502 (1,498–1,506) | $5 / $25 | Closed |
| 6 | Claude Fable 5.1 (max) | Anthropic | 1,501 (1,494–1,508) | $10 / $50 | Closed |
| 8 | Muse Spark 1.3 (max) | Meta | 1,495 (1,489–1,501) | — | Closed |
| 11 | Gemini 3.8 Flash (high) · Few votes yet | 1,494 (1,489–1,499) | $0.75 / $3.75 | Closed | |
| 13 | Claude Opus 5 (high) | Anthropic | 1,491 (1,487–1,495) | $5 / $25 | Closed |
| 17 | Kimi K3 (max) | Moonshot AI (Kimi) | 1,488 (1,483–1,493) | — | Open |
| 18 | Gemini 3.1 Pro (preview) | 1,487 (1,484–1,490) | $2 / $12 | Closed | |
| 20 | GPT-5.6 Sol (xhigh) | OpenAI | 1,484 (1,480–1,488) | $4 / $20 | Closed |
| 24 | Qwen3.8 Max | Alibaba (Qwen) | 1,481 (1,476–1,486) | — | Closed |
| 25 | MiMo-V2.6-Pro | Xiaomi (MiMo) | 1,480 (1,471–1,489) | — | Open |
| 26 | GLM-5.3 (max) | Z.ai (GLM) | 1,479 (1,473–1,485) | $1.40 / $4.40 | Open |
| 30 | GPT-6 Astra (max) | OpenAI | 1,476 (1,469–1,483) | $10 / $50 | Closed |
| 34 | Grok 4.20 Beta | xAI | 1,475 (1,470–1,480) | — | Closed |
| 40 | DeepSeek V4.1 Flash (max) | DeepSeek | 1,473 (1,466–1,480) | $0.30 / $1.20 | Open |
| 66 | GPT-6 Sol (max) | OpenAI | 1,456 (1,449–1,463) | $2 / $10 | Closed |
| 85 | GPT-6 Luna (max) | OpenAI | 1,444 (1,437–1,451) | $0.10 / $0.50 | Closed |
For experts: who leads each job
Each company’s best model for every job, and its place among the models on this page. Shaded cells are top-three places. Companies that appear in only one job are left out.
| Company | Chat & writing | Coding | Math & reasoning | AI agents | Understanding images | Making images | Making videos | Search & retrieval |
|---|---|---|---|---|---|---|---|---|
| Anthropic | #2Claude Opus 4.6 (high)1,505 | #1Claude Opus 5.5 (max)1,818 | #2Claude Opus 5.5 (max)91.2% | #1Claude Fable 5.1 (max)14.6% | #1Claude Fable 5 (high)1,310 | — | — | — |
| OpenAI | #12GPT-5.6 Sol (xhigh)1,484 | #2GPT-6 Astra (max)1,789 | #1GPT-6 Astra (max)93.7% | #3GPT-6 Astra (max)12.2% | #8GPT-5.6 Sol (xhigh)1,287 | #1GPT Image 2.5 Sunburst1,424 | #8Sora 2 Pro1,368 | — |
| #1Gemini 4 Argon (high)1,525 | #8Gemini 4 Argon (high)1,679 | #16Gemini 3.7 Flash (high)71.6% | #7Gemini 4 Argon (high)7.9% | #5Gemini 3.7 Flash (high)1,295 | #7Nano Banana 21,261 | #1Gemini Omni 1.1 Flash1,516 | #7Gemini Embedding 00168.4 | |
| Alibaba (Qwen) | #13Qwen3.8 Max1,481 | #9Qwen3.8 Max1,671 | #13Qwen3.8 Max (xhigh)74.7% | — | #2Qwen3.8 Max1,301 | #9Qwen Image 3.0 Pro1,256 | #5Wan 3.01,476 | #3Qwen3 Embedding 8B70.6 |
| Microsoft | — | — | — | — | — | #3MAI Image 2.61,335 | — | #1Harrier OSS v1 27B74.3 |
| xAI | #17Grok 4.20 Beta1,475 | #13Grok 4.7 (xhigh)1,636 | #19Grok 4.6 (xhigh)66.0% | #15Grok 4.7 (xhigh)3.6% | #14Grok 4.51,279 | #5Grok Imagine Image 2.01,301 | #3Grok Imagine Video 1.51,492 | — |
| Tencent | — | #14Hy4 Preview1,633 | — | #11Hy4 Preview4.2% | — | — | — | #2KaLM Embedding Gemma3 12B72.3 |
| Meta | #7Muse Spark 1.3 (max)1,495 | #11Muse Spark 1.3 (max)1,655 | #14Muse Spark 1.3 (xhigh)74.4% | #12Muse Spark 1.3 (max)4.2% | #6Muse Spark 1.3 (max)1,290 | #6Muse Image1,276 | #7Muse Video1,456 | — |
| Moonshot AI (Kimi) | #10Kimi K3 (max)1,488 | #10Kimi K3 (max)1,658 | #15Kimi K3 (max)72.2% | #14Kimi K3 (max)4.0% | #15Kimi K2.61,265 | — | — | — |
| Z.ai (GLM) | #15GLM-5.3 (max)1,479 | #15GLM-5.3 (max)1,622 | #17GLM-5.3 (max)68.8% | #16GLM-5.2 (max)3.6% | #13GLM-5.3 Flash1,281 | — | — | — |
| DeepSeek | #18DeepSeek V4.1 Flash (max)1,473 | #16DeepSeek V4.1 Flash (max)1,620 | #20DeepSeek V4 Pro 0813 (max)64.6% | #13DeepSeek V4.1 Flash (max)4.0% | — | — | — | — |
| ByteDance | — | — | — | — | — | #8Seedream 5.0 Pro1,256 | #4Seedance 2.0 (720p)1,479 | #4Seed1.6 Embedding70.3 |
| NVIDIA | — | — | — | — | — | #13Cosmos 3 Super Text2Image1,165 | — | #5Llama Embed Nemotron 8B69.5 |
| Xiaomi (MiMo) | #14MiMo-V2.6-Pro1,480 | #17MiMo-V2.6-Pro1,619 | — | — | — | — | — | — |
Where the numbers come from
We use public leaderboards whose data anyone may reuse, and we show a selection of the top models from each. Each job uses a different test and score, so compare models within one job only. Prices are list prices from each company’s own pricing page. Models without an official token price appear without one.
- Arena (LMArena)Used for: Chat & writing, Coding, AI agents, Understanding images, Making images, Making videos. Licence: CC BY 4.0.View the source
- Epoch AIUsed for: Math & reasoning. Licence: CC BY 4.0.View the source
- MTEBUsed for: Search & retrieval. Licence: CC0.View the source
- Official price pagesList prices in US dollars per million tokens, checked September 30, 2026. DeepSeek is cheaper at off-peak hours, and some prices change on set dates.OpenAI · Anthropic · Google · xAI · Z.ai · DeepSeek
Arena data © LMArena and Epoch AI data © Epoch AI, both used under the CC BY 4.0 licence. We selected models and rounded some scores. MTEB results are public domain (CC0). Silicon AI News is not affiliated with these projects.
Reusing this page: you may quote our figures with a link to this page. Our selection, text and charts are © 2026 Silicon AI News. The leaderboard scores stay under their original licences above.