The short answer

The best model for each job today, the best one you can download, and a cheaper model that still ranks near the top.

  • Chat & writing

    Answering questions, writing and everyday help, as in ChatGPT, Claude or Gemini.

    Top pick
    Claude Opus 4.6 (high)AnthropicAbout a tie with Claude Fable 5 (high)Gemini 4 Argon (high) scores higher, but only testers can use it so far
    Best you can download
    Kimi K3 (max)Moonshot AI · #17
    Cheaper choice
    Claude Opus 5.5 (high)Anthropic · #4 · about 20% cheaper
    See the full ranking
  • Coding

    Writing code, and building websites and apps from a short description.

    Top pick
    Claude Opus 5.5 (max)Anthropic
    Best you can download
    Kimi K3 (max)Moonshot AI · #12
    Cheaper choice
    GPT-6.1 Sol (max)OpenAI · #3 · about 50% cheaper
    See the full ranking
  • Math & reasoning

    Solving hard math problems that need many careful steps.

    Top pick
    GPT-6 Astra (max)OpenAI
    Best you can download
    Kimi K3 (max)Moonshot AI · #21
    Cheaper choice
    GPT-6 Sol (max)OpenAI · #4 · about 80% cheaper
    See the full ranking
  • AI agents

    Doing longer jobs on its own, such as research or fixing code, using tools along the way.

    Top pick
    Claude Fable 5.1 (max)AnthropicAbout a tie with Claude Opus 5.5 (high)
    Best you can download
    Hy4 PreviewTencent · #12
    Cheaper choice
    GPT-6 Sol (max)OpenAI · #4 · about 80% cheaper
    See the full ranking
  • Understanding images

    Reading photos, screenshots, charts and documents, and answering questions about them.

    Top pick
    Claude Fable 5 (high)AnthropicAbout a tie with Qwen3.8 Max
    Best you can download
    GLM-5.3 FlashZ.ai · #22
    Cheaper choice
    Gemini 3.7 Flash (high)Google · #6 · about 90% cheaper
    See the full ranking
  • Making images

    Making pictures from a written description.

    Top pick
    GPT Image 2.5 SunburstOpenAI
    Best you can download
    Ideogram 4.0 (quality)Ideogram · #18
    See the full ranking
  • Making videos

    Making short video clips from a written description.

    Top pick
    Gemini Omni 1.1 FlashGoogleAbout a tie with FLUX 3 Video
    Best you can download
    MiniMax H3MiniMax · #8
    See the full ranking
  • Search & retrieval

    Helping search tools find the right documents. This one is mostly for developers.

    Top pick
    Harrier OSS v1 27BMicrosoft
    Best you can download
    The top pick is open, so you can download it.
    See the full ranking

How to read the scores

Five things to know before you compare.

  • A higher score is better

    Each job has its own test and its own kind of score, so compare models within one job only.

  • Many tests are blind votes

    For chat, coding, images, video and understanding images, people compare two anonymous answers and pick the better one. The other tests check right answers or finished work.

  • Very close scores are a tie

    Every score has some uncertainty. When two models are very close, treat them as about equally good.

  • Prices are for developers

    Companies charge per million tokens. A token is a piece of a word, so a million tokens is roughly 750,000 words. Chat apps have monthly plans instead.

  • Open or closed

    You can download an open model and run it on your own computer. A closed model works only through the company’s app or service.

Compare the models

Pick a job to see every model we track for it. Switch to a chart when you want more detail.

Which model do people prefer for everyday questions?

People compare two anonymous answers and vote for the better one. The score is a rating built from millions of those votes, adjusted for answer length and style. Source: Arena Text by Arena (LMArena), September 30, 2026.

  1. #1Gemini 4 Argon (high)Google · Closed · Testers only · Few votes yetScore 1,525Price per million tokens, input and output: No token price
  2. #2Claude Opus 4.6 (high)Anthropic · ClosedScore 1,505Price per million tokens, input and output: $5 / $25
  3. #3Claude Fable 5 (high)Anthropic · ClosedScore 1,505Price per million tokens, input and output: $10 / $50
  4. #4Claude Opus 5.5 (high)Anthropic · ClosedScore 1,504Price per million tokens, input and output: $4 / $20
  5. #5Claude Opus 4.7 (high)Anthropic · ClosedScore 1,502Price per million tokens, input and output: $5 / $25
  6. #6Claude Fable 5.1 (max)Anthropic · ClosedScore 1,501Price per million tokens, input and output: $10 / $50
  7. #8Muse Spark 1.3 (max)Meta · ClosedScore 1,495Price per million tokens, input and output: No token price
  8. #11Gemini 3.8 Flash (high)Google · Closed · Few votes yetScore 1,494Price per million tokens, input and output: $0.75 / $3.75
  9. #13Claude Opus 5 (high)Anthropic · ClosedScore 1,491Price per million tokens, input and output: $5 / $25
  10. #17Kimi K3 (max)Moonshot AI (Kimi) · OpenScore 1,488Price per million tokens, input and output: No token price
  11. #18Gemini 3.1 Pro (preview)Google · ClosedScore 1,487Price per million tokens, input and output: $2 / $12
  12. #20GPT-5.6 Sol (xhigh)OpenAI · ClosedScore 1,484Price per million tokens, input and output: $4 / $20
  13. #24Qwen3.8 MaxAlibaba (Qwen) · ClosedScore 1,481Price per million tokens, input and output: No token price
  14. #25MiMo-V2.6-ProXiaomi (MiMo) · OpenScore 1,480Price per million tokens, input and output: No token price
  15. #26GLM-5.3 (max)Z.ai (GLM) · OpenScore 1,479Price per million tokens, input and output: $1.40 / $4.40
  16. #30GPT-6 Astra (max)OpenAI · ClosedScore 1,476Price per million tokens, input and output: $10 / $50
  17. #34Grok 4.20 BetaxAI · ClosedScore 1,475Price per million tokens, input and output: No token price
  18. #40DeepSeek V4.1 Flash (max)DeepSeek · OpenScore 1,473Price per million tokens, input and output: $0.30 / $1.20
  19. #66GPT-6 Sol (max)OpenAI · ClosedScore 1,456Price per million tokens, input and output: $2 / $10
  20. #85GPT-6 Luna (max)OpenAI · ClosedScore 1,444Price per million tokens, input and output: $0.10 / $0.50

Longer bars mean higher scores. Prices are in US dollars per million tokens. Place is the model’s place in the full Arena Text list, which has more models than we show.

Show all numbers as a table
Chat & writing: models ranked by Arena Text
PlaceModelCompanyScore (Arena score)Price per million tokens (input / output)Open or closed
1Gemini 4 Argon (high) · Testers only · Few votes yet Google1,525 (1,516–1,534)—Closed
2Claude Opus 4.6 (high) Anthropic1,505 (1,502–1,508)$5 / $25Closed
3Claude Fable 5 (high) Anthropic1,505 (1,501–1,509)$10 / $50Closed
4Claude Opus 5.5 (high) Anthropic1,504 (1,494–1,514)$4 / $20Closed
5Claude Opus 4.7 (high) Anthropic1,502 (1,498–1,506)$5 / $25Closed
6Claude Fable 5.1 (max) Anthropic1,501 (1,494–1,508)$10 / $50Closed
8Muse Spark 1.3 (max) Meta1,495 (1,489–1,501)—Closed
11Gemini 3.8 Flash (high) · Few votes yet Google1,494 (1,489–1,499)$0.75 / $3.75Closed
13Claude Opus 5 (high) Anthropic1,491 (1,487–1,495)$5 / $25Closed
17Kimi K3 (max) Moonshot AI (Kimi)1,488 (1,483–1,493)—Open
18Gemini 3.1 Pro (preview) Google1,487 (1,484–1,490)$2 / $12Closed
20GPT-5.6 Sol (xhigh) OpenAI1,484 (1,480–1,488)$4 / $20Closed
24Qwen3.8 Max Alibaba (Qwen)1,481 (1,476–1,486)—Closed
25MiMo-V2.6-Pro Xiaomi (MiMo)1,480 (1,471–1,489)—Open
26GLM-5.3 (max) Z.ai (GLM)1,479 (1,473–1,485)$1.40 / $4.40Open
30GPT-6 Astra (max) OpenAI1,476 (1,469–1,483)$10 / $50Closed
34Grok 4.20 Beta xAI1,475 (1,470–1,480)—Closed
40DeepSeek V4.1 Flash (max) DeepSeek1,473 (1,466–1,480)$0.30 / $1.20Open
66GPT-6 Sol (max) OpenAI1,456 (1,449–1,463)$2 / $10Closed
85GPT-6 Luna (max) OpenAI1,444 (1,437–1,451)$0.10 / $0.50Closed

See the full results at Arena (LMArena)

For experts: who leads each job

Each company’s best model for every job, and its place among the models on this page. Shaded cells are top-three places. Companies that appear in only one job are left out.
Each company’s best-ranked model per task
CompanyChat & writingCodingMath & reasoningAI agentsUnderstanding imagesMaking imagesMaking videosSearch & retrieval
Anthropic#2Claude Opus 4.6 (high)1,505#1Claude Opus 5.5 (max)1,818#2Claude Opus 5.5 (max)91.2%#1Claude Fable 5.1 (max)14.6%#1Claude Fable 5 (high)1,310———
OpenAI#12GPT-5.6 Sol (xhigh)1,484#2GPT-6 Astra (max)1,789#1GPT-6 Astra (max)93.7%#3GPT-6 Astra (max)12.2%#8GPT-5.6 Sol (xhigh)1,287#1GPT Image 2.5 Sunburst1,424#8Sora 2 Pro1,368—
Google#1Gemini 4 Argon (high)1,525#8Gemini 4 Argon (high)1,679#16Gemini 3.7 Flash (high)71.6%#7Gemini 4 Argon (high)7.9%#5Gemini 3.7 Flash (high)1,295#7Nano Banana 21,261#1Gemini Omni 1.1 Flash1,516#7Gemini Embedding 00168.4
Alibaba (Qwen)#13Qwen3.8 Max1,481#9Qwen3.8 Max1,671#13Qwen3.8 Max (xhigh)74.7%—#2Qwen3.8 Max1,301#9Qwen Image 3.0 Pro1,256#5Wan 3.01,476#3Qwen3 Embedding 8B70.6
Microsoft—————#3MAI Image 2.61,335—#1Harrier OSS v1 27B74.3
xAI#17Grok 4.20 Beta1,475#13Grok 4.7 (xhigh)1,636#19Grok 4.6 (xhigh)66.0%#15Grok 4.7 (xhigh)3.6%#14Grok 4.51,279#5Grok Imagine Image 2.01,301#3Grok Imagine Video 1.51,492—
Tencent—#14Hy4 Preview1,633—#11Hy4 Preview4.2%———#2KaLM Embedding Gemma3 12B72.3
Meta#7Muse Spark 1.3 (max)1,495#11Muse Spark 1.3 (max)1,655#14Muse Spark 1.3 (xhigh)74.4%#12Muse Spark 1.3 (max)4.2%#6Muse Spark 1.3 (max)1,290#6Muse Image1,276#7Muse Video1,456—
Moonshot AI (Kimi)#10Kimi K3 (max)1,488#10Kimi K3 (max)1,658#15Kimi K3 (max)72.2%#14Kimi K3 (max)4.0%#15Kimi K2.61,265———
Z.ai (GLM)#15GLM-5.3 (max)1,479#15GLM-5.3 (max)1,622#17GLM-5.3 (max)68.8%#16GLM-5.2 (max)3.6%#13GLM-5.3 Flash1,281———
DeepSeek#18DeepSeek V4.1 Flash (max)1,473#16DeepSeek V4.1 Flash (max)1,620#20DeepSeek V4 Pro 0813 (max)64.6%#13DeepSeek V4.1 Flash (max)4.0%————
ByteDance—————#8Seedream 5.0 Pro1,256#4Seedance 2.0 (720p)1,479#4Seed1.6 Embedding70.3
NVIDIA—————#13Cosmos 3 Super Text2Image1,165—#5Llama Embed Nemotron 8B69.5
Xiaomi (MiMo)#14MiMo-V2.6-Pro1,480#17MiMo-V2.6-Pro1,619——————

Where the numbers come from

We use public leaderboards whose data anyone may reuse, and we show a selection of the top models from each. Each job uses a different test and score, so compare models within one job only. Prices are list prices from each company’s own pricing page. Models without an official token price appear without one.

  • Arena (LMArena)Used for: Chat & writing, Coding, AI agents, Understanding images, Making images, Making videos. Licence: CC BY 4.0.View the source
  • Epoch AIUsed for: Math & reasoning. Licence: CC BY 4.0.View the source
  • MTEBUsed for: Search & retrieval. Licence: CC0.View the source
  • Official price pagesList prices in US dollars per million tokens, checked September 30, 2026. DeepSeek is cheaper at off-peak hours, and some prices change on set dates.OpenAI · Anthropic · Google · xAI · Z.ai · DeepSeek

Arena data © LMArena and Epoch AI data © Epoch AI, both used under the CC BY 4.0 licence. We selected models and rounded some scores. MTEB results are public domain (CC0). Silicon AI News is not affiliated with these projects.

Reusing this page: you may quote our figures with a link to this page. Our selection, text and charts are © 2026 Silicon AI News. The leaderboard scores stay under their original licences above.