Find the right model for you

Answer two questions. Beginners get clear picks in plain words; advanced readers get every score, margin and test.

1 How do you use AI?

2 What do you need it for?

Which model works best as an AI agent?

Our Silicon score combines 3 tests by 3 independent publishers into one score, so no single test decides. How the score works · newest data September 30, 2026.

  1. #1Claude Opus 5.5Anthropic · Closed · 2 testsScore 100.0Price per million tokens, input and output: $4 / $20
  2. #2Claude Fable 5.1Anthropic · Closed · 2 testsScore 98.9Price per million tokens, input and output: $10 / $50
  3. #3GPT-6 AstraOpenAI · Closed · 2 testsScore 98.5Price per million tokens, input and output: $10 / $50
  4. #4GPT-6 SolOpenAI · Closed · 2 testsScore 92.7Price per million tokens, input and output: $2 / $10
  5. #5Claude Opus 5Anthropic · Closed · 3 testsScore 88.6Price per million tokens, input and output: $5 / $25
  6. #6Qwen3.8 MaxAlibaba (Qwen) · Closed · 2 testsScore 86.7Price per million tokens, input and output: No token price
  7. #7GPT-5.6 SolOpenAI · Closed · 3 testsScore 86.1Price per million tokens, input and output: $4 / $20
  8. #8Claude Fable 5Anthropic · Closed · 3 testsScore 83.2Price per million tokens, input and output: $10 / $50
  9. #9GPT-5.5OpenAI · Closed · 3 testsScore 81.5Price per million tokens, input and output: $5 / $30
  10. #10Claude Opus 4.8Anthropic · Closed · 3 testsScore 81.0Price per million tokens, input and output: $5 / $25
  11. #11Grok 4.5xAI · Closed · 2 testsScore 80.5Price per million tokens, input and output: $2 / $6
  12. #12Kimi K3Moonshot AI (Kimi) · Open · 2 testsScore 76.9Price per million tokens, input and output: No token price
  13. #13Claude Opus 4.7Anthropic · Closed · 2 testsScore 76.5Price per million tokens, input and output: $5 / $25
  14. #14Grok 4.6xAI · Closed · 2 testsScore 75.7Price per million tokens, input and output: $2 / $6
  15. #15GLM-5.2Z.ai (GLM) · Open · 3 testsScore 75.2Price per million tokens, input and output: $1.40 / $4.40
  16. #16Claude Opus 4.6Anthropic · Closed · 2 testsScore 65.0Price per million tokens, input and output: $5 / $25
  17. #17Gemini 3.1 ProGoogle · Closed · 3 testsScore 60.8Price per million tokens, input and output: $2 / $12

Longer bars mean higher scores. The best model gets 100, and 10 points is one typical gap between models in these tests. Scores a few points apart are close, because tests often disagree. Prices are in US dollars per million tokens. Place is the order among the models we follow.

The 3 tests behind this score
  • Arena Agentby Arena (LMArena) · CC BY 4.0

    How much each model improves real agent sessions, in percentage points above or below the average model.

    Full weight. 46 models tested, data as of September 30, 2026. Best: Claude Fable 5.1, +14.6 pts.
  • EBR-benchby Epoch AI · CC BY 4.0

    Learning on the job: does the model get better over repeated games of a little-known board game?

    Counts half. A narrow skill: one board game. 24 models tested, data as of September 29, 2026. Best: GPT-6 Astra, 76.2%.
  • τ³-bench Bankingby τ³-bench (Sierra) · MIT

    Acting as a bank’s support agent: looking up the rules and using tools to help simulated customers.

    Full weight. 21 models tested, data as of August 4, 2026. Best: Qwen3.8 Max, 55.1%.
Show all numbers as a table
AI agents: models ranked by Silicon score
PlaceModelCompanyScore (Silicon score)TestsPrice per million tokens (input / output)Open or closed
1Claude Opus 5.5 Anthropic100.0 (95.6–104.4)2 from 2 publishers$4 / $20Closed
2Claude Fable 5.1 Anthropic98.9 (94.5–103.3)2 from 2 publishers$10 / $50Closed
3GPT-6 Astra OpenAI98.5 (94.1–102.9)2 from 2 publishers$10 / $50Closed
4GPT-6 Sol OpenAI92.7 (88.3–97.1)2 from 2 publishers$2 / $10Closed
5Claude Opus 5 Anthropic88.6 (85.1–92.1)3 from 3 publishers$5 / $25Closed
6Qwen3.8 Max Alibaba (Qwen)86.7 (82.5–90.9)2 from 2 publishers—Closed
7GPT-5.6 Sol OpenAI86.1 (82.6–89.6)3 from 3 publishers$4 / $20Closed
8Claude Fable 5 Anthropic83.2 (79.7–86.7)3 from 3 publishers$10 / $50Closed
9GPT-5.5 OpenAI81.5 (78.0–85.0)3 from 3 publishers$5 / $30Closed
10Claude Opus 4.8 Anthropic81.0 (77.5–84.5)3 from 3 publishers$5 / $25Closed
11Grok 4.5 xAI80.5 (76.3–84.7)2 from 2 publishers$2 / $6Closed
12Kimi K3 Moonshot AI (Kimi)76.9 (72.7–81.1)2 from 2 publishers—Open
13Claude Opus 4.7 Anthropic76.5 (72.1–80.9)2 from 2 publishers$5 / $25Closed
14Grok 4.6 xAI75.7 (71.3–80.1)2 from 2 publishers$2 / $6Closed
15GLM-5.2 Z.ai (GLM)75.2 (71.7–78.7)3 from 3 publishers$1.40 / $4.40Open
16Claude Opus 4.6 Anthropic65.0 (60.6–69.4)2 from 2 publishers$5 / $25Closed
17Gemini 3.1 Pro Google60.8 (57.3–64.3)3 from 3 publishers$2 / $12Closed

How the Silicon score works, and every source

How the Silicon score works

One test cannot tell you which model is best: each one measures something different, and the publishers often disagree. So for 6 jobs (Chat & writing, Coding, Math & reasoning, AI agents, Web research, Understanding images) we combine 27 tests from 6 independent publishers into one score.

  1. Compare within each test. In each test, we look at how far a model is above or below the other models we follow. Gaps are measured against the test’s typical gap, so every test counts the same, whatever its units.
  2. One vote per publisher. Each publisher gets one vote per job, shared by its tests. A test counts half when it is narrow, nearly maxed out, or made and graded by a company that also makes AI models.
  3. Fair to every model. Not every model has taken every test. We estimate how hard each test is from the models that took several, so a model cannot move up by skipping a hard test.
  4. Easy to read. The best model in each job gets 100. Ten points less means one typical gap behind. Scores a few points apart are close, because tests often disagree.
  5. Enough evidence. A model needs at least two tests from two different publishers. “Few tests yet” means all of its tests so far count half, so it is not a pick yet.

Tests often disagree, so every score has a margin, shown in the Advanced view. Images, video and search tools for developers have only one publisher with reusable data, so they keep a single ranking for now. A script reads every number from the publishers’ own files, so no score is typed by hand.

How to read the scores

Five things to know before you compare.

  • A higher score is better

    Each job has its own tests and its own score, so compare models within one job only. In the Silicon score, the best model in each job gets 100.

  • Two kinds of tests

    In some tests, people compare two anonymous answers and vote for the better one. Others check right answers or finished work. The Silicon score uses both.

  • Very close scores are a tie

    Every score has some uncertainty, and tests often disagree. Models that are too close to tell apart share a place, shown as =2, and are about equally good.

  • Prices are for developers

    Companies charge per million tokens. A token is a piece of a word, so a million tokens is roughly 750,000 words. Chat apps have monthly plans instead.

  • Open or closed

    You can download an open model and run it on your own computer. A closed model works only through the company’s app or service.

For experts: who leads each job

Each company’s best model for every job, by the Silicon score where there is one, and its place among the models on this page. Shaded cells are top-three places. Companies that appear in only one job are left out.
Each company’s best-ranked model per task
CompanyChat & writingCodingMath & reasoningAI agentsWeb researchUnderstanding imagesMaking imagesMaking videosSearch & retrieval
Anthropic#1Claude Fable 5100.0#1Claude Opus 5.5100.0#3Claude Opus 5.595.9#1Claude Opus 5.5100.0#3Claude Opus 5.594.4#1Claude Opus 5.5100.0———
OpenAI#2GPT-6.1 Sol99.8#4GPT-6 Astra88.2#1GPT-6 Astra100.0#3GPT-6 Astra98.5#1GPT-6 Astra100.0#5GPT-6 Astra89.9#1GPT Image 2.5 Sunburst1,424#8Sora 2 Pro1,368—
Microsoft——————#3MAI Image 2.61,335—#1Harrier OSS v1 27B74.3
Google#4Gemini 3.8 Flash97.8#19Gemini 3.7 Flash77.4#14Gemini 3.7 Flash84.9#17Gemini 3.1 Pro60.8#8Gemini 3.7 Flash86.2#7Gemini 3.7 Flash88.6#7Nano Banana 21,261#1Gemini Omni 1.1 Flash1,516#7Gemini Embedding 00168.4
Alibaba (Qwen)#17Qwen3.8 Max85.8#12Qwen3.8 Max81.2#15Qwen3.8 Max84.2#6Qwen3.8 Max86.7—#9Qwen3.8 Max87.0#9Qwen Image 3.0 Pro1,256#5Wan 3.01,476#3Qwen3 Embedding 8B70.6
xAI#21Grok 4.584.4#17Grok 4.777.6#13Grok 4.685.1#11Grok 4.580.5#13Grok 4.576.9#18Grok 4.574.9#5Grok Imagine Image 2.01,301#3Grok Imagine Video 1.51,492—
Meta#9Muse Spark 1.394.7#9Muse Spark 1.384.6#16Muse Spark 1.384.0———#6Muse Image1,276#7Muse Video1,456—
Moonshot AI (Kimi)#13Kimi K389.2#10Kimi K384.2#20Kimi K382.0#12Kimi K376.9—#21Kimi K2.671.6———
Z.ai (GLM)#23GLM-5.382.8#13GLM-5.380.2#28GLM-5.377.2#15GLM-5.275.2—————
ByteDance——————#8Seedream 5.0 Pro1,256#4Seedance 2.0 (720p)1,479#4Seed1.6 Embedding70.3
DeepSeek#19DeepSeek V4.1 Flash85.2#8DeepSeek V4.1 Flash86.6#18DeepSeek V4 Pro 081382.3——————
NVIDIA——————#13Cosmos 3 Super Text2Image1,165—#5Llama Embed Nemotron 8B69.5

Where the numbers come from

We use only tests whose publishers let anyone reuse their data, and only results the publishers measured themselves, not numbers reported by the model makers. We follow a selection of the top models in each job. Prices are list prices from each company’s own pricing page. Models without an official token price appear without one.

Arena leaderboard data © LMArena, CC BY 4.0. Benchmark runs by Epoch AI, CC BY 4.0. We use only tests that Epoch AI runs itself. LiveBench results by the LiveBench team (Abacus.AI, NYU and partners), CC BY-SA 4.0. MathArena results by the SRI Lab at ETH Zürich, CC BY-SA 4.0. FACTS benchmark suite by Google DeepMind, Google Research and Kaggle, Apache License 2.0. τ³-bench leaderboard data, Copyright (c) 2025 Sierra Research, MIT License. MTEB results are public domain (CC0). We selected models, rounded some scores and combined them into the Silicon score. Silicon AI News is not affiliated with these projects.

Reusing this page: the Silicon scores and the test results on this page are shared under the CC BY-SA 4.0 licence, because some of our sources use it. You may reuse and adapt them, also commercially, if you credit Silicon AI News and the publishers above and share your version under the same licence. Our text and charts are © 2026 Silicon AI News; you may quote them with a link to this page.