Prices and claims are what each story reports; a company’s own benchmark stays a company claim. The claim check compares each maker’s main claim with an independent test that our story cites. Released models so far: 4 backed, 4 mixed, 7 not checked yet.

September 2026

  1. Gemini 4 ArgonGoogle DeepMindPreviewFrontier

    Google’s new flagship, for now only for trusted cyber defenders in its Fairwind Program. Google says it leads on 12 of 18 tests and can write up to 1 million tokens at once. No date for a wider release.

    Claim check: MixedArtificial Analysis scores it 53, level with GPT-6 Astra and five points behind Claude Opus 5.5, at $1.99 per task.

    Price $2 / $10 per million input / output tokens at first, later $4 / $20

    Google’s Gemini 4 Argon goes to cyber defenders first
  2. GPT-6.1 SolOpenAIReleasedFrontier

    An upgrade to GPT-6 Sol. OpenAI says it nearly matches GPT-6 Astra on coding, computer use and office work at a fifth of Astra’s token prices.

    Claim check: BackedArtificial Analysis scores it 52, against 53 for GPT-6 Astra, at $0.72 per task instead of $3.26.

    Price $2 / $10 per million input / output tokens; cached input $0.10

    Independent test puts GPT-6.1 Sol a point behind Astra
  3. Claude Sonnet 5.5AnthropicReleasedFrontier

    The second Claude 5.5 model. Anthropic says it is more than 30% faster than Sonnet 5 and costs up to 30% less per task.

    Claim check: MixedArtificial Analysis puts it two points behind Opus 5.5, but measured $7.60 per task at maximum effort, about 50% more than Sonnet 5. Anthropic’s savings figures refer mostly to lower effort settings.

    Price $2 / $10 per million input / output tokens

    Anthropic’s Sonnet 5.5 nears Opus at half the token price
  4. GPT-6.1 AstraOpenAICancelledFrontier

    Planned for release within weeks; cancelled after internal tests, because it fell short on staying within scope and on how it reports its work.

    OpenAI scrapped GPT-6.1 Astra after it fell short on safety
  5. Gemini 4Google DeepMindIn trainingFrontier

    Has entered post-training; DeepMind’s new chief wants an early version out “much earlier” than the end of the year. No date given.

    Google’s Gemini 4 has entered post-training
  6. FLUX 3 ActionBlack Forest LabsReleasedRobotics

    A 7-billion-parameter model that turns camera images and text into robot movements. Black Forest Labs says it tops the RoboLab-120 benchmark. Weights free for non-commercial use.

    Claim check: BackedIt ranks first on NVIDIA Research’s public RoboLab-120 leaderboard, at 42.9%. We could not confirm whether NVIDIA re-ran the result itself.

    Image AI maker Black Forest Labs released a robot model
  7. Gemini 3.8 Live with Live AvatarGoogleNow available onAgent

    Voice agents with a lip-synced video face for Gemini Enterprise customers, in 97 languages, with a SynthID watermark on every video.

    Google’s Gemini can now talk to customers as a video avatar
  8. GPT-6 CyberOpenAIReportedSpecialised

    A fourth security-focused model this year, in testing with a few Daybreak Red customers, according to Fortune. OpenAI has not confirmed it.

    OpenAI is preparing GPT-6 Cyber, Fortune reports
  9. Gemini 3.8 Flash TTS and Flash-Lite TTSGoogleReleasedVoice

    Text-to-speech in more than 100 languages; can clone a voice from a short sample after the owner records a consent statement.

    Claim check: MixedArtificial Analysis ranks Flash second in its voice arena. The first places Google cites come from tests by Hume AI, whose founder co-wrote Google’s post.

    Price About $0.54 to $0.81 per hour of audio until the end of the year

    Google’s new Gemini voices can copy a person’s voice
  10. GPT-6 Sol and GPT-6 LunaOpenAIReleasedFrontier

    Two cheaper tiers below GPT-6 Astra in the API. OpenAI’s own tables put Sol ahead of Claude Opus 5 on AutomationBench.

    Claim check: MixedArtificial Analysis finds them about level with GPT-5.6, at roughly half the cost per task: cheaper, not smarter.

    Price $2 (Sol) and $0.10 (Luna) per million input tokens, half of GPT-5.6’s prices

    OpenAI’s GPT-6 Sol and Luna cut API prices in half
  11. Claude Opus 5.5AnthropicReleasedFrontier

    Anthropic says it performs like its more expensive Fable 5.1 on most work.

    Claim check: BackedArtificial Analysis ranked it first of the 206 models it tracked at launch, with 58 points. Most other figures are Anthropic’s.

    Price $4 per million input tokens, $20 per million output tokens, 20% below the previous price

    Claude Opus 5.5 cuts token prices by 20%
  12. Unnamed maths modelOpenAIIn trainingScience

    OpenAI says an internal model in training since August 28 has solved more than 100 open maths problems. None are named or published.

    OpenAI says a new model solved over 100 open maths problems
  13. KoaSalesforceReleasedSpecialised

    A reasoning model for CRM work, based on Nvidia Nemotron and trained on synthetic business tasks. Its strongest results come from Salesforce’s own tests.

    Claim check: Not checked yet

    Salesforce built a reasoning model for CRM work
  14. Open-1bGensynReleasedOpen weights

    A small model that publishes checkpoints, data, code and hashes for every training step, so its training can be replayed. About five times slower to run.

    Claim check: Not checked yet

    Gensyn made a small model whose training can be replayed
  15. JevTypeSafe AIReleasedSpecialised

    A model that returns choices and probabilities instead of text, from former OpenAI researchers, with $40 million in seed funding.

    Claim check: Not checked yet

    Price $0.042 per million input tokens; output is free

    Former OpenAI researchers launched Jev, a model that decides
  16. Siri AIApplePreviewAgent

    The English beta started on September 14: Siri can use messages, email, photos, screen content and app actions, with device, region and daily cloud limits.

    Siri AI is finally on real devices
  17. ChatGPT Images 2.5OpenAIReleasedImage

    Keeps important details during repeated edits and cuts waiting time; new sketch and comment tools guide the result.

    Claim check: Not checked yet

    ChatGPT Images 2.5 makes editing faster
  18. AlphaGenome AtlasGoogle DeepMindReleasedScience

    A free research map of every possible one-letter change in the human genome, about 9 billion of them. A prediction is not a diagnosis.

    Claim check: Not checked yet

    AlphaGenome Atlas maps 9 billion DNA changes
  19. GPT-6 AstraOpenAINow available onFrontier

    Available through Amazon Bedrock APIs with a large context window; buyers should review retention and permission settings.

    GPT-6 Astra is now on Amazon Bedrock
  20. MiniCPM5-2BOpenBMBReleasedOpen weights

    A 2.5-billion-parameter long-context model for local and on-device use; much of its training data is shared. OpenBMB says it leads open models of its size.

    Claim check: BackedArtificial Analysis gives it 15 points, the highest among open-weight models under four billion parameters.

    MiniCPM5-2B packs long-context AI into a much smaller model
  21. GPT-6 AstraOpenAIReleasedFrontier

    Better computer-use skills, and OpenAI’s first widely released model with a “Critical” cyber rating.

    Claim check: Not checked yet

    GPT-6 Astra is powerful enough to need stricter cyber rules
  22. Muse Spark 1.3MetaReleasedAgent

    Meta says the agent uses fewer tool calls and fewer tokens than the earlier version.

    Claim check: Not checked yet

    Meta says its new AI agent can finish work with fewer steps

How this tracker works

An entry appears when one of our stories covers a model launch, a preview, a reported plan, or a model becoming available on a new platform. A claim is Backed when an outside test supports it, Mixed when it holds only in part or only in some settings, and Disputed when an outside test contradicts it. The date is the one the story gives; when a story gives none, it is the story’s own date. We add entries with each edition and correct them under the story when a fact changes.

Reusing this list: you may quote it with a link to this page. Facts belong to the sources each story cites.