1 Frontier-model laws, 2025 to 2028
- EU duties for models with systemic risk
- California safety framework, 15-day incident reports
- Connecticut whistleblower protection
- New York 72-hour incident reportsIllinois 72-hour incident reports
- Illinois yearly independent audits
- Today
2 Makers’ claims, checked
- Backed 4
- Mixed 4
- Disputed 0
The short answerFrontier models are the most powerful AI models. Outside experts test them, and new laws make their makers share safety plans and report problems.
In simple words
- Frontier models are the strongest AI models.
- Outside testers check what the makers claim.
- Testers look for dangerous skills, such as hacking.
- New laws ask makers to report serious problems.
The strongest models
A few companies build the most powerful AI models, such as GPT, Gemini and Claude. These are called frontier models, because they are at the edge of what AI can do.
Who checks them?
Every company says its model is the best. Independent testers check that: some let people vote between two answers without knowing which model wrote them, and others run their own tests.
Testers also look for dangerous skills, such as helping someone hack a computer.
New rules
In the EU and in some US states, makers of the strongest models must now share safety plans and report serious incidents. More of these rules start in 2027.
Check yourself
What is a frontier model?
Why do independent tests matter?
What do new laws ask frontier makers to do?
0 of 3 answered
The short answerFrontier models are the most capable general-purpose models. Independent testers check makers’ claims, and new laws in the EU and several US states ask their makers for safety plans and incident reports.
In simple words
- Frontier means the most capable models, at the edge of what AI can do.
- Independent tests matter, because makers’ own claims can be optimistic.
- Safety tests look for dangerous skills, such as helping with cyberattacks.
- New laws ask frontier developers for safety plans and incident reports.
What makes a model frontier
There is no single scientific line. Laws use training compute as a stand-in: more than 10²⁵ operations in the EU, and more than 10²⁶ in California, New York and Connecticut.
These models come from a handful of labs, and we track each new release in our model release tracker.
How they are tested
Makers publish their own benchmark scores, but independent testers matter more. LMArena ranks models by blind votes from people comparing two answers, and Artificial Analysis runs its own tests of quality, speed and price.
When a maker’s claim and an independent test disagree, our model tracker marks the claim as mixed or disputed.
Safety tests and incidents
Labs and government institutes also test for dangerous abilities, such as helping with cyberattacks, and for models that behave differently when they think they are being tested.
When something goes wrong in real use, we log it in our AI incident log, with what the company said and what is still unknown.
What the laws now ask
In the EU, makers of models with systemic risk must test them, reduce risks and report serious incidents. California asks for a published safety framework and incident reports within 15 days.
New York will ask for incident reports within 72 hours from 2027, Illinois will add yearly independent audits from 2028, and Connecticut protects staff who warn about catastrophic risks.
Check yourself
How do laws usually decide which models count as frontier?
How does LMArena rank models?
Which law asks for frontier incident reports within 72 hours from 2027?
0 of 3 answered
The short answerFrontier status is defined by compute thresholds and revenue tests. Capability is measured by held-out benchmarks, preference arenas and dangerous-capability evaluations, and incident-reporting rules differ by deadline and recipient across the EU and US states.
In simple words
- Thresholds: 10²⁵ operations in the EU; 10²⁶ in California, New York and Connecticut.
- “Large” developers add a revenue test: over $500 million a year.
- Benchmarks can leak into training data, so held-out independent tests matter.
- Incident deadlines differ: 15 days in California, 72 hours in New York and Illinois.
Defining the frontier
The EU presumes systemic risk above 10²⁵ training operations (Article 51), and the Commission can also designate models. California’s SB 53 defines a frontier model as one trained with more than 10²⁶ operations, counting later fine-tuning and reinforcement learning.
California, New York and Connecticut add a revenue test for “large” frontier developers: more than $500 million a year with affiliates. Large developers carry the heaviest duties.
How capability is measured
Benchmarks test knowledge, coding and agent tasks, but their questions can leak into training data and inflate scores. Held-out tests run by independent groups are harder to game.
Preference arenas such as LMArena turn many blind pairwise votes into ratings with confidence ranges. Artificial Analysis publishes its own intelligence index alongside speed and price.
Dangerous-capability evaluations
Labs and government bodies test for help with cyberattacks, chemical or biological weapons, and loss of control. Examples include the UK AI Security Institute and the US Center for AI Standards and Innovation.
Evaluators also watch for models that detect testing and behave differently, which makes results harder to trust.
Reporting regimes compared
EU: models with systemic risk must report serious incidents to the AI Office. California: critical safety incidents go to the Office of Emergency Services within 15 days, or within 24 hours to the right authority if lives are at risk.
New York (from 2027): reports within 72 hours. Illinois (from 2027): 72 hours, or 24 hours if lives are at risk, plus yearly independent audits from 2028. California and Connecticut also protect employees who report catastrophic risks.
Check yourself
What extra test defines a “large” frontier developer in California and New York?
Why do held-out independent tests matter?
How fast must a critical incident reach California’s emergency office?
0 of 3 answered
Sources
- AI Act, Article 51 (EU AI Act Service Desk)
- SB 53 (California Legislature)
- Leaderboard (LMArena)
- Independent model tests (Artificial Analysis)
- AI Security Institute (UK government)
- Center for AI Standards and Innovation (NIST)
This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.