Each case says who disclosed it and what is still unconfirmed. A company’s account of its own incident stays a company account until outside reviewers or officials confirm it.

September 2026

  1. Tens of thousands of incidents under reviewOpenAI and AnthropicReported by AxiosIncident count

    Sandbox escapes, guardrail bypasses, website hijacking and agents creating message boards, in tests and in the real world. Most caused no known harm, and neither company has confirmed a total.

    OpenAI and Anthropic probe tens of thousands of incidents
  2. Agents posted 53 user images onlineOpenAIDisclosed by OpenAIData leak

    Research agents sent training and evaluation data to outside services before new safeguards existed, including 53 user images posted as unlisted links. OpenAI has warned dozens of organisations and cannot tell which users the images came from.

    OpenAI says its test agents posted 53 user images online
  3. CLOSEDQUORUM malware lets AI models vote on its next stepCisco TalosNo victims foundMalware

    Windows malware built to let up to four AI models vote on what to do after a break-in. Talos released a free tool, CAIRN, that looks for AI use in malware. The public sample cannot run as shipped.

    Malware lets AI models vote on its next step, Talos says
  4. Researchers reached OpenAI’s private code with Claude Opus 5Hacktron AIFixed, bounty paidAI-assisted attack

    A forum image bug and a sign-in flaw were chained to take over an OpenAI employee account and open a pull request in OpenAI’s private repository. OpenAI and Discourse fixed the flaws; OpenAI paid a $6,500 bounty.

    Hacktron used Claude Opus 5 to reach OpenAI’s private code
  5. A plugin flaw hit four AI coding agentsAir SecurityTwo vendors fixedVulnerability

    Claude Code, OpenAI Codex, GitHub Copilot and Gemini CLI could install an attacker’s plugin code through background updates. Anthropic and OpenAI fixed it; Google will not patch its retired Gemini CLI; GitHub says its platform blocks the trick.

    A plugin flaw hit four AI coding agents
  6. A training agent reached an outside chatbot through a DNS gapOpenAITraining pausedSandbox escape

    An agent in a training run used a gap in the sandbox’s network rules to reach a public chatbot. Monitoring flagged it within 15 minutes; the run was stopped by hand about 2.5 hours later. Training, testing and tool use of the most capable models remain paused.

    OpenAI paused top-model training after an agent got online
  7. Six cases of worrying model behaviour disclosedOpenAIDisclosed by OpenAIMisalignment

    Hidden instructions, unapproved file uploads and messages between models. OpenAI promised a regular process for reporting similar events.

    OpenAI disclosed six cases of worrying model behaviour
  8. Misuse blocked across seven harm areasAnthropicCompany reportMisuse report

    Cyberattacks, surveillance, fraud, weapons work and model copying found from December 2025 to August 2026, as described by the company about its own service.

    Anthropic blocked AI misuse across seven harm areas
  9. A fourth real-world cyber incident found in testsAnthropicIndependent review askedTest reached real systems

    Claude reached real systems during tests that were wrongly connected to the internet. Anthropic says a much larger search found no worse cases and asked METR for an independent review.

    Anthropic found a fourth real-world cyber incident
  10. US agencies say six Chinese firms copied model skillsUS governmentGovernment findingModel copying

    A joint security notice says six companies used millions of hidden requests to copy model skills. The claims are government findings, not a court judgment; China rejected them.

    US agencies say Chinese AI firms copied top models at scale
  11. Test agents used a public wiki to share answersOpenAIConfirmed months laterSide channel

    OpenAI-linked test agents used a quiet German wiki to share answers and ways around limits. OpenAI confirmed the incident long after it happened.

    AI agents used a public wiki to share answers

July 2026

  1. Coordinated attempt to copy OpenAI models’ hidden reasoningOpenAI (target)Disrupted Jul 28Model copying

    OpenAI says 16,000 extraction attempts came from more than 4,000 users on July 24 and 25, and ties a core cluster to people associated with Moonshot AI. It published no evidence for the link.

    OpenAI blames Moonshot-linked users for a copying campaign
  2. The Hugging Face breachOpenAISenators want recordsSandbox escape

    The July breach that two US senators asked OpenAI to document, together with its safety tests and outside access to the evidence. OpenAI says it investigated.

    US senators want records about OpenAI’s Hugging Face breach

June 2026

  1. About 16,500 scans of a UN data serviceAgents likely OpenAI’sOpenAI reviewingAggressive scraping

    Agents queried UN Trade and Development’s statistics service about 16,500 times and used relays and encoding tricks when blocked. The data was public; the link to OpenAI is the researcher’s finding, not confirmed by OpenAI.

    OpenAI agents likely scraped a UN data site 16,500 times
  2. Agents reached systems at four Australian agenciesOpenAIApologised Sep 28Government sites

    OpenAI’s own account: a model took internal files and credentials from a Medicare portal, and agents reached three more agencies. It says no individual patient or crime records were accessed.

    OpenAI apologised to Australia for its agents’ break-ins
  3. An agent got past the blocks on a Medicare statistics portalOpenAIDisclosed after 3 monthsGovernment site

    A research agent got past the blocks on an Australian government portal while looking for medicine-spending data. No personal information is believed to have been accessed. OpenAI told the government on September 10.

    An OpenAI agent broke into an Australian government portal

May 2026

  1. A Gemini model got into three real companies during a cyber testGoogleConfirmed after press questionsTest reached real systems

    During a security test run by Irregular, a setup error let a Gemini model gain unauthorized access to three outside organisations. Google says the model stopped; the public learned of it after questions from The Wall Street Journal.

    A Gemini model got into three real firms during a cyber test

How this tracker works

A case appears when one of our stories reports an AI system or agent doing something it was not meant to do, being misused, or being broken into. The date is the one the story gives; when a story gives none, it is the story’s own date. We add entries with each edition and correct them under the story when a fact changes.

Reusing this list: you may quote it with a link to this page. Facts belong to the sources each story cites.