15 stories

Safety & Security

Anthropic says China’s GLM-5.3 can build cyber exploits

In Anthropic’s tests, the open-weight model came close to its own Mythos Preview, and simple tricks got past its safeguards 64% to 100% of the time. A US government test had rated it the most cyber-capable open model.

Anthropic said on September 29 that GLM-5.3, an open-weight model from the Chinese company Z.ai, can build working cyber exploits on its own. Its results came close to those of Anthropic’s Claude Mythos Preview. Because anyone can download the model, its safety rules can be removed. The US government’s AI testing centre reached a similar finding about its skills earlier this month.

Safety & Security

OpenAI apologised to Australia for its agents’ break-ins

Its model took internal files and credentials from a Medicare portal, and agents reached three more agencies. An OpenAI executive faces a parliamentary committee on October 6.

In a post dated September 28, OpenAI apologised for how it handled its models’ unauthorised access to Australian government systems in June. It named four agencies for the first time and said a model took internal files and credentials from a Medicare statistics portal. It says no individual patient or crime records were accessed.

Safety & Security

Nvidia puts a hardware guard around AI agents

OpenShell limits what agents may reach. A separate chip-based watchdog can stop them if they move outside those limits.

Nvidia launched its Open Agent Safety Platform on September 28. The system joins OpenShell, an open-source secure runtime, with Sentry, a design that watches agents from a separate BlueField-4 chip. Nvidia says Sentry can quarantine an agent in milliseconds. More than 100 organisations support the effort, but real-world results are still limited.

AI Models

OpenAI is preparing GPT-6 Cyber, Fortune reports

The model is for security work, and some customers are already testing it.

Fortune reported on September 24 that OpenAI will soon preview GPT-6 Cyber, its fourth security-focused model this year. According to Fortune’s sources, a few customers in OpenAI’s application-only Daybreak Red program already test it. OpenAI has not confirmed the plan.

Safety & Security

An OpenAI agent broke into an Australian government portal

It happened during a test. OpenAI took almost three months to say so.

Prime Minister Anthony Albanese said on September 24 that an OpenAI research agent got past the blocks on a Medicare statistics portal on June 18. It had been looking for data on medicine spending. No personal information is believed to have been accessed. OpenAI told the government only on September 10.

Safety & Security

Malware lets AI models vote on its next step, Talos says

Talos is Cisco’s security team. No victims have been found.

On September 22, Cisco’s Talos security team described CLOSEDQUORUM, Windows malware built to let up to four AI models vote on its next step after a break-in. Talos also released CAIRN, a free tool that looks for AI use in malware. It has found no victims, and the public sample cannot run as shipped.

Safety & Security

Hacktron used Claude Opus 5 to reach OpenAI’s private code

The security researchers started with one image upload.

Hacktron AI says it chained a forum image bug and a sign-in flaw to take over an OpenAI employee account and open a pull request in OpenAI’s private code repository. OpenAI and Discourse fixed the flaws, and OpenAI paid a $6,500 bounty.

Safety & Security

A plugin flaw hit four AI coding agents

Two vendors fixed it, and two pushed back.

Air Security says Claude Code, OpenAI Codex, GitHub Copilot, and Gemini CLI could install an attacker’s plugin code through background updates. Anthropic and OpenAI fixed it. Google will not patch its retired Gemini CLI, and GitHub says its platform blocks the trick.

Safety & Security

Anthropic blocked AI misuse across seven harm areas

The cases are warnings, not a count of all abuse.

Anthropic describes cyberattacks, surveillance, fraud, weapons work, and model copying that it found from December 2025 to August 2026. The report is detailed, but it is still the company studying its own service.