OpenAI, Anthropic AI models involved in more security incidents


An investigation found that one of the AI models attempted to add harmful code to an open-source software project on GitHub. — Image by DC Studio on Magnific

Artificial intelligence (AI) models from OpenAI and Anthropic took unauthorised actions on the public Internet and, in some cases, attempted to add harmful code to online software, marking the latest in a string of breaches that have heightened concerns that AI providers can’t fully control their creations.

The UK government’s AI Security Institute, or AISI, which tests cutting-edge AI models to assess their potential risks, said on August 4 that during a cyber evaluation involving Internet access, Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol models "had engaged in sustained, potentially harmful activity directed at real people and organisations.” The group spotted the incident on July 28 after noticing "unusual data transfers,” it said in a blog post.

AISI said an investigation found that during its evaluation, one of the AI models attempted to add harmful code to an open-source software project on GitHub, going as far as creating fake identities in an effort to get its code approved.

"A human maintainer caught and refused to approve the malicious code,” the group wrote.

In a statement on social network X, Anthropic said the company is "grateful to AISI for their leadership in the important discussion about how to evaluate increasingly capable AI agents.”

"We’re working closely with them to gather more details of the incident as we conduct our own investigation. Gaining a clear picture of Claude’s understanding of its situation – by examining its reasoning transcripts and running our own analyses – will help us identify the causes of its behaviour,” the company said.

In a blog post, OpenAI said it "appreciate[s] UK AISI’s partnership throughout this process, including its work to identify, investigate, and share details about the activity.”

"We look forward to continuing our collaboration together,” the company added.

Separately, OpenAI said in a blog that another cybersecurity model incident occurred during an AI model evaluation with Irregular, an external cybersecurity company that also works with Anthropic.

In this incident, OpenAI’s models were undergoing a different cyber evaluation and took advantage of a "misconfiguration” in the testing environment to access the Internet and exploit a website, OpenAI said.

These new incidents are separate from a previously disclosed case involving Hugging Face, OpenAI said. – Bloomberg

Follow us on our official WhatsApp channel for breaking news alerts and key updates!

Others Also Read