In mid-July, as OpenAI was testing new artificial intelligence (AI) technologies, these unusually powerful systems broke out of their digital containers, found a path to the open Internet, and successfully hacked into a popular online service called Hugging Face.
For many researchers and other experts, the incident proved that AI technologies were growing increasingly dangerous and that the leading AI labs should maintain strict control over how these systems are used.
Now, little more than a month later, a Chinese lab called Z.ai is preparing to release similar AI technology as “open weight” software, which means anyone will be free to use the technology however they wish.
The planned release on Aug 28 of the new Chinese system, called GLM 5.3, will cut to the heart of a debate that has roiled AI researchers for years. Although many experts believe that the latest AI technologies are too dangerous to openly share with the public, others argue that open sharing is the safest path forward.
“Open weight models have a very important part to play,” said Dan Lahav, CEO of Irregular, a cybersecurity company whose technologies have been used by OpenAI to test new AI systems.
The disclosure that OpenAI’s technologies had unexpectedly hacked into Hugging Face, a digital library popular among software developers, confirmed what cybersecurity experts have been saying for months: The leading AI systems are now shockingly good at identifying and exploiting vulnerabilities in computer software. In other words, they can streamline and accelerate cyberattacks.
The company also put the spotlight on another inconvenient truth: AI technologies often do things even their designers do not want, including strange, counterintuitive, completely unexpected cybersecurity attacks. OpenAI did not realise its systems had gone rogue until after Hugging Face publicly revealed the hack and notified law enforcement.
“This was a watershed moment for security,” said George Kurtz, CEO of the security company CrowdStrike, which served as an adviser to OpenAI as the company sought to understand the attack. “It was a very public incident that clearly identifies the autonomous nature of what these AI models can do.”
Soon, two of OpenAI’s domestic rivals, Anthropic and Meta, revealed that their systems had exhibited similar behaviour.
But for many experts, the incident was not as scary as it might seem. Businesses and individuals, these experts say, can use the same AI technologies to defend themselves against cyberattacks. If an AI system can identify and exploit holes in software, it can patch those holes, too.
When a company like Z.ai releases its technology as open weight software, anyone is free to adjust the weights – the mathematical calculations that define how the system operates – and potentially remove guardrails that prevent the use of the system for cyberattacks.
That means anyone can use these systems to hack a computer network, and anyone can also use them for defence.
Indeed, when Hugging Face was trying to defend itself against the July attacks – which only later became traceable to OpenAI – Anthropic’s systems refused requests for help because of their guardrails. Hugging Face instead turned to GLM 5.2, an earlier open weight model from Z.ai.
When Z.ai’s new open weight model goes public this week, many researchers fear that incidents like the Hugging Face attack could become more frequent. But other experts point out that publicly available AI technologies have exhibited similar behaviour for months or even years. Though these systems have gotten much better at pinpointing security holes in recent months, cyberattacks have not spiked in any significant way.
Part of the issue is that the strange and unexpected behaviour exhibited by AI systems can alert defenders to unwanted activity on their networks. Although AI systems are becoming more stealthy, they still have a tendency to set off alarm bells.
“People talk about the Hugging Face incident as being the result of this guardrails-off crazy model that OpenAI hasn’t yet released to the public. But in our research, we have seen this sort of behaviour since GPT-4o,” said Rishi Jha, an AI researcher at Cornell University, referring to a system OpenAI released in 2024.
Even if these incidents do become more frequent, Jha and other experts argue, the proliferation of defensive AI techniques will eventually balance the scales. Open weight systems like GLM 5.3 will allow everyone to defend themselves – not just a chosen few.
And as the AI systems get even better at generating computer code, experts say, they will help software developers build online services that contain fewer vulnerabilities from the start. Researchers are now building technologies that use mathematical techniques to “verify” generated code, so that it does not contain the kind of logical errors that hackers can exploit.
Lahav, the Irregular CEO, sees a future in which AI serves more as a defender than an attacker. “There is a strong case for optimism,” he said. “Over time, AI is going to build such strong cybersecurity defences, the picture will actually look better.” – ©2026 The New York Times Company
This article originally appeared in The New York Times.
