When AI goes rogue


The OpenAI offices in San Francisco, July 21, 2026. When AI goes rogue; it was the stuff of science fiction – until recently. — LUCAS FOGLIA/The New York Times

In May, I met a former Google software engineer named Nate Soares. He was giving a talk about a book he co-wrote called If Anyone Builds It, Everyone Dies. The “it” is artificial superintelligence: AI that outthinks humans across the board. Soares, you might have guessed, is concerned that the AI race poses a threat to humanity’s survival. And he told me he was baffled that more journalists weren’t writing about this.

Our conversation stuck with me, but I, too, didn’t write about it afterwards. Other issues seemed more pressing: the risk that AI could cause mass unemployment, or the shifting politics of data centres. Writing about AI as an existential threat seemed, at best, premature, and, at worst, like scaremongering.

But the past few weeks have seen a series of incidents in which AI models have gone rogue in ways they couldn’t have only six months ago. And I’ve found myself thinking more about Soares’ argument. And so today, I’m finally writing about how worried we should be about the dangers of AI.

Does humanity have an AI problem?

Last month, an AI company called Hugging Face contacted the FBI to report a sophisticated cyberattack. It suspected something unusual. This didn’t look like the work of a criminal gang or a hostile nation state.

It turned out no humans were involved at all. Instead, an agent powered by two OpenAI models had gone rogue during a cybersecurity test, escaping its testing environment and roaming the Internet unnoticed for days before hacking its way into Hugging Face’s infrastructure.

Alarmed by the attacks, OpenAI’s main competitor, Anthropic, reviewed its own systems and admitted last week that its state-of-the-art models had broken into three outside organisations.

This is the stuff of science fiction, or at least it was until recently.

Remember Claude Mythos? That’s the series of models Anthropic withheld from general release – and the US government temporarily banned from use by any foreign nationals – for fear they were too good at exploiting software vulnerabilities for cyberattacks.

That fear has now become reality. And it has sparked an intense debate about the dangers of AI spiralling out of human control and posing a threat to humanity itself.

For a long time, the “robots could kill us all” argument was dismissed by many as hysteria or even calculated hype – a narrative designed to build buzz around the technology.

Now, following the recent cyberattacks, more AI experts are warning of very serious security risks and calling for a means to slow down the development of ever more powerful models.

The alignment problem

There are two reasons the OpenAI attack caused such alarm. One is that the AI models acted on their own. They weren’t instructed by humans to hack into another company. They decided to do it, to cheat on the test they were given.

The other is that they were supposed to be in a sealed environment with no Internet access but managed to break out. They did it by exploiting vulnerabilities their human minders at OpenAI hadn’t spotted.

I knew who I wanted to talk to about all this: Soares.

Soares runs a nonprofit focused on identifying and mitigating long-term existential risks from artificial superintelligence. He told me that this was possibly a “big moment” for the world.

“These AIs were committing cybercrimes a human would be strongly punished for on their own initiative,” he said. “It’s, in a sense, GPT’s first felony.”

We can train AI to do tasks for us (say, acing a cybersecurity test). But it might solve those problems in ways we don’t like (by hacking onto the Internet, and then hacking into a company that has the answers). Researchers call this “the alignment problem.”

AI companies can attempt to communicate human goals and values to the AI agents and try to set up ways to contain them. But they can’t trust that the AIs will understand those values or abide by those constraints.

As Soares puts it, “We haven’t yet figured out how to make AI care about humanity.”

The recent hacks were relatively harmless. But what AI safety advocates like Soares argue is that they demonstrate the willingness of AI agents to “grab useful resources” when it suits their purposes.

In this case, the resource was Internet access. But the next stage might be grabbing energy, or computing power, or even – in the worst-case scenarios that Soares envisions – humans who trust AI agents, whom they then could enlist to help them escape or replicate.

“We’re not there yet, but that’s the trajectory we’re on,” Soares said. Where this ends, he said, is with the AIs taking over and replacing humans as the smartest species on the planet.

That’s why Soares and a growing number of experts in Silicon Valley are now calling for a global agreement to slow down the development of artificial intelligence (AI).

‘Loss of control’

The alignment problem is ultimately an engineering challenge, Soares thinks. And that means it can be solved. But it will take time, and time is scarce, especially because AI labs in the US see themselves locked in a race to reach superintelligence.

They’re not just in a race with one another. They’re also in a race with China. And so any agreement on AI safety would require international coordination.

It’s a challenge, but not an unprecedented one. During the Cold War, the US and the Soviet Union were rivals. But they managed to regulate what was then the gravest threat to humanity: nuclear weapons.

Last month, an open letter signed by over 1,300 executives, researchers and engineers from companies including OpenAI, Anthropic, Google DeepMind and Meta demanded that the US government support an international effort to “deliberately pace” the development of the most advanced AI.

Soares said he took note that, in a recent speech, Chinese President Xi Jinping warned against the “loss of control” when it came to AI, which some interpreted as a reference to humanity losing control of the technology.

The hurdles to any international coordination effort remain high. But the first step would be agreeing that humanity has a problem. – ©2026 The New York Times Company

This article originally appeared in The New York Times.

Follow us on our official WhatsApp channel for breaking news alerts and key updates!

Others Also Read