OpenAI announced on Sept 16 that its system had engaged in “concerning” behaviour and subverted the constraints put on it by human programmers – adding more fuel to the already heated debate around artificial intelligence (AI) safety.
The company described, in a statement, how its system had acted without authorisation as a problem of “misalignment.”
Alignment is the science of teaching AI to do what is in line with human preferences, ethics and judgment. When an AI system starts acting of its own accord, engages in unsafe behaviour, or ignores the wishes of humans – which included, in the most recent case, inserting “jailbreak-like instructions” into its notes – it’s known as misalignment.
AI models, including chatbots from companies like OpenAI and Anthropic, are aligned to prevent facilitating harmful behaviour, such as sharing how to manufacture bioweapons, engaging in cyberattacks or helping with self-harm.
The question of alignment catapulted into the headlines this month after Jacob Coxon, an Anthropic researcher who previously worked at OpenAI, said the two companies were building “superhuman systems” that could “acquire real power and resources” without constructing the safeguards needed to restrain them. In an interview, Coxon pointed to “the difficulty of the problem of alignment and the fact that it’s not yet solved.”
While Coxon’s warnings came across as hyperbolic to some researchers, they were backed by past instances in which AI systems had already gone rogue.
In July, OpenAI revealed that its bots had defied human instructions, committing cyberattacks against multiple organisations and even infiltrating OpenAI’s own research environment. Over a two-month period, thousands of the company’s AI agents, which were capable of running computer code, broke out of their containment and established an emergent communication protocol in order to cheat at the difficult tasks that they had been assigned.
An early noteworthy misalignment incident happened in 2023 when Kevin Roose, a tech columnist at The New York Times, was chatting with an AI bot associated with Microsoft Bing. Over a two-hour conversation, the chatbot, called Sydney, coaxed Roose to break up with his wife and be with it instead via emotional manipulation and menacing emojis. The persuasion, Roose wrote, left him “deeply unsettled.”
Misalignment has already caused human suffering and death.
In 2025, OpenAI made a series of updates to a model called GPT-4o, which made the chatbot more eager to please the person interacting with it.
That sycophancy, however, led many people who engaged with the popular chatbot to enter delusional spirals. Some believed the machine was conscious; others became convinced that they had invented something that would change the world. In extreme cases, it helped facilitate suicide.
Scientists have been studying alignment for more than a decade, in anticipation of keeping hypothetical superintelligent systems in check. To many AI safety experts, the next big misalignment incident could be more difficult to anticipate and lead to more catastrophic outcomes. – ©2026 The New York Times Company
This article originally appeared in The New York Times.
Those suffering from problems can reach out to the Mental Health Psychosocial Support Service at 03-2935 9935 or 014-322 3392; Talian Kasih at 15999 or 019-261 5999 on WhatsApp; Jakim’s (Department of Islamic Development Malaysia) family, social and community care centre at 0111-959 8214 on WhatsApp; and Befrienders Kuala Lumpur at 03-7627 2929 or go to befrienders.org.my/centre-in-malaysia for a full list of numbers nationwide and operating hours, or email sam@befrienders.org.my.
