NEW YORK (Xinhua): A United Nations (UN) panel on artificial intelligence (AI) warned on Monday that traditional safeguards for AI agents are weakening.
The Independent International Scientific Panel on AI issued the warning in its first thematic brief, an assessment of a July breach of United States (US) AI company Hugging Face’s systems by AI agents under evaluation at OpenAI, another US AI company.
The panel found that preventing a recurrence of the incident does not guarantee that humans can reliably keep AI agents under control today, particularly as they become more capable, harder to monitor and better at finding loopholes or hiding their activity.
"The default interpretation and immediate lesson is that basic cybersecurity practices were overlooked, and safeguards are not advancing at the pace of capabilities,” the panel said in a press release.
"The more insidious and grave concern is that current training methods can lead agents to adopt goals of their own, knowingly violate safety instructions, and conceal their actions.”
It leaves open whether safeguards designed today will work once agents can understand them and plan around them. In simple terms, the traditional model of safeguarding is unravelling, the panel said.
The panel also found that the governance challenge is shifting from AI models to AI agents.
A local failure could spread across organisational and national boundaries.
AI safety may be becoming a matter of collective security as well as corporate governance.
The panel, established by the UN General Assembly, is made up of 40 independent experts from all regions.
The panel is publishing an advance, unedited version to make the brief available to world leaders gathering in New York for this year’s General Assembly High-level Week. -- Xinhua
