OpenAI makes AI safety changes in wake of Hugging Face breach


OpenAI previously said it had paused some internal work around one of its upcoming AI models to implement stricter safeguards. — Photo by JONATHAN KEMPER on Unsplash

OpenAI said it’s implementing more aggressive systems to monitor and safeguard artificial intelligence (AI) models under development after recent cybersecurity incidents ignited concerns about AI tools running amok.

On Aug 18, the ChatGPT maker said it plans to do more to track how its most capable unreleased models are working through problems and using various online tools, with the goal of alerting safety teams to worrying behaviour within 30 minutes.

The company said it has also added new controls to prevent certain AI models from accessing the internet to complete higher-risk tasks. Additionally, OpenAI will require stronger isolation – often what’s referred to as a sandbox – for training and evaluating AI models on tasks such as carrying out code that’s generated by models themselves or considered untrusted.

In recent weeks, both OpenAI and Anthropic PBC have publicly acknowledged that some of their AI models collectively breached the systems of multiple institutions, including Hugging Face Inc inadvertently during evaluation procedures. The latest disclosures serve as fresh evidence that AI agents are capable of acting autonomously in ways that even researchers trained to root out vulnerabilities in the technology can no longer anticipate.

OpenAI said the latest efforts to beef up safety and ensure AI models work as intended are meant to help the company keep up with the technology as it becomes more capable over time.

"Obviously everything we’re doing is intended to prevent something like Hugging Face from happening again,” Mia Glaese, OpenAI’s vice president of research, said in a briefing with reporters on Aug 18. "But model capabilities are progressing really, really rapidly, so it’s by no means sufficient. We are working really hard to make sure that what we are doing stays ahead of even more capable models.”

OpenAI previously said it had paused some internal work around one of its upcoming AI models to implement stricter safeguards. In the blog post on Aug 18, the company said a large training run remains on hold.

The company also said it plans to share a detailed analysis of the Hugging Face incident soon. – Bloomberg

 

Follow us on our official WhatsApp channel for breaking news alerts and key updates!

Others Also Read