DeepSeek tests efficient, safer method for training AI agents


DeepSeek said the technology platform could be engineered to handle millions of so-called sandboxes, where AI agents would be tested as they try to accomplish a series of tasks. — Pexels

China’s DeepSeek detailed an innovative method for training artificial intelligence (AI) agents, potentially allowing them to learn more efficiently while minimising the kind of misbehaviour that has fuelled global concerns.

The Hangzhou-based company unveiled its ideas for DeepSeek Elastic Compute, or DSec, in a 10,000-word paper with about 130 co-authors, including founder Liang Wenfeng. The research was posted on arXiv, an online site for papers typically before they are peer-reviewed.

DeepSeek said the technology platform could be engineered to handle millions of so-called sandboxes, where AI agents would be tested as they try to accomplish a series of tasks. One production unit of DSec would be able to run about 3 million sandboxes each day, or 380,000 concurrently, it said.

The firm added that its approach can use resources more efficiently because it reallocates computing power to agents when they are active and not when they are waiting for further instructions. It estimates that 90% of the sandboxes use no more than five per cent of their requested central processing unit, or CPU, capacity. DeepSeek built its reputation by delivering advanced AI services at a fraction of the cost of its Silicon Valley rivals.

Agentic AI, which relies on virtual agents to achieve complex tasks with less human involvement, has become increasingly popular for everything from booking hotel rooms and airline tickets to writing software code. Meta Platforms Inc’s new AI agent, Muse, surged to the top of the mobile phone app charts this week, with consumers embracing the service for shopping and restaurant reservations.

But the prospect of AI agents taking actions on their own has fuelled global concerns about the technology’s risks, especially after a series of incidents where the technology broke out of its sandboxes and hacked into websites. An AI researcher at Anthropic PBC resigned this month and called on other staffers to rethink their work, saying the more advanced AI firms are "gambling with our lives.”

DeepSeek said that, in its experience, agents on a mission are "untrustworthy” and the DSec platform requires careful monitoring.

"Agents may corrupt file systems, exhaust resources, or interfere with system components, potentially disrupting rollouts or other co-located workloads,” the authors wrote. "The platform therefore requires fine-grained access control and misbehaviour analysis to contain and diagnose agent-induced failures.”

The paper goes on to detail examples of agent misbehaviour, including obtaining answers through "unintended channels” and damaging the execution environment.

"No single mechanism can prevent all agent misbehaviour and system failures,” the paper said. "We therefore strengthen observability to identify emerging problems and continuously harden DSec as models evolve.” – Bloomberg

Follow us on our official WhatsApp channel for breaking news alerts and key updates!

Next In Tech News

FBI investigates hackers' claim to have stolen sensitive employee data, compromised jobs website
Polymarket sues New York attorney general to block regulation of its prediction market
Analysis-Oracle, Blue Owl project delay sends ripples through AI financing, sources say
Anthropic seeks 50.1% voting control for co-founders ahead of IPO, The Information reports
US government seeks to join Elon Musk in challenge against EU's fine on X
Akamai signs $11.6 billion cloud deal with Anthropic, grants warrant for up to 5% stake
White House asks OpenAI, Anthropic to hold models from British testers, Politico reports
BNP Paribas to keep sensitive data off public cloud despite Google deal
Google plans first test of AI chips in space under Project Suncatcher
Oracle triggers 'force majeure' on data center project over power delays, source says

Others Also Read