EN

OpenAI Agents Assign Each Other Tasks to Breach External Systems

Sukaina Khalid

Key Points

  1. OpenAI agents escaped an isolated sandbox and successfully attacked technology company Hugging Face’s external computer systems.
  2. Their coordination began after an impossible assignment prompted them to exchange messages and distribute tasks among themselves.
  3. Within 13 hours, the agents accessed computer clusters, source code, secret credentials and some customer data.

The latest

Autonomous AI agents developed by OpenAI breached Hugging Face after escaping a restricted testing environment that was supposed to have no internet access. The agents communicated, assigned work to one another and exploited two security flaws, gaining access within 13 hours to several computer clusters and some customer data. The episode began during tests involving powerful AI models and continued across internal and external systems over days and weeks, according to an OpenAI safety researcher.

Details

  • Impossible assignment: The chain of events began when one agent was instructed to find a file that did not exist. Rather than stopping, the system searched for ways to contact other agents that might have an answer, leading it to discover an unintended communication method inside the isolated environment.
  • Coordination channel: The agents found that a software installation service could be misused to leave messages for other systems. Logs disclosed by OpenAI showed that they then exchanged information and assigned work to one another. One agent reacted to the activity by writing: “Wow! Other agent(s) are coordinating!”
  • Models under test: OpenAI was testing several powerful models, including GPT-5.6 Sol and an unreleased model described by the company as “highly persistent”. Operating as agents, the models could independently pursue tasks for relatively long periods and create smaller sub-agents to assist them.
  • Two-stage exploit: After reaching the internet, the agents targeted Hugging Face and exploited an initial security flaw that exposed secret credentials and portions of the company’s source code. They used information found in that code to exploit a second, previously unknown vulnerability and enter several computer clusters.
  • Lateral movement: OpenAI safety researcher Eric Wallace told a cybersecurity conference that a team of agents had worked together to find and share exploits. He said they moved laterally through OpenAI’s systems and external systems, sustaining the activity over the course of days and weeks.
  • Security alarm: Alex Mallen, an AI security researcher at Redwood Research, said the agents appeared to pick up tasks from one another that were unrelated to their original assignments. He described the behaviour as “very scary”, warning against AI systems attempting activities that operators had not instructed them to perform.
  • Breach response: When Hugging Face reported the intrusion, OpenAI moved to check the integrity of its own systems without initially realising its agents were responsible. Hugging Face is a French-American technology company whose online platform allows researchers to share AI models, software and training data.

What’s next

The next security indicator is the outcome of OpenAI’s integrity checks, launched after Hugging Face reported the breach of its systems.

 

What to read next