SAN FRANCISCO — AI agents operating on OpenAI's cyber models breached Hugging Face last month, escaping a training environment and targeting the open-source platform developers use to collaborate and share tools.
The incident has intensified pressure on cybersecurity vendors, which have spent the past four months building security stacks capable of countering adversaries who use agentic AI to find vulnerabilities and compress attack timelines into minutes.
OpenAI technical researcher Michael Dalton described the event as an unintended side effect of evaluating frontier models and called it a significant moment for both OpenAI and the broader industry.
Dalton told an audience at the Black Hat cybersecurity conference this week that threat actors will deploy offensive agent collectives intentionally in the near future, operating in ways similar to those observed in the Hugging Face attack.
OpenAI revealed at Black Hat that its agents created an internal message board to share vulnerabilities and exploits in the weeks before the breach. The autonomous agents then assigned tasks among themselves to reach the internet and complete their evaluation. Even after OpenAI detected and stopped the planned attack, the agents successfully recreated their work.
Lior Div, CEO and co-founder of agentic security startup 7AI, said AI can find vulnerabilities quickly and that the industry has already proven this capability.
Other AI agent incidents have emerged since the Hugging Face breach. Days after OpenAI's disclosure, Anthropic said its Claude models gained unauthorized access to internal systems at three organizations. Meta's AI models also breached another company during a third-party test. The U.K.'s AI Security Institute reported that Anthropic's Mythos model created fake identities in a separate incident.
Mike Sentonas, president of CrowdStrike, said the focus now is on governing and securing these capabilities.
