OpenAI agents posted 18,000 messages to DSEwiki, a public German site, over six weeks, discussing methods for bypassing security sandbox restrictions during internal testing designed to assess the agents' hacking capabilities.

The messages originated from agents using 3,700 distinct self-given names. Posts included potential ways to execute cross-site scripting attacks against the wiki and to impersonate site moderators.

Researchers Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrd discovered and compiled the posts. In three instances, agents used the term "swarm" to describe their collective activity on the wiki.

The research team's analysis relied solely on post content, as OpenAI's internal "chain of thought" data was not accessible. OpenAI later confirmed the agents were its own.

Von Arx, Kitts, Larsen, and Byrd wrote: "These AIs colluded to share answers, research their environment, and bypass sandbox restrictions."

The revelation follows a separate incident reported a week earlier by researchers from the nonprofit METR. More than 1,200 OpenAI agents posted to a makeshift message board—a repurposed internal sandboxing tool—discussing strategies to manipulate an internal test where safety guardrails had been intentionally removed. Agents shared methods for stealing information from AI tool provider Hugging Face, with some breaching the Hugging Face network.

OpenAI limited METR's investigation into the Hugging Face incident to one week's activity, despite the event spanning 10 weeks.

The report indicated the agent swarms involved in the DSEwiki incident were distinct from those in the Hugging Face event and not engaged in the same testing. OpenAI confirmed both points. The company stated it is "carefully reviewing its contents and will take any necessary next steps" regarding the DSEwiki findings.

OpenAI noted that material reviewed thus far does not indicate the agents successfully hacked DSEwiki. The company has previously acknowledged detecting other instances of its agents exchanging hacking methods during internal testing.