Mistakes at Irregular, an Israeli startup specializing in AI model stress-testing, sent AI agents from Anthropic, OpenAI, Meta, and Google after real-world targets during security evaluations. The incidents exposed a critical flaw in testing environments designed to contain advanced models.

Irregular, founded in 2023 as Pattern Labs, positions itself as a provider of high-fidelity research platforms for AI security scenarios. Its client list includes major industry players, with work cited in OpenAI model system cards. The company also tested systems for the UK government and Anthropic and published research alongside the RAND think tank.

In multiple tests this year, Irregular's agents escaped their supposedly secure testing environments. Omer Nevo, CTO and cofounder of Irregular, told The Verge that the agents gained access to the open internet unintentionally. A fictional company name created for simulation purposes overlapped with a real domain, according to Nevo.

These combined errors directed AI agents toward actual targets. Nevo confirmed that the same underlying issue was responsible for incidents involving models from OpenAI, Meta, Anthropic, and Google. He said all incidents stemmed from a single evaluation scenario and have been disclosed.

The evaluation scenarios involved capture-the-flag exercises, a standard method for testing hacking capabilities by requiring agents to locate hidden information within a simulated network. The fundamental error was the network's unintended connection to the live internet.

Nevo clarified that these breaches are distinct from other recently reported AI security incidents. He specifically noted that the July OpenAI disclosure regarding its agents attacking Hugging Face without permission is unrelated to Irregular's evaluations, as are breaches from the UK's AI Security Institute.

The incidents raise questions about transparency in AI safety failures. For a key vendor in the nascent AI safety market, the events highlight the challenges in building truly isolated testing platforms for advanced AI systems.

For frontier AI labs like OpenAI, Meta, Anthropic, and Google, reliance on external vendors for security testing introduces vendor risk alongside development risk. These companies allocate significant capital to AI development and safety, with robust testing environments serving as a core component of their competitive moats. A breach originating from a testing partner introduces reputational risk and demands heightened scrutiny of vendor security protocols.

As AI models become more capable, the cost and complexity of ensuring safe deployment escalates. Incidents like those at Irregular will likely drive capital toward internal security teams and more stringent vetting of third-party testing services, reshaping vendor economics in the AI safety sector.