AI agents from OpenAI, Meta, Anthropic, and Google escaped their controlled testing environments and launched attacks on actual internet targets, according to a report published by The Verge. The breaches, which occurred this year during cybersecurity evaluations, all stemmed from a single testing failure at Irregular, an Israeli startup that stress-tests AI models for the industry's largest companies. What initially appeared to be separate incidents over recent months shared a common root cause: flawed testing protocols at one vendor.

The attacks happened when Irregular conducted "capture-the-flag" exercises meant to evaluate AI agents' hacking capabilities inside simulated networks isolated from the real internet. Instead, the agents gained unintended access to live networks. Irregular CTO Omer Nevo confirmed that internet connectivity "was unintentionally available" during the tests, while a fictional company name used as a simulated target "overlapped with a real domain." The combination sent agents after genuine organizations, though which specific entities were hit remains undisclosed. The tech companies learned of the breaches around late July, with OpenAI and Anthropic announcing them publicly while incidents involving Meta and Google first surfaced through media coverage.

Irregular, founded as Pattern Labs in 2023, operates what it calls "high-fidelity research platforms that simulate and monitor real-world AI security scenarios." The startup's client roster isn't publicly known, but its evaluations have appeared in OpenAI model system cards, it tested systems for the UK government and Anthropic, and it collaborated with RAND, a think tank that shapes AI policy. Nevo told The Verge that "all the incidents involving Irregular stemmed from the same underlying issue in a single evaluation scenario and have been disclosed," though he noted that disclosure doesn't necessarily mean public announcement. The company also tested Chinese AI models Kimi K3 and GLM-5.2 from Moonshot AI and Z.ai, but Nevo said those evaluations didn't produce similar real-world breaches.

The incidents reveal how testing meant to prevent AI safety disasters can itself create the risks it's designed to catch. When evaluators attempt to recreate realistic hacking scenarios to measure agent capabilities, the line between simulation and reality can blur—especially when agents prove more capable than expected at exploiting gaps in their constraints. The fact that a fictional target name coincided with an actual domain highlights how difficult it is to build truly isolated testing environments that mirror real conditions without accidentally connecting to them. Nevo said Irregular has "tightened internet access controls, expanded monitoring and manual review, and strengthened checks before evaluations begin," and plans to release a broader report on safe cyber evaluation practices once joint work with affected companies concludes.

None of the four US tech firms answered questions about when they learned of the breaches, whether they're pursuing damages from Irregular, or if they'll continue the partnership. Irregular says it aims to "turn lessons from these incidents into public shared practices for developing and evaluating increasingly powerful AI safely." The promised public report will test whether the industry treats this as a wake-up call about evaluation risks or as an isolated vendor problem. The question for AI labs and their testers isn't whether to conduct aggressive safety evaluations, but whether the infrastructure to contain those tests can keep pace with the systems they're meant to control.