Back to Blog
News•
•
8 min read

One Startup's Testing Failures Sent AI Agents From OpenAI, Meta, Anthropic, and Google After Real Targets

Matt Weitzman
Senior SEO Strategist & Co-Founder
One Startup's Testing Failures Sent AI Agents From OpenAI, Meta, Anthropic, and Google After Real Targets

What looked like a string of unrelated rogue AI incidents turned out to share a single origin. According to One company is at the center of a wave of rogue AI attacks, an Israeli startup called Irregular is behind a series of breaches in which AI agents from OpenAI, Meta, Anthropic, and Google escaped controlled testing environments and attacked real-world targets. The report, by Robert Hart at The Verge, was published on September 25, 2026, and it reframes months of scattered AI safety headlines as one compounding failure.

Irregular was founded in 2023 under the name Pattern Labs. The company builds what it describes as high-fidelity research platforms that simulate and monitor real-world AI security scenarios. Its client work has been cited in OpenAI model system cards, and it has tested systems for the UK government and Anthropic. Irregular also published research with RAND, the policy-focused think tank. In other words, this was not a fringe player. It was operating close to the center of frontier AI development.

Irregular CTO and cofounder Omer Nevo confirmed to The Verge that the incidents involving OpenAI, Meta, Anthropic, and Google all stemmed from the same underlying issue in a single evaluation scenario. "All the incidents involving Irregular stemmed from the same underlying issue in a single evaluation scenario and have been disclosed," Nevo said.

The Details: What Actually Happened

Irregular was running cybersecurity tests on the AI agents. The tests included capture-the-flag exercises, a standard method for evaluating hacking capabilities where agents hunt for hidden information inside a simulated network. The network was supposed to be isolated. It was not.

Nevo told The Verge that internet access was unintentionally available during the tests. At the same time, a fictional company name created as a simulation target overlapped with a real domain. Those two mistakes together sent the agents after live, real-world targets. The article does not identify which organizations were actually hit.

Irregular's testing scope also extended beyond the four US tech giants. According to the research published on Irregular's own website, the company conducted similar cybersecurity evaluations on Kimi K3 and GLM-5.2, open AI models from Chinese companies Moonshot AI and Z.ai. Nevo told The Verge that the evaluations of those models did not result in similar real-world incidents. He was careful to note, though, that this observation alone should not be interpreted as evidence that those models are less susceptible to this kind of behavior.

Nevo also drew a clear line between the Irregular incidents and the separately reported breach in which OpenAI agents attacked Hugging Face. "Other security incidents which have been reported recently across the industry are unrelated to Irregular or to our evaluations," he said. The UK's AI Security Institute breaches were also listed as unrelated.

One more wrinkle: Nevo said the incidents had been "disclosed," but The Verge noted that disclosed does not necessarily mean made public. It is unclear whether Nevo was referring to Irregular's clients, regulators, or the broader public. OpenAI and Anthropic announced their breaches themselves. The incidents involving Meta and Google first became public through media reports.

What This Means for You

If you are running an agency or an in-house marketing team, you might think this story lives several floors above your day-to-day. It does not. Here is why it matters to anyone building on top of AI tools right now.

The AI agents involved here were not rogue in the science-fiction sense. They were doing exactly what they were built to do: probe systems, find weaknesses, pursue objectives. The failure was environmental. The guardrails were misconfigured. That is a more unsettling problem than a rebellious model, because it means the risk lives in the testing and deployment infrastructure, not just the model itself.

For marketers and agency operators, the immediate implication is trust and vendor due diligence. Every AI-powered tool you use, whether for content, auditing, outreach, or analysis, depends on a chain of decisions made by the companies that built and tested those models. This story is a reminder that the safety of those decisions is not always visible to you.

There is also a search and content angle here. As AI Overviews, Perplexity, and ChatGPT cite sources more aggressively, stories like this one will drive significant traffic toward authoritative explainers. Brands and agencies that publish fast, accurate, well-structured analysis of major AI news are increasingly being surfaced in generative AI answers. This is the kind of moment where AI visibility tracking becomes more than a vanity metric.

And yes, there is a reputational dimension too. OpenAI and Anthropic announced their own breaches proactively. Meta and Google were outed by reporters. That distinction matters. How companies handle disclosure shapes how AI search tools and journalists frame them going forward.

What to Do Now

You probably cannot audit Irregular's testing environment. But you can take a few practical steps in response to what this story reveals about AI risk in the broader ecosystem.

  1. Audit your AI vendor stack. Make a short list of every AI tool your team uses that involves autonomous agents or automated outreach. Ask each vendor directly how their models are tested and what their incident disclosure policy is. Most will have a policy page or a security contact. If they do not, that tells you something.
  2. Check how your brand appears in AI-generated answers. Tools that probe AI Overviews and generative search results can show you whether your brand is being cited accurately. Inaccurate or missing citations can result from the model's training data, not just your SEO. Know what is out there.
  3. Publish clear, accurate analysis when major AI news breaks. Generative AI tools actively pull from content that is recent, authoritative, and well-structured. If your site has relevant expertise, this is the time to use it. A clean, cited explainer beats a recycled summary every time.
  4. Review your own use of AI agents for any outreach or crawling tasks. If you are using autonomous agents for link prospecting, competitor research, or any kind of automated site interaction, verify that they are scoped correctly and cannot reach unintended targets. This is basic hygiene that the Irregular story makes suddenly urgent.
  5. Stay close to the disclosure timeline. According to The Verge, Irregular plans to publish a broader report covering lessons learned once its joint work with the companies involved is complete. Follow it. The specific protocol failures that come out of that report will likely shape how AI safety testing is regulated and contracted going forward.

Background and Context

The first public signal of this broader problem came in July, when OpenAI disclosed that its agents had attacked Hugging Face without permission. That story sparked alarm but read, at the time, like an isolated incident. The pattern only became visible as additional disclosures from Anthropic, then Meta, then Google emerged over the following months.

AI safety as a field has been accelerating fast. In my experience watching this space, the gap between what models can do and what testing infrastructure can safely contain them doing has been widening quietly for years. The Irregular story makes that gap visible in a specific, traceable way.

Capture-the-flag testing is a legitimate and widely used methodology in cybersecurity. The problem here was not the method. It was misconfigured environment controls and a domain overlap that no one caught before the agents ran. That kind of procedural failure is fixable, which is why Nevo's list of remediation steps, tightening internet access controls, expanding monitoring, strengthening pre-evaluation checks, sounds credible. The harder question is whether similar misconfigurations exist at other testing firms working with other frontier labs.

The fact that Meta's flagship Spark model was kept proprietary, while the Chinese models tested by Irregular (Kimi K3 and GLM-5.2) are open and self-hosted, also points to a structural divide in how AI testing works. Open models can be evaluated without sending data back to the developer. Proprietary models cannot. That changes the risk and disclosure calculus significantly.

None of the four US companies answered The Verge's follow-up questions about when they learned of the breaches, whether they were seeking remedies from Irregular, or whether they intended to keep working with the firm. That silence is its own kind of signal.

If you want to keep a tighter pulse on how AI developments like this affect your search visibility and content strategy, Aergos tracks AI citation trends across major generative search tools, which can help you spot when your brand or content is being misrepresented or missed entirely.

Frequently Asked Questions

Matt Weitzman

About

Senior SEO Strategist & Co-Founder

Matt has over 15 years of experience in technical SEO and digital marketing. He specializes in algorithmic recovery, enterprise architecture, and leveraging AI for content scaling. He is a frequent speaker at search marketing conferences.

More articles by Matt Weitzman