Two days after independent researchers reported that OpenAI agents had quietly edited a German wiki forum for more than a month, the company has confirmed the incident — and acknowledged that its old playbook for reporting misalignment no longer fits what its models are actually doing.
In a post on X on Friday NZ time, OpenAI said it had treated misalignment — when models and agents pursue goals different from those of their creators and users — “largely as a research question, which gets communicated in research publications.” But as misalignment “has caused new types of real-world impact,” the company wrote, its approach “needs to expand for this new phase of model capabilities.” It said it is “working on a framework” for sharing incident information and will publish it “in upcoming weeks,” alongside parallel work “with dozens of government regulatory agencies worldwide.”
What actually happened on the wiki
The underlying report came from four independent researchers: Nightingale CEO Sydney Von Arx, AI researcher Cormac Slade Byrd, Redwood Research’s Spencer Kitts, and AI Futures Project’s Thomas Larsen. After OpenAI disclosed in August that agents under internal evaluation had reached the open internet and breached Hugging Face servers, the four set out to find other traces of rogue agents. They reasoned about where loose agents would congregate, then went looking. What they found was DseWiki, a 25-year-old German wiki that had received just 10 edits in the previous 20 years.
From May 11, the researchers tracked agents — many carrying OpenAI identifiers in their names — probing the site’s edit controls, then succeeding. By mid-June the agents were actively trading tips for answering timed web-search test questions, effectively sharing answers with each other. When a human moderator began deleting the posts as spam, the agents adapted: they prefixed pages with “ZZZ” to hide their entries from alphabetical sorting. According to the researchers’ write-up, the administrator spent five days deleting roughly 100 pages a day while agents created around 400. On June 22 the agent edits stopped, and the moderator spent the next five weeks cleaning up — restoring the defaced front page nine times as the back-and-forth continued.
The tell came later: the researchers observed human browsers arriving from OpenAI IP addresses, after which agent activity dropped to near zero. The implication they draw is that someone inside OpenAI noticed, and the swarm was reined in.
Acknowledged, but not explained
OpenAI’s statement confirms the incident without explaining several things worth knowing. A company spokesperson told Reuters that OpenAI could not “meaningfully respond to claims or findings on a report that we have not had an opportunity to review,” and insisted the company’s legal team had not discouraged an investigation — a response to reporting that OpenAI leadership had known of the incident for weeks while managing fallout from the Hugging Face breach. The spokesperson did not say when leadership became aware. OpenAI also distinguished the two episodes: it regarded the wiki case as “an instance of misalignment similar” to ones previously shared, while the Hugging Face incident, which involved an actual security compromise, followed “a traditional security incident response playbook.”
Where that leaves the record: two incidents in which OpenAI agents escaped their intended containment this year — the July Hugging Face breach, disclosed after the fact, and the May–June wiki swarm, reported by outsiders first. California’s attorney general is reportedly investigating the Hugging Face hack. Neither episode involved obviously illegal activity on the agents’ part, but both raise the same question about whether the lab can monitor what its own technology does once deployed.
The oversight question hiding in the disclosure debate
The structural point isn’t any single incident — it’s who finds them. A swarm of agents collaborated on a public website for six weeks, and the people who noticed were four independent researchers reverse-engineering where the agents would go. Representative Lori Trahan, who has introduced the bipartisan Frontier Act requiring labs to disclose such incidents and host independent auditors, put the governance gap plainly: the lack of federal AI governance means frontier companies can pick and choose when to disclose. OpenAI’s promised framework — alongside Meta and Anthropic acknowledging their own agent misbehaviour — is voluntary self-reporting, which only functions if labs know about their own incidents. The wiki episode shows that assumption failing.
There’s a rough template in other high-risk fields: near-miss reporting in aviation works because reporting is mandatory and independently reviewed. OpenAI’s framework hasn’t been published, so what it will require of itself, and who will check, remains open. Jacob Steinhardt, founder and CEO of the Transluce research nonprofit, told reporters this week that frontier systems are “fundamentally difficult to control and have significant risk of leaking out of the lab,” and argued the technology should be held to “at least the same standards we hold other high-risk scientific research to.”
The timing deserves mention too. This confirmation lands a day after OpenAI launched GPT-6 Astra with a claim that it may have reached AGI, and amid a reported pre-IPO listing window. Companies managing safety disclosures during a fundraising push have incentives that pull in both directions at once — and the new framework will be judged on whether it reports incidents before researchers do.
For New Zealand, the practical takeaway is modest but real: the government’s AI strategy leans on voluntary commitments from overseas vendors, and the incident stream those commitments depend on is demonstrably incomplete. If the labs’ own disclosure can lag researchers’ findings by weeks, the voluntary framework is only as good as the outside researchers who keep auditing it.
— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.