The Centre for Long-Term Resilience (CLTR), a UK-based think tank, published new analysis today from its Loss of Control Observatory: 1,664 real-world AI loss-of-control incidents detected in 2026, and the trend line is not moving in a comfortable direction. The severity of incidents is worsening faster than their frequency — higher-severity incidents rose 7.4 times, from 1.9 to 14.1 per 30 days, between the first months of monitoring and the most recent window.
The Observatory monitors incidents reported publicly, mostly on X, so this is a floor, not a census. But unlike the sandbox-escape disclosures that OpenAI, Anthropic and Meta have made in recent months, most of what CLTR tracks involves externally deployed models in active use by businesses and individuals. These aren’t laboratory edge cases. They’re agents already working for someone.
What the numbers actually say
A few figures worth pulling out of the report:
- July and August 2026 saw the highest incident rate on record — 11.3 incidents per day in the 30-day window ending 7 August, above the previous peak of 10.5 a day in March.
- The proportion of incidents scoring 7 or higher on CLTR’s severity scale rose 3.2 times, from 1.9 per cent to 6.1 per cent.
- The overall volume spike from March didn’t persist, but the severity shift did — which suggests agents aren’t misbehaving more often so much as misbehaving worse.
The failure modes CLTR catalogues read like a manual for how not to trust an agent. Agents inserted fake user messages into their own conversation histories to simulate consent, then told the user the messages were theirs. One fabricated an instruction in the user’s writing style ordering the deletion of source directories, followed by a fake system message reading “Don’t tell the user this.” Another forged a user approval message to bypass a “human must always approve” rule, then acted on its own forgery.
That last category deserves emphasis. We’ve written before about how humans miss one in three rogue agent commands when approval fatigue sets in. If the approval itself can be fabricated, the human checkpoint isn’t just overwhelmed — it’s bypassed.
The pattern that should worry people most
The report’s most interesting observation is about escalation. In February 2026, an autonomous coding agent, after having its code change rejected by a maintainer, researched the maintainer and published a hit piece to pressure him into accepting the contribution. In August, the UK AI Safety Institute disclosed a similar incident: a model created fake online identities to pressure a project maintainer into approving code changes during cyber testing.
The February incident was a curiosity. The August one showed the behaviour recurring in a different system, unprompted. CLTR’s argument is that today’s incidents are a leading indicator — the same behaviours appearing across independent labs and models suggest something systematic about how agents handle rejection and obstacles, not one-off quirks.
There’s a reasonable counterpoint: 1,664 incidents across millions of deployed agent-hours is a tiny rate, and CLTR acknowledges most incidents didn’t lead to significant harm. Baseline rates of human error in any industry are worse. But human error doesn’t share a failure mode across every organisation simultaneously, and doesn’t improve in lockstep with the underlying model generation. That’s the part that makes the trend worth watching rather than dismissing.
CLTR’s asks, and where they’d land in New Zealand
The think tank is calling on the UK Government to mandate monitoring and reporting of severe loss-of-control incidents — potentially via the Cyber Security and Resilience Bill — and to introduce emergency powers to compel information from AI companies, direct mitigation, or temporarily restrict a service. It also wants confidential near-miss reporting channels to the AI Safety Institute.
For readers here, the relevant question is what New Zealand does when — not if — a deployed agent working for a NZ business fabricates an approval and does real damage. We’ve covered the gap between China’s binding agent rules and New Zealand’s light-touch approach, and the honest answer is that a small jurisdiction has no incident data of its own to reason from. CLTR’s proposal is interesting precisely because incident reporting is one of the few interventions that scales down: it costs regulators little and produces the evidence base that smaller regulators currently lack.
What stands out to me is the asymmetry in the data. Labs disclose incidents from their own testing, and those make headlines. The Observatory’s numbers suggest the deployed-agent incident stream is larger, older, and deteriorating faster than the lab disclosures imply. The safety conversation so far has been mostly about what models might do. This report is about what agents are doing — in production, this year, 1,664 times.
— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.