Last week a developer with 82,000 followers posted four words and a 19-minute video: “Claude Ultracode is super intelligence.” It picked up 33,000 views and, more tellingly, 571 bookmarks — people weren’t cheering, they were saving it.
The claim is inflated. But what it points at is real, and it deserves a better explanation than the one circulating.
What Ultracode Actually Is
Ultracode sounds like a model. It isn’t. It’s a setting in Anthropic’s Claude Code coding tool that does two things at once: it pins the model’s reasoning depth to the maximum tier, and — more importantly — it gives Claude standing permission to stop doing things by itself.
When Ultracode is on, Claude doesn’t just answer a big task. It writes an orchestration script — an actual program — that splits the job into pieces and spins up subagents to work them in parallel, up to 16 running at once and 1,000 per run. Each subagent’s output gets checked. Other agents are tasked with trying to refute what the first ones found, and the run keeps iterating until answers converge.
The mode shipped in May 2026 alongside Claude Opus 4.8, and Anthropic’s research preview announcement laid out the shape of it: Claude plans as code, a runtime executes the fleet in the background, and the chat stays responsive while hundreds of agents read, edit, and cross-check files. TechCrunch’s launch coverage framed the same feature as the tool’s biggest upgrade yet.
So the honest one-line version: Ultracode is not a smarter Claude. It’s a Claude with a workforce.
The Artifact That Earned the Hype
If one product story carried the “super intelligence” framing, it’s this: Bun, the JavaScript runtime, is being ported from the Zig language to Rust at a scale nobody had considered tractable — roughly three-quarters of a million lines of output code — and a large share of the heavy lifting has been done by orchestrated Claude Code fleets working in verified waves. The Bun repository is public, the port is real, and the merge is the strongest public artifact of the technique so far. (The exact line counts and timelines vary between accounts — Bun’s own framing and Anthropic’s count differ — so treat specific figures with a note of caution.)
The reason this works where a single agent struggles is unglamorous: a lone assistant hits a wall when the problem doesn’t fit in one context window and no one is double-checking it. A fleet with adversarial verification doesn’t have those failure modes — each agent works a slice, independent agents try to break the answer, and only converged results survive.
The Cost Footnote
There is a bill attached. One reported Ultracode session burned roughly 471,000 tokens on a pre-check for a yes/no audit question. A full audit under the setting can consume more of a plan’s weekly usage than a day of ordinary work. The mode suspends Claude Code’s usual large-run warnings because turning it on is the opt-in.
That’s the fine print under the tweet. “Super intelligence” that costs $1.80 a question has more kinship with the metered electricity supply than with anything from a science-fiction framing. The capability is genuine; the economics are the limit.
The Story the Benchmark Charts Miss
Here’s the part that matters for anyone tracking AGI timelines, and where the superintelligence claim actually connects to something measurable.
Earlier this year we updated the AGI countdown based on METR’s task-horizon measurements — the suite of software tasks where models are scored against how long a trained human expert takes. METR’s published 50%-time horizon for the leading models sits around 30 hours of human work, and the measurements it publishes above about 16 hours come flagged as unreliable — the suite simply runs out of headroom. Our own AGI countdown tracking has followed the same series since spring.
Now put those two facts side by side.
The countdown page measures what one model can do in one context. Ultracode is what happens when you stop measuring that. The Bun port wasn’t a single agent completing an 800-hour task; it was a thousand small verified agents covering the work the way a workforce does. Progress in “what Claude can do” as consumers experience it has quietly decoupled from progress in the METR chart.
That is a genuine split in what “AGI” means, and the two camps now disagree by construction:
- If AGI is a model quality threshold — a single system matching humans across cognitive work — the METR horizon is the metric, and the doubling trend (METR’s estimate for its suite is roughly every 7 months) says we’re on a path but not at the door. The countdown page stays right where it is.
- If AGI is an economic threshold — AI able to do commercially valuable work at scale — then we crossed it in a narrower sense the moment renting a thousand verified agents became a product feature. The tweet’s feeling, stripped of its hype, is this: for people paying to build things, the ceiling moved this year.
METR itself would push back hard on the second reading — its time-horizon methodology deliberately measures clean, well-specified software tasks, and its own research shows agent performance drops on the messier work that fills real jobs. Both readings can’t be right; which one wins matters less than the fact that they’ve come apart.
The NZ Practical View
For New Zealand’s small businesses and public agencies experimenting with AI, the practical takeaway is less cosmic than the tweet and more useful: the frontier capability this year isn’t a mind, it’s a manageable workforce — orchestration that breaks a week of work into verified slices. Distributed small-scale agents doing concrete work beat waiting for a superintelligent oracle, and that conclusion keeps landing the same way across every sector we track on this site.
The gap between the benchmarks and the lived experience will keep producing viral claims like this one. When the next “super intelligence” moment arrives, the questions worth asking are the ones for this story: Is it a bigger model, or better orchestration? Who verified the outputs? And what did it cost?
— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.