Two mathematicians working through equations at a whiteboard in a university office
AI & Singularity

Claude Formalised Fermat's Last Theorem in 11 Days. The Maths Is Old. The Method Is Not.

Claude spent 11 days and 13 million lines of Lean formalising Wiles's 1995 proof of Fermat's Last Theorem. Kevin Buzzard has checked it. What it means for how mathematics gets done.

AnthropicClaudeMathematicsLeanFormal Verification

Kevin Buzzard had a plan to formalise Fermat’s Last Theorem. He had a blueprint, a multi-year community project, and Imperial College behind him. Then Anthropic’s researchers went away for 11 days and came back with the finished article.

On September 4, Anthropic published the first complete computer-checked proof of Fermat’s Last Theorem, written in the Lean proof assistant by Claude working largely autonomously. The numbers are the kind you read twice: 13 million lines of Lean, roughly 30,300 intermediate theorems proven (29,500 of which made the final cut), and about six billion output tokens from a research model Anthropic describes as roughly comparable to Claude Fable 5.1. The formal proof is over five times the size of Mathlib, the entire community proof library it builds on.

The maths is 30 years old. That’s the point.

Fermat’s Last Theorem needs no introduction, but its AI angle needs a careful framing, because this was not discovery. Andrew Wiles proved the theorem in 1995, in a 129-page paper that took a year of patching after a reviewer found a critical gap two months into checking it. What Claude produced is a formalisation: a translation of that argument into Lean, a language where a computer checks every logical step mechanically, with no reliance on a human referee’s stamina or judgment.

The distinction matters more than it first appears. A formal proof uses nothing but the axioms of mathematics — Buzzard, who reviewed the result, confirmed it proves the theorem “with no assumptions other than the axioms of mathematics.” A comparator tool also confirmed the statement matches Mathlib’s own statement of FLT. Nobody has to trust Anthropic’s word for anything. Lean did the trusting.

For years, the assumption was that a full FLT formalisation would take a coordinated community effort years to complete. Buzzard’s own project, kicked off in 2024, was building toward exactly that. His reaction, on the Xena project blog, was less competitive than delighted: “Congratulations to Anthropic!” He noted the last item on Freek Wiedijk’s famous list of 100 formalisation challenges is now closed, wrapping up a twenty-year benchmark.

What actually broke the problem

The interesting engineering detail is that Claude’s first attempts failed. Agents made early progress, then lost track of the project’s state and stopped collaborating effectively. Failed runs still contributed around 7 per cent of the non-boilerplate lines in the final proof.

The breakthrough wasn’t a smarter model. It was shared state. Anthropic researcher Tianyi Peng and collaborators at Columbia University built Prove2Me, an open collaborative platform that maintains a directed acyclic graph of theorem statements. Agents consult the graph to decide what to prove next, which solves two problems at once: memory degradation over long tasks, and parallel agents tripping over each other. Human input was limited to occasional high-level nudges — “push the Mazur theorem to be done soon” — which is about as close to managing a team of PhD students as it is to programming.

The proof itself follows the Darmon–Diamond–Taylor exposition of the Wiles–Taylor–Wiles argument, running through Langlands–Tunnell and Ribet’s level-lowering theorem. The repository develops Fontaine theory and enough of Mazur’s work on the Eisenstein ideal to rule out any Frey curve. If none of that means anything to you, that’s rather the point: the machine now does the parts that take mathematicians years to learn, and hands back a certificate Lean has already checked.

Why verification beats discovery

A proof of the Riemann hypothesis from an AI would be a new result, and we’d spend a decade deciding whether to believe it. This is different, and quietly more consequential. As Anthropic put it: AI produces more and more claims; the bottleneck in mathematics is that verifying a new result can take months or years of expert effort. If formalisation becomes routine, the checking gets cheap and automatic, and the trusted body of mathematics grows faster than it decays.

There’s a reasonable objection: the model burned six billion tokens on a theorem humanity already possessed. That’s a fair description of a demo. But the demo was the point — this was the hardest item on the most famous formalisation checklist in mathematics, and it fell in under two weeks. The cost of that compute buys a template every mathematics department in the world can now borrow.

What stands out to me is what this says about agent workflows. The model didn’t get smarter between the failed attempts and the successful one. The scaffolding did the work — a shared graph of what’s proven, what’s pending, and what to attack next. That’s the same lesson every serious agent deployment is learning, whether the domain is mathematics or warehouse logistics.

For New Zealand’s universities, the signal is less about Fermat and more about the toolchain: Lean skills, formal methods, and agent orchestration are about to be in demand well beyond pure maths. A country with a small research workforce should read that as an opportunity rather than a threat.

— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.

Sources: https://www.anthropic.com/research/formalizing-fermats-last-theorem, https://github.com/anthropics/fermats-last-theorem, https://xenaproject.wordpress.com/2026/09/04/flt-anthropic-has-beaten-me-to-it/, https://www.techtimes.com/articles/326745/20260905/fermats-last-theorem-machine-checked-claude-completes-11-day-formalization.htm