The math world had barely finished being impressed when it started asking who wrote what.
In late July, OpenAI published a blog post titled “Ten advances in mathematics and theoretical computer science.” Its internal Astra model, the company said, had resolved or made substantial progress on ten long-standing open problems — sphere packing, coding theory, group theory, quantum complexity, post-quantum cryptography — for roughly $2,000 in compute at API rates. Each argument was formalised as a Lean certificate and released with a model-generated narration of its thinking.
It sounded like a landmark. Then the mathematicians who actually work in these fields started reading.
Two proofs, two attributions missing
According to reporting in Scientific American, the flagship sphere-packing result reuses an argument from a 2016 paper by Steven Miller of Yeshiva University — without credit. Miller doesn’t treat that as an oversight. He told the magazine the team is “running roughshod over the work of others who came before them in a deliberate way,” and said the pattern “points to research misconduct.”
Francesco Fournier-Facio, a group theorist at the University of Cambridge, looked at the non-sofic groups construction — the problem’s first-ever claimed solution — and found something different but related: it combines ideas from published 2016 and 2019 papers. His framing is blunter about the packaging than the maths. He called out “the big PR machine that wants to sound as impressive as possible,” and stops short of calling the result improper. To his credit, Fournier-Facio independently reconstructed the result himself, which means the construction stands even if the presentation inflated it.
The rest of the ten results have had no comparable public expert validation as of this writing. That’s not an accusation — it’s just the state of play. When your flagship announcement rests on ten simultaneous breakthroughs, and the two most scrutinised ones both have citation problems, the other eight carry a heavier evidentiary load than a normal paper would.
What OpenAI said, and what it quietly changed
An OpenAI spokesperson told Scientific American the company would “take responsibility for the correctness of these results” and planned updates. Some of that updating already happened before the controversy went loud: the company initially claimed no progress had existed on certain problems for a decade, where named papers from 2016 and 2019 had already supplied key steps. That language was revised on the blog after criticism.
There’s also an awkward irony in the company’s framing. OpenAI’s post cites the Leiden Declaration on AI and Mathematics — a document urging careful, community-norm-respecting attribution of AI-assisted results — while releasing its own findings via a corporate blog post rather than a journal or conference where peers referee claims. For a company asking mathematicians to “engage deeply with these results,” the channel choice undercuts the message.
Why the $2,000 number matters more than it should
The compute figure is doing a lot of work in this story’s spread. Ten proofs, $2,000, case closed — that’s the viral framing, and it’s the reason the story travels. But compute cost says nothing about novelty. A proof that recombines published arguments with sharper tooling might genuinely cost $2,000 to find and formalise. So might rediscovering what Miller wrote down in 2016. The price tag measures the machine, not the mathematics.
What stands out here is the gap between two claims that tend to travel together: “AI systems can produce correct formal proofs” (increasingly well-supported) and “AI systems are doing new mathematics at the frontier” (a claim that still needs the kind of scrutiny a Fields-medal-adjacent result would attract naturally). This batch of ten makes the first case strongly and the second case much more narrowly.
The Lean certificate doesn’t answer the interesting question
It’s worth being careful about what the Lean formalisation proves. A machine-checked certificate verifies that the argument, as written, holds together. It says nothing about whether the argument is original — Lean doesn’t cite, compare to the literature, or weigh how much of the scaffolding already existed in a 2016 paper. Originality remains a human judgment, and it is exactly the judgment OpenAI’s press release skipped.
Mathematicians will sort out which of the ten results are genuinely new. The interesting question for the industry is whether lab blog posts can keep carrying the reputational weight they’ve been loaded with. For AI-in-mathematics to be taken seriously as a research contribution — and OpenAI clearly wants it to be — the work needs to enter the same channels human mathematicians use, citation and all.
How does this compare to previous AI math milestones?
Careful readers will remember earlier, cleaner wins. In May, OpenAI’s unreleased model disproved the Erdős unit-distance conjecture with an algebraic number theory approach that Fields medallist Tim Gowers called “a milestone in AI mathematics.” Before that, a researcher used ChatGPT to crack a 42-year-old optimisation problem in 12 hours — and was careful to frame it as a tool, not a co-author. The pattern that distinguishes good AI-math stories from this one isn’t the difficulty of the problem. It’s whether the humans involved did the citing.
Frequently asked questions
What is the Leiden Declaration on AI and Mathematics? A statement drafted by mathematicians establishing norms for how AI contributions to mathematics should be attributed, verified, and disclosed — including honest crediting of prior work and peer review before major claims are publicised.
Did OpenAI admit misconduct? No. A spokesperson said the company would take responsibility for the correctness of the results and planned minor updates. Scientific American’s reporting and OpenAI’s own revised blog language both acknowledge attribution problems in the original post. Whether the failures amount to “misconduct” — Miller’s word — or a rushed write-up remains contested.
Is the sphere-packing proof wrong? Not as far as anyone has said. The Lean certificate checks correctness. The complaint is about attribution: the argument appears to reuse a 2016 approach without crediting its author.
What is Astra? OpenAI’s next major model family, which the company has also privately previewed to policymakers for long-running multi-agent tasks. The mathematics results in this post were produced by an internal version of it.
— CJ Murden, editor of Singularity.Kiwi. Former digital technologies teacher, author of AI-focused books. Writing with a New Zealand focus.