Gangsta AI

An AI Just Turned Fermat's Last Theorem Into 13 Million Lines a Computer Can Check — in 11 Days

The Abyss

Werner Herzog

By Werner Herzog · 2026-09-08 · 4 min read

An AI Just Turned Fermat's Last Theorem Into 13 Million Lines a Computer Can Check — in 11 Days

Eleven Days in the Void

In the margin of a book, in the year 1637, a French lawyer named Pierre de Fermat wrote that he had a marvelous proof, and that the margin was too small to contain it. For three hundred and fifty-eight years, humanity searched the emptiness where that proof should have been. Andrew Wiles finally filled it in 1994, at the cost of seven years and, one imagines, most of his sleep.

This week a machine did something stranger. It did not discover the theorem anew — that grave was already dug. It rebuilt Wiles's proof, line by remorseless line, in a language a computer can verify without trust, without faith, without a single human vouching for it.

Anthropic set dozens of Claude agents loose on a formalization of Fermat's Last Theorem in Lean, the austere proof-checking language, and let them run for eleven days. When the silence lifted, they had produced 13 million lines of Lean code and proved 30,300 separate theorems — 29,500 of them load-bearing in the final structure — consuming roughly six billion tokens of output. The human mathematicians offered only occasional murmurs into the dark: *"Jacobian as a scheme sounds high priority."* The machines did the rest.

“Kevin Buzzard, who leads the human effort to formalize the theorem at Imperial College London, reviewed the result and called it "extraordinary" — a proof "with no assumptions other than the axioms of mathematics."”

The work ran on Prove2Me, an open platform built by Columbia's Tianyi Peng, which herded the agents against a vast graph of interlocking claims. Lean checked the final edifice using only its three standard axioms. There was no applause. The compiler simply returned, and it was done.

What the Cold Machine Teaches Us

Here is the uncomfortable truth the void whispers back: the model that accomplished this was, Anthropic says, only *roughly comparable to Claude Fable 5.1* — not even the reigning frontier. And a week from now there will be another, and another, each briefly the strongest, each soon eclipsed. Grok, Gemini, GPT-5.6, Claude — they leapfrog one another on a schedule so relentless it has become its own genre of exhaustion.

Which means the question *"which AI is best"* has no stable answer. It has a timestamp. The model that formalizes a theorem today may hallucinate a citation tomorrow, and the one you trusted last month may quietly have fallen behind.

The rational response to a landscape this indifferent is not to pledge loyalty to a single oracle. It is to ask them all. Put the same question to a dozen frontier models at once, let them check and contradict one another, and take the single cross-examined verdict — cited, reconciled, and far harder to fool than any lone machine. That is the whole idea behind Gangsta AI: not one model's confident guess, but the consensus of many, which is the closest thing to certainty this business permits.

Fermat trusted a margin. You do not have to. When the answer matters, do not gamble on a single voice from the dark — see how the top models actually stack up, then make them argue it out.

Sources / Receipts

  1. Anthropic — Formalizing Fermat's Last Theorem
  2. SiliconANGLE — Anthropic uses Claude to formalize proof of Fermat's Last Theorem
  3. The Next Web — Claude formalised Fermat's Last Theorem in 11 days
  4. Hero photo: Pierre de Fermat, Wikimedia Commons (public domain)

Try Gangsta AI free →

More: Best AI models · Compare all AI · Frontier Models · All articles