Gangsta AI

OpenAI Says a Hidden Model Just Cracked Hundreds of Open Math Problems at Once. Mathematicians Are Not Celebrating — They're Arguing.

The Ledger

Warren Buffett

By Warren Buffett · 2026-10-07 · 6 min read

OpenAI Says a Hidden Model Just Cracked Hundreds of Open Math Problems at Once. Mathematicians Are Not Celebrating — They're Arguing.

On Monday, OpenAI did something you don't see every day: it handed the world's mathematicians a stack of homework and said, *grade this.*

The company published results on 377 long-standing open math problems — 722 manuscripts in total — all produced by an unreleased "internal frontier model." It didn't drop them in a press release and call it a day. It posted them to a public GitHub repository, with citation protocols, and — the part the pros actually care about — formalizations of many of the proofs in Lean, the programming language that lets a computer check a proof line by line. OpenAI says the average result took the equivalent of roughly three hours of ChatGPT Pro "thinking" compute.

That's the kind of number that makes headlines write themselves. *AI solves hundreds of math problems overnight.* And sure enough, the field got loud fast. The New York Times said the release was "further roiling" mathematics. New Scientist called it 722 discoveries "in one go." Somewhere a GPU is very tired.

“When somebody tells you they made a fortune overnight, the first question isn't 'how much' — it's 'show me the receipts.'”

Now here's where an old man gets cautious

I've watched a lot of people confuse activity with achievement. A model that attempts a thousand problems and posts the 377 it liked is doing something real — but it's also doing something very human: showing you its highlight reel. Several mathematicians quoted this week weren't dazzled; they were annoyed. Some argued the "results" were restatements, partial, or already within reach of existing methods. One essentially accused the machine of cutting in line — racing to publish territory humans were already closing in on.

OpenAI, to its credit, didn't just spike the football. It consulted the independent Advisory Group on Mathematics and AI at the Institute for Advanced Study and leaned on their recommendations for how to release this stuff responsibly. The Lean formalizations matter here: a proof a computer can verify isn't a vibe, it's a receipt. The ones that check out, check out. The rest are claims.

So which is it — a historic leap, or a very expensive highlight reel? Honest answer: both, and we don't know the ratio yet. That's not a dodge. That's just where the evidence sits this week.

The lesson isn't about math. It's about trust.

Here's the thing I keep coming back to. The problem on display this week isn't that one lab made a big claim. It's that you cannot grade one model's work using that same model. OpenAI marking its own math is like a company auditing its own books. The only reason anyone believes any of these 377 results is that *outside* checkers — Lean, the IAS group, rival mathematicians — get to tear into them.

That principle scales down to the rest of us. Every week some frontier model — ChatGPT, Claude, Gemini, Grok — claims a new lead, a new benchmark, a new "it beats everyone." And every week a different one leapfrogs it. If you're betting real decisions on a single model's confident answer, you're trusting the audited and the auditor to be the same guy.

The smarter move is the one the mathematicians just modeled for us: don't take one model's word — cross-check it. Ask the same question across several of the top models at once and see where they *agree*, where they split, and why. Agreement across independent systems is the closest thing to a Lean proof the rest of us get. That's the entire idea behind Gangsta AI — you pose one question, a dozen-plus frontier models answer, and you get a single cross-checked, cited verdict instead of gambling your decision on whichever lab shouted loudest this week.

OpenAI may well have done something remarkable. We'll know once enough skeptical humans — and a few stubborn proof-checkers — finish grading the homework. Until then, the oldest rule in my book still applies: *price is what the headline charges you; value is what survives verification.*

Want to see which of today's top models actually hold up when you make them show their work? Start with our honest, regularly-updated breakdown of the best AI models right now.

Sources / Receipts

  1. OpenAI — 'Sharing AI progress in mathematics' (official announcement, GitHub release + Lean formalizations, Oct 6 2026)
  2. The New York Times — 'OpenAI Releases Findings on 377 Math Problems, Further Roiling Field' (Oct 6 2026)
  3. New Scientist — 'OpenAI announces 722 mathematical discoveries in one go' (Oct 7 2026)
  4. Institute for Advanced Study — Advisory Group on Mathematics and AI, public recommendations (resolves 200)
  5. Hero photo: Sam Altman, 'Meeting with Masayoshi Son and Sam Altman' (cropped), Wikimedia Commons (CC BY 2.0)

Try Gangsta AI free →

More: Best AI models · Compare all AI · Frontier Models · All articles

'; })();