AI theorem proving hits industrial scale — with a lesson about credit

Anthropic's Lean formalization of Fermat's Last Theorem landed without dispute, while OpenAI's 10,000-agent Navier–Stokes result awaits independent review. Multi-agent theorem proving is now routine — and so is the scrutiny.

Formal mathematics has become the proving ground where AI labs measure their research models, and this week's results mark the moment it became routine. Anthropic reports that a Claude research model produced a complete Lean formalization of Fermat's Last Theorem in 11 days, using roughly 6 billion output tokens, building on — and crediting — existing formalization work. Unlike the recent Navier–Stokes episode, this result drew no dispute.

The contrast with OpenAI's Navier–Stokes announcement is instructive. There, 10,000 agents reached a result in 88 hours, with Lean formalization taking another 17 hours across 4.9 million messages and around 300 billion output tokens — striking scale, but with independent review still pending and a credit and data-use dispute unresolved. Both efforts show the same pattern: massive parallel agent fleets, machine-checked proofs as the arbiter, and enormous token budgets. What separates a durable result from a contested one is not compute — it is verification, attribution, and whether the broader community can check the work. Fermat passes that bar cleanly; Navier–Stokes is still waiting.