On September 4, 2026, Anthropic published an unusual milestone: an advanced prototype of Claude produced the first complete, computer-verified proof of Fermat's Last Theorem , written in the Lean proof assistant. This is not a "seems right" paper: Lean checks the logic step by step. In 11 days, the model wrote approximately 13 million lines and proved tens of thousands of intermediate theorems.
For product teams, AI startups, and SMEs already delegating reasoning to models, the useful message isn't "AI is already Wiles." It's this: when the volume of generated results exceeds what humans can manually review, formalization (translating the argument into verifiable code) becomes a layer of trust. This article summarizes the official announcement, the Prove2Me framework, and what questions to ask before including "auditable AI" in your roadmap.

According to Anthropic's research post:
Kevin Buzzard (Imperial College London), who leads the community effort to formalize FLT, called the result an extraordinary achievement of self-formalization and a step towards tools that ease the burden on referees and detect errors in the mathematical corpus.

Anthropic is explicit: the first multi-agent attempts stalled. Agents lost the project status and stopped collaborating effectively. The breakthrough came with the use of Prove2Me , an open collaborative formalization platform (Peng et al., Columbia), which:
With that framework and a Claude Code-style multi-agent harness, the team closed the campaign in less than two weeks. Anthropic estimates that the internal research model, comparable to Claude Fable 5.1, generated around six billion tokens . Previous failed attempts contributed approximately 7% of the non-boilerplate lines in the final artifact.

Three practical readings:
Brief checklist if you sell or buy "verifiable AI":

You don't need to formalize Fermat. You do need to decide which parts of your system (payments, permissions, pricing, compliance) deserve a more rigorous approach than simply "the LLM said it's okay." Start with a narrow pilot: a formal specification or a set of properties with automated checkers; measure QA acceptance rate and time to green artifact. If you need to finalize that pipeline—product, landing page, and integration—at Presticorp we work by stage: startups , SMEs , and e-commerce .
Treat the formalization of FLT as a sign of trust infrastructure , not as "superintelligence" marketing. This week, define: (1) which claim of your product would require a checker, (2) who writes the statement, and (3) what token and human review budget you're willing to accept. When AI generates more than it can read, the winner is the one with verification, not the one with the longest prompt.
Editorial note: Line counts, theorems, tokens, and timelines are quoted as they appear in Anthropic's post of September 4, 2026. Lean/Mathlib and the public repo status may change; please check the GitHub linked from the announcement on the day you cite numbers in a business brief.
Enviando comentario…
Si tu proyecto requiere una solución más enfocada, entra directo a la landing ideal para tu negocio y envíanos tu información en el formulario correspondiente.
0 Comentarios