<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Kevin Buzzard | guIA</title><link>https://guia.desdeelsur.org/en/tags/kevin-buzzard/</link><atom:link href="https://guia.desdeelsur.org/en/tags/kevin-buzzard/index.xml" rel="self" type="application/rss+xml"/><description>Kevin Buzzard</description><generator>HugoBlox Kit (https://hugoblox.com)</generator><language>en-us</language><lastBuildDate>Sat, 05 Sep 2026 00:00:00 +0000</lastBuildDate><image><url>https://guia.desdeelsur.org/media/sharing.png</url><title>Kevin Buzzard</title><link>https://guia.desdeelsur.org/en/tags/kevin-buzzard/</link></image><item><title>Thirteen million lines in the margin: what makes this the good case</title><link>https://guia.desdeelsur.org/en/blog/2026-09-05-trece-millones-de-lineas-en-el-margen/</link><pubDate>Sat, 05 Sep 2026 00:00:00 +0000</pubDate><guid>https://guia.desdeelsur.org/en/blog/2026-09-05-trece-millones-de-lineas-en-el-margen/</guid><description>&lt;p&gt;&lt;strong&gt;On:&lt;/strong&gt; &amp;ldquo;Formalizing Fermat&amp;rsquo;s Last Theorem&amp;rdquo; — Anthropic, 4 September 2026.
·
&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;Eleven days, more than thirteen million four hundred thousand lines of Lean, twenty-nine thousand five hundred intermediate theorems and some six billion output tokens. With that, several dozen Claude agents produced the first end-to-end computer-checked proof of Fermat&amp;rsquo;s Last Theorem, closing out along the way the list of a hundred theorems Freek Wiedijk has kept open for twenty years. The mathematician who runs the project to formalize that same theorem reviewed the result and wrote that, mathematically, this tells us essentially nothing. He is right, and that is exactly why it is worth looking at.&lt;/p&gt;
&lt;h2 id="what-is-inside-the-artifact"&gt;What is inside the artifact&lt;/h2&gt;
&lt;p&gt;The version doing the rounds — the machine cracked in eleven days what took Wiles seven years — is false in both halves, and Anthropic itself does not claim it. Andrew Wiles proved the theorem in 1995, in a hundred and twenty-nine pages that took months to review and needed a patch a year later for a hole found during that review. What the agents did was not find a proof but transcribe one: they followed the simplified exposition by Darmon, Diamond and Taylor and poured it into a language whose compiler takes nothing on trust. The eleven days are wall-clock time for dozens of agents in parallel, coordinated through Prove2Me — an open platform from Tianyi Peng&amp;rsquo;s group at Columbia that maintains the dependency graph between theorems and tells each agent what is missing — with human intervention limited to occasional very high-level instructions, on the order of &lt;em&gt;the Jacobian as a scheme sounds high priority&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;The result is also narrower than the headline suggests: the formalization covers prime exponents greater than or equal to seventeen, and the rest was already done by humans.&lt;sup id="fnref:1"&gt;&lt;a href="#fn:1" class="footnote-ref" role="doc-noteref"&gt;1&lt;/a&gt;&lt;/sup&gt; Buzzard, who had gone on record before any of this saying he was 99.9% sure Fermat&amp;rsquo;s proof was correct, did not revise that figure after reading the repository. There is no new mathematical knowledge here. There is an engineering artifact at a scale that did not exist a month ago, and what it tells us is not about Fermat but about what these systems can do when someone puts something on the other side that can tell them no.&lt;/p&gt;
&lt;h2 id="the-kernel-does-not-ask-who-wrote-it"&gt;The kernel does not ask who wrote it&lt;/h2&gt;
&lt;p&gt;Anthropic writes, in a sentence that in institutional prose amounts to a confession, that formalization is a place where they feel unambiguously good about the role of AI. It is worth reconstructing why before arguing with it, because the reason is a good one and has nothing to do with the model&amp;rsquo;s virtues. Lean checks every step against a tiny kernel that admits only three axioms, and that kernel is indifferent to the provenance of what it checks: it does not know whether the lines were written by a doctoral student, a swarm of agents or a very lucky random generator, and its verdict costs a few hours of compute where producing what it checks cost six billion tokens.&lt;sup id="fnref:2"&gt;&lt;a href="#fn:2" class="footnote-ref" role="doc-noteref"&gt;2&lt;/a&gt;&lt;/sup&gt; An unreliable producer coupled to a cheap, independent, hostile verifier yields a reliable system. That is the whole trick, and it is a property of the domain, not of the model.&lt;/p&gt;
&lt;p&gt;What survives of the standard objection — that nobody read the thirteen million lines and therefore nobody knows what they say — is smaller and sharper than it looks. The kernel certifies that the lines entail the statement from the axioms; it does not certify that the statement is the one you wanted to prove. That is why the team ran a comparator tool against Mathlib&amp;rsquo;s own formulation of the theorem, and why Buzzard went and looked at the statement. That step, the shortest in the whole process, is the only irreducibly human one, and no quantity of tokens shortens it: someone has to answer whether the line sitting on top of thirteen million others says what the world means by &lt;em&gt;no positive integers satisfy the equation&lt;/em&gt;. The entire trustworthiness of the artifact rests on a three-line reading.&lt;/p&gt;
&lt;h2 id="a-gift-the-commons-cannot-lift"&gt;A gift the commons cannot lift&lt;/h2&gt;
&lt;p&gt;None of this happens without Mathlib. The community library the proof stands on was written over years by hundreds of mathematicians, largely unpaid for it; Anthropic&amp;rsquo;s repository credits a hundred and six files taken from Buzzard&amp;rsquo;s FLT project at Imperial College — funded by the UK&amp;rsquo;s EPSRC with a million pounds over five years — and from the flt-regular project. The resulting proof is over five times the size of Mathlib and takes nearly twenty times as long to compile, on a ninety-six-core machine. It is on GitHub, openly licensed, and it will probably stay there.&lt;/p&gt;
&lt;p&gt;That is the part worth looking at slowly, because it is not an enclosure and calling it one would be more comfortable than accurate. Nothing was taken: Mathlib is intact, the code is published, and in a year when model releases have consisted of choosing which layer to open and which to charge for, a complete and auditable repository is more than usually shows up. The problem is of a different kind. Mathlib is a commons because it can be maintained: someone reads a file, understands what it does, generalizes it, refactors it, argues about the name of a lemma on Zulip. Thirteen million four hundred thousand lines written to satisfy a compiler do not admit that treatment, and Buzzard&amp;rsquo;s guess — a well-founded one — is that Anthropic will not do the work of turning them into something the community can absorb. His two stated goals, contributing the fundamental objects of modern number theory to Mathlib and building a dynamic document that lets a human walk through the proof, remain his and remain pending. A commons is measured not by what can be downloaded but by what someone can sustain; by that measure, what came back to the commons is not the proof but the news that the proof is possible.&lt;/p&gt;
&lt;h2 id="three-subscriptions"&gt;Three subscriptions&lt;/h2&gt;
&lt;p&gt;There is a second experiment in the announcement that matters more to a reader in this region than the thirteen million lines: a group of agents running on three personal Claude Max subscriptions, coordinated by the same platform, formalized Vinogradov&amp;rsquo;s three primes theorem in three days. That is not a show of force; it is a change in the entry price, and it deserves to be conceded in full. Until this week, formalization was the one branch of contemporary mathematics whose barrier was time rather than capital: Lean runs on a laptop, Mathlib is free, and the community accepts contributions from anyone whose code compiles. For departments that will not be buying a GPU cluster this decade, it was the open door — it is still open, and now it takes less time to walk through.&lt;/p&gt;
&lt;p&gt;The objection is not about access to the tool but about the unit of measurement. If what used to be a doctoral thesis&amp;rsquo; worth of formalization becomes three days of agents, then what counts as &lt;em&gt;a project&lt;/em&gt; in the field is set by whoever can pay for the largest swarm, and the distance between three subscriptions and dozens of agents on an internal model is not one of price but of availability: the model that did Fermat is not for sale.&lt;sup id="fnref:3"&gt;&lt;a href="#fn:3" class="footnote-ref" role="doc-noteref"&gt;3&lt;/a&gt;&lt;/sup&gt; It is also worth naming what is &lt;em&gt;not&lt;/em&gt; here, because critical vocabulary wears out when used reflexively: there was no data extraction from the South, no ghost annotation labour, no local corpus turned into raw material. The asymmetry is of another kind, simpler and harder to reverse, and it consists in the fact that the capacity that produced the result can neither be hosted here nor bought abroad.&lt;/p&gt;
&lt;h2 id="what-else-has-a-kernel"&gt;What else has a kernel&lt;/h2&gt;
&lt;p&gt;The transferable lesson of these eleven days is not that machines can do mathematics. It is a criterion, and a more demanding one than it looks: AI works unsupervised where there exists a verifier that is cheap relative to production, independent of the producer, and public. It is worth walking through the deployments actually under discussion in the region with that criterion in hand. A benefits-allocation system has no cheap verifier: checking that a denial was correct costs more than issuing it, and the affected person finds out when the transfer does not arrive. An automated exam grader has no independent verifier: it is audited by whoever bought it. Medical triage has no public verifier: the truth arrives months later, scattered across records nobody cross-references. All three are being deployed anyway, and the difference from Fermat is not one of risk or ambition but of epistemic infrastructure.&lt;/p&gt;
&lt;p&gt;Lean and Mathlib have been under construction for over a decade, largely on volunteer labour and public funding, and without that kernel this week&amp;rsquo;s result would not be a result but a thirteen-million-line file nobody would have reason to believe. The question it leaves open is not whether the machine proves theorems. It is how long a kernel takes to build, on whose money and under whose responsibility, when what needs verifying is not theorems but case files.&lt;/p&gt;
&lt;div class="footnotes" role="doc-endnotes"&gt;
&lt;hr&gt;
&lt;ol&gt;
&lt;li id="fn:1"&gt;
&lt;p&gt;The proof Claude formalized holds for prime exponents p ≥ 17, which is as far as the route through Fontaine&amp;rsquo;s theory and Mazur&amp;rsquo;s work on the Eisenstein ideal reaches. The small exponents were already covered: Best, Birkbeck, Brasca, Rodriguez, van der Velde and Yang had formalized the case of odd regular primes, and the smallest irregular prime is 37, so 3, 5, 7, 11 and 13 come in that way. The union of the two pieces closes the theorem, which means the complete result is, in this respect too, a collective object.&amp;#160;&lt;a href="#fnref:1" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:2"&gt;
&lt;p&gt;Lean&amp;rsquo;s kernel is the piece of software whose correctness you have to assume in order to believe anything else, and it is deliberately tiny; the three axioms are propositional extensionality, classical choice and soundness of quotients. The strategy — a small, auditable checker verifying proofs produced by anything at all, including heuristics with no guarantees — is known as the de Bruijn criterion and is half a century old. That it turns out to be exactly the architecture that makes a probabilistic producer tolerable today is a historical coincidence deserving more attention than it got this week.&amp;#160;&lt;a href="#fnref:2" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li id="fn:3"&gt;
&lt;p&gt;An order of magnitude, with every caveat attached: six billion output tokens, at the public rate of a model comparable to the one used (fifty dollars per million), come to about three hundred thousand dollars. The actual run used an internal model and was not billed at that rate, so the figure is not the cost but the price of buying it, and it serves only for the comparison that matters: roughly a third of the five-year grant funding the equivalent human project, spent in eleven days.&amp;#160;&lt;a href="#fnref:3" class="footnote-backref" role="doc-backlink"&gt;&amp;#x21a9;&amp;#xfe0e;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;/div&gt;</description></item></channel></rss>