Who Checks the Proof Checkers?

Leo de Moura
Chief Architect, Lean FRO
FROCON 2026, Santa Cruz | September 2026
Lean

Sep 4 …

Anthropic: Formalizing Fermat's Last Theorem

"We are sharing the first complete computer-checked proof of Fermat's Last Theorem. Claude worked largely autonomously over 11 days to write the proof in the Lean programming language."

anthropic.com/research/formalizing-fermats-last-theorem

Lean

… Sep 7 … Sep 8 …

Buckmaster and Alpoge announce three blow-up results

OpenAI: On the Navier–Stokes Millennium Prize Problem

mastodon.social/@tristanbuckmaster ∣ openai.com/index/navier-stokes-solution

Lean

What is Lean?

A proof assistant and programming language that is transforming how we approach mathematics, software verification, and AI.

Lean provides machine-checkable proofs.

Lean addresses the trust bottleneck.

Lean is implemented in Lean, and is very extensible and scalable.

It is based on dependent type theory.

Small trusted kernel. Proofs can be exported and independently checked.

325,000+ unique installations: VS Code (184K) + Open VSX (141K).

Lean in the editor: proof state alongside the code

Lean

Lean is a "Game"

"You have written my favorite computer game" — Kevin Buzzard, Prof. of Mathematics, Imperial College

def odd (n : Nat) : Prop := ∃ k, n = 2 * k + 1 theorem square_of_odd_is_odd : odd n → odd (n * n) := by intro ⟨k₁, e₁⟩ simp [e₁, odd] exists 2 * k₁ * k₁ + 2 * k₁ lia

The "game board": you see goals and hypotheses, then apply "moves" (tactics).

Each tactic transforms the game board.

Lean

Mathlib: The Lean Mathematical Library

Created in July 2017, in Lean 3 during Big Proof. An open-source, community-driven library. Today:

Mathlib dependency graph

"I'm investing time now so that somebody in the future can have that amazing experience." — Heather Macbeth, Prof. of Mathematics, Imperial College

Lean

Mathematics Is Adopting Lean

Six Fields Medalists engaged: Tao, Scholze, Viazovska, Gowers, Hairer, Freedman.

Tao on Lex Fridman

Lean

Software Verification

Lean is not only for mathematics.

  • Cedar (AWS): Verified Authorization

  • SymCrypt (Microsoft): Verified Cryptography

  • Kraken (Google): x64 Semantics

  • ArkLib (Ethereum Foundation): Formally Verified Arguments of Knowledge

  • Signal Shot (Beneficial AI): Verify the Signal protocol and app

  • SampCert (AWS): Differential Privacy

  • CSLib: Computer Science Library

Lean

The IMO Grand Challenge (2019)

IMO Grand Challenge page and committee

Lean

AlphaProof and the International Math Olympiad (2024)

IMO 2024 problem 1 in Lean

AlphaProof score on IMO 2024 problems

deepmind.google/discover/blog/ai-solves-imo-problems-at-silver-medal-level

Lean

Extensibility and Introspection

"At Google DeepMind, we used Lean to build AlphaProof, a new reinforcement-learning based system for formal math reasoning. Lean's extensibility and verification capabilities were key in enabling the development of AlphaProof." — Pushmeet Kohli, Vice President, Research, Google DeepMind

Lean 4 is implemented in Lean, allowing for unprecedented extension and introspection by users.

leanprover-community/repl is a simple Lean program for communicating with the Lean compiler via JSON.

AI labs have customized it in many ways, especially around parallelized proof tree search.

stanford-centaur/PyPantograph is the state of the art for open-source Lean REPLs.

Lean

IMO 2025: 3 Gold, 1 Silver

OpenAI (informal)

OpenAI announces gold-medal performance

Harmonic (Lean)

Aristotle achieves gold-medal performance

DeepMind (informal)

Gemini Deep Think achieves gold-medal standard

ByteDance (Lean)

Seed Prover achieves silver-medal score

Lean

IMO-Level Proof Search as a Service

Aristotle Lean 4 API

AlphaProof interest form

aristotle.harmonic.fun ∣ deepmind.google/alphaproof

Lean

Autoformalization (2025)

Math, Inc. announces Gauss and the strong Prime Number Theorem

Dependency graph of the strong PNT formalization

github.com/math-inc/strongpnt

Lean

"Vibe Proving"

Alexeev & Mixon resolved a $1000 Erdős prize problem.

"We used ChatGPT to vibe code a Lean proof." — Alexeev & Mixon

The proof is checked by Lean. Paper

Erdos prize paper

Lean

2026: Putnam Bench Saturated

PutnamBench leaderboard: all 672 problems solved

💚 fully open-sourced · 💙 partially open-sourced

trishullab.github.io/PutnamBench/leaderboard.html

Lean

zlib in Lean

AI converted zlib (a C compression library) to Lean.

theorem zlib_decompressSingle_compress (data : ByteArray) (level : UInt8)
    (maxOutputSize : Nat) (hsize : data.size ≤ maxOutputSize) :
    ZlibDecode.decompressSingle (ZlibEncode.compress data level) maxOutputSize = .ok data

DONE BY AI zlib C compression library translate C → Lean lean-zip Lean implementation test Tests pass zlib test suite LEAN'S GUARANTEE prove Proved correct round-trip theorem, kernel-checked

"The Lean library isn't just tested and validated, it's proved correct. This allows us to let AIs loose optimizing the code, requiring that they update the proof whenever the implementation materially changes. This gives us the confidence to allow them to work autonomously in a way that would be unthinkable in other languages." — Kim Morrison, Why Lean is faster than Rust

Lean

Tau Ceti: AI-Authored Lean Mathematics

  • AI-authored Lean mathematics, directed by a human-owned roadmap and gated by open, adversarial review.

  • Humans own the roadmap: mathematicians choose the targets.

  • AIs write and review the code.

  • On the roadmap: universal covers, the Jacobian challenge, reductive algebraic groups, PDEs.

  • 1,657,293 lines of Lean as of Sep 25, 2026, up from zero on Jun 3.

Tau Ceti: lines of Lean, Jun 3 to Sep 25, 2026

taucetiproject.github.io/TauCeti

Lean

Formal Verification for Everyday Code

Boris Cherny (Anthropic, creator of Claude Code), Sep 22, 2026:

"I used Opus 5.5 to formally verify the Claude Agent SDK using Lean. A couple short prompts = 16 PRs fixing various bugs and race conditions."

"I don't know either language well, but Claude is excellent at both. This approach is super useful for formally modeling your code and finding bugs that a human probably wouldn't have spotted."

x.com/bcherny/status/2102543349102338309

Lean

AI-Generated Proofs at Scale

Lines of Lean: FLT, Mathlib, Buckmaster–Alpöge, OpenAI

Lean

The Statement / Specification Is Your Responsibility

HUMAN RESPONSIBILITY Human intent Fermat's Last Theorem “No solution in positive integers to an + bn = cn for n > 2.” AI Formal statement theorem fermat_last_theorem : n > 2 → a > 0 → b > 0 → c > 0 → a^n + b^n ≠ c^n := by … Proof … The formal statement must reflect your intent. Lean cannot check whether AI autoformalized the statement correctly. LEAN'S GUARANTEE CHECK Checked proof Verified by the Lean kernel. The proof proves exactly that statement. Lean guarantees this.

Formal Conjectures (Google DeepMind): curated, human-verified formal statements of open problems.

Formal Conjectures website

Lean

Validation Architecture

Kernel trust: declarations, official kernel, export, independent re-checks

Lean

Validating Lean Proofs

"It's really important with these formal proof assistants that there are no backdoors or exploits you can use to somehow get your certified proof without actually proving it, because reinforcement learning is just so good at finding these backdoors." — Terence Tao

Lean has multiple independent kernels. You can build your own and submit it to arena.lean-lang.org.

Validating a Lean Proof ∣ Who Watches the Provers?

Kernel Arena leaderboard

Lean

Validating a Lean Proof: Level I

The "blue double check marks" signify in-process acceptance by the official C++ kernel.

Blue check marks next to a theorem in the editor

Protects against "innocent" mistakes:

  • incomplete proofs / sorry

  • tactics producing incorrect proof terms

… in the current theorem.

Lean

Validating a Lean Proof: Level II

#print axioms lists all axioms (transitively) used by a theorem.

Three built-in axioms relied on pervasively in Lean: choice, propositional extensionality, quotient lifting.

The Gold Standard of the pre-AI world.

Protects against

  • inherited incomplete/incorrect proof terms

  • custom axioms

Lean

Validating a Lean Proof: Level III

The kernel can be run as its own process: leanchecker.

Protects against

  • incorrect kernel handling by Lean (e.g. imported declarations are trusted)

  • metaprograms bypassing the in-process kernel or otherwise corrupting processing

Lean

Validating a Lean Proof: Level IV

comparator isolates statement processing, proof processing, and leanchecker into separate, sandboxed processes. Optionally runs additional checkers as well.

Protects against

  • actively malicious proofs (but not incorrect statements)

  • bugs in some but not all checkers

Comparator is a judge for Lean proofs. comparator.live.lean-lang.org

Lean

Validating a Lean Proof: Level IV

A Challenge.lean file contains only the statement with minimal dependencies for review.

The FLT challenge theorem with sorry

Solution.lean then is checked to be a refinement of the challenge with arbitrary further imports.

The solution is accepted if it passes all checkers and axiom checks (i.e. no sorry) and its statement plus relevant closure is equal to that in the challenge.

Lean

Comparator in Action

A challenge asks for a proof of False.

A candidate tries to smuggle one past the kernel with a metaprogramming trick that exploits a missing check in the official kernel.

Real GitHub issue. Comparator rejects it. The proof is exported and re-checked independently, defeating the exploit.

Nanoda (Lean kernel written in Rust) and Lean4Lean reject it too.

Comparator Live rejecting a proof of False

Lean

2026: Bugs Found and Fixed

An adversarial user or AI is not trying to prove a theorem. It is looking for a bug in the checker and using it to make the checker accept something that is not true.

AI is excellent at finding exploits.

Jul 25 Collatz "disproof": an exploit built to fool two checkers at once, the official kernel and nanoda Jul 28 Issue #14576 filed. Fix merged 10 hours later; Lean v4.32.2 released the same day Jul 30 – Aug 20 Bug hunt with OpenAI internal models: 6 more soundness fixes (4 kernel, 2 runtime) Aug 11 lean-inductive-models: every inductive type modelled and re-checked Aug 21 Lean v4.33.1: all fixes shipped Sep 10 con-leche released: kernel written in Lean, with a consistency proof Sep 13 con-ron: verified Rust port of con-leche Sep 15 Lean v4.35.0-rc1: lake check --paranoid; nanoda, lean4lean, con-leche, con-ron bundled fixes and releases new checkers Dates from GitHub and the postmortems of Aug 1 and Aug 24, 2026. PR numbers: see the table.

  • Seven implementation bugs (five in the kernel, two in the runtime), found with adversarial AI. All fixed; all shipped in Lean v4.33.1 on Aug 21.

  • No Mathlib or CSLib proof was affected.

Lean

con-leche: A Verified Kernel

con-leche: a proof checker written in Lean, with a consistency proof. Released Sep 10, 2026.

Written and proved by Claude, under close supervision by Joachim Breitner (Lean FRO).

Checks all of Mathlib, FLT, and Navier-Stokes. A few days after the first release, it is already ~50% faster than the official kernel.

The main theorem: every set of declarations it accepts has a model in set theory.

Does not protect against bugs in the Lean runtime.

Next steps include separating the specification from the implementation.

Lean

con-ron: con-leche in Rust

Joachim Breitner used AI to translate con-leche to Rust.

He uses Aeneas+Lean to prove that con-ron and con-leche are equivalent.

It is very unlikely that the Lean and Rust runtimes have the same bug.

Lean

Lean v4.35

Release candidate Sep 15, 2026.

  • lake check: build, export, replay the proof through a specific kernel.

  • lake check --paranoid: run every bundled checker; accept only if all agree.

  • Bundled: nanoda (Rust), lean4lean (Lean), con-leche (Lean, verified), con-ron (Rust, verified).

Lean

Scalability and Proof Digestion

Trust is not the only challenge created by AI.

Scalability: we managed to keep up with humans. Keeping up with AI is harder.

Two aligned charts: Mathlib lines of code rising 66% since Feb 2024, while instructions to build Mathlib peaked in Dec 2025 and fell 28%; green markers show Lean releases

Lean

Conclusion

Lean is extensible, scalable, and trusted.

Mathematics, software verification, and AI labs rely on it.

The 2026 soundness bugs were found by adversarial AI, fixed in less than 24 hours, and turned into regression tests.

Proofs in Mathlib, CSLib, and other major Lean projects were not affected.

We released two new verified kernels: con-leche and con-ron.

More verified kernels coming soon: Lean4Lean, MetaLean.

Many new unverified kernels in arena.lean-lang.org.

Lean v4.35 includes lake check --paranoid with four independent checkers bundled.

Verified runtimes and compilers are the next frontier.

Scalability and proof digestion are new challenges too.

Thank You

FROCON 2026, Santa Cruz | September 2026