dyb

17 of 32

אִם יִרְצֶה הַשֵּׁם

Peer review is a social process. Verification is a mechanical one. They are not the same thing, and I can now put a number on the difference.

For the last few months I have been auditing published mathematical claims the hard way: not by reading the proofs and nodding, but by writing small programs — gates — that replay each claim's key step and report pass or fail. Thirty-two gates, covering published claims across number theory, PDEs, combinatorics, and dynamical systems. Every gate is code you can run. Every verdict is pinned in a file called EXPECTED_VERDICT, and a CI check fails the build if any gate drifts from its recorded verdict in either direction.

The lock, as of October 5: 17 BREAK, 15 PASS.

Read that again. More than half of the audited proof-routes failed to replay.

What a BREAK means — and what it doesn't

Here is the doctrine, stated plainly, because it matters more than the number: gates refute routes, not theorems. A BREAK means the published proof-route — the specific argument the paper offered — did not survive mechanical replay. It does not mean the theorem is false. Mathematics is full of true statements with broken proofs and broken statements with salvageable cores. The gate tells you which route died and where, and stops there. It does not editorialize.

This is what makes the number honest instead of sensational. I am not claiming seventeen theorems are false. I am claiming seventeen published arguments could not be replayed by a machine that was trying to help them succeed.

One example, since you asked

Of the 32 gates, only one is a claimed-proof refutation: Zhi-Wei Sun's claimed proof that Catalan's constant is irrational. The audit hash-pinned the source PDF, confirmed authorship from the title page, and found the route does not close. That is the campaign's single BREAK of a full claimed proof — and it took a machine to find it, not a referee's afternoon.

On the other side of the ledger: Jude Gomila's Λ-bound audit passed all four replay lanes. Yucai Su's 2D-Jacobian claim passed the computational gates and the logical-gap pass alike. A PASS from this machine means something, precisely because the machine is not in the habit of passing things.

Check it yourself

That is the whole point. The thesis of this post is not "trust me." It is this:

A published proof that cannot be mechanically replayed has not finished being published.

And it is verifiable in the most literal sense. The repository is public. The lock lives in scripts/gates/check.py and its expected-verdict file. Clone it, run the gates, watch 17 break and 15 pass. If you find a gate that is wrong — a route I broke that shouldn't have broken, or a pass I granted that wasn't earned — the CI will tell you, and so will I. That is the deal peer review never quite offered: the audit itself is auditable.

Seventeen of thirty-two. The number will move as the campaign continues. What will not move is the standard: replay it, or it isn't done.

← Previous
○ Ricky polyglot software developer
Next →