
working paper: cheap-science by Nic Fishman (Harvard Statistics), Gabriel Sekeres (Cornell Economics)
אִם יִרְצֶה הַשֵּׁם
Every evidence norm is secretly a cost norm. When automation collapses the cost of generating evidence, you must change the unit of evidence, not just the standards or the thresholds. Selection pressure can't be eliminated, only displaced. So commit the whole search space, not a single hypothesis, and audit the ledger, not the survivor.
1. TL;DR - By Audience
For the expert (econometrician/theorist): The paper builds a commitment-model of editorial screening in which researchers sequentially search over a correlated evidence stream at per-test cost γ, stop optimally, and selectively disclose. It proves a mechanism-independent Bernoulli-KL information bound on the screening frontier (Thm 4.3), shows short/sublinear disclosure collapses discrimination at the effective-sample-size scale (Thm 4.5), and constructs a robustness-check mechanism (require m significant passes out of a testing-capacity-scaled window) that attains the information-theoretic envelope up to constants (Thm 5.2, Cor 5.3). Empirically, they build an audited "specification surface" pipeline, validate it against the I4R (Institute for Replication) benchmark on 39 AEA papers, then fit a three-type folded-Gaussian mixture and AR(1) dependence model on 103 papers to calibrate γ, obtaining λ ≈ 1/172 and Δ̂ ≈ 0.849.
For the practitioner (empirical economist / journal editor): If specification search is now ~170x cheaper because of LLM coding agents, the traditional "show us a handful of robustness checks" norm no longer screens out false positives. To hold the false-discovery rate at a conventional level, you'd need roughly 7,000 qualifying robustness checks per paper instead of ~50. That is obviously unworkable. The actionable takeaway isn't "demand more tables," it's "the unit of evidence must change": editors and referees should move toward auditing a pre-registered, machine-executed universe of specifications rather than a curated handful.
For the general public: Scientists testing an idea used to face a real cost. Each new analysis took real time and effort, which naturally limited how many chances they had to get a lucky (but false) result. AI coding tools have made running these alternate analyses almost free. This paper shows mathematically that "cheap science" breaks the old honor-system of listing a few robustness checks, and that fixing this isn't about running even more manual checks. It's about auditing the entire menu of things a researcher could have tried, before they try it.
For the skeptic: Reasonable pushback: the theory rests on stylized assumptions (bounded per-test information, "type sandwich" conditions, a specific AR(1)/folded-Gaussian evidence model) that may not generalize past their calibration sample. The "7,000 checks" headline number is explicitly presented as illustrative/non-literal by the authors themselves. It's a reductio, not a policy proposal, but it will likely be quoted out of that context. The three-type mixture model, while diagnostically well-fit (PP/QQ plots, AIC/BIC), is a modeling convenience, not a claim about what any individual paper "really is."
For the decision-maker (journal editor, funder, policy-maker): This is a rigorous argument, with both theory and empirical calibration, that current robustness-check norms in empirical economics (and by extension, any cheap-testing observational science) are becoming an inadequate screening mechanism as AI agents lower the cost of specification search. The recommended institutional response is not "demand thousands of checks" but "shift editorial infrastructure toward auditing pre-committed, machine-executed specification universes." In essence, build the equivalent of a pre-registration plus automated-audit pipeline into peer review.
2. The Core Problem: Checking Was Never About Truth
Empirical economics (and observational social science broadly) relies on researchers reporting a handful of robustness checks to convince editors/referees/readers that a result isn't an artifact of arbitrary analytical choices (which controls to include, which sample restriction, which functional form, etc.).
Signal ≠ table. Signal = (table × cost-to-fake-table). Collapse the second factor and the first becomes free-floating decoration.
Peer review was quietly running an implicit proof-of-work scheme without ever writing the protocol down. The hidden collateral backing any table was the labor it cost to fabricate it, and labor was the cost that made "this looked robust" worth believing.
3. The Disruption: The Calculator Got 172x Faster
The disruption: LLM-based coding agents (the authors used Claude Opus 4.5/4.6 via Claude Code) can now execute an entire "robustness universe" (dozens to hundreds of estimator/sample/control variants) in minutes instead of days. The paper's own benchmark: an automated baseline reproduction took 14 minutes versus a conservative human benchmark of 40 hours (fastest completed reanalysis in a 41-paper replication study), implying a ~172x cost reduction.
That 1/172 ratio is the paper's load-bearing scalar, and the authors are honest that it's a single-trial point estimate deliberately biased toward the fastest human denominator, so the true ratio is likely larger. All the downstream numbers inherit that fragility. It's a cost ratio, not a capability claim, but every number downstream jumps off it.
This breaks the informal screening value of "the author showed several robustness checks," because a researcher (or, more subtly, an unscrupulous automated pipeline) can now cheaply search until a favorable-looking result appears and disclose only that, without doing anything that looks like fraud.
4. Surprising / Counterintuitive Findings
- Tightening standards doesn't help, and can actively backfire. The intuitive fix ("just require a lower p-value threshold") is proven to be strictly dominated by simply requiring more disclosed passes. Standards alone induce a "search race" without producing verifiable discrimination.
- The two editorial levers are fundamentally asymmetric. "Tighten the bar" and "force more disclosure" look similar but have completely different asymptotic behavior as testing gets cheap. Only forced disclosure that scales with testing capacity (Θ(1/γ) checks) recovers exponential screening. Raising the bar → researcher searches harder → same FDR, more wasted compute. Widening the report → exponential screening recovered.
- Effort that isn't emitted is not evidence. Even if a researcher does capacity-scale search internally, if they only disclose a bounded/sublinear number of results, the editor's ability to discriminate high-quality from null results provably collapses (Theorem 4.5). The bottleneck is how much got transmitted, not honesty. A diligent researcher who reports a small slice is indistinguishable from a p-hacker, structurally, not morally.
- There's an information speed-limit on editorial cleverness (Thm 4.3). No accept/reject rule, however clever, can separate good from null papers faster than the KL information the disclosed evidence physically carries. Editorial cleverness is a constant factor; disclosure volume is the exponent.
- The dependency structure matters as much as the raw count. Because specifications within a paper are correlated (AR(1) persistence φ̂ ≈ 0.151), the "effective" number of independent tests is only ~85% of the raw count (Δ̂ ≈ 0.849). Dependence doesn't save you from the cheap-testing problem, but it isn't catastrophic either.
- The implied fix is almost absurd on its face (deliberately). Restoring current false-discovery targets requires ~7,000 robustness checks. That's a number the authors themselves call undeliverable, precisely to motivate their real recommendation: audit the universe, don't count checks.
5. Key Terms, Translated
| Term | Plain-language translation |
|---|---|
| Specification surface | A pre-registered, machine-readable "menu" of every reasonable way to re-run an analysis (different controls, samples, models). Defined and locked in before anyone looks at results, so nobody can cherry-pick after the fact. |
| Specification search / p-hacking | Trying many versions of an analysis until one gives a "significant" result, then reporting only that one. |
| Robustness check | Re-running the same analysis a slightly different way to see if the finding survives. Like re-measuring something with a different ruler. |
| False discovery rate (FDR) | Out of all accepted findings, what fraction are actually false (null) results dressed up to look real. |
| Screening frontier | The best possible trade-off a journal can achieve between publishing lots of papers (throughput) and not publishing false ones (purity). A speed-vs-accuracy trade-off for an editor. |
| Effective sample size (n_eff) | How many truly independent pieces of evidence you have once you account for correlated (non-independent) analyses. |
| Bernoulli-KL divergence / information budget | A mathematical "speed limit" on how much an accept/reject decision can tell apart a good paper from a bad one, given how much real evidence went in. |
| Optimal stopping | The researcher's decision of when to stop running more analyses and just submit. Like knowing when to stop shuffling and deal. |
| Selective disclosure | Only reporting the analyses that came out favorably (a résumé that only lists your best grades). |
| AR(1) / persistence coefficient (φ) | How much one analysis's result "carries over" to predict the next one. High φ means the specs are all telling roughly the same story, not independent ones. |
| Folded-Gaussian mixture | A statistical way of sorting evidence strength (t-statistics) into three buckets: "probably nothing," "modest real effect," "unmistakably large effect." |
| Witness window | The zone (e.g., p < 0.05) that a policy treats as "counts as a pass" when tallying robustness checks. |
6. The Actual Proposal: Pre-Commit the Universe
The real recommendation is not an unchecked escalation. It's a concrete architectural shift:
Pre-run specification surface → independent verifier agent edits → whitelisted execution, contract-checked, hash-audited → post-run drift verifier (no re-run).
That's a locked, machine-readable, hash-audited ledger of every reasonable way to run the analysis, committed before touching outcomes, with provenance. It shifts the trust object from the author's curated table to the pipeline's recorded ledger.
How this differs from pre-registration: pre-registration commits to a single hypothesis; the specification surface commits to the entire search space. Pre-registration relocates selection pressure (search happens before the plan; rejection pressure migrates from author to field via selective publication of pre-registered hits). The surface contains it, because the whole space is on the ledger, not just the survivor.
7. Methodology: What They Actually Did
(A) Theory: a commitment / mechanism-design model - A journal commits ex ante to an accept/reject rule mapping disclosed evidence to acceptance probability. - A researcher (type = null, moderate, or extreme, unknown to her at first) sequentially observes correlated p-values at cost γ per test, chooses when to stop (optimal-stopping via Snell envelope), and discloses whichever sub-multiset of results maximizes acceptance (omission is unverifiable). - They derive: a mechanism-independent upper bound on discrimination (Thm 4.3), a short-disclosure collapse result (Thm 4.5), and an achievability result (Thm 5.2) for a simple "require m passes in window B" mechanism that hits the envelope up to constants. - ≈60 pages of appendices formalize regularity conditions and verify them in a running Gaussian AR(1) three-type example.
(B) Empirics: an agentic replication pipeline - Built an LLM-agent workflow (Claude Opus 4.5/4.6 via Claude Code CLI) that: (1) builds a pre-run "specification surface," (2) has a separate LLM agent verify/edit it before execution, (3) executes only whitelisted specs with contract-checked, hash-audited outputs, (4) has a post-run verification agent flag drift without re-running anything. - Validation (Sample A, n=39 AEA papers): |t|-statistics track the Institute for Replication's (I4R) independent human reanalysis closely (Figure 2); matched reproductions agree claim-by-claim (most within |Δt| < 0.5). - Calibration (Sample B, n=103 papers, 5,793 specs): fit a 3-component folded-Gaussian mixture (null ≈62%/μ≈1.6, moderate ≈31%/μ≈3.9, extreme ≈8%/μ≈7.9), within-paper AR(1) dependence (φ̂≈0.151 → Δ̂≈0.849), and the cost ratio λ≈1/172. - Counterfactual: plugged estimates into the mechanism to compute required passes (m) to hold FDR=0.05 post-shift, given baseline m_old=50.
8. Key Quantitative Results
| Quantity | Estimate | Context |
|---|---|---|
| Cost ratio (λ = γ_new/γ_old) | ≈ 1/172 (baseline); 1/103 conservative | 14 min automated vs. 40 hr (fastest I4R reanalysis; avg. I4R time was 13 days) |
| Effective-independence rate (Δ̂ = 1−φ̂) | 0.849 | Preferred AR(1) ordering, R²=0.036; robust 0.849-0.859 across five non-random orderings |
| Baseline disclosure requirement (m_old) | 50 | Median author-reported regressions (mean 81, max 693) |
| Required disclosure to hold FDR=0.05 (m_new) | ≈ 6,994 | ≈140x increase over m_old=50 at the calibration point |
| Mixture: null / moderate / extreme | 62% / 31% / 8% | Folded-Gaussian, σ=1 fixed; μ̂≈1.6 / 3.9 / 7.9 |
| Sample A validation (I4R match) | Most claims within | Δt |
| Full spec dataset | 5,793 specs / 97 papers | Mean ≈60 specs/paper |
| Within-paper variance share | 35.2% | ICC=0.648 (between-paper); spec search within a paper really has room to move results |
Authors also report: bootstrap CIs on mixture parameters (Fig. 16), AR(1) across six orderings (Table 7), AIC/BIC model selection (K=3 preferred over K=2, diminishing at K=4), folded-vs-truncated normal robustness, AER vs. non-AER subsample replication, and a full sensitivity grid over λ, windows, and m_old (Table 9).
A warning about the "7,000" number: it's a reductio, not a quota, engineered to be absurd. Its function is to prove the unit of evidence itself is wrong, not to set a target. Predictably, the headline will detach from the caveat on the way into media/policy gravity. Read it as "count checks is the wrong unit," not "demand 7,000 tables."
9. Practical Deployment Considerations
- Who would implement it? Journal editorial offices, the Institute for Replication (I4R), funders requiring pre-registration, or third-party auditing services. The pipeline (prompts, validators, contracts) is described as open/reproducible.
- Integration pathway: the natural point is the pre-registration/pre-analysis-plan stage. A specification surface is essentially a much-more-granular, machine-checkable pre-analysis plan, verified by a second LLM agent before data are touched.
- Cost/latency: ~14 minutes per paper. The bottleneck shifts from compute to institutional adoption and trust in agentic audits. Editors must trust an LLM-mediated verification layer, a real governance/liability question the paper doesn't fully settle.
- Failure modes to watch: two of 41 papers couldn't be handled by the pipeline (structural macro; restricted-access data). Coverage isn't universal, especially structural/theory-heavy or non-public-data work.
- Human-in-the-loop: the "surface verifier" step is a natural candidate for human referee oversight rather than full automation at production stakes.
10. Where the Argument Is Soft
The authors shoot at their own feet. It's worth reading the caveats as more than hedging:
- Theory rests on stylized regularity conditions (bounded per-test KL, "type sandwich," Bayes-factor tails) only verified in the Gaussian AR(1) running example.
- The three-type mixture is a measurement instrument, not a taxonomy. Their point: extreme type = automatically "true" for FDR accounting, a hidden assumption that large |t| may itself be artifact. This is exactly where LLM search could manufacture spurious large effects.
- Scope is thin. 103 papers, all AEA, observational economics with public replication packages. Only one RCT. The strong claims carry a lot for a single-domain observation.
- λ is a fragile point estimate. One automated-vs-fastest-human timing comparison on one vendor's model. Numbers will shift with other agents.
- Pre-registration is explicitly an incomplete fix. Selection pressure is conserved and displaced, so the surface's scope of commitment is what matters, not the commitment itself.
11. What's Next
- Multiverse/specification-curve interpretation that scales: reading sets of specifications jointly, extending Simonsohn et al. (2020) and Steegen et al. (2016).
- Third-party "clearinghouse" auditing infrastructure for specification surfaces, rather than per-journal stacks.
- Beyond economics to any cheap-testing observational science (psychology, epidemiology) facing the same LLM-driven cost collapse.
- Capacity-rationing mechanisms (Appendix A.3.4) for journals with hard throughput caps rather than just FDR targets.
- Better dependence primitives. AR(1) orderings range 0.849-0.859, so graph-based models could sharpen the effective-sample-size estimate.
12. Conflicts of Interest & Potential Biases
- Direct use of Anthropic Claude models as the empirical engine. The λ≈1/172 figure is tied to this model/tooling combination; another agent would vary (the authors appropriately caveat this).
- No disclosed funding/industry sponsorship visible in the provided text. Check the published version for an acknowledgments section.
- Institutional stake in a "problem is real" framing. As methods researchers proposing a new auditing methodology, they have a natural incentive to frame cheap-testing as significant and surface-auditing as necessary. Expected for a methods paper, not disqualifying.
- No ideological/political valence. Institutionally and methodologically focused on peer-review mechanics rather than any substantive economic question.