03 / CHECK THE CONSEQUENCE
A false positive, every time.
Now let each of the 70 assignments play the observed role. For each one, select its winner and apply each procedure.
70/70Frozen-selection procedure rejects
100% false rejection at a 5% threshold.
0/70Global procedure rejects
0% rejection in this constructed example.
This panel was deliberately constructed to expose a failure. The values stay fixed under every treatment assignment, so every rejection is false here.
The global result is no evidence against the sharp null in this example. It does not establish that a real treatment has no effect.
Why is there always a perfect-looking feature?
The panel contains one binary vector from each pair of complementary assignments. There are 35 such pairs among 70 assignments.
Each assignment therefore has a feature that exactly matches its treatment labels or their complement. Its absolute mean difference is 1.
For that feature alone, only its matching assignment and complement have difference 1. Repeating the search finds a perfect separator for every assignment.
A feature fixed in advance can have the same numerical p-value, 2/70. The failure comes from selecting the winner first.
Does the issue also appear in less structured synthetic counts?
The retained research run generated 100 Poisson and 100 gamma-Poisson matrices. Each had eight units and 100 features.
Each matrix was held fixed while all 70 assignments were enumerated. The table averages the exact conditional rejection fractions.
On narrow screens, scroll the table sideways.
Assignment probabilities are exact. The reported averages depend on the 200 generated matrices and the recorded seed, 104729.
Two-sided symmetry and 70 assignments limit the valid rejection fraction to at most 2/70 at this threshold. This discreteness explains the conservative averages.
All 200 matrix results · Generator and experiment code
04 / REVIEW THE WHOLE CLAIM
What an expert review catches.
Correct calculations can answer the wrong question. A review connects the scientific claim, selection history, randomization procedure and implementation.
Trace the search into the test
The mean differences, maximum and tail counts can each be correct. Freezing a label-dependent winner changes the composed procedure.
Rerun selection over the fixed panel for this global test. A different conditional method needs its own justification.
Keep the claim at the right level
The global sharp null says treatment changes no measured feature for any unit. Rejection does not identify a particular affected gene.
Selected-gene inference, partial nulls and false discovery rate control need separate arguments.
The premises travel with the result.
A fixed panel and procedureThe 35-feature panel is constructed without an observed assignment. The same selection rule runs across all assignments.
The declared assignment lawAll 70 four-of-eight assignments are equally likely. The real experiment must justify the randomization law being used.
The stated scientific claimThe test concerns a global sharp null. These eight units are not automatically interchangeable with eight cells from a study.
A recorded history is not a verified history. If someone first selects the winning column, a correct one-column analysis can still inherit that hidden selection.
The retained boundary check accepted all 70 false fixed-panel declarations. It refused all 70 truthful declarations of observed-assignment selection.
What does the formal result add—and what stays outside it?
The research includes a Lean-checked finite counting theorem for inclusive ranks, including ties. It establishes a precise property of finite score lists.
A mathematical argument connects that property to the error bound under a uniform assignment law and the global sharp null.
The proof does not cover the full Python/R workflow, panel-selection history, actual assignment mechanism, measurement quality or biological interpretation.
Strong differential tests and runtime checks each caught all eight seeded defects in the retained comparison. This establishes no superiority over strong testing.
This page replays exact precomputed results. It does not run Lean or certify a visitor’s dataset. The source receipts concern hypothetical assignment tables.
Inspect the exact method and source artifacts
For each feature, compute the absolute treated-minus-control mean difference. The retained implementation scales this by 4 × 4 = 16.
The global statistic is the maximum across the fixed panel. Count reference statistics greater than or equal to the observed statistic, including ties.
Divide that inclusive tail count by 70. No assignments are sampled, and no random numbers are generated in this page.
The displayed matrix and both R output tables are copied byte-for-byte. A separate Python calculation reproduces all 140 p-value rows.
Build-generated data also records all 2,450 feature-by-assignment scores. The browser selects from that table for the graphics.
Full scientific contract · Global R output · Frozen R output
Selection-history boundary · Testing comparison · Source identities and SHA-256 hashes