Understanding Statistics and Experimental Design: How to Not Lie with Statistics
Michael H. Herzog (EPFL), Gregory Francis (Purdue) and Aaron Clarke (Bilkent) — Springer, Learning Materials in Biosciences, 2019, 142 pages, open access under CC BY-NC 4.0. The spoke’s second statistics text and its first free one, which is what growth edge #5 asked for. It is also unlike anything else in this corpus: a textbook whose last four chapters are an argument that the field it teaches is being used wrongly.
What it covers
Twelve chapters. The first eight are a standard applied sequence — basic probability, experimental design and signal detection, the core concept of statistics, variations on the t-test, the multiple testing problem, ANOVA, model fits and power, correlation. The last four are the book’s real subject: meta-analysis (9), understanding replication (10), magnitude of excess success (11), and suggested improvements and challenges (12).
Set against larsen-marx-mathematical-statistics, the difference is one of altitude and purpose. Larsen & Marx derives the machinery — likelihood, sufficiency, the Cramér–Rao bound, the generalized likelihood ratio — and treats testing as a body of theory. Herzog et al. assume far less mathematics, skip the derivations, and spend the recovered space on what the machinery does to a research literature when many people use it at once. Neither replaces the other. The corpus now holds the theory and the failure mode of its practice, which is a better pair than two more theory texts would have been.
The Test for Excess Success
The book’s distinctive contribution, and the reason it earns a page rather than a line. The argument is arithmetic rather than statistical sophistication:
If an effect is real but power is moderate, random sampling alone guarantees that some experiments fail. So a set of studies should succeed at roughly the rate their power predicts. Sum the power values across the set; that is how many significant results you should see. If a paper reports many more than that, the results are too good to be true — not necessarily fraudulent, but inconsistent with the effect size the paper itself reports.
The worked example is a ten-experiment precognition set. Pooled effect size g* = 0.1855; summing power across the ten experiments gives 6.27 expected successes against 9 reported. The probability of getting 9 or more, if the effect is exactly what the author claims, is 0.058 — so the author’s own data imply about a 6% chance of the author’s own success rate.
The control case matters as much. Applying the same test to nineteen bystander-effect studies gives 10.77 expected successes against 10 observed, and a probability of 0.76. The test does not simply flag everything; it distinguishes a believable literature from an implausible one.
Why the book says replication cannot save you
The authors’ sharpest claim, and the one that generalizes beyond statistics: “replication cannot be the final arbiter for science when hypothesis testing is used, unless experimental power is very high.” If power is moderate, a failed replication is the expected outcome some of the time and proves nothing, while a suspiciously successful series is evidence of a problem rather than of truth.
They cite the Open Science Collaboration’s 2015 replication of 97 studies from three top journals — 36% replicated — and note the part usually skipped: there was hardly any relationship between the original p-value and the replication p-value, and many replications had larger samples than the originals, so they should have produced smaller p-values. The problem is not confined to one field: Amgen could not reproduce findings in 47 of 53 landmark cancer papers.
Chapter 11 then shows, by simulation, that excess success arises from publication bias and from optional stopping without anyone intending to cheat — which is why the book treats this as a structural property of the method rather than a story about bad actors.
The chapter that declines to sell a fix
Chapter 12 evaluates the standard remedies and endorses none of them cleanly. On publishing every experiment: can an experimenter reliably tell a methodological failure from a sampling failure, and if not, does publishing everything clutter the literature or clarify it? Preregistration, alternative analyses and replication each get the same treatment. The authors state plainly that they “do not have a specific proposal that will address all of these problems.”
That refusal is worth recording as the book’s stance rather than as a gap. A textbook that ends by declining to resolve the problem it spent four chapters establishing is making a claim about the difficulty, and it is a more honest ending than a checklist.
Standing and limits
T1 — peer-reviewed Springer textbook, open access, authors at EPFL and Purdue. Two things to hold against it. The licence is CC BY-NC, so it is free to read and share but not fully open in the way hefferon-linear-algebra or trench-real-analysis are — the non-commercial clause is a real restriction, and edge #5 asked for an open-licensed text without specifying how open. And the mathematical level is deliberately low: there are no derivations here, so it does not deepen the spoke’s theory of inference at all. It is a text about the use of statistics, filed here because it teaches the subject from first principles, with its examples drawn from one field.
Cross-spoke context. Chapters 10–12 are substantially about the psychology replication crisis and its canonical cases (Open Science Collaboration, Bem’s precognition studies, the bystander effect), so they bear directly on psychology-wiki’s standing replication caveat — that spoke was the runner-up in routing. The Test for Excess Success is also a general-purpose tool for reading any literature that reports many significant results, which touches the evidence problems tracked in agentic-tooling-wiki and quant-trading-wiki (where a reported edge’s plausibility is the standing open question). Linked as context, not duplicated.