The replication crisis
Psychology’s discovery that a large share of its published findings do not reproduce. It is the
reason this spoke’s CLAUDE.md carries a standing evidence caveat, and it is not a footnote to the
corpus — it is doing active work inside two of the founding arguments.
Where it shows up here
In a celebrated book’s own bibliography. Studies cited in Chapter 4 of thinking-fast-and-slow carry a replicability index of 14, and broader analyses find much of the book resting on shaky literature, worst in the social-priming material. daniel-kahneman conceded it: “I placed too much faith in underpowered studies.”
As a premise in a live argument. free-will-belief-effects rests part of its case on the claim that the older findings — free-will scepticism increasing cheating and aggression — “have often failed to be replicated.” The crisis is not being confessed there; it is being used, to clear away the evidence on the other side.
That second use is the one to watch. “It didn’t replicate” is a legitimate reason to discount a finding and also a convenient one, and a reader has no way to tell which is happening without checking. The asymmetry is worth naming: the failures are cited against the opposing literature and rarely against one’s own. Walton’s two experiments (260 and 262 participants) are themselves the shape of study the crisis taught the field to distrust.
The complaint that isn’t non-replication (added 2026-08-05)
magnus-peresetsky-2022 brings a challenge this page had no slot for. Its argument against the dunning-kruger-effect is not that the result stops appearing. A doubly-censored tobit model — noise plus a random floor and ceiling, no psychological parameter anywhere in it — fits 665 students’ predictions of their own exam grades almost perfectly. Their conclusion: “there is an effect, but it does not reflect human nature.” Krueger & Mueller (2002) made the milder version of the point with regression to the mean.
Keep the two complaints apart, because the remedies differ. A replication failure is fixed by running the study again, bigger. An artifact objection survives every replication you can run: the finding is reproducible and the mechanism may still be absent. Larger samples make it worse, not better — they make the artifact cleaner.
And an artifact objection can itself be under-powered against the evidence. Reading both primaries turns this into a two-sided methodological fight rather than a debunking. kruger-dunning-1999 anticipated the regression objection and built Study 4 against it — train the bottom quartile for ten minutes and their self-estimates drop while an untrained control group’s do not, which censoring cannot produce. Magnus and Peresetsky model the cross-sectional shape and never mention that experiment. The lesson for this page is symmetric: a statistical critique needs to reach the experimental evidence, and this one addresses the chart.
The corpus also now holds its first named successful replication in any subject: a 2021 Nature study of roughly 4,000 participants each in two experiments, supporting the metacognitive reading for grammar and logic — reported by dunning-kruger-misread, which names neither its authors nor its title. So the wiki knows a replication succeeded and still cannot look at it.
A failure that shrank a finding instead of deleting it (added 2026-08-12)
handwriting-vs-typing-notes brings the corpus its first direct replication failure with the numbers attached. Mueller & Oppenheimer (2014) — handwritten notes beat typed ones — is one of the most repeated findings in popular writing about learning. A 2021 direct replication with 74 laptop and 68 longhand note-takers found that “on the immediate quiz, neither group performed better.” A 2026 EEG study of 33 undergraduates found no significant difference either. And a 2024 meta-analysis of 24 studies put the surviving advantage at Hedges’ g = 0.248.
This is a third pattern, distinct from the two above, and the most common one in practice: the
effect neither replicated cleanly nor turned out to be an artifact — it got smaller. A famous
finding became a small true one, which is the outcome least well served by how findings travel. There
is no headline for g = 0.248, so the 2014 version keeps circulating.
Note what this does to the asymmetry named earlier on this page. Here the failures are cited against the popular claim the article’s own readers hold, not against an opponent’s literature — the cleaner use of the two.
The same evidential hole, again. The meta-analysis carrying the headline number is reported with no authors and no institution, by the same publisher as dunning-kruger-misread. Two of the corpus’s most load-bearing replication facts now rest on studies it cannot look up. Recorded on silicon-canals as a property of the outlet.
What it means for reading this wiki
- A famous finding is not a reliable one. Fame and replication are close to uncorrelated in this field, and priming is the standard example.
- A claim’s presence in a bestseller is not evidence for it (thinking-fast-and-slow).
- Record the study, not the summary: sample, method, year, venue — or say plainly that the source gave none, as kahneman-focusing-illusion didn’t.
- The psychophysics end of the corpus is less exposed than the social-cognition end. Effects measured in a species comparison with a physical stimulus are a different evidential animal from priming studies on undergraduates.
The credit where it’s due
Kahneman’s public retraction is the standard the field needed and mostly didn’t get, and it came from the person with the most reputational exposure. Recorded on daniel-kahneman. The uncomfortable version — a book about systematic errors of judgment that made one — is also the clearest demonstration its thesis ever got.
Related
thinking-fast-and-slow · daniel-kahneman · heuristics-and-biases · dunning-kruger-effect · kruger-dunning-1999 · magnus-peresetsky-2022 · dunning-kruger-misread · free-will-belief-effects · perceptual-learning · handwriting-vs-typing-notes · note-taking-and-learning · silicon-canals · synthesis