What checking did to published numbers, setting by setting
Five settings, four different quantities, no shared scale. Anyone selling you a single correction factor has built a statistic out of other people's conditions.
| Setting | What was done | What happened to the published numbers | What that quantity is |
|---|---|---|---|
| Psychology: 100 studies, three journals | Re-run with high-powered designs and original materials where available | "Replication effects were half the magnitude of original effects." 97% of originals were significant; of the replications, 36% were significant, 47% of original effects sat inside the replication's 95% confidence interval, 39% were rated as having replicated, and 68% stayed significant with original and replication evidence combined | Effect-size ratio plus four success measures: the authors report four precisely so that no single verdict exists |
| Laboratory economics: 18 studies, two top journals | Re-run at 90%-plus power under pre-defined plans | Replicated effect sizes averaged 66% of the original; 61% showed a significant effect in the original direction; four further replicability indicators ran 67–78% | Effect-size ratio and pass rates: a floor-condition decay, with almost nothing about the setting changed |
| The most-cited clinical research: 49 studies examined | Held against later, larger or better-controlled studies | Of the 45 the paper counts: 7 contradicted (16%), 7 initially stronger than what followed (16%), 20 replicated (44%), 11 never seriously re-tested (24%) | Verdict counts on famous findings, selected for fame, not a base rate. Most were never overturned |
| Preclinical cancer biology | Independent re-run of published experiments | Median replication effect 85% smaller than the original; 92% of replication effects came out smaller than their originals | The extreme of the range, in the setting furthest from a commercial reader |
| Nudges: a published sample against practice | A full census of two US units' 126 trials, held against a sample from two meta-analyses | 8.7 percentage points in the published sample; 1.4 across everything run; about 70% of the gap is attributed to selective publication with low power | A selected sample against an honest denominator, not a discount rate |
Swipe or scroll horizontally if the table is wider than your screen.
Reference & Evidence
Source: Open Science Collaboration (2015), Science 349(6251), aac4716, abstract; Camerer et al. (2016), Science 351(6280), 1433–1436, abstract; Ioannidis (2005), JAMA 294(2), 218–228; Errington et al. (2021), eLife 10:e71601; DellaVigna & Linos (2020), NBER Working Paper No. 27594, abstract. Each row reports a different quantity; no shared scale exists, and that is the finding.
Each line is a claim from the register this journal publishes against, resolved from the register at build time.
- B Ninety-seven percent of the originals were significant: "ninety sevenpercent of original studies had significant results p 05", against "thirty six percent of replications" Open Science Collaboration (2015), abstract ·
OSC15-C1 - B The headline, verbatim: "Replication effects were half the magnitude of original effects", over a set the authors describe as "We conducted replications of 100 experimental and correlational studies published in three psychology journals" Open Science Collaboration (2015), abstract ·
OSC15-C2 - B The original-side figure in the abstract's own words: "Ninety-seven percent of original studies had statistically significant results" Open Science Collaboration (2015), abstract ·
OSC15-C3 - B Four measures, not one, and they disagree: "thirty six percent of replications had significant results", "47 of original effect sizes were in the 95 confidence interval of the replication effect size", "39 of effects were subjectively rated to have replicated the original result", and combining both "left 68 with statistically significant effects" Open Science Collaboration (2015), abstract ·
OSC15-C4 - B The average replicated size, in the authors' words: "on average, the replicated effect size is 66% of the original", in a set where "we replicated 18 studies published in the American Economic Review and the Quarterly Journal of Economics between 2011 and 2014" Camerer et al. (2016), abstract ·
CAMERER16-C1 - B The direction test, verbatim: "We found a significant effect in the same direction as in the original study for 11 replications (61%)" Camerer et al. (2016), abstract ·
CAMERER16-C2 - B The design, and it is the reason the result carries: "they all have a statistical power of at least 90% to detect the original effect size at the 5% significance level" Camerer et al. (2016), abstract ·
CAMERER16-C3 - B The other indicators, in the abstract's own words: "The replicability rate varies between 67% and 78% for four additional replicability indicators" Camerer et al. (2016), abstract ·
CAMERER16-C4 - B The counts, verbatim from the abstract: "Of 49 highly cited original clinical research studies, 45 claimed that the inter"vention was effective, of which "7 (16%) were contradicted by subsequent studies, 7 others (16%) had found effects that were stronger than those of subsequent studies", while 20 "(44%) were replicated, and 11 (24%) remained largely unchallenged" Ioannidis (2005), JAMA abstract and results ·
IOANNIDIS05J-C1 - B The design contrast, verbatim: five of six highly "cited nonrandomized studies had been contradicted or had found stronger effects vs 9 of 39 randomized controlled trials" Ioannidis (2005), JAMA abstract and results ·
IOANNIDIS05J-C2 - A The size comparison, verbatim: "the median effect size in the replications was 85% smaller than the median effect size in the original experiments" Errington et al. (2021), eLife, results ·
ERRINGTON21-C1 - A The direction of nearly all of them: "92% of replication effect sizes were smaller than the original" Errington et al. (2021), eLife, results ·
ERRINGTON21-C2 - A The published-sample figure, verbatim: "an 8.7 percentage point take-up effect, which is a 33.4% increase over the average control", where "the average impact of a nudge is very large" DellaVigna & Linos (2022), Econometrica 90(1), 81–116, abstract ·
DELLAVIGNA22-C1 - A The decomposition, verbatim: "selective publication in the Academic Journals sample, exacerbated by low statistical power, explains about 70 percent of the difference in effect sizes between the two samples. Different nudge characteristics account for most of the residual difference." A model result, never a raw measurement DellaVigna & Linos (2022), Econometrica 90(1), 81–116, abstract ·
DELLAVIGNA22-C2 - A The at-scale average, verbatim: "still sizable and highly statistically significant, but smaller at 1.4 percentage points, an 8.0% increase" DellaVigna & Linos (2022), Econometrica 90(1), 81–116, abstract ·
DELLAVIGNA22-C3 - A The census, verbatim: "We assemble a unique data set of 126 RCTs covering 23 million individuals, including all trials run by two of the largest Nudge Units in the United States" DellaVigna & Linos (2022), Econometrica 90(1), 81–116, abstract ·
DELLAVIGNA22-C4
Grades: A, verified against the printed page of the primary source · B, primary source, text layer only · C, authoritative secondary · D, reported.
Related exhibits
-
Causal claim specification sheet
From the essay An uplift claim needs a specification before it needs a number
-
The synthetic reconstruction worksheet
From the essay What is reproducibility? Can another analyst reconstruct the result?
-
Three questions before a number enters the decision
From the essay Evidence over anecdote: what a number has to survive.