The Rules of the Game Called Psychological Science

Marjan Bakker, Annette van Dijk, Jelte M. Wicherts

Perspectives on Psychological Science · 2012 · 847 citations · 48 references

DOIFull text

Open access

Concepts

TL;DR

Psychological science routinely reports statistically significant findings, yet reviews show that about 96 % of null‑hypothesis significance tests are significant despite many studies being underpowered, leading to inflated effects, high false‑positive rates, and questionable practices such as adding participants after interim testing. The authors aim to explain why multiple small, underpowered samples can be a more efficient strategy for finding significant results than a single larger sample. They demonstrate this by comparing the probability of obtaining p < .05 from multiple small samples versus a single large sample, showing the former is more efficient. Simulations and an analysis of 13 meta‑analyses covering 281 primary studies reveal severe biases and excess significant results in seven, underscoring the need for adequately powered replications and changes to journal policies.

Abstract

If science were a game, a dominant rule would probably be to collect results that are statistically significant. Several reviews of the psychological literature have shown that around 96% of papers involving the use of null hypothesis significance testing report significant outcomes for their main results but that the typical studies are insufficiently powerful for such a track record. We explain this paradox by showing that the use of several small underpowered samples often represents a more efficient research strategy (in terms of finding p < .05) than does the use of one larger (more powerful) sample. Publication bias and the most efficient strategy lead to inflated effects and high rates of false positives, especially when researchers also resorted to questionable research practices, such as adding participants after intermediate testing. We provide simulations that highlight the severity of such biases in meta-analyses. We consider 13 meta-analyses covering 281 primary studies in various fields of psychology and find indications of biases and/or an excess of significant results in seven. These results highlight the need for sufficiently powerful replications and changes in journal policies.

References

48