Statistical Significance vs. Practical Significance: Analyzing the Roots of the Replicability Crisis in Modern Scientific Research

Authors

  • Jiacheng Mao

DOI:

https://doi.org/10.61173/4ca6jq05

Keywords:

Reproducibility crisis, Statistical significance, Practical significance, P-value manipulation

Abstract

Modern scientific research is facing a severe crisis of reproducibility, with numerous published findings failing to replicate in subsequent independent verification, severely undermining the reliability of scientific knowledge. This paper systematically argues that the methodological imbalance between statistical significance and practical significance constitutes the core root cause of this crisis. We delve into how the misuse of p-values in current research practices—through p-value manipulation and selective reporting—spawns fragile findings. We also reveal the deep-seated impacts of academic incentive structures, publication bias, cognitive biases, and insufficient scientific literacy. Building on this analysis, we propose multidimensional solutions: at the research culture level, reforming incentive mechanisms to prioritize replicable studies and open science practices; and at the educational level, strengthening the cultivation of statistical thinking. The scientific community must achieve a paradigm shift from pursuing statistical significance to evaluating evidence strength and practical relevance, thereby building a more resilient and credible research ecosystem.

References

[1] Wasserstein, R. L. (2016). ASA statement on statistical significance and P-values. American Statistician, 70(2), 131- 133.

[2] Cohen, J. (1994). The earth is round (p<. 05). American psychologist, 49(12), 997.

[3] Cumming, G. (2014). The new statistics: Why and how. Psychological science, 25(1), 7-29.

[4] Goodman, S. (2008, July). A dirty dozen: twelve p-value misconceptions. In Seminars in hematology (Vol. 45, No. 3, pp. 135-140). WB Saunders.

[5] Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., & Altman, D. G. (2016). Statistical tests, P values, confidence intervals, and power: a guide to misinterpretations. European journal of epidemiology, 31(4), 337-350.

[6] Halsey, L. G., Curran-Everett, D., Vowler, S. L., & Drummond, G. B. (2015). The fickle P value generates irreproducible results. Nature methods, 12(3), 179-185.

[7] Ioannidis, J. P. (2005). Why most published research findings are false. PLoS medicine, 2(8), e124.

[8] Muff, S., Nilsen, E. B., O’Hara, R. B., & Nater, C. R. (2022). Rewriting results sections in the language of evidence. Trends in ecology & evolution, 37(3), 203-210.

[9] Májovský, M., Černý, M., Kasal, M., Komarc, M., & Netuka, D. (2023). Artificial intelligence can generate fraudulent but authentic-looking scientific medical articles: Pandora’s box has been opened. Journal of medical Internet research, 25(1), e46924.

[10] Nosek, B. A., Hardwicke, T. E., Moshontz, H., Allard, A., Corker, K. S., Dreber, A., ... & Vazire, S. (2022). Replicability, robustness, and reproducibility in psychological science. Annual review of psychology, 73(1), 719-748.

[11] Althubaiti, A. (2023). Sample size determination: A practical guide for health researchers. Journal of general and family medicine, 24(2), 72-78.

[12] Kim, J. H. (2022). Moving to a world beyond p-value< 0.05: a guide for business researchers. Review of managerial science, 16(8), 2467-2493.

Downloads

Published

2025-12-19