The null hypothesis significance test in health sciences research (1995-2006): statistical analysis and interpretation
- PMID: 20482841
- PMCID: PMC2886084
- DOI: 10.1186/1471-2288-10-44
The null hypothesis significance test in health sciences research (1995-2006): statistical analysis and interpretation
Abstract
Background: The null hypothesis significance test (NHST) is the most frequently used statistical method, although its inferential validity has been widely criticized since its introduction. In 1988, the International Committee of Medical Journal Editors (ICMJE) warned against sole reliance on NHST to substantiate study conclusions and suggested supplementary use of confidence intervals (CI). Our objective was to evaluate the extent and quality in the use of NHST and CI, both in English and Spanish language biomedical publications between 1995 and 2006, taking into account the International Committee of Medical Journal Editors recommendations, with particular focus on the accuracy of the interpretation of statistical significance and the validity of conclusions.
Methods: Original articles published in three English and three Spanish biomedical journals in three fields (General Medicine, Clinical Specialties and Epidemiology - Public Health) were considered for this study. Papers published in 1995-1996, 2000-2001, and 2005-2006 were selected through a systematic sampling method. After excluding the purely descriptive and theoretical articles, analytic studies were evaluated for their use of NHST with P-values and/or CI for interpretation of statistical "significance" and "relevance" in study conclusions.
Results: Among 1,043 original papers, 874 were selected for detailed review. The exclusive use of P-values was less frequent in English language publications as well as in Public Health journals; overall such use decreased from 41% in 1995-1996 to 21% in 2005-2006. While the use of CI increased over time, the "significance fallacy" (to equate statistical and substantive significance) appeared very often, mainly in journals devoted to clinical specialties (81%). In papers originally written in English and Spanish, 15% and 10%, respectively, mentioned statistical significance in their conclusions.
Conclusions: Overall, results of our review show some improvements in statistical management of statistical results, but further efforts by scholars and journal editors are clearly required to move the communication toward ICMJE advices, especially in the clinical setting, which seems to be imperative among publications in Spanish.
Figures
Similar articles
-
Statistical inference in abstracts of major medical and epidemiology journals 1975-2014: a systematic review.Eur J Epidemiol. 2017 Jan;32(1):21-29. doi: 10.1007/s10654-016-0211-1. Epub 2016 Nov 17. Eur J Epidemiol. 2017. PMID: 27858205 Review.
-
The reporting of p values, confidence intervals and statistical significance in Preventive Veterinary Medicine (1997-2017).PeerJ. 2021 Nov 24;9:e12453. doi: 10.7717/peerj.12453. eCollection 2021. PeerJ. 2021. PMID: 34900418 Free PMC article.
-
Meta-analysis: Problems with Russian Publications.Int J Risk Saf Med. 2015;27 Suppl 1:S89-90. doi: 10.3233/JRS-150702. Int J Risk Saf Med. 2015. PMID: 26639728
-
In support of null hypothesis significance testing.Proc Biol Sci. 2004 Feb 7;271 Suppl 3(Suppl 3):S82-4. doi: 10.1098/rsbl.2003.0105. Proc Biol Sci. 2004. PMID: 15101426 Free PMC article.
-
The Use and Misuse of Statistical Methods in Cytopathology Studies: Review of 6 Journals.Lab Med. 2019 Jan 1;50(1):8-15. doi: 10.1093/labmed/lmy036. Lab Med. 2019. PMID: 30020520 Review.
Cited by
-
Tracing scientific reasoning in psychiatry: Reporting of statistical inference in abstracts of top journals 1975-2015.Int J Methods Psychiatr Res. 2018 Aug 2;27(3):e1735. doi: 10.1002/mpr.1735. Online ahead of print. Int J Methods Psychiatr Res. 2018. PMID: 30073728 Free PMC article.
-
Publication bias in psychology: a diagnosis based on the correlation between effect size and sample size.PLoS One. 2014 Sep 5;9(9):e105825. doi: 10.1371/journal.pone.0105825. eCollection 2014. PLoS One. 2014. PMID: 25192357 Free PMC article.
-
Breaking free of sample size dogma to perform innovative translational research.Sci Transl Med. 2011 Jun 15;3(87):87ps24. doi: 10.1126/scitranslmed.3001628. Sci Transl Med. 2011. PMID: 21677197 Free PMC article.
-
What possibly affects nighttime heart rate? Conclusions from N-of-1 observational data.Digit Health. 2022 Aug 24;8:20552076221120725. doi: 10.1177/20552076221120725. eCollection 2022 Jan-Dec. Digit Health. 2022. PMID: 36046637 Free PMC article.
-
Recommendations for a Complete Reporting of Statistical Methods in Veterinary Pharmacology.J Vet Pharmacol Ther. 2025 Jul;48(4):221-233. doi: 10.1111/jvp.70001. Epub 2025 Jun 6. J Vet Pharmacol Ther. 2025. PMID: 40474839 Free PMC article. Review.
References
-
- Fisher RA. Statistical Methods for Research Workers. Edinburgh: Oliver & Boyd; 1925.
-
- Neyman J, Pearson E. On the use and interpretation of certain test criteria for purposes of statistical inference. Biometrika. 1928;20:175–240.
-
- Silva LC. Los laberintos de la investigación biomédica. En defensa de la racionalidad para la ciencia del siglo XXI. Madrid: Díaz de Santos; 2009.
-
- Berkson J. Test of significance considered as evidence. J Am Stat Assoc. 1942;37:325–335. doi: 10.2307/2279000. - DOI
MeSH terms
LinkOut - more resources
Full Text Sources