Skip to main page content
U.S. flag

An official website of the United States government

Dot gov

The .gov means it’s official.
Federal government websites often end in .gov or .mil. Before sharing sensitive information, make sure you’re on a federal government site.

Https

The site is secure.
The https:// ensures that you are connecting to the official website and that any information you provide is encrypted and transmitted securely.

Access keys NCBI Homepage MyNCBI Homepage Main Content Main Navigation
. 2015 Jun;68(6):627-36.
doi: 10.1016/j.jclinepi.2014.12.014. Epub 2015 Jan 22.

The number of subjects per variable required in linear regression analyses

Affiliations
Free article

The number of subjects per variable required in linear regression analyses

Peter C Austin et al. J Clin Epidemiol. 2015 Jun.
Free article

Abstract

Objectives: To determine the number of independent variables that can be included in a linear regression model.

Study design and setting: We used a series of Monte Carlo simulations to examine the impact of the number of subjects per variable (SPV) on the accuracy of estimated regression coefficients and standard errors, on the empirical coverage of estimated confidence intervals, and on the accuracy of the estimated R(2) of the fitted model.

Results: A minimum of approximately two SPV tended to result in estimation of regression coefficients with relative bias of less than 10%. Furthermore, with this minimum number of SPV, the standard errors of the regression coefficients were accurately estimated and estimated confidence intervals had approximately the advertised coverage rates. A much higher number of SPV were necessary to minimize bias in estimating the model R(2), although adjusted R(2) estimates behaved well. The bias in estimating the model R(2) statistic was inversely proportional to the magnitude of the proportion of variation explained by the population regression model.

Conclusion: Linear regression models require only two SPV for adequate estimation of regression coefficients, standard errors, and confidence intervals.

Keywords: Bias; Explained variation; Linear regression; Monte Carlo simulations; Regression; Statistical methods.

PubMed Disclaimer

Similar articles

Cited by

Publication types

LinkOut - more resources