Power and sample size calculations for cluster randomized trials with binary outcomes when intracluster correlation coefficients vary by treatment arm
- PMID: 34879711
- PMCID: PMC8883478
- DOI: 10.1177/17407745211059845
Power and sample size calculations for cluster randomized trials with binary outcomes when intracluster correlation coefficients vary by treatment arm
Abstract
Background/aims: Generalized estimating equations are commonly used to fit logistic regression models to clustered binary data from cluster randomized trials. A commonly used correlation structure assumes that the intracluster correlation coefficient does not vary by treatment arm or other covariates, but the consequences of this assumption are understudied. We aim to evaluate the effect of allowing variation of the intracluster correlation coefficient by treatment or other covariates on the efficiency of analysis and show how to account for such variation in sample size calculations.
Methods: We develop formulae for the asymptotic variance of the estimated difference in outcome between treatment arms obtained when the true exchangeable correlation structure depends on the treatment arm and the working correlation structure used in the generalized estimating equations analysis is: (i) correctly specified, (ii) independent, or (iii) exchangeable with no dependence on treatment arm. These formulae require a known distribution of cluster sizes; we also develop simplifications for the case when cluster sizes do not vary and approximations that can be used when the first two moments of the cluster size distribution are known. We then extend the results to settings with adjustment for a second binary cluster-level covariate. We provide formulae to calculate the required sample size for cluster randomized trials using these variances.
Results: We show that the asymptotic variance of the estimated difference in outcome between treatment arms using these three working correlation structures is the same if all clusters have the same size, and this asymptotic variance is approximately the same when intracluster correlation coefficient values are small. We illustrate these results using data from a recent cluster randomized trial for infectious disease prevention in which the clusters are groups of households and modest in size (mean 9.6 individuals), with intracluster correlation coefficient values of 0.078 in the control arm and 0.057 in an intervention arm. In this application, we found a negligible difference between the variances calculated using structures (i) and (iii) and only a small increase (typically ) for the independent correlation structure (ii), and hence minimal effect on power or sample size requirements. The impact may be larger in other applications if there is greater variation in the ICC between treatment arms or with an additional covariate.
Conclusion: The common approach of fitting generalized estimating equations with an exchangeable working correlation structure with a common intracluster correlation coefficient across arms likely does not substantially reduce the power or efficiency of the analysis in the setting of a large number of small or modest-sized clusters, even if the intracluster correlation coefficient varies by treatment arm. Our formulae, however, allow formal evaluation of this and may identify situations in which variation in intracluster correlation coefficient by treatment arm or another binary covariate may have a more substantial impact on power and hence sample size requirements.
Keywords: Cluster randomized trial; generalized estimating equations; intracluster correlation coefficient; logistic regression; sample size.
Conflict of interest statement
Declaration of conflicting interests
The authors declare no conflicting interests.
Figures



Similar articles
-
swdpwr: A SAS macro and an R package for power calculations in stepped wedge cluster randomized trials.Comput Methods Programs Biomed. 2022 Jan;213:106522. doi: 10.1016/j.cmpb.2021.106522. Epub 2021 Nov 12. Comput Methods Programs Biomed. 2022. PMID: 34818620 Free PMC article.
-
A readily available improvement over method of moments for intra-cluster correlation estimation in the context of cluster randomized trials and fitting a GEE-type marginal model for binary outcomes.Clin Trials. 2019 Feb;16(1):41-51. doi: 10.1177/1740774518803635. Epub 2018 Oct 8. Clin Trials. 2019. PMID: 30295512
-
Appropriate statistical methods for analysing partially nested randomised controlled trials with continuous outcomes: a simulation study.BMC Med Res Methodol. 2018 Oct 11;18(1):105. doi: 10.1186/s12874-018-0559-x. BMC Med Res Methodol. 2018. PMID: 30314463 Free PMC article.
-
A review of current practice in the design and analysis of extremely small stepped-wedge cluster randomized trials.Clin Trials. 2025 Feb;22(1):45-56. doi: 10.1177/17407745241276137. Epub 2024 Oct 8. Clin Trials. 2025. PMID: 39377196 Free PMC article. Review.
-
Adherence to key recommendations for design and analysis of stepped-wedge cluster randomized trials: A review of trials published 2016-2022.Clin Trials. 2024 Apr;21(2):199-210. doi: 10.1177/17407745231208397. Epub 2023 Nov 21. Clin Trials. 2024. PMID: 37990575 Free PMC article. Review.
Cited by
-
Characterisation of between-cluster heterogeneity in malaria cluster randomised trials to inform future sample size calculations.Nat Commun. 2025 Jul 18;16(1):6615. doi: 10.1038/s41467-025-61502-w. Nat Commun. 2025. PMID: 40681482 Free PMC article.
-
The Effect of School-Linked Module-Based Friendly-Health Education on Adolescents' Sexual and Reproductive Health Knowledge, Guji Zone, Ethiopia - Cluster Randomized Controlled Trial.Adolesc Health Med Ther. 2024 Jan 23;15:5-18. doi: 10.2147/AHMT.S441957. eCollection 2024. Adolesc Health Med Ther. 2024. PMID: 38282688 Free PMC article.
-
Effect of a school-linked life skills intervention on adolescents' sexual and reproductive health skills in Guji zone, Ethiopia (CRT)-A generalized linear model.Front Public Health. 2023 Oct 23;11:1203376. doi: 10.3389/fpubh.2023.1203376. eCollection 2023. Front Public Health. 2023. PMID: 37937073 Free PMC article. Clinical Trial.
References
-
- Hayes RJ and Moulton LH. Cluster Randomised Trials. 2nd ed. Boca Raton, FL: CRC Press, 2017.
-
- Campbell MK, Fayers PM and Grimshaw JM. Determinants of the intracluster correlation coefficient in cluster randomized trials: the case of implementation research. Clin Trials 2005; 2: 99–107. - PubMed
-
- Choudhry NK, Avorn J, Glynn RJ, et al. Full coverage for preventive medications aftermyocaridal infarction. N Engl J Med 2011; 365: 2088–2097. - PubMed
-
- Shih WJ. Sample size and power calculations for periodontal and other studies with clustered samples using the method of generalized estimating equations. Biom J 1997; 39: 899–908.
Publication types
MeSH terms
Grants and funding
LinkOut - more resources
Full Text Sources