. 2022 Feb;19(1):42-51.

doi: 10.1177/17407745211059845. Epub 2021 Dec 8.

Power and sample size calculations for cluster randomized trials with binary outcomes when intracluster correlation coefficients vary by treatment arm

Lee Kennedy-Shaffer^{1

2}, Michael D Hughes¹

Affiliations

¹ Department of Biostatistics, Harvard T. H. Chan School of Public Health, Boston, MA, USA.
² Department of Mathematics and Statistics, Vassar College, Poughkeepsie, NY, USA.

PMID: 34879711
PMCID: PMC8883478
DOI: 10.1177/17407745211059845

Power and sample size calculations for cluster randomized trials with binary outcomes when intracluster correlation coefficients vary by treatment arm

Lee Kennedy-Shaffer et al. Clin Trials. 2022 Feb.

. 2022 Feb;19(1):42-51.

doi: 10.1177/17407745211059845. Epub 2021 Dec 8.

Authors

Lee Kennedy-Shaffer^{1

2}, Michael D Hughes¹

Affiliations

¹ Department of Biostatistics, Harvard T. H. Chan School of Public Health, Boston, MA, USA.
² Department of Mathematics and Statistics, Vassar College, Poughkeepsie, NY, USA.

PMID: 34879711
PMCID: PMC8883478
DOI: 10.1177/17407745211059845

Abstract

Background/aims: Generalized estimating equations are commonly used to fit logistic regression models to clustered binary data from cluster randomized trials. A commonly used correlation structure assumes that the intracluster correlation coefficient does not vary by treatment arm or other covariates, but the consequences of this assumption are understudied. We aim to evaluate the effect of allowing variation of the intracluster correlation coefficient by treatment or other covariates on the efficiency of analysis and show how to account for such variation in sample size calculations.

Methods: We develop formulae for the asymptotic variance of the estimated difference in outcome between treatment arms obtained when the true exchangeable correlation structure depends on the treatment arm and the working correlation structure used in the generalized estimating equations analysis is: (i) correctly specified, (ii) independent, or (iii) exchangeable with no dependence on treatment arm. These formulae require a known distribution of cluster sizes; we also develop simplifications for the case when cluster sizes do not vary and approximations that can be used when the first two moments of the cluster size distribution are known. We then extend the results to settings with adjustment for a second binary cluster-level covariate. We provide formulae to calculate the required sample size for cluster randomized trials using these variances.

Results: We show that the asymptotic variance of the estimated difference in outcome between treatment arms using these three working correlation structures is the same if all clusters have the same size, and this asymptotic variance is approximately the same when intracluster correlation coefficient values are small. We illustrate these results using data from a recent cluster randomized trial for infectious disease prevention in which the clusters are groups of households and modest in size (mean 9.6 individuals), with intracluster correlation coefficient values of 0.078 in the control arm and 0.057 in an intervention arm. In this application, we found a negligible difference between the variances calculated using structures (i) and (iii) and only a small increase (typically $< 5 %$ ) for the independent correlation structure (ii), and hence minimal effect on power or sample size requirements. The impact may be larger in other applications if there is greater variation in the ICC between treatment arms or with an additional covariate.

Conclusion: The common approach of fitting generalized estimating equations with an exchangeable working correlation structure with a common intracluster correlation coefficient across arms likely does not substantially reduce the power or efficiency of the analysis in the setting of a large number of small or modest-sized clusters, even if the intracluster correlation coefficient varies by treatment arm. Our formulae, however, allow formal evaluation of this and may identify situations in which variation in intracluster correlation coefficient by treatment arm or another binary covariate may have a more substantial impact on power and hence sample size requirements.

Keywords: Cluster randomized trial; generalized estimating equations; intracluster correlation coefficient; logistic regression; sample size.

PubMed Disclaimer

Conflict of interest statement

Declaration of conflicting interests

The authors declare no conflicting interests.

Figures

**Figure 1.**
Asymptotic standard error of estimated ${\hat{β}}_{1}$ (SE, panels A and B) and asymptotic relative efficiency (ARE, C and D) vs. ratio of intracluster correlation coefficient (ICC) among treated compared to control clusters ( $ρ_{1} / ρ_{0}$ , A and C) and ratio of outcome prevalence among treated compared to control clusters ( $π_{1} / π_{0}$ , B and D) by working correlation structure, for fixed ρ₀, π₀, π₁ (A and C), ρ₁ (B and D), cluster size distribution, and K. The vertical line indicates the parameters observed in the trial by Lin *et al*.

**Figure 2.**
Asymptotic standard error of estimated ${\hat{β}}_{1}$ (SE, panels A and B) and asymptotic relative efficiency (ARE, C and D) vs. mean cluster size (A and C) and coefficient of variation of the cluster size distribution (CV, B and D) by working correlation structure, for fixed ρ₀, ρ₁, π₀, π₁, CV (A and C), mean cluster size (B and D), and K. The vertical line indicates the parameters observed in the trial by Lin *et al*.

**Figure 3.**
Power to detect effect vs. ratio of ICC among treated compared to control clusters ( $ρ_{1} / ρ_{0}$ , panel A), ratio of outcome prevalence among treated compared to control clusters ( $π_{1} / π_{0}$ , B), mean cluster size (C), and coefficient of variation of the cluster size distribution (D) by combination of working correlation structure, assumed underlying correlation structure, and *a priori* estimate of the common ICC used to calculate sample size (SS), for fixed ρ₀, ρ₁ (B,C,D), π₀, π₁ (A,C,D), CV of cluster size distribution (A,B,C), mean cluster size (A,B,D), significance level of 0.05, desired 80% power, and null and alternative hypothesis values (A,C,D). The vertical line indicates the parameters observed in the trial by Lin *et al*.

See this image and copyright information in PMC

Cited by

Characterisation of between-cluster heterogeneity in malaria cluster randomised trials to inform future sample size calculations.
Biggs J, Challenger JD, Dee D, Elobolobo E, Chaccour C, Saute F, Staedke SG, Vilakati S, Chung JB, Hsiang MS, Dabira ED, Erhart A, D'Alessandro U, Tripura R, Peto TJ, Von Seidlein L, Mukaka M, Mosha J, Protopopoff N, Accrombessi M, Hayes R, Churcher TS, Cook J. Biggs J, et al. Nat Commun. 2025 Jul 18;16(1):6615. doi: 10.1038/s41467-025-61502-w. Nat Commun. 2025. PMID: 40681482 Free PMC article.
The Effect of School-Linked Module-Based Friendly-Health Education on Adolescents' Sexual and Reproductive Health Knowledge, Guji Zone, Ethiopia - Cluster Randomized Controlled Trial.
Boku GG, Garoma Abeya S, Ayers N, Abera Wordofa M. Boku GG, et al. Adolesc Health Med Ther. 2024 Jan 23;15:5-18. doi: 10.2147/AHMT.S441957. eCollection 2024. Adolesc Health Med Ther. 2024. PMID: 38282688 Free PMC article.
Effect of a school-linked life skills intervention on adolescents' sexual and reproductive health skills in Guji zone, Ethiopia (CRT)-A generalized linear model.
Godana G, Garoma S, Ayers N, Abera M. Godana G, et al. Front Public Health. 2023 Oct 23;11:1203376. doi: 10.3389/fpubh.2023.1203376. eCollection 2023. Front Public Health. 2023. PMID: 37937073 Free PMC article. Clinical Trial.

References

1. Hayes RJ and Moulton LH. Cluster Randomised Trials. 2nd ed. Boca Raton, FL: CRC Press, 2017.
1. Campbell MK, Fayers PM and Grimshaw JM. Determinants of the intracluster correlation coefficient in cluster randomized trials: the case of implementation research. Clin Trials 2005; 2: 99–107. - PubMed
1. Choudhry NK, Avorn J, Glynn RJ, et al. Full coverage for preventive medications aftermyocaridal infarction. N Engl J Med 2011; 365: 2088–2097. - PubMed
1. Benger JR, Kirby K, Black S, et al. Effect of a strategy of a supraglottic airway device vs tracheal intubation during out-of-hospital cardiac arrest on functional outcome: the AIRWAYS-2 randomized clinical trial. JAMA 2018; 320: 779–791. - PMC - PubMed
1. Shih WJ. Sample size and power calculations for periodontal and other studies with clustered samples using the method of generalized estimating equations. Biom J 1997; 39: 899–908.

Publication types

Actions

MeSH terms

Actions
Actions
Actions
Actions
Actions
Actions

Grants and funding

LinkOut - more resources

Full Text Sources

Save citation to file

Email citation

Add to Collections

Add to My Bibliography

Your saved search

Create a file for external citation management software

Your RSS Feed

Power and sample size calculations for cluster randomized trials with binary outcomes when intracluster correlation coefficients vary by treatment arm

Affiliations

Power and sample size calculations for cluster randomized trials with binary outcomes when intracluster correlation coefficients vary by treatment arm

Authors

Affiliations

Abstract

Conflict of interest statement

Figures

Similar articles

Cited by

References

Publication types

MeSH terms

Grants and funding

LinkOut - more resources

Full Text Sources

Abstract

Conflict of interest statement

Figures

Similar articles

Cited by

References

Publication types

MeSH terms

Related information

Grants and funding

LinkOut - more resources

Full Text Sources