Diagnostic Accuracy and Agreement
Estimate the number of subjects needed to obtain a specified confidence-interval width for an intraclass correlation coefficient when each subject is measured by the same number of raters, instruments, or repeated measurements. The calculation uses the large-sample Bonett approximation used by agreement/reliability sample-size procedures such as AOC3.
This calculator uses the precision-based intraclass correlation sample-size approximation described by Douglas G. Bonett (2002). The objective is to choose the number of subjects so that the planned confidence interval for the ICC has a specified total width. Bonett derived an approximation for the required number of subjects for ICCs estimated under one-way and certain two-way ANOVA models.
Let ρ be the anticipated ICC, k be the number of measurements or raters per subject, w be the desired total confidence-interval width, and z1-α/2 be the corresponding standard-normal quantile. The approximate number of subjects is:
The calculated value is rounded upward because a study cannot enroll a fractional subject. The formula uses the total confidence-interval width. Therefore, if the desired distance from the estimated ICC to either confidence limit is 0.10, the corresponding total width is 0.20.
The intraclass correlation describes the degree of similarity among measurements made on the same subject relative to the variation between subjects. In a reliability study, the repeated measurements may represent raters, instruments, observers, or measurement occasions.
Important: This is a confidence-interval precision calculation, not a hypothesis-test power calculation. It determines the sample size needed to target a specified confidence-interval width around an anticipated ICC.
The AOC3 documentation gives a worked reliability example with four measurements/raters per subject, an expected ICC of 0.85, and a two-sided 95% confidence interval extending approximately 0.10 from the observed ICC. Thus the desired total confidence-interval width is 0.20.
The calculator reproduces this result: the unrounded calculation is approximately 19.15 subjects, which rounds upward to 20 subjects.
A larger anticipated ICC generally reduces the required number of subjects, while a narrower desired confidence interval increases it. Increasing the number of measurements per subject also changes the precision through the k(k − 1) term and the ICC-dependent variance expression. The number of measurements should therefore be selected based on the actual reliability design rather than simply maximizing the number of raters.
Bonett, D. G. (2002). Sample size requirements for estimating intraclass correlations with desired precision. Statistics in Medicine, 21(9), 1331–1335. DOI: 10.1002/sim.1108.
the software. Tests for Intraclass Correlation. this method Sample Size Software documentation. the relevant methodological literature describes the ICC reliability design with N subjects and K observations per subject and identifies Walter, Eliasziw, and Donner (1998) as a source for its formulation.
Statistical Solutions / Advisor. Confidence Interval for Intraclass Correlation for k Measurements (AOC3). The documentation describes the procedure as a confidence-interval sample-size calculation using the number of measurements/raters and the expected ICC.
Walter, S. D., Eliasziw, M., & Donner, A. (1998). Sample size and optimal designs for reliability studies. Statistics in Medicine, 17(1), 101–110. DOI: 10.1002/(SICI)1097-0258(19980115)17:1<101::AID-SIM727>3.0.CO;2-E.