Multiplicity and Multiple Testing
Sample-size planning for two-family hierarchical multiple-testing designs. Specify primary and secondary hypothesis families, serial or parallel gatekeeping, standardized treatment effects, the family-wise significance level, and the desired probability of reaching a successful secondary objective.
Gatekeeping procedures are used when hypotheses are organized into ordered families and later hypotheses may be tested only after a prespecified success condition in an earlier family. Dmitrienko, Millen, Brechenmacher, and Paux describe gatekeeping strategies for confirmatory clinical trials with hierarchically ordered objectives and multiple hypothesis families. The framework is intended to preserve the overall family-wise Type I error rate while using the logical hierarchy among clinical objectives.
This calculator implements a transparent two-family Bonferroni gatekeeping model. Within each family, the global significance level is divided equally:
For an equal-size two-group comparison with standardized effect size d, the normal-theory noncentrality parameter is d√(n/2), where n is the number of evaluable participants per treatment group. The individual-hypothesis power is calculated from the corresponding normal distribution using the Bonferroni-adjusted critical value.
In a serial gate, the secondary family opens only when every primary hypothesis rejects. If the individual primary-hypothesis power is π₁, the probability of opening the gate under the independent planning model is:
The probability that at least one secondary hypothesis rejects after the gate opens is:
Therefore the planning probability of successfully reaching at least one secondary objective is the product of the gate-opening probability and the conditional secondary-family success probability.
In a parallel gate, the secondary family opens when at least one primary hypothesis rejects. The corresponding gate-opening probability is:
The same secondary-family calculation is then applied after the gate opens. Thus, changing from serial to parallel gatekeeping changes the probability that the secondary family becomes available without changing the per-hypothesis primary power.
The calculator searches integer values of evaluable participants per group and returns the smallest value for which the calculated secondary success probability is at least the requested target. If a dropout rate is specified, the evaluable sample size is inflated using:
The implementation is deliberately restricted to equal allocation, equal standardized effects within each hypothesis family, two hypothesis families, and independent normal-theory test statistics. More elaborate gatekeeping structures, unequal effects, correlated endpoints, graphical alpha recycling, truncated Hochberg procedures, or endpoint-specific covariance matrices require a more general multivariate power calculation.
A fixed numerical validation example uses two primary hypotheses and two secondary hypotheses, a serial gate, two-sided α = 0.05, primary effect size d₁ = 0.50, secondary effect size d₂ = 0.40, and a target probability of secondary success of 80%. No dropout inflation is applied.
At 86 evaluable participants per group, the calculated secondary success probability is approximately 79.45%. At 87 per group it rises to approximately 80.08%, so 87 per group is the first integer sample size meeting the 80% target. The implementation was independently evaluated against these calculations before presentation.
Dmitrienko, A., Millen, B. A., Brechenmacher, T., & Paux, G. (2011). Development of gatekeeping strategies in confirmatory clinical trials. Biometrical Journal, 53(6), 875–893. doi:10.1002/bimj.201100036.
Bretz, F., Maurer, W., Brannath, W., & Posch, M. (2009). A graphical approach to sequentially rejective multiple test procedures. Statistics in Medicine, 28, 586–604. This paper provides the broader sequentially rejective multiple-testing framework in which gatekeeping procedures can be represented.
/ Statistical Solutions. Sample Size and Power Calculation. documentation and validation materials describe the software's sample-size and power calculation framework and its validation of procedure-specific computational tables.
Clinical-trial applications illustrate the practical use of serial gatekeeping for preserving the study-wide Type I error rate across primary and secondary endpoints; for example, published protocols have used serial gatekeeping to stop downstream confirmatory testing after a failed higher-priority endpoint.