Statistical Calculators › Group Sequential, Adaptive, and Interim Analysis › Multi-Arm Multi-Stage (MAMS) Design Sample Size
← All Calculators

Group Sequential, Adaptive, and Interim Analysis

Multi-Arm Multi-Stage (MAMS) Design Sample Size

Sample-size planning for a multi-arm, multi-stage trial with several experimental treatments sharing one control arm. The calculator uses the generalized Dunnett MAMS framework: ineffective treatments may be dropped at interim analyses, while an efficacy boundary can stop the trial early. For continuous outcomes, the calculation uses the shared-control correlation structure and estimates the probability that the best treatment is correctly selected under the specified alternative.

MAMS Design Parameters

Continuous/normal endpoint — MAMS group-sequential design for means
Allocation: a ratio of 1 means equal allocation to each experimental arm and the shared control. A ratio of 2 means twice as many experimental patients as control patients at each stage.
Efficacy & futility boundaries
Enter one upper and one lower boundary for each stage. The upper boundary represents efficacy; the lower boundary represents lack of benefit/futility. Boundaries should be calibrated in advance to control the desired familywise type-I error rate. The calculator does not independently certify an arbitrary custom boundary set.

Required Sample Size

Simulation-based operating-characteristic calculation using the shared-control Dunnett correlation structure.
Enter design parameters and click Calculate Sample Size.

Methodology

The MAMS framework extends group-sequential testing to several experimental treatments compared with a common control. At each interim analysis, each active treatment has a test statistic comparing it with control. Treatments crossing the futility boundary are discontinued, while crossing the efficacy boundary permits early stopping for efficacy. Treatments remaining between the boundaries continue to the next stage.

Shared-control test statistics

For a continuous outcome with a common standard deviation, the standardized treatment-versus-control statistic has an approximate normal distribution. With an experimental-to-control allocation ratio of r, the same-stage correlation between two treatment-versus-control statistics is induced by the shared control group:

ρ = r / (1 + r)

Thus, with equal allocation, the correlation is 0.5. This is the key Dunnett-type feature of a MAMS design: the treatment comparisons are not independent because they share the same control observations.

Stagewise information

The calculator assumes equal numbers of patients per arm at each stage. Independent stage increments are generated with the shared-control correlation, and cumulative Z-statistics are formed at each analysis. For an effect difference δ, common SD σ, experimental-to-control allocation ratio r, and n patients per experimental arm and stage, the mean of a stagewise standardized treatment-control increment is:

E(Zincrement) = δ √[ n / { σ²(1 + 1/r) } ]

Cumulative statistics therefore acquire information as patients accumulate across stages while retaining the appropriate group-sequential correlation.

Least-favourable configuration

The power calculation uses a least-favourable configuration in which one experimental treatment has the clinically relevant effect δ1, while the remaining experimental treatments have the smaller effect δ0. This reflects the MAMS objective of correctly identifying a genuinely effective treatment while controlling erroneous treatment selection.

Selection rule

At each stage, the treatment with the largest Z-statistic is the candidate for efficacy stopping. If the largest statistic exceeds that stage's upper boundary, that treatment is selected and the trial stops. Any treatment whose statistic is at or below the lower boundary is dropped. If at least one treatment remains and efficacy has not been demonstrated, the surviving arms continue.

The reported power is therefore the probability that the prespecified best treatment is the treatment ultimately selected. This is the relevant selection-power interpretation for the MAMS design used in the validation example.

Sample-size search

Candidate maximum sample sizes are evaluated sequentially. For each candidate, the design is simulated under the least-favourable configuration using a fixed deterministic random-number seed. The smallest candidate satisfying the target power is reported, subject to the specified minimum stage-1 sample size.

Maximum evaluable N = (patients per arm at each stage) × (number of stages) × (number of experimental arms + 1 control arm)

Validation example

The calculator's default example reproduces the published TAILoR MAMS configuration: three experimental treatments plus control; two stages; standardized effects of 0.545 for the clinically relevant treatment and 0.178 for the other treatments; common SD of 1; equal allocation; interim efficacy boundary 2.782; final efficacy boundary 2.086; and a futility boundary of 0. The published design specified 42 patients per arm at the interim and a maximum of 84 evaluable patients per arm, giving 336 evaluable patients across four arms.

Stage 1: 42 / arm  →  168 total
Stage 2: +42 / continuing arm
Maximum: 84 / arm  →  336 total

With the calculator's deterministic validation simulation, the 84-per-arm configuration produces a best-treatment selection probability above the requested 90% target, while the imposed minimum stage-1 requirement is satisfied. The published TAILoR protocol reports the same 42-per-arm interim and 84-per-arm maximum design.

Important interpretation

MAMS sample size is not simply the sample size from an ordinary two-group comparison multiplied by the number of treatments. The shared control, correlated treatment-control statistics, treatment dropping, interim efficacy stopping, and treatment-selection objective all affect the operating characteristics. Consequently, a final design should use the exact prespecified stopping boundaries and, for a regulatory trial, should normally be independently verified with the intended validated statistical software.

References

Magirr D, Jaki T, Whitehead J. (2012). A generalized Dunnett test for multi-arm multi-stage clinical studies with treatment selection. Biometrika, 99(2), 494–501. doi:10.1093/biomet/ass002.

Jaki T, Pallmann P, Magirr D. (2019). The R Package MAMS for Designing Multi-Arm Multi-Stage Clinical Trials. Journal of Statistical Software, 88(4), 1–25.

Statsols / nQuery. nQuery Advanced User Manual. Multi-Arm Multi-Stage (MAMS) Group Sequential Tests: MGT6 for means and PGT4 for proportions.

Pushpakom SP, et al. TAILoR protocol: a dose-ranging Phase II randomized trial of telmisartan for reduction of insulin resistance in HIV-positive individuals. The published design used three active treatments, one control, two stages, critical values 2.782 and 2.086, 42 patients per arm at interim, and 336 maximum evaluable patients.