← Clinical Trials
Chronic Lymphocytic Leukemia Phase 3 Time-to-Event Analysis NCT01886872

ALLIANCE A041202: Complete Statistical Analysis of Ibrutinib in Previously Untreated Chronic Lymphocytic Leukemia

An independent statistical review of the randomized phase 3 ALLIANCE A041202 trial comparing rituximab and bendamustine hydrochloride, ibrutinib, and ibrutinib plus rituximab in older patients with previously untreated chronic lymphocytic leukemia.

Trial status: Active, not recruiting  ·  Enrollment: 547  ·  Primary completion: August 7, 2018
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

ALLIANCE A041202 is a randomized phase 3 trial evaluating three treatment strategies in older patients with previously untreated chronic lymphocytic leukemia: rituximab plus bendamustine hydrochloride, ibrutinib, and ibrutinib plus rituximab. The registered primary endpoint is progression-free survival (PFS), analyzed as a time-to-event outcome.

547
Enrollment
Randomized phase 3 trial
3
Treatment arms
A, B, and C
3
Primary analyses
Pairwise PFS comparisons
0.39
PFS HR
Ibrutinib vs Arm A
FeatureALLIANCE A041202
TrialALLIANCE A041202
ClinicalTrials.gov identifierNCT01886872
PhasePhase 3
Therapeutic areaHematology
ConditionStage I, Stage II, Stage III, and Stage IV chronic lymphocytic leukemia
AllocationRandomized
Design modelCrossover
MaskingNone
Primary purposeTreatment
Enrollment547
Primary endpointProgression Free Survival (PFS)
Primary endpoint typeTime-to-event
Primary statistical methodLog-rank test
Primary effect measureHazard ratio
Hypothesis typeSuperiority
Lead sponsorNational Cancer Institute (NCI)
Sponsor typeNIH
StartJanuary 15, 2014
Primary completionAugust 7, 2018
Results postedYes

2. Clinical Question

The registered clinical question concerns progression-free survival in older patients with previously untreated chronic lymphocytic leukemia. The trial compares three treatment strategies rather than a simple two-arm intervention-versus-control design.

Population

Older patients with previously untreated chronic lymphocytic leukemia. The registered conditions include Stage I, Stage II, Stage III, and Stage IV chronic lymphocytic leukemia.

Arm A

Rituximab and bendamustine hydrochloride.

Arm B

Ibrutinib.

Arm C

Ibrutinib and rituximab.

The primary statistical question is whether the observed PFS distributions differ between the randomized treatment groups, with the registry reporting three pairwise primary analyses and superiority hypotheses.

3. Trial Design

01
Randomize547 enrolled
02
Three armsA, B, and C
03
Study treatmentRituximab, bendamustine, ibrutinib
04
Follow-upPFS events and censoring
05
Primary analysisLog-rank and HR
ARM A

Rituximab + Bendamustine Hydrochloride

  • Rituximab
  • Bendamustine hydrochloride
  • Serious adverse events reported: 87 affected / 176 at risk
ARM B

Ibrutinib

  • Ibrutinib
  • Serious adverse events reported: 112 affected / 180 at risk
  • Crossover serious adverse events: 16 affected / 27 at risk
ARM C

Ibrutinib + Rituximab

  • Ibrutinib
  • Rituximab
  • Serious adverse events reported: 108 affected / 181 at risk

The registry identifies the design model as crossover and masking as none. The ClinicalTrials.gov record does not provide additional details about the crossover rule, timing, eligibility criteria for crossover, or how crossover was incorporated into the primary PFS analysis. Those details therefore should not be inferred from the design label alone.

4. Trial Timeline and Registry Status

January 15, 2014

Trial start

The registry lists January 15, 2014 as the study start date.

August 7, 2018

Primary completion

The registry lists August 7, 2018 as the primary completion date.

Current registry status

Active, not recruiting

The ClinicalTrials.gov record classifies the study as ACTIVE_NOT_RECRUITING.

5. Primary Endpoint

EndpointRegistry definition / time frameAnalysis
Progression Free Survival (PFS) Time from study entry to the time of documented disease progression or death. Kaplan-Meier estimation with log-rank comparison; hazard ratio reported for pairwise comparisons.

The registry describes the PFS analysis as event driven. The registry time-frame field continues beyond the available text, but the reported core definition is time from study entry to documented disease progression or death.

Why this endpoint matters: PFS is a time-to-event endpoint. Patients who have not experienced documented progression or death by the time of their last evaluable follow-up can contribute censored follow-up rather than being treated as if the event occurred. This is one reason ordinary comparisons of proportions are not the appropriate primary method for this endpoint.

6. Statistical Analysis Plan Reflected in the Registry

The ClinicalTrials.gov record identifies the log-rank test as the primary normalized statistical method and the hazard ratio as the effect measure. All three posted primary analyses use the same general framework.

ComponentRegistry information
EndpointProgression Free Survival (PFS)
Endpoint typeTime-to-event
Analysis familySurvival analysis
Comparison methodLog-rank test
Effect measureHazard ratio
Hypothesis typeSuperiority
Confidence intervals95%, two-sided
Primary analyses posted3

Three primary analyses are reported because each of the three treatment arms participates in a pairwise PFS comparison. The registry data do not state an adjustment procedure for multiplicity among these three primary comparisons. Consequently, the reported P-values should be presented exactly as posted rather than assigning them an additional multiplicity interpretation that is not contained in the ClinicalTrials.gov record.

7. Results: Progression-Free Survival

ClinicalTrials.gov contains three statistical analyses for the registered primary endpoint. Each compares a pair of randomized treatment groups using a log-rank test and reports a hazard ratio with a two-sided 95% confidence interval.

Arm B: Ibrutinib vs Arm A: Rituximab + Bendamustine Hydrochloride

Hazard ratio for progression-free survival

0.39

95% CI: 0.26–0.58   ·   P < 0.001

Comparison: Arm B (Ibrutinib) versus Arm A (Rituximab, Bendamustine Hydrochloride)

Clinical Biostats interpretation

An HR of 0.39 means that, within the time-to-event framework used for this comparison, the estimated instantaneous rate of the PFS event in the ibrutinib group was 0.39 times that of Arm A. Expressed as a simple relative-hazard interpretation, this corresponds to an estimated 61% lower hazard because \(1-0.39=0.61\).

The HR does not mean that 61% of patients avoided progression or death, nor does it mean that every patient experienced exactly a 61% reduction in risk. It is a relative time-to-event measure summarizing the treatment comparison under the statistical model.

The 95% confidence interval of 0.26–0.58 describes uncertainty around the estimated hazard ratio. It does not describe the range of PFS times for individual patients. The interval also remains below 1, which is consistent with the direction of the reported superiority comparison.

The P < 0.001 value addresses evidence against the null hypothesis in the reported statistical test. It is not a measure of the magnitude of the treatment effect and does not tell us how clinically important an effect is on its own.

As with other hazard-ratio analyses, interpretation of a single HR is most straightforward when the proportional-hazards assumption is reasonably appropriate. The ClinicalTrials.gov record does not provide a proportional-hazards diagnostic or a time-varying treatment-effect analysis.

Arm C: Ibrutinib + Rituximab vs Arm A: Rituximab + Bendamustine Hydrochloride

Hazard ratio for progression-free survival

0.38

95% CI: 0.25–0.59   ·   P < 0.001

Comparison: Arm C (Ibrutinib, Rituximab) versus Arm A (Rituximab, Bendamustine Hydrochloride)

Clinical Biostats interpretation

An HR of 0.38 indicates an estimated instantaneous PFS-event rate equal to 0.38 times that of Arm A under the reported analysis. In simple relative terms, \(1-0.38=0.62\), so the estimated hazard is approximately 62% lower for the Arm C comparison.

This does not mean that 62% of patients are progression-free, that 62% of patients are cured, or that an individual patient's PFS is expected to be 62% longer. The hazard ratio summarizes a relative comparison of event rates over time.

The 95% CI of 0.25–0.59 quantifies uncertainty around the estimated HR. It remains below 1 throughout the reported interval, while the associated P-value is < 0.001.

Again, the P-value should not be interpreted as an effect-size statistic. The effect magnitude is conveyed by the HR and its confidence interval; the P-value addresses the hypothesis test associated with the comparison.

The registry does not provide a formal interaction analysis comparing the treatment effect in Arm C with the treatment effect in Arm B. Therefore, the similarity of the two HR estimates versus Arm A should not by itself be interpreted as proof that the combination and ibrutinib-alone strategies have identical effects.

Arm C: Ibrutinib + Rituximab vs Arm B: Ibrutinib

Hazard ratio for progression-free survival

1.00

95% CI: 0.62–1.62   ·   P = 0.49

Comparison: Arm C (Ibrutinib, Rituximab) versus Arm B (Ibrutinib)

Clinical Biostats interpretation

An HR of 1.00 is the point estimate of equal instantaneous PFS-event rates between the two groups under the reported analysis. The confidence interval is considerably wider than the point estimate: 0.62–1.62.

The interval spans both values below and above 1. Therefore, the data represented by this confidence interval are compatible with a lower estimated hazard in Arm C as well as a higher estimated hazard in Arm C, within the uncertainty represented by the model and sample data.

The reported P = 0.49 does not establish that the two treatments are equivalent. A nonsignificant superiority test means that the reported analysis did not provide sufficient statistical evidence for a difference under that test; it is not the same as a prespecified equivalence or non-inferiority demonstration.

The P-value also does not measure the size of any potential treatment difference. The confidence interval is more informative about the range of effects compatible with the analysis than the P-value alone.

8. Primary PFS Results Side by Side

ComparisonMethodHR95% CIP-valueHypothesis
Arm B: Ibrutinib vs Arm A: Rituximab + bendamustine hydrochloride Log-rank 0.39 0.26–0.58 <0.001 Superiority
Arm C: Ibrutinib + rituximab vs Arm A: Rituximab + bendamustine hydrochloride Log-rank 0.38 0.25–0.59 <0.001 Superiority
Arm C: Ibrutinib + rituximab vs Arm B: Ibrutinib Log-rank 1.00 0.62–1.62 0.49 Superiority

The three comparisons show why a multi-arm trial needs to be read as a set of prespecified pairwise questions rather than as one overall treatment-versus-control comparison. Two comparisons use Arm A as the reference, while the third directly compares the two ibrutinib-containing strategies.

9. Understanding the Three Pairwise Comparisons

Arm B vs Arm A

The reported HR of 0.39 describes the relative PFS hazard for ibrutinib compared with rituximab plus bendamustine hydrochloride. The 95% CI is 0.26–0.58.

Arm C vs Arm A

The reported HR of 0.38 describes the relative PFS hazard for ibrutinib plus rituximab compared with rituximab plus bendamustine hydrochloride. The 95% CI is 0.25–0.59.

Arm C vs Arm B

The reported HR of 1.00 compares ibrutinib plus rituximab directly with ibrutinib. The 95% CI is 0.62–1.62.

Why the direct comparison matters

Comparing each active strategy with Arm A does not answer whether the two active strategies differ from one another. That question is addressed directly by the Arm C versus Arm B analysis.

This structure is statistically important. If two treatments each differ from a common reference, the two treatment effects cannot automatically be compared by visually inspecting their separate confidence intervals or P-values. A direct comparison, such as the reported Arm C versus Arm B analysis, addresses the treatment contrast itself.

10. Kaplan-Meier Estimation

Because PFS is a time-to-event endpoint, Kaplan-Meier estimation is the natural descriptive framework identified in the registry's endpoint definition. It estimates the probability of remaining event-free over time while allowing patients to be censored when their available follow-up ends without a documented event.

Conceptual Kaplan-Meier form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at an event time and ni represents the number at risk immediately before that time. The resulting curve estimates the event-free survival function over time.

The ClinicalTrials.gov record does not include the underlying patient-level event and censoring times or a digitized Kaplan-Meier curve. It is therefore inappropriate to manufacture median PFS values, landmark PFS probabilities, or a reconstructed survival curve from the hazard ratios alone.

11. Log-Rank Testing

The reported comparison method is the log-rank test. Conceptually, the log-rank test compares the observed and expected numbers of events between treatment groups over the observed event times.

What the log-rank test asks
Observed events vs expected events under the null hypothesis

The null hypothesis for a standard two-group log-rank comparison is that the event-time distributions do not differ between groups. The test uses the ordering and timing of events rather than reducing PFS to a single fixed-time proportion.

For ALLIANCE A041202, the registry reports log-rank analyses for all three pairwise primary comparisons. The reported P-values are therefore associated with time-to-event comparisons, not with ordinary comparisons of percentages.

12. Hazard Ratios and Confidence Intervals

The hazard ratio is the effect measure reported for each primary PFS analysis. Its interpretation requires care because a hazard is an instantaneous event rate conditional on having remained event-free up to that point.

HRGeneral interpretation
Less than 1The first-named treatment has a lower estimated instantaneous event rate than the comparison treatment.
Equal to 1The estimated instantaneous event rates are equal at the point estimate.
Greater than 1The first-named treatment has a higher estimated instantaneous event rate than the comparison treatment.

For the two comparisons against Arm A, the reported HRs are below 1. For the direct comparison of Arm C versus Arm B, the point estimate is exactly 1.00.

A confidence interval adds information that a point estimate cannot provide. For example, the Arm C versus Arm B interval of 0.62–1.62 is substantially broader around the point estimate than would be implied by reporting 1.00 alone. This makes the uncertainty surrounding the direct comparison central to interpretation.

A common mistake

A hazard ratio should not be translated automatically into a percentage of patients who benefited. An HR of 0.39 is not a statement that 39% of patients remained progression-free. It is a relative time-to-event effect measure.

13. P-Values: What They Do and Do Not Tell You

The registry reports P < 0.001 for the comparisons of Arm B versus Arm A and Arm C versus Arm A, and P = 0.49 for Arm C versus Arm B.

These P-values answer a narrower question than the hazard ratios. They quantify the compatibility of the observed data with the null hypothesis used for the corresponding test, under the assumptions of that statistical framework.

P-value is not effect size

The P-value does not tell you whether an effect is large or small. The HR and its confidence interval provide the principal effect-size information here.

P-value is not probability of the null

A P-value is not the probability that the null hypothesis is true or false.

P = 0.49 is not equivalence

The Arm C versus Arm B analysis was registered as a superiority hypothesis. A nonsignificant superiority test is not automatically evidence of equivalence or non-inferiority.

Multiple comparisons matter

Three primary pairwise analyses are posted. The ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure, so the P-values should not be assigned an unreported adjustment.

14. Crossover Design

ClinicalTrials.gov identifies the design model as crossover. Crossover creates an important distinction between treatment assignment and subsequent treatment exposure.

In a crossover trial, patients may contribute information under more than one treatment exposure depending on the protocol. This can complicate causal interpretation if outcomes occurring after crossover are attributed to the original randomized assignment without considering the timing and mechanism of crossover.

For ALLIANCE A041202, the ClinicalTrials.gov record identifies the crossover design and separately report serious adverse events for an Arm B crossover group: 16 affected / 27 at risk. The ClinicalTrials.gov record does not provide the crossover timing, crossover eligibility rule, or a specific statistical adjustment used to account for crossover in the PFS analyses.

Interpretation caution: the existence of a crossover design does not by itself tell us how much crossover affected the reported hazard ratios. That requires the protocol, analysis population definitions, crossover timing, and an analysis method specifically addressing treatment switching. Those details are not contained in the ClinicalTrials.gov record.

15. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected patients divided by patients at risk.

GroupSerious adverse eventsPopulation at risk
Arm A: Rituximab + bendamustine hydrochloride87 affected176 at risk
Arm B: Ibrutinib112 affected180 at risk
Arm C: Ibrutinib + rituximab108 affected181 at risk
Arm B crossover16 affected27 at risk

These counts should be kept separate from the PFS efficacy analyses. Serious adverse events and progression-free survival answer different questions and use different statistical frameworks. A treatment arm's safety experience cannot be inferred from its PFS hazard ratio, and a PFS effect cannot be inferred from serious-adverse-event counts.

Data presentation rule: the affected/at-risk figures above are reproduced exactly from the ClinicalTrials.gov record. No additional safety percentages, event classifications, severity grades, or comparative safety tests have been added because they were not provided.

16. Non-Inferiority, Superiority, and the Role of the Null Value

The registry identifies the hypothesis type for all three primary analyses as superiority. This is important for interpreting the HR of 1.00 in the Arm C versus Arm B comparison.

For a superiority analysis of a hazard ratio, the conventional null value is 1. An HR below 1 is in the direction of a lower event hazard for the first-named group, while an HR above 1 is in the opposite direction.

QuestionWhat would be required?
SuperiorityEvidence against the null value of 1 in the prespecified direction and statistical framework.
Non-inferiorityA prespecified non-inferiority margin and an analysis demonstrating that the confidence interval excludes effects worse than that margin.
EquivalencePrespecified lower and upper equivalence margins and an analysis demonstrating that the confidence interval lies entirely within those margins.

No non-inferiority margin or equivalence margin is provided in the ClinicalTrials.gov record. Therefore, the Arm C versus Arm B result should be interpreted as a reported superiority comparison, not converted into a non-inferiority or equivalence claim.

17. Multiplicity Across Three Primary Analyses

One of the most important statistical features of this trial is that the registry reports three primary-Endpoint analyses for the same registered primary endpoint. They correspond to the three pairwise treatment comparisons:

  1. Arm B versus Arm A.
  2. Arm C versus Arm A.
  3. Arm C versus Arm B.

When multiple hypothesis tests are performed within a confirmatory framework, the probability of obtaining at least one false-positive result can depend on how the tests are handled jointly. Common strategies include a prespecified multiplicity adjustment, hierarchical testing, or an explicitly defined alpha-allocation scheme.

The ClinicalTrials.gov record does not specify such an adjustment. Consequently, the page reports the three P-values exactly as posted but does not reinterpret them as adjusted or unadjusted familywise-error probabilities.

Why this matters

Multiplicity is not a reason to discard the reported results; it is a reason to distinguish the individual reported tests from a formal claim about the entire family of three hypotheses. The statistical strength of any familywise conclusion depends on the prespecified testing strategy, which is not contained in the ClinicalTrials.gov record.

18. Interim Analysis

The ClinicalTrials.gov record does not provide an interim-analysis description, alpha-spending plan, stopping boundary, information fraction, or interim efficacy decision rule.

Because these details are absent, no interim-monitoring method is attributed to ALLIANCE A041202 on this page. In general, an interim analysis requires special control of the type I error if repeated looks at accumulating efficacy data are part of a confirmatory design.

Important distinction: the absence of an interim-analysis method in the ClinicalTrials.gov record does not establish that no interim analysis occurred. It means only that the ClinicalTrials.gov record do not document one sufficiently to describe it here.

19. Missing Data and Censoring

PFS is inherently subject to censoring because not every patient will necessarily experience documented disease progression or death during the period in which their outcome is observed. Kaplan-Meier and log-rank methods are designed for time-to-event data containing such right-censored observations under their relevant assumptions.

The ClinicalTrials.gov record does not specify the detailed censoring rules, missing-data handling, sensitivity analyses, or imputation strategy for PFS.

Censoring

Censoring means that the exact event time is not observed within the available follow-up, while the patient can still contribute information up to the censoring point.

Imputation

No specific PFS imputation method is reported in the ClinicalTrials.gov record. It would therefore be inappropriate to claim that a particular imputation procedure was used.

20. Stratification

The ClinicalTrials.gov record identifies the method as a log-rank test but do not provide stratification factors. The registry data therefore do not support adding a stratified log-rank test or a stratified Cox model to the formal description of the posted analyses.

This distinction matters because a trial can be randomized and multi-arm without every analysis necessarily being described as stratified. The method actually reported for these analyses is the appropriate basis for this page: log-rank, with hazard ratio as the effect measure.

21. Bayesian Methods

No Bayesian statistical method is identified in the ClinicalTrials.gov record. The reported primary analysis uses frequentist log-rank testing with hazard-ratio estimates and two-sided 95% confidence intervals.

Accordingly, this page does not interpret the reported results as posterior probabilities, credible intervals, Bayesian predictive probabilities, or Bayesian decision criteria.

22. Statistical Methods Explained

Why was a log-rank test used?

PFS is a time-to-event endpoint, so patients can have different follow-up times and can be censored. The log-rank test compares event-time distributions while accounting for the timing of events and censoring rather than treating every patient as a simple binary event/no-event observation at one fixed time.

What does an HR of 0.39 mean?

For the Arm B versus Arm A comparison, an HR of 0.39 means that the estimated instantaneous rate of the PFS event for Arm B was 0.39 times the corresponding rate for Arm A under the reported analysis. The simple relative-hazard interpretation is a 61% lower estimated hazard, calculated as \(1-0.39\).

Why is an HR of 1.00 not evidence of equivalence?

The Arm C versus Arm B analysis was specified as a superiority hypothesis. Its HR point estimate is exactly 1.00, but the 95% CI spans 0.62–1.62. Equivalence or non-inferiority requires prespecified margins and a corresponding confidence-interval decision rule. Neither is provided in the ClinicalTrials.gov record.

Why is the confidence interval important?

A point estimate alone does not show how uncertain the estimated treatment effect is. The confidence interval supplies a range of values compatible with the statistical analysis under its assumptions. For Arm C versus Arm B, the interval is 0.62–1.62, illustrating substantially more uncertainty around the direct comparison than the point estimate of 1.00 alone would convey.

Why can the P-value not replace the hazard ratio?

The P-value addresses evidence against the null hypothesis; it does not quantify the magnitude of the treatment effect. Two studies can have similar P-values with different effect sizes, and a statistically small P-value can occur with a relatively modest effect when information is abundant. Here, the HR and its confidence interval are needed to describe the estimated PFS treatment effect.

Why does the three-arm structure matter?

There are three distinct pairwise treatment questions. The two comparisons against Arm A establish how Arms B and C compare with the same reference, while the Arm C versus Arm B analysis directly addresses the difference between the two ibrutinib-containing strategies. Separate comparisons against a common control should not be substituted for the direct comparison.

Why does crossover complicate interpretation?

Crossover can cause subsequent treatment exposure to differ from the originally randomized treatment assignment. That distinction matters particularly for causal interpretation of outcomes after treatment switching. The registry identifies the design as crossover, but the ClinicalTrials.gov record does not provide enough information to determine the magnitude or statistical handling of crossover in the reported PFS analyses.

23. What the Confidence Intervals Say About Precision

ComparisonPoint estimate95% CIPrecision interpretation
Arm B vs Arm A 0.39 0.26–0.58 The interval remains below 1 and is relatively concentrated around a point estimate below 1.
Arm C vs Arm A 0.38 0.25–0.59 The interval remains below 1 and is centered around a point estimate below 1.
Arm C vs Arm B 1.00 0.62–1.62 The interval spans 1 and includes values representing both lower and higher hazards for Arm C.

The intervals should not be interpreted as a range of individual patient outcomes. They quantify uncertainty in the estimated population-level treatment contrast under the statistical model and confidence-interval procedure.

24. Limitations

25. Why This Trial Matters Statistically

ALLIANCE A041202 is a useful statistical teaching case because it combines a randomized three-arm design, a time-to-event primary endpoint, pairwise hazard-ratio comparisons, and a crossover design. Those features create several distinct statistical questions that should not be collapsed into one summary number.

ConceptHow it appears in ALLIANCE A041202
RandomizationThe trial uses randomized allocation.
Three-arm designArm A, Arm B, and Arm C generate three pairwise primary comparisons.
Time-to-event endpointPFS is defined from study entry to documented disease progression or death.
Kaplan-Meier estimationThe registered endpoint definition specifies Kaplan-Meier estimation of PFS distributions.
Log-rank testThe posted statistical method for all three primary analyses is log-rank.
Hazard ratioHR is the reported effect measure for each primary comparison.
Confidence intervalEach HR has a two-sided 95% confidence interval.
Superiority testingAll three primary analyses are identified as superiority hypotheses.
CrossoverThe registry identifies the design model as crossover and reports a crossover safety group.
MultiplicityThree primary pairwise analyses are posted; the ClinicalTrials.gov record does not state the multiplicity procedure.
Safety analysisSerious adverse events are reported separately by treatment arm and crossover group.

26. Interpreting the Trial as a Statistical Story

The most important feature of the reported results is the contrast between the two comparisons against Arm A and the direct comparison between Arms B and C.

Against Arm A, the reported PFS hazard ratios are 0.39 for Arm B and 0.38 for Arm C, with two-sided 95% confidence intervals of 0.26–0.58 and 0.25–0.59, respectively. Both corresponding P-values are reported as < 0.001.

The direct comparison of Arm C versus Arm B is different: the HR is 1.00, the 95% CI is 0.62–1.62, and the P-value is 0.49.

These three results illustrate a fundamental principle of multi-arm clinical-trial analysis: treatment effects are defined by specific contrasts. A treatment's HR versus one reference treatment does not, by itself, establish the HR versus another treatment. The direct comparison is the appropriate statistical object when the question is whether two active strategies differ.

What the reported evidence does not establish

The ClinicalTrials.gov record does not provide median PFS, absolute PFS probabilities at specific time points, subgroup treatment effects, a formal multiplicity procedure, or detailed crossover-adjusted analyses. Those quantities should not be inferred from the reported HRs and P-values.

27. Results and Statistical Interpretation: A Compact Summary

QuestionReported resultStatistical reading
Does Arm B differ from Arm A in PFS? HR 0.39; 95% CI 0.26–0.58; P < 0.001 The estimated PFS hazard is lower for Arm B under the reported superiority analysis.
Does Arm C differ from Arm A in PFS? HR 0.38; 95% CI 0.25–0.59; P < 0.001 The estimated PFS hazard is lower for Arm C under the reported superiority analysis.
Does Arm C differ from Arm B in PFS? HR 1.00; 95% CI 0.62–1.62; P = 0.49 The reported superiority test does not provide evidence of a difference; the confidence interval remains compatible with both lower and higher hazards for Arm C.

28. Related Tutorials

Learn more about the methods used in this trial:

29. Related Calculators

30. Sources

Continue through Clinical Biostats

Explore statistical tutorials, calculators, and additional clinical-trial analyses organized around the methods used in real randomized studies.

31. Record Summary

ALLIANCE A041202 provides a useful example of statistical analysis for a randomized three-arm clinical trial with a time-to-event primary endpoint. The registered endpoint is progression-free survival, and the posted analyses use log-rank tests with hazard ratios and two-sided 95% confidence intervals under superiority hypotheses.

The reported comparisons show HR 0.39 for Arm B versus Arm A and HR 0.38 for Arm C versus Arm A, while the direct Arm C versus Arm B comparison reports HR 1.00. The corresponding confidence intervals and P-values provide the uncertainty and hypothesis-testing context for those estimates.

The statistical interpretation must also account for the trial's three-arm structure, crossover design, and the fact that three primary pairwise analyses are posted without a multiplicity procedure being reported in the ClinicalTrials.gov record used for this page. The result is best understood not as one overall number, but as a set of treatment contrasts answering distinct questions about PFS.

Clinical Biostats methodology: A trial-results page should distinguish reported statistical evidence from educational interpretation. For ALLIANCE A041202, that means preserving the posted hazard ratios, confidence intervals, P-values, endpoint definition, and analysis method while avoiding unsupported reconstruction of median survival, subgroup effects, crossover adjustments, multiplicity procedures, or other quantities not contained in the ClinicalTrials.gov record.