← Clinical Trials
Advanced NSCLC Phase 3 ALK-Positive NCT02075840

ALEX: Complete Statistical Analysis of Alectinib in Advanced ALK-Positive Non-Small Cell Lung Cancer

An independent statistical analysis of the randomized phase 3 ALEX trial comparing alectinib with crizotinib in treatment-naive participants with anaplastic lymphoma kinase-positive advanced non-small cell lung cancer.

Completed  ·  Randomized parallel design  ·  303 participants  ·  Hoffmann-La Roche
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results presented here are limited to the ClinicalTrials.gov record. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

ALEX was a randomized, open-label, parallel phase 3 trial comparing alectinib with crizotinib in treatment-naive participants with ALK-positive advanced non-small cell lung cancer. The registry reports 303 enrolled participants, two treatment arms, two registered primary endpoints, and 28 posted outcome measures.

Independent analysis notice: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
303
Enrolled
Total participants
2
Treatment Arms
Alectinib vs crizotinib
0.47
Primary PFS HR
95% CI 0.34–0.65
<0.0001
Primary PFS P-value
Two-sided
FeatureALEX
Trial nameALEX
NCT identifierNCT02075840
PhasePhase 3
ConditionNon-Small Cell Lung Cancer
PopulationTreatment-naive participants with anaplastic lymphoma kinase-positive advanced non-small cell lung cancer
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment303.0
InterventionsAlectinib; crizotinib
Primary endpointsProgression-Free Survival by Investigator Assessment; Percentage of Participants With PFS Event by Investigator Assessment
Results postedYes
Outcome measures posted28
Statistical analyses posted12
Lead sponsorHoffmann-La Roche
Sponsor typeIndustry

2. Clinical Question

The central question was whether treatment with alectinib compared with crizotinib affected progression-free survival and other clinical outcomes in treatment-naive participants with ALK-positive advanced non-small cell lung cancer.

Population

Treatment-naive participants with anaplastic lymphoma kinase-positive advanced non-small cell lung cancer.

Intervention

Alectinib.

Comparator

Crizotinib.

Primary question

Does alectinib produce a different progression-free survival outcome than crizotinib in the randomized comparison?

3. Trial Design

01
Randomize 303 participants
02
Parallel arms Alectinib vs crizotinib
03
Open label No masking
04
Assess outcomes PFS, CNS, response, OS, QoL
05
Analyze Survival and categorical methods
ARM A · ALECTINIB

Alectinib

  • Drug intervention
  • Randomized treatment assignment
  • Evaluated against crizotinib
ARM B · CRIZOTINIB

Crizotinib

  • Drug intervention
  • Randomized treatment assignment
  • Comparator for the alectinib group
Design characteristicRegistry description
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Phase3
Enrollment303.0

Randomization is the central design feature supporting the treatment comparison. Because the registry describes the trial as randomized and parallel, the principal efficacy comparison is naturally framed between the randomized alectinib and crizotinib groups rather than as a comparison of independently selected patient cohorts.

4. Trial Timeline

August 19, 2014

Trial start

The registry lists the study start date as 2014-08-19.

February 9, 2017

Primary completion

The registry lists the primary completion date as 2017-02-09.

Completed

Registry status

The trial is listed as completed, with results posted for 28 outcome measures and 12 statistical analyses.

5. Endpoints

The registry lists two primary endpoints: a time-to-event progression-free survival endpoint and a binary endpoint describing the percentage of participants with a progression-free survival event. The registry also reports statistical analyses for several secondary outcomes.

EndpointRegistry definition / time frameType
Progression-Free Survival (PFS) by Investigator Assessment Randomization to first documented disease progression or death, whichever occurs first (assessed every 8 weeks up to 33 months) Time-to-event
Percentage of Participants With PFS Event by Investigator Assessment Randomization to first documented disease progression or death, whichever occurs first (assessed every 8 weeks up to 33 months) Binary
PFS Independent Review Committee (IRC)-Assessed Randomization to first documented disease progression or death, whichever occurs first (assessed every 8 weeks up to 33 months) Time-to-event
Percentage of Participants With CNS Progression as Determined by IRC Using RECIST V1.1 Criteria Randomization to CNS PD as first occurrence of disease progression (assessed every 8 weeks up to 33 months) Time-to-event
Objective Response Rate (CR or PR) as Determined by Investigators According to RECIST V1.1 Criteria Randomization to first documented disease progression or death, whichever occurs first (assessed every 8 weeks up to 33 months) Binary
Overall Survival (OS) From randomization until death (up to 10.5 years) Time-to-event
Time to Deterioration by EORTC Quality of Life Questionnaire Core 30 (C30) Baseline, every 4 weeks until disease progression (up to 33 months) Time-to-event
Time to Deterioration by EORTC Quality of Life Questionnaire Lung Cancer Module 13 (LC13) Baseline, every 4 weeks until disease progression (up to 33 months) Time-to-event
Endpoint-definition discipline: The registry's registry-reported time-frame text for several endpoints ends at "up to 33" rather than providing a complete unit. This page preserves that registry wording rather than completing or inferring the missing text.

6. Primary Endpoint Results

Progression-Free Survival by Investigator Assessment

The registry reports a formal statistical analysis for the primary investigator-assessed PFS endpoint. The analysis population was the ITT population, defined in the registry as including all randomized participants in the study, with the number analyzed corresponding to participants evaluable for the outcome.

Stratified hazard ratio for progression or death

0.47

95% CI: 0.34–0.65   ·   P <0.0001

Log-rank test; superiority hypothesis.

FeatureReported analysis
EndpointProgression-Free Survival (PFS) by Investigator Assessment
ComparisonAlectinib vs Crizotinib
PopulationITT population
MethodLog-rank test
Effect measureHazard Ratio, stratified
Estimate0.47
95% CI0.34–0.65
P-value<0.0001
HypothesisSuperiority
StratificationRace (Asian vs Non-Asian) and CNS metastases at baseline by IRC
Clinical Biostats interpretation

The hazard ratio of 0.47 means that the estimated instantaneous rate of progression or death was approximately 47% of the corresponding rate in the comparator group under the fitted time-to-event comparison. Equivalently, 1 − 0.47 = 0.53, so the estimate corresponds to an approximately 53% lower estimated hazard for the alectinib group relative to crizotinib.

The hazard ratio does not mean that 53% of participants avoided progression, that every participant experienced exactly a 53% reduction in risk, or that the median PFS differed by 53%. It is a relative time-to-event measure rather than an absolute probability.

The 95% confidence interval of 0.34–0.65 describes statistical uncertainty around the estimated hazard ratio. It does not describe the range of effects that individual participants experienced.

The p-value of <0.0001 addresses the statistical evidence against the null hypothesis under the specified analysis. It is not a measure of the magnitude or clinical importance of the treatment effect. The magnitude is described by the hazard ratio and its confidence interval.

The registry identifies the analysis as stratified by race and baseline CNS metastases. Because the effect measure is a stratified hazard ratio, its interpretation depends on the specified stratified analysis rather than an unadjusted comparison.

Percentage of Participants With PFS Event by Investigator Assessment

The registry posts this second primary endpoint as a binary outcome, but the ClinicalTrials.gov record does not contain a corresponding formal statistical-analysis record for this endpoint. Therefore, a treatment-group difference, confidence interval, and p-value are not reported here.

How this endpoint would normally be analyzed

The endpoint records whether each participant experienced a progression-free survival event during the specified assessment period. A binary comparison would ordinarily compare the proportions with events between randomized groups, with an appropriate confidence interval for the between-group difference or another prespecified effect measure.

For ALEX, the ClinicalTrials.gov record identifies the endpoint and indicates that results are posted, but it does not provide a formal statistical-analysis result for this primary binary endpoint.

7. Secondary Endpoint Results

Independent Review Committee-Assessed PFS

Stratified hazard ratio

0.50

95% CI: 0.36–0.70   ·   P <0.0001

Log-rank test; superiority hypothesis.

Clinical Biostats interpretation

The IRC-assessed PFS hazard ratio of 0.50 corresponds to an estimated hazard approximately half that of the comparator group in the reported stratified analysis. The corresponding 95% CI, 0.36–0.70, quantifies uncertainty around that estimate.

The confidence interval does not establish the probability that the true hazard ratio lies inside the interval, nor does the p-value of <0.0001 quantify the size of the treatment effect. The p-value addresses evidence against the specified null hypothesis; the hazard ratio and confidence interval describe the estimated effect and its precision.

The analysis was stratified by race and baseline CNS metastases by IRC. This should be kept in mind when comparing the IRC result with the investigator-assessed primary PFS result.

CNS Progression

Cause-specific hazard ratio

0.16

95% CI: 0.10–0.28   ·   P <0.0001

IRC, RECIST v1.1 stratified analysis.

Clinical Biostats interpretation

The cause-specific hazard ratio of 0.16 indicates a substantially lower estimated cause-specific hazard for the CNS progression endpoint in the alectinib group relative to crizotinib under the reported analysis.

The 95% CI of 0.10–0.28 indicates the statistical uncertainty around that estimate. The endpoint is a time-to-event outcome, so the hazard ratio should not be interpreted as a simple ratio of the final percentages of participants with CNS progression.

The registry describes the analysis as an IRC RECIST v1.1 stratified analysis by race and baseline CNS metastases. The p-value of <0.0001 provides evidence against the specified null hypothesis, but it does not measure the magnitude of the CNS effect.

Objective Response Rate

Difference in overall response rates

7.40

95% CI: -1.71–16.50   ·   P = 0.0936

Cochran-Mantel-Haenszel analysis; superiority hypothesis.

Clinical Biostats interpretation

The reported effect measure is the difference in overall response rates, estimated as 7.40 percentage points for alectinib versus crizotinib under the reported analysis.

The 95% CI extends from -1.71 to 16.50. Thus, the interval includes zero, meaning that the reported confidence interval includes both a possible negative difference and positive differences of varying magnitude.

The p-value of 0.0936 is not an estimate of how large the response difference is. It quantifies the evidence against the specified null hypothesis under the Cochran-Mantel-Haenszel analysis. The effect estimate and confidence interval are needed to understand the magnitude and precision.

This endpoint is binary, so its interpretation differs from the hazard-ratio endpoints. It summarizes whether a participant achieved complete or partial response according to the specified RECIST v1.1 investigator assessment rather than the timing of progression or death.

Overall Survival

Hazard ratio for death

0.78

95% CI: 0.56–1.08   ·   P = 0.1320

Log-rank test; superiority hypothesis.

Clinical Biostats interpretation

The reported OS hazard ratio of 0.78 corresponds to an estimated instantaneous hazard of death approximately 78% of that in the crizotinib group under the reported time-to-event analysis.

The 95% CI of 0.56–1.08 crosses 1.00. Consequently, the interval includes the possibility of no difference in hazard as well as a range of lower and higher relative hazards.

The p-value of 0.1320 is evidence against the prespecified null hypothesis only within the framework of this analysis; it is not a probability that the treatment has no effect and it is not an effect-size measure.

The ClinicalTrials.gov record does not report median OS or an absolute survival percentage, so those quantities are not inferred here.

Time to Deterioration: EORTC QLQ-C30 Fatigue

Hazard ratio

0.74

95% CI: 0.46–1.19   ·   P = 0.2079

Log-rank test; fatigue-stratified analysis.

Clinical Biostats interpretation

The hazard ratio of 0.74 represents the estimated relative hazard of the fatigue deterioration event for alectinib compared with crizotinib under the reported analysis. The 95% CI of 0.46–1.19 includes 1.00, so the estimate is compatible with a range of relative effects.

The p-value of 0.2079 should not be interpreted as the probability that the treatment has no effect. It is a hypothesis-test quantity and does not replace the effect estimate or its confidence interval.

Time to Deterioration: EORTC QLQ-C30 Dyspnea

Hazard ratio

1.66

95% CI: 0.88–3.15   ·   P = 0.1137

Log-rank test.

Clinical Biostats interpretation

A hazard ratio of 1.66 means that the estimated hazard of the specified deterioration event was higher in the alectinib group relative to crizotinib under this analysis. Because the 95% CI extends from 0.88 to 3.15, substantial uncertainty remains around the estimate.

The direction of a hazard ratio must always be interpreted in relation to the event definition. Here, the endpoint is time to deterioration for dyspnea, so a hazard ratio above 1 does not automatically mean that the overall treatment effect was unfavorable; it describes this particular deterioration endpoint.

Time to Deterioration: EORTC QLQ-LC13 Coughing

Hazard ratio

0.88

95% CI: 0.44–1.74   ·   P = 0.7042

Log-rank test.

Clinical Biostats interpretation

The reported hazard ratio of 0.88 is close to 1.00, while the 95% CI of 0.44–1.74 is relatively broad and includes 1.00. The interval therefore allows for a range of possible relative effects.

The p-value of 0.7042 does not measure the magnitude of the observed estimate. It reflects the statistical evidence against the specified null hypothesis for this endpoint.

Time to Deterioration: EORTC QLQ-LC13 Dyspnea

Hazard ratio

1.76

95% CI: 1.05–2.92   ·   P = 0.0285

Log-rank test; dyspnea analysis.

Clinical Biostats interpretation

The reported hazard ratio of 1.76 indicates a higher estimated hazard of the specified dyspnea deterioration event in the alectinib group relative to crizotinib in this analysis. The 95% CI of 1.05–2.92 lies above 1.00.

The endpoint is specifically a time-to-deterioration outcome for dyspnea. Its interpretation should therefore remain tied to the endpoint definition rather than being generalized to overall efficacy or overall quality of life.

The p-value of 0.0285 provides evidence against the null hypothesis for this individual analysis, but the ClinicalTrials.gov record contains multiple secondary time-to-deterioration analyses. Those multiple comparisons are relevant when interpreting isolated p-values.

Time to Deterioration: EORTC QLQ-LC13 Pain in Arm and Shoulder

Hazard ratio

1.43

95% CI: 0.79–2.61   ·   P = 0.2377

Log-rank test.

Clinical Biostats interpretation

The estimated hazard ratio of 1.43 indicates a higher estimated hazard of the specified deterioration event in the alectinib group, but the 95% CI of 0.79–2.61 includes 1.00. The confidence interval therefore does not isolate a single direction of effect with high precision.

The p-value of 0.2377 should be interpreted as a hypothesis-test result, not as a probability that the observed treatment effect is absent.

Time to Deterioration: EORTC QLQ-LC13 Pain in Chest

Hazard ratio

0.51

95% CI: 0.24–1.10   ·   P = 0.0796

Log-rank test.

Clinical Biostats interpretation

The hazard ratio of 0.51 is below 1.00, indicating a lower estimated hazard of the specified chest-pain deterioration event in the alectinib group. However, the 95% CI of 0.24–1.10 crosses 1.00, so the precision of the estimate does not exclude no difference.

The p-value of 0.0796 should be kept in the context of the confidence interval and the fact that this was one of several secondary time-to-deterioration analyses.

Time to Deterioration: EORTC QLQ-LC13 Composite Score

Hazard ratio

1.10

95% CI: 0.72–1.68   ·   P = 0.6435

Log-rank test.

Clinical Biostats interpretation

The hazard ratio of 1.10 is close to 1.00, while the 95% CI of 0.72–1.68 spans both sides of 1.00. The estimate therefore provides limited precision about the direction or magnitude of the relative hazard.

The p-value of 0.6435 is a test result for this endpoint and should not be interpreted as a measure of effect size.

8. Statistical Methodology

Log-rank test

The registry identifies the log-rank test as the principal method for the reported time-to-event comparisons. The log-rank framework compares the observed and expected event patterns between randomized groups over follow-up, taking censoring into account.

Conceptual comparison
H0: the survival distributions are equivalent under the specified comparison

The test uses information accumulated over event times rather than reducing the outcome to a single fixed-time proportion.

Hazard ratios

The principal time-to-event effect measure reported for ALEX is the hazard ratio. A hazard ratio compares instantaneous event rates between groups within the specified survival-analysis framework.

Interpretation of the reported primary estimate
HR = 0.47  →  estimated hazard approximately 47% of the comparator hazard

Equivalently, the estimate corresponds to an approximately 53% lower estimated hazard. This is not the same as a 53% reduction in the probability of ever experiencing the event.

Stratified analysis

The primary investigator-assessed PFS analysis reports a stratified hazard ratio and p-value. The registry specifies stratification by race (Asian vs Non-Asian) and CNS metastases at baseline by IRC.

Stratification allows the analysis to account for prespecified categorical factors when comparing treatment groups. It is particularly relevant when those factors are associated with prognosis or were incorporated into the randomized design or analysis framework.

Intention-to-treat analysis

The registry defines the ITT population as including all randomized participants in the study. The reported efficacy analyses use this population, with the number analyzed described as the total participants evaluable for the relevant outcome measure.

The ITT principle preserves the treatment assignment created by randomization. In statistical terms, this helps maintain the comparability established by the randomized design rather than redefining groups according to treatment received after randomization.

Cochran-Mantel-Haenszel analysis

The objective-response endpoint was analyzed using a Cochran-Mantel-Haenszel test, reported in the registry as "Mantel Haenszel." This method is designed for categorical comparisons while allowing stratification when appropriate.

Binary endpoint framework
Difference in response rates = Response rateAlectinib − Response rateCrizotinib

The registry's reported effect measure for ORR was the difference in overall response rates, estimated as 7.40 with a 95% CI of -1.71–16.50.

9. Statistical Methods Explained

Why was a log-rank test used?

PFS, CNS progression, OS, and the quality-of-life deterioration endpoints are time-to-event outcomes. A participant's follow-up can end before the event occurs, producing censoring. The log-rank test is designed to compare event-time distributions between groups while using the available follow-up information.

What does a hazard ratio of 0.47 mean?

For the primary investigator-assessed PFS analysis, an HR of 0.47 means that the estimated instantaneous rate of progression or death was 47% of the comparator rate under the reported analysis. The complementary calculation, 1 − 0.47, gives an approximately 53% lower estimated hazard. This does not mean that 53% of patients were protected from progression.

Why is the confidence interval important?

The estimate alone does not communicate its statistical precision. The primary PFS estimate is 0.47, but the 95% CI is 0.34–0.65. The interval shows the range of parameter values compatible with the analysis under the stated confidence framework. A narrow interval generally conveys more precision than a wide interval; the width here is therefore part of the evidence, not an optional supplement to the point estimate.

Why doesn't the p-value measure effect size?

A p-value evaluates evidence against a null hypothesis under a specified statistical model and sampling framework. It does not say how large the treatment effect is. For ALEX, the primary PFS p-value is <0.0001, while the effect size is represented by the hazard ratio of 0.47 and its 95% CI of 0.34–0.65.

Why does stratification matter?

The primary PFS analysis was stratified by race and baseline CNS metastases by IRC. A stratified analysis allows the treatment comparison to account for these specified categories instead of treating all participants as though those factors were irrelevant to the comparison.

Why is the ORR analysis different from the PFS analysis?

ORR is a binary endpoint: a participant is classified according to whether a complete or partial response occurred. PFS is a time-to-event endpoint: the timing of progression or death matters, and participants can be censored. Consequently, the registry uses a Cochran-Mantel-Haenszel method for ORR and a log-rank framework for PFS.

Why should the secondary p-values be interpreted carefully?

The registry reports multiple secondary endpoints, including CNS progression, overall survival, and several quality-of-life deterioration outcomes. When many hypotheses are tested, the chance of observing at least one apparently unusual p-value can increase. The ClinicalTrials.gov record does not provide a multiplicity-adjustment scheme for these secondary analyses, so individual p-values should not automatically be treated as independent confirmatory findings.

10. Understanding the Primary PFS Result

The primary PFS result contains three different statistical pieces of information that answer different questions.

Effect size

HR 0.47 describes the estimated relative event hazard in the alectinib group compared with crizotinib.

Precision

95% CI 0.34–0.65 describes uncertainty around the hazard-ratio estimate.

Evidence against the null

P <0.0001 quantifies the statistical evidence under the reported hypothesis test.

Analysis structure

The estimate and p-value were derived from a stratified log-rank analysis of the ITT population.

Keeping these concepts separate prevents a common statistical error: treating the p-value as though it were the effect estimate. The p-value does not tell the reader whether the hazard ratio is 0.47, 0.80, or 0.20. That information comes from the effect estimate itself and its confidence interval.

11. Secondary Endpoint Pattern

Secondary endpointEffect95% CIP-valueMethod
IRC-assessed PFSHR 0.500.36–0.70<0.0001Log-rank
CNS progressionCause-specific HR 0.160.10–0.28<0.0001Log-rank
Objective response rateDifference 7.40-1.71–16.500.0936Cochran-Mantel-Haenszel
Overall survivalHR 0.780.56–1.080.1320Log-rank
C30 fatigue deteriorationHR 0.740.46–1.190.2079Log-rank
C30 dyspnea deteriorationHR 1.660.88–3.150.1137Log-rank
LC13 coughing deteriorationHR 0.880.44–1.740.7042Log-rank
LC13 dyspnea deteriorationHR 1.761.05–2.920.0285Log-rank
LC13 pain in arm and shoulder deteriorationHR 1.430.79–2.610.2377Log-rank
LC13 pain in chest deteriorationHR 0.510.24–1.100.0796Log-rank
LC13 composite-score deteriorationHR 1.100.72–1.680.6435Log-rank

The secondary results illustrate why a clinical-trial analysis should not reduce every endpoint to a simple "significant" versus "not significant" classification. The estimates differ in direction and precision, and the endpoints themselves measure different phenomena. PFS and CNS progression are disease-control time-to-event outcomes; ORR is a binary response outcome; OS measures death; and the EORTC endpoints measure time to specified quality-of-life deterioration.

12. Safety

The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm. The numbers are presented as affected participants divided by participants at risk.

Safety measureAlectinibCrizotinib
Serious adverse events70/15248/151
Serious adverse events: affected / at risk
Alectinib
70/152
Crizotinib
48/151
Safety interpretation: The ClinicalTrials.gov record reports affected and at-risk counts rather than a formal statistical comparison, confidence interval, or p-value for serious adverse events. This page therefore does not construct an inferential treatment comparison that is not present in the registry analysis data.

Safety and efficacy also answer different statistical questions. The efficacy analyses preserve randomized treatment assignment through the ITT framework, whereas safety summaries are fundamentally connected to treatment exposure. The registry-reported ALEX data does not provide enough detail to construct a complete exposure-adjusted safety analysis.

13. Censoring and Time-to-Event Interpretation

Several ALEX endpoints are time-to-event outcomes. For PFS, the event is the first documented disease progression or death, whichever occurs first. For OS, the event is death. For the quality-of-life endpoints, the event is deterioration as defined by the corresponding questionnaire endpoint.

Why censoring matters
Observed follow-up = event time, or information available until censoring

A participant who has not experienced the specified event during observed follow-up is not equivalent to a participant who will never experience the event. Survival methods use the available follow-up while accounting for censoring under their underlying assumptions.

This is why a time-to-event analysis cannot generally be replaced by simply calculating the percentage of participants who experienced an event at the end of the study. The timing of events and censoring both contribute information to the analysis.

14. Proportional-Hazards Considerations

A hazard ratio is a model-based relative measure of event hazard. Its most straightforward interpretation as a stable relative hazard over time relies on the proportional-hazards concept. The registry-reported ALEX data reports hazard ratios but does not provide a diagnostic assessment of the proportional-hazards assumption.

Interpretation caution: A single hazard ratio summarizes a time-to-event comparison, but it does not show the complete shape of the survival curves. Without the underlying event and censoring data, this page does not reconstruct Kaplan-Meier curves or independently test proportional hazards.

This distinction is important because two treatment groups can have complex time-varying differences even when a single hazard ratio is reported. The registry's reported estimate should therefore be understood as the result of the specified survival-analysis framework rather than as a complete description of every feature of the underlying event-time distributions.

15. Stratification and Covariate Adjustment

The primary investigator-assessed PFS analysis reports that the stratified hazard ratio and p-value were stratified for race (Asian vs Non-Asian) and CNS metastases at baseline by IRC. The registry also identifies covariate adjustment and stratified analysis among the concepts appearing in the analysis text.

FeatureReported role
RaceAsian vs Non-Asian
Baseline CNS metastasesStratification factor by IRC
Analysis populationITT
Primary time-to-event methodLog-rank test
Effect measureStratified hazard ratio

Stratification is not equivalent to creating separate treatment effects for each subgroup. A stratified hazard ratio produces an overall treatment comparison that accounts for the specified strata. Formal claims that treatment effects differ between strata require an appropriate interaction or heterogeneity analysis, which is not posted on ClinicalTrials.gov for the primary PFS result here.

16. Multiplicity and Multiple Endpoints

The registry reports two primary endpoints and numerous secondary endpoints. The ClinicalTrials.gov record identifies the hypothesis type for the reported analyses as superiority, but it does not provide a complete multiplicity-adjustment scheme for all posted secondary analyses.

Endpoint categoryReported statusStatistical implication
Investigator-assessed PFSPrimary; formal analysis postedPrimary time-to-event result
Percentage with PFS eventPrimary; no formal statistical analysis reportedResults are not accompanied here by a formal comparison
IRC-assessed PFSSecondary; formal analysis postedSecondary time-to-event result
CNS progressionSecondary; formal analysis postedSecondary time-to-event result
ORRSecondary; formal analysis postedSecondary binary result
OSSecondary; formal analysis postedSecondary time-to-event result
Quality-of-life deterioration endpointsSecondary; multiple analyses postedIndividual p-values should be interpreted in the context of multiple testing

A p-value such as 0.0285 for one secondary endpoint should not automatically be interpreted as having the same confirmatory status as the primary PFS analysis. Whether a secondary endpoint is confirmatory depends on the prespecified hierarchy, multiplicity strategy, and alpha allocation. Those details are not reported in the ClinicalTrials.gov record.

17. What the PFS Hazard Ratio Does — and Does Not — Mean

Effect size

The primary PFS hazard ratio of 0.47 means that the estimated instantaneous rate of progression or death in the alectinib group was approximately 47% of the comparator rate under the reported stratified survival analysis.

What it does not mean

It does not mean that exactly 47% of participants progressed, that 53% of participants were protected from progression, or that every individual participant experienced a 53% reduction in risk.

Precision

The 95% CI of 0.34–0.65 quantifies uncertainty around the estimated hazard ratio. It is not a prediction interval for individual patients and does not describe the variability of treatment effects across individual participants.

P-value

The p-value of <0.0001 measures evidence against the specified null hypothesis under the reported analysis. It does not measure the size, clinical importance, or patient-level benefit of the effect.

18. Comparison of the Main Reported Effects

OutcomeEffect estimate95% CIP-valueEndpoint type
Investigator-assessed PFSHR 0.470.34–0.65<0.0001Time-to-event
IRC-assessed PFSHR 0.500.36–0.70<0.0001Time-to-event
CNS progressionCause-specific HR 0.160.10–0.28<0.0001Time-to-event
Objective response rateDifference 7.40-1.71–16.500.0936Binary
Overall survivalHR 0.780.56–1.080.1320Time-to-event

The pattern demonstrates why the endpoint definition matters. The PFS estimates are hazard ratios for progression or death; CNS progression is reported as a cause-specific hazard ratio; ORR is a difference in response rates; and OS is a hazard ratio for death. These estimates cannot be placed on a common numerical scale and interpreted as though they were interchangeable.

19. Limitations

20. Why This Trial Matters Statistically

ALEX is a useful statistical teaching case because it combines randomized treatment comparison, time-to-event endpoints, stratified survival analysis, a binary response endpoint, CNS-specific progression, overall survival, and repeated patient-reported quality-of-life deterioration outcomes.

ConceptHow it appears in ALEX
RandomizationRandomized parallel comparison of alectinib and crizotinib
Intention-to-treat analysisThe reported efficacy analyses use the ITT population
Time-to-event analysisPFS, CNS progression, OS, and quality-of-life deterioration endpoints
Hazard ratioPrimary and secondary time-to-event effect measure
Confidence intervalQuantifies uncertainty around reported effect estimates
Log-rank testReported method for the time-to-event comparisons
Stratified analysisPrimary PFS analysis stratified by race and baseline CNS metastases
Cochran-Mantel-Haenszel testReported method for the binary ORR analysis
Binary endpointObjective response and the percentage with a PFS event
Multiple endpointsTwo primary endpoints plus numerous secondary analyses
Open-label designNo masking was used

The trial also demonstrates why statistical interpretation should remain connected to endpoint definitions. A hazard ratio for PFS, a cause-specific hazard ratio for CNS progression, a difference in response rates, and a hazard ratio for OS are not interchangeable measures. Each describes a different aspect of the randomized comparison.

21. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The reported primary investigator-assessed PFS analysis produced a stratified HR of 0.47 with a 95% CI of 0.34–0.65 and P <0.0001. The registry also reports secondary time-to-event and binary analyses with their corresponding estimates and uncertainty intervals.

Endpoint-specific interpretation

The different estimates describe different outcomes. PFS concerns progression or death, CNS progression concerns CNS disease progression as defined by the registry, ORR concerns complete or partial response, and OS concerns death.

A statistically strong result for one endpoint should not automatically be transferred to another endpoint. For example, the primary PFS estimate and the OS estimate answer different questions and have different reported confidence intervals and p-values.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Calculators

24. Sources

Continue through Clinical Biostats

Explore statistical tutorials, calculators, and additional clinical-trial analyses covering the methods used in randomized clinical research.

25. Record Summary

ALEX is a randomized phase 3 comparison of alectinib and crizotinib in treatment-naive participants with ALK-positive advanced non-small cell lung cancer. The registry reports a primary investigator-assessed PFS hazard ratio of 0.47 with a 95% CI of 0.34–0.65 and P <0.0001, using a stratified log-rank analysis of the ITT population. The ClinicalTrials.gov record also reports secondary analyses for IRC-assessed PFS, CNS progression, objective response, overall survival, and multiple quality-of-life deterioration endpoints.

The statistical story is therefore broader than a single p-value. The primary PFS result is a stratified time-to-event comparison; the ORR analysis uses a Cochran-Mantel-Haenszel method; secondary survival outcomes use hazard ratios and log-rank tests; and the quality-of-life endpoints demonstrate why endpoint definition, event direction, confidence intervals, and multiplicity all matter when interpreting a clinical-trial evidence set.

Clinical Biostats methodology: A trial-results page should distinguish the reported statistical evidence from the interpretation of that evidence. For ALEX, this means preserving the registry's endpoint definitions and reported estimates while avoiding unsupported median survival values, event counts, subgroup results, or additional statistical assumptions.