This page separates reported trial results from statistical interpretation. The numerical results presented here are limited to the ClinicalTrials.gov record. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
ALEX was a randomized, open-label, parallel phase 3 trial comparing alectinib with crizotinib in treatment-naive participants with ALK-positive advanced non-small cell lung cancer. The registry reports 303 enrolled participants, two treatment arms, two registered primary endpoints, and 28 posted outcome measures.
| Feature | ALEX |
|---|---|
| Trial name | ALEX |
| NCT identifier | NCT02075840 |
| Phase | Phase 3 |
| Condition | Non-Small Cell Lung Cancer |
| Population | Treatment-naive participants with anaplastic lymphoma kinase-positive advanced non-small cell lung cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 303.0 |
| Interventions | Alectinib; crizotinib |
| Primary endpoints | Progression-Free Survival by Investigator Assessment; Percentage of Participants With PFS Event by Investigator Assessment |
| Results posted | Yes |
| Outcome measures posted | 28 |
| Statistical analyses posted | 12 |
| Lead sponsor | Hoffmann-La Roche |
| Sponsor type | Industry |
2. Clinical Question
The central question was whether treatment with alectinib compared with crizotinib affected progression-free survival and other clinical outcomes in treatment-naive participants with ALK-positive advanced non-small cell lung cancer.
Population
Treatment-naive participants with anaplastic lymphoma kinase-positive advanced non-small cell lung cancer.
Intervention
Alectinib.
Comparator
Crizotinib.
Primary question
Does alectinib produce a different progression-free survival outcome than crizotinib in the randomized comparison?
3. Trial Design
Alectinib
- Drug intervention
- Randomized treatment assignment
- Evaluated against crizotinib
Crizotinib
- Drug intervention
- Randomized treatment assignment
- Comparator for the alectinib group
| Design characteristic | Registry description |
|---|---|
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Phase | 3 |
| Enrollment | 303.0 |
Randomization is the central design feature supporting the treatment comparison. Because the registry describes the trial as randomized and parallel, the principal efficacy comparison is naturally framed between the randomized alectinib and crizotinib groups rather than as a comparison of independently selected patient cohorts.
4. Trial Timeline
Trial start
The registry lists the study start date as 2014-08-19.
Primary completion
The registry lists the primary completion date as 2017-02-09.
Registry status
The trial is listed as completed, with results posted for 28 outcome measures and 12 statistical analyses.
5. Endpoints
The registry lists two primary endpoints: a time-to-event progression-free survival endpoint and a binary endpoint describing the percentage of participants with a progression-free survival event. The registry also reports statistical analyses for several secondary outcomes.
| Endpoint | Registry definition / time frame | Type |
|---|---|---|
| Progression-Free Survival (PFS) by Investigator Assessment | Randomization to first documented disease progression or death, whichever occurs first (assessed every 8 weeks up to 33 months) | Time-to-event |
| Percentage of Participants With PFS Event by Investigator Assessment | Randomization to first documented disease progression or death, whichever occurs first (assessed every 8 weeks up to 33 months) | Binary |
| PFS Independent Review Committee (IRC)-Assessed | Randomization to first documented disease progression or death, whichever occurs first (assessed every 8 weeks up to 33 months) | Time-to-event |
| Percentage of Participants With CNS Progression as Determined by IRC Using RECIST V1.1 Criteria | Randomization to CNS PD as first occurrence of disease progression (assessed every 8 weeks up to 33 months) | Time-to-event |
| Objective Response Rate (CR or PR) as Determined by Investigators According to RECIST V1.1 Criteria | Randomization to first documented disease progression or death, whichever occurs first (assessed every 8 weeks up to 33 months) | Binary |
| Overall Survival (OS) | From randomization until death (up to 10.5 years) | Time-to-event |
| Time to Deterioration by EORTC Quality of Life Questionnaire Core 30 (C30) | Baseline, every 4 weeks until disease progression (up to 33 months) | Time-to-event |
| Time to Deterioration by EORTC Quality of Life Questionnaire Lung Cancer Module 13 (LC13) | Baseline, every 4 weeks until disease progression (up to 33 months) | Time-to-event |
6. Primary Endpoint Results
Progression-Free Survival by Investigator Assessment
The registry reports a formal statistical analysis for the primary investigator-assessed PFS endpoint. The analysis population was the ITT population, defined in the registry as including all randomized participants in the study, with the number analyzed corresponding to participants evaluable for the outcome.
Stratified hazard ratio for progression or death
95% CI: 0.34–0.65 · P <0.0001
Log-rank test; superiority hypothesis.
| Feature | Reported analysis |
|---|---|
| Endpoint | Progression-Free Survival (PFS) by Investigator Assessment |
| Comparison | Alectinib vs Crizotinib |
| Population | ITT population |
| Method | Log-rank test |
| Effect measure | Hazard Ratio, stratified |
| Estimate | 0.47 |
| 95% CI | 0.34–0.65 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
| Stratification | Race (Asian vs Non-Asian) and CNS metastases at baseline by IRC |
The hazard ratio of 0.47 means that the estimated instantaneous rate of progression or death was approximately 47% of the corresponding rate in the comparator group under the fitted time-to-event comparison. Equivalently, 1 − 0.47 = 0.53, so the estimate corresponds to an approximately 53% lower estimated hazard for the alectinib group relative to crizotinib.
The hazard ratio does not mean that 53% of participants avoided progression, that every participant experienced exactly a 53% reduction in risk, or that the median PFS differed by 53%. It is a relative time-to-event measure rather than an absolute probability.
The 95% confidence interval of 0.34–0.65 describes statistical uncertainty around the estimated hazard ratio. It does not describe the range of effects that individual participants experienced.
The p-value of <0.0001 addresses the statistical evidence against the null hypothesis under the specified analysis. It is not a measure of the magnitude or clinical importance of the treatment effect. The magnitude is described by the hazard ratio and its confidence interval.
The registry identifies the analysis as stratified by race and baseline CNS metastases. Because the effect measure is a stratified hazard ratio, its interpretation depends on the specified stratified analysis rather than an unadjusted comparison.
Percentage of Participants With PFS Event by Investigator Assessment
The registry posts this second primary endpoint as a binary outcome, but the ClinicalTrials.gov record does not contain a corresponding formal statistical-analysis record for this endpoint. Therefore, a treatment-group difference, confidence interval, and p-value are not reported here.
The endpoint records whether each participant experienced a progression-free survival event during the specified assessment period. A binary comparison would ordinarily compare the proportions with events between randomized groups, with an appropriate confidence interval for the between-group difference or another prespecified effect measure.
For ALEX, the ClinicalTrials.gov record identifies the endpoint and indicates that results are posted, but it does not provide a formal statistical-analysis result for this primary binary endpoint.
7. Secondary Endpoint Results
Independent Review Committee-Assessed PFS
Stratified hazard ratio
95% CI: 0.36–0.70 · P <0.0001
Log-rank test; superiority hypothesis.
The IRC-assessed PFS hazard ratio of 0.50 corresponds to an estimated hazard approximately half that of the comparator group in the reported stratified analysis. The corresponding 95% CI, 0.36–0.70, quantifies uncertainty around that estimate.
The confidence interval does not establish the probability that the true hazard ratio lies inside the interval, nor does the p-value of <0.0001 quantify the size of the treatment effect. The p-value addresses evidence against the specified null hypothesis; the hazard ratio and confidence interval describe the estimated effect and its precision.
The analysis was stratified by race and baseline CNS metastases by IRC. This should be kept in mind when comparing the IRC result with the investigator-assessed primary PFS result.
CNS Progression
Cause-specific hazard ratio
95% CI: 0.10–0.28 · P <0.0001
IRC, RECIST v1.1 stratified analysis.
The cause-specific hazard ratio of 0.16 indicates a substantially lower estimated cause-specific hazard for the CNS progression endpoint in the alectinib group relative to crizotinib under the reported analysis.
The 95% CI of 0.10–0.28 indicates the statistical uncertainty around that estimate. The endpoint is a time-to-event outcome, so the hazard ratio should not be interpreted as a simple ratio of the final percentages of participants with CNS progression.
The registry describes the analysis as an IRC RECIST v1.1 stratified analysis by race and baseline CNS metastases. The p-value of <0.0001 provides evidence against the specified null hypothesis, but it does not measure the magnitude of the CNS effect.
Objective Response Rate
Difference in overall response rates
95% CI: -1.71–16.50 · P = 0.0936
Cochran-Mantel-Haenszel analysis; superiority hypothesis.
The reported effect measure is the difference in overall response rates, estimated as 7.40 percentage points for alectinib versus crizotinib under the reported analysis.
The 95% CI extends from -1.71 to 16.50. Thus, the interval includes zero, meaning that the reported confidence interval includes both a possible negative difference and positive differences of varying magnitude.
The p-value of 0.0936 is not an estimate of how large the response difference is. It quantifies the evidence against the specified null hypothesis under the Cochran-Mantel-Haenszel analysis. The effect estimate and confidence interval are needed to understand the magnitude and precision.
This endpoint is binary, so its interpretation differs from the hazard-ratio endpoints. It summarizes whether a participant achieved complete or partial response according to the specified RECIST v1.1 investigator assessment rather than the timing of progression or death.
Overall Survival
Hazard ratio for death
95% CI: 0.56–1.08 · P = 0.1320
Log-rank test; superiority hypothesis.
The reported OS hazard ratio of 0.78 corresponds to an estimated instantaneous hazard of death approximately 78% of that in the crizotinib group under the reported time-to-event analysis.
The 95% CI of 0.56–1.08 crosses 1.00. Consequently, the interval includes the possibility of no difference in hazard as well as a range of lower and higher relative hazards.
The p-value of 0.1320 is evidence against the prespecified null hypothesis only within the framework of this analysis; it is not a probability that the treatment has no effect and it is not an effect-size measure.
The ClinicalTrials.gov record does not report median OS or an absolute survival percentage, so those quantities are not inferred here.
Time to Deterioration: EORTC QLQ-C30 Fatigue
Hazard ratio
95% CI: 0.46–1.19 · P = 0.2079
Log-rank test; fatigue-stratified analysis.
The hazard ratio of 0.74 represents the estimated relative hazard of the fatigue deterioration event for alectinib compared with crizotinib under the reported analysis. The 95% CI of 0.46–1.19 includes 1.00, so the estimate is compatible with a range of relative effects.
The p-value of 0.2079 should not be interpreted as the probability that the treatment has no effect. It is a hypothesis-test quantity and does not replace the effect estimate or its confidence interval.
Time to Deterioration: EORTC QLQ-C30 Dyspnea
Hazard ratio
95% CI: 0.88–3.15 · P = 0.1137
Log-rank test.
A hazard ratio of 1.66 means that the estimated hazard of the specified deterioration event was higher in the alectinib group relative to crizotinib under this analysis. Because the 95% CI extends from 0.88 to 3.15, substantial uncertainty remains around the estimate.
The direction of a hazard ratio must always be interpreted in relation to the event definition. Here, the endpoint is time to deterioration for dyspnea, so a hazard ratio above 1 does not automatically mean that the overall treatment effect was unfavorable; it describes this particular deterioration endpoint.
Time to Deterioration: EORTC QLQ-LC13 Coughing
Hazard ratio
95% CI: 0.44–1.74 · P = 0.7042
Log-rank test.
The reported hazard ratio of 0.88 is close to 1.00, while the 95% CI of 0.44–1.74 is relatively broad and includes 1.00. The interval therefore allows for a range of possible relative effects.
The p-value of 0.7042 does not measure the magnitude of the observed estimate. It reflects the statistical evidence against the specified null hypothesis for this endpoint.
Time to Deterioration: EORTC QLQ-LC13 Dyspnea
Hazard ratio
95% CI: 1.05–2.92 · P = 0.0285
Log-rank test; dyspnea analysis.
The reported hazard ratio of 1.76 indicates a higher estimated hazard of the specified dyspnea deterioration event in the alectinib group relative to crizotinib in this analysis. The 95% CI of 1.05–2.92 lies above 1.00.
The endpoint is specifically a time-to-deterioration outcome for dyspnea. Its interpretation should therefore remain tied to the endpoint definition rather than being generalized to overall efficacy or overall quality of life.
The p-value of 0.0285 provides evidence against the null hypothesis for this individual analysis, but the ClinicalTrials.gov record contains multiple secondary time-to-deterioration analyses. Those multiple comparisons are relevant when interpreting isolated p-values.
Time to Deterioration: EORTC QLQ-LC13 Pain in Arm and Shoulder
Hazard ratio
95% CI: 0.79–2.61 · P = 0.2377
Log-rank test.
The estimated hazard ratio of 1.43 indicates a higher estimated hazard of the specified deterioration event in the alectinib group, but the 95% CI of 0.79–2.61 includes 1.00. The confidence interval therefore does not isolate a single direction of effect with high precision.
The p-value of 0.2377 should be interpreted as a hypothesis-test result, not as a probability that the observed treatment effect is absent.
Time to Deterioration: EORTC QLQ-LC13 Pain in Chest
Hazard ratio
95% CI: 0.24–1.10 · P = 0.0796
Log-rank test.
The hazard ratio of 0.51 is below 1.00, indicating a lower estimated hazard of the specified chest-pain deterioration event in the alectinib group. However, the 95% CI of 0.24–1.10 crosses 1.00, so the precision of the estimate does not exclude no difference.
The p-value of 0.0796 should be kept in the context of the confidence interval and the fact that this was one of several secondary time-to-deterioration analyses.
Time to Deterioration: EORTC QLQ-LC13 Composite Score
Hazard ratio
95% CI: 0.72–1.68 · P = 0.6435
Log-rank test.
The hazard ratio of 1.10 is close to 1.00, while the 95% CI of 0.72–1.68 spans both sides of 1.00. The estimate therefore provides limited precision about the direction or magnitude of the relative hazard.
The p-value of 0.6435 is a test result for this endpoint and should not be interpreted as a measure of effect size.
8. Statistical Methodology
Log-rank test
The registry identifies the log-rank test as the principal method for the reported time-to-event comparisons. The log-rank framework compares the observed and expected event patterns between randomized groups over follow-up, taking censoring into account.
The test uses information accumulated over event times rather than reducing the outcome to a single fixed-time proportion.
Hazard ratios
The principal time-to-event effect measure reported for ALEX is the hazard ratio. A hazard ratio compares instantaneous event rates between groups within the specified survival-analysis framework.
Equivalently, the estimate corresponds to an approximately 53% lower estimated hazard. This is not the same as a 53% reduction in the probability of ever experiencing the event.
Stratified analysis
The primary investigator-assessed PFS analysis reports a stratified hazard ratio and p-value. The registry specifies stratification by race (Asian vs Non-Asian) and CNS metastases at baseline by IRC.
Stratification allows the analysis to account for prespecified categorical factors when comparing treatment groups. It is particularly relevant when those factors are associated with prognosis or were incorporated into the randomized design or analysis framework.
Intention-to-treat analysis
The registry defines the ITT population as including all randomized participants in the study. The reported efficacy analyses use this population, with the number analyzed described as the total participants evaluable for the relevant outcome measure.
The ITT principle preserves the treatment assignment created by randomization. In statistical terms, this helps maintain the comparability established by the randomized design rather than redefining groups according to treatment received after randomization.
Cochran-Mantel-Haenszel analysis
The objective-response endpoint was analyzed using a Cochran-Mantel-Haenszel test, reported in the registry as "Mantel Haenszel." This method is designed for categorical comparisons while allowing stratification when appropriate.
The registry's reported effect measure for ORR was the difference in overall response rates, estimated as 7.40 with a 95% CI of -1.71–16.50.
9. Statistical Methods Explained
Why was a log-rank test used?
PFS, CNS progression, OS, and the quality-of-life deterioration endpoints are time-to-event outcomes. A participant's follow-up can end before the event occurs, producing censoring. The log-rank test is designed to compare event-time distributions between groups while using the available follow-up information.
What does a hazard ratio of 0.47 mean?
For the primary investigator-assessed PFS analysis, an HR of 0.47 means that the estimated instantaneous rate of progression or death was 47% of the comparator rate under the reported analysis. The complementary calculation, 1 − 0.47, gives an approximately 53% lower estimated hazard. This does not mean that 53% of patients were protected from progression.
Why is the confidence interval important?
The estimate alone does not communicate its statistical precision. The primary PFS estimate is 0.47, but the 95% CI is 0.34–0.65. The interval shows the range of parameter values compatible with the analysis under the stated confidence framework. A narrow interval generally conveys more precision than a wide interval; the width here is therefore part of the evidence, not an optional supplement to the point estimate.
Why doesn't the p-value measure effect size?
A p-value evaluates evidence against a null hypothesis under a specified statistical model and sampling framework. It does not say how large the treatment effect is. For ALEX, the primary PFS p-value is <0.0001, while the effect size is represented by the hazard ratio of 0.47 and its 95% CI of 0.34–0.65.
Why does stratification matter?
The primary PFS analysis was stratified by race and baseline CNS metastases by IRC. A stratified analysis allows the treatment comparison to account for these specified categories instead of treating all participants as though those factors were irrelevant to the comparison.
Why is the ORR analysis different from the PFS analysis?
ORR is a binary endpoint: a participant is classified according to whether a complete or partial response occurred. PFS is a time-to-event endpoint: the timing of progression or death matters, and participants can be censored. Consequently, the registry uses a Cochran-Mantel-Haenszel method for ORR and a log-rank framework for PFS.
Why should the secondary p-values be interpreted carefully?
The registry reports multiple secondary endpoints, including CNS progression, overall survival, and several quality-of-life deterioration outcomes. When many hypotheses are tested, the chance of observing at least one apparently unusual p-value can increase. The ClinicalTrials.gov record does not provide a multiplicity-adjustment scheme for these secondary analyses, so individual p-values should not automatically be treated as independent confirmatory findings.
10. Understanding the Primary PFS Result
The primary PFS result contains three different statistical pieces of information that answer different questions.
Effect size
HR 0.47 describes the estimated relative event hazard in the alectinib group compared with crizotinib.
Precision
95% CI 0.34–0.65 describes uncertainty around the hazard-ratio estimate.
Evidence against the null
P <0.0001 quantifies the statistical evidence under the reported hypothesis test.
Analysis structure
The estimate and p-value were derived from a stratified log-rank analysis of the ITT population.
Keeping these concepts separate prevents a common statistical error: treating the p-value as though it were the effect estimate. The p-value does not tell the reader whether the hazard ratio is 0.47, 0.80, or 0.20. That information comes from the effect estimate itself and its confidence interval.
11. Secondary Endpoint Pattern
| Secondary endpoint | Effect | 95% CI | P-value | Method |
|---|---|---|---|---|
| IRC-assessed PFS | HR 0.50 | 0.36–0.70 | <0.0001 | Log-rank |
| CNS progression | Cause-specific HR 0.16 | 0.10–0.28 | <0.0001 | Log-rank |
| Objective response rate | Difference 7.40 | -1.71–16.50 | 0.0936 | Cochran-Mantel-Haenszel |
| Overall survival | HR 0.78 | 0.56–1.08 | 0.1320 | Log-rank |
| C30 fatigue deterioration | HR 0.74 | 0.46–1.19 | 0.2079 | Log-rank |
| C30 dyspnea deterioration | HR 1.66 | 0.88–3.15 | 0.1137 | Log-rank |
| LC13 coughing deterioration | HR 0.88 | 0.44–1.74 | 0.7042 | Log-rank |
| LC13 dyspnea deterioration | HR 1.76 | 1.05–2.92 | 0.0285 | Log-rank |
| LC13 pain in arm and shoulder deterioration | HR 1.43 | 0.79–2.61 | 0.2377 | Log-rank |
| LC13 pain in chest deterioration | HR 0.51 | 0.24–1.10 | 0.0796 | Log-rank |
| LC13 composite-score deterioration | HR 1.10 | 0.72–1.68 | 0.6435 | Log-rank |
The secondary results illustrate why a clinical-trial analysis should not reduce every endpoint to a simple "significant" versus "not significant" classification. The estimates differ in direction and precision, and the endpoints themselves measure different phenomena. PFS and CNS progression are disease-control time-to-event outcomes; ORR is a binary response outcome; OS measures death; and the EORTC endpoints measure time to specified quality-of-life deterioration.
12. Safety
The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm. The numbers are presented as affected participants divided by participants at risk.
| Safety measure | Alectinib | Crizotinib |
|---|---|---|
| Serious adverse events | 70/152 | 48/151 |
Safety and efficacy also answer different statistical questions. The efficacy analyses preserve randomized treatment assignment through the ITT framework, whereas safety summaries are fundamentally connected to treatment exposure. The registry-reported ALEX data does not provide enough detail to construct a complete exposure-adjusted safety analysis.
13. Censoring and Time-to-Event Interpretation
Several ALEX endpoints are time-to-event outcomes. For PFS, the event is the first documented disease progression or death, whichever occurs first. For OS, the event is death. For the quality-of-life endpoints, the event is deterioration as defined by the corresponding questionnaire endpoint.
A participant who has not experienced the specified event during observed follow-up is not equivalent to a participant who will never experience the event. Survival methods use the available follow-up while accounting for censoring under their underlying assumptions.
This is why a time-to-event analysis cannot generally be replaced by simply calculating the percentage of participants who experienced an event at the end of the study. The timing of events and censoring both contribute information to the analysis.
14. Proportional-Hazards Considerations
A hazard ratio is a model-based relative measure of event hazard. Its most straightforward interpretation as a stable relative hazard over time relies on the proportional-hazards concept. The registry-reported ALEX data reports hazard ratios but does not provide a diagnostic assessment of the proportional-hazards assumption.
This distinction is important because two treatment groups can have complex time-varying differences even when a single hazard ratio is reported. The registry's reported estimate should therefore be understood as the result of the specified survival-analysis framework rather than as a complete description of every feature of the underlying event-time distributions.
15. Stratification and Covariate Adjustment
The primary investigator-assessed PFS analysis reports that the stratified hazard ratio and p-value were stratified for race (Asian vs Non-Asian) and CNS metastases at baseline by IRC. The registry also identifies covariate adjustment and stratified analysis among the concepts appearing in the analysis text.
| Feature | Reported role |
|---|---|
| Race | Asian vs Non-Asian |
| Baseline CNS metastases | Stratification factor by IRC |
| Analysis population | ITT |
| Primary time-to-event method | Log-rank test |
| Effect measure | Stratified hazard ratio |
Stratification is not equivalent to creating separate treatment effects for each subgroup. A stratified hazard ratio produces an overall treatment comparison that accounts for the specified strata. Formal claims that treatment effects differ between strata require an appropriate interaction or heterogeneity analysis, which is not posted on ClinicalTrials.gov for the primary PFS result here.
16. Multiplicity and Multiple Endpoints
The registry reports two primary endpoints and numerous secondary endpoints. The ClinicalTrials.gov record identifies the hypothesis type for the reported analyses as superiority, but it does not provide a complete multiplicity-adjustment scheme for all posted secondary analyses.
| Endpoint category | Reported status | Statistical implication |
|---|---|---|
| Investigator-assessed PFS | Primary; formal analysis posted | Primary time-to-event result |
| Percentage with PFS event | Primary; no formal statistical analysis reported | Results are not accompanied here by a formal comparison |
| IRC-assessed PFS | Secondary; formal analysis posted | Secondary time-to-event result |
| CNS progression | Secondary; formal analysis posted | Secondary time-to-event result |
| ORR | Secondary; formal analysis posted | Secondary binary result |
| OS | Secondary; formal analysis posted | Secondary time-to-event result |
| Quality-of-life deterioration endpoints | Secondary; multiple analyses posted | Individual p-values should be interpreted in the context of multiple testing |
A p-value such as 0.0285 for one secondary endpoint should not automatically be interpreted as having the same confirmatory status as the primary PFS analysis. Whether a secondary endpoint is confirmatory depends on the prespecified hierarchy, multiplicity strategy, and alpha allocation. Those details are not reported in the ClinicalTrials.gov record.
17. What the PFS Hazard Ratio Does — and Does Not — Mean
The primary PFS hazard ratio of 0.47 means that the estimated instantaneous rate of progression or death in the alectinib group was approximately 47% of the comparator rate under the reported stratified survival analysis.
It does not mean that exactly 47% of participants progressed, that 53% of participants were protected from progression, or that every individual participant experienced a 53% reduction in risk.
The 95% CI of 0.34–0.65 quantifies uncertainty around the estimated hazard ratio. It is not a prediction interval for individual patients and does not describe the variability of treatment effects across individual participants.
The p-value of <0.0001 measures evidence against the specified null hypothesis under the reported analysis. It does not measure the size, clinical importance, or patient-level benefit of the effect.
18. Comparison of the Main Reported Effects
| Outcome | Effect estimate | 95% CI | P-value | Endpoint type |
|---|---|---|---|---|
| Investigator-assessed PFS | HR 0.47 | 0.34–0.65 | <0.0001 | Time-to-event |
| IRC-assessed PFS | HR 0.50 | 0.36–0.70 | <0.0001 | Time-to-event |
| CNS progression | Cause-specific HR 0.16 | 0.10–0.28 | <0.0001 | Time-to-event |
| Objective response rate | Difference 7.40 | -1.71–16.50 | 0.0936 | Binary |
| Overall survival | HR 0.78 | 0.56–1.08 | 0.1320 | Time-to-event |
The pattern demonstrates why the endpoint definition matters. The PFS estimates are hazard ratios for progression or death; CNS progression is reported as a cause-specific hazard ratio; ORR is a difference in response rates; and OS is a hazard ratio for death. These estimates cannot be placed on a common numerical scale and interpreted as though they were interchangeable.
19. Limitations
- Incomplete binary primary analysis: the ClinicalTrials.gov record identifies the percentage of participants with a PFS event as a primary endpoint and indicates that results are posted, but it does not supply a formal statistical-analysis result for that endpoint.
- Registry-level detail: the ClinicalTrials.gov record does not provide median PFS, median OS, absolute event rates, Kaplan-Meier coordinates, or detailed baseline characteristics. Those quantities are therefore not inferred.
- Multiple secondary endpoints: the registry reports numerous secondary analyses. Individual secondary p-values require context concerning multiplicity and endpoint hierarchy.
- Hazard-ratio assumptions: hazard ratios summarize time-to-event comparisons but do not provide the full survival distribution and may be sensitive to departures from proportional hazards.
- Censoring: time-to-event estimates depend on how follow-up and censoring are handled. The ClinicalTrials.gov record does not provide participant-level censoring information.
- Analysis population: the registry defines the ITT population as all randomized participants, while the number analyzed for an outcome corresponds to participants evaluable for that outcome. The ClinicalTrials.gov record does not provide the detailed participant disposition needed to reconcile every outcome-specific denominator.
- Open-label design: masking is listed as none. This is particularly relevant when interpreting investigator-assessed and patient-reported outcomes because knowledge of treatment assignment can potentially influence assessment or reporting.
- Safety detail: serious adverse-event counts are reported by arm, but the data does not provide a formal comparative statistical analysis for those counts.
20. Why This Trial Matters Statistically
ALEX is a useful statistical teaching case because it combines randomized treatment comparison, time-to-event endpoints, stratified survival analysis, a binary response endpoint, CNS-specific progression, overall survival, and repeated patient-reported quality-of-life deterioration outcomes.
| Concept | How it appears in ALEX |
|---|---|
| Randomization | Randomized parallel comparison of alectinib and crizotinib |
| Intention-to-treat analysis | The reported efficacy analyses use the ITT population |
| Time-to-event analysis | PFS, CNS progression, OS, and quality-of-life deterioration endpoints |
| Hazard ratio | Primary and secondary time-to-event effect measure |
| Confidence interval | Quantifies uncertainty around reported effect estimates |
| Log-rank test | Reported method for the time-to-event comparisons |
| Stratified analysis | Primary PFS analysis stratified by race and baseline CNS metastases |
| Cochran-Mantel-Haenszel test | Reported method for the binary ORR analysis |
| Binary endpoint | Objective response and the percentage with a PFS event |
| Multiple endpoints | Two primary endpoints plus numerous secondary analyses |
| Open-label design | No masking was used |
The trial also demonstrates why statistical interpretation should remain connected to endpoint definitions. A hazard ratio for PFS, a cause-specific hazard ratio for CNS progression, a difference in response rates, and a hazard ratio for OS are not interchangeable measures. Each describes a different aspect of the randomized comparison.
21. Statistical Interpretation vs Clinical Interpretation
Statistical interpretation
The reported primary investigator-assessed PFS analysis produced a stratified HR of 0.47 with a 95% CI of 0.34–0.65 and P <0.0001. The registry also reports secondary time-to-event and binary analyses with their corresponding estimates and uncertainty intervals.
Endpoint-specific interpretation
The different estimates describe different outcomes. PFS concerns progression or death, CNS progression concerns CNS disease progression as defined by the registry, ORR concerns complete or partial response, and OS concerns death.
A statistically strong result for one endpoint should not automatically be transferred to another endpoint. For example, the primary PFS estimate and the OS estimate answer different questions and have different reported confidence intervals and p-values.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Calculators
24. Sources
- ClinicalTrials.gov: ALEX — NCT02075840.
- PubMed: PMID 39661917.
- PubMed: PMID 35275991.
- PubMed: PMID 32418886.
- PubMed: PMID 30215676.
- PubMed: PMID 29656744.
Continue through Clinical Biostats
Explore statistical tutorials, calculators, and additional clinical-trial analyses covering the methods used in randomized clinical research.
25. Record Summary
ALEX is a randomized phase 3 comparison of alectinib and crizotinib in treatment-naive participants with ALK-positive advanced non-small cell lung cancer. The registry reports a primary investigator-assessed PFS hazard ratio of 0.47 with a 95% CI of 0.34–0.65 and P <0.0001, using a stratified log-rank analysis of the ITT population. The ClinicalTrials.gov record also reports secondary analyses for IRC-assessed PFS, CNS progression, objective response, overall survival, and multiple quality-of-life deterioration endpoints.
The statistical story is therefore broader than a single p-value. The primary PFS result is a stratified time-to-event comparison; the ORR analysis uses a Cochran-Mantel-Haenszel method; secondary survival outcomes use hazard ratios and log-rank tests; and the quality-of-life endpoints demonstrate why endpoint definition, event direction, confidence intervals, and multiplicity all matter when interpreting a clinical-trial evidence set.