← Clinical Trials
Non-Squamous Lung Cancer Phase 3 Completed NCT01154140

PROFILE 1014: Complete Statistical Analysis of Crizotinib in ALK-Positive Non-Squamous Lung Cancer

An independent statistical review of the randomized phase 3 PROFILE 1014 trial comparing crizotinib with standard chemotherapy consisting of pemetrexed plus cisplatin or carboplatin in patients with ALK-positive non-squamous lung cancer.

Trial start: 2011-01-13  ·  Primary completion: 2013-11-30  ·  Enrollment: 343
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

PROFILE 1014 was a randomized, parallel-group, open-label phase 3 treatment trial enrolling 343 participants with non-squamous lung cancer. The statistical record compares crizotinib with chemotherapy and contains a formal time-to-event analysis for the primary endpoint of progression-free survival based on independent radiologic review.

343
Enrolled
2 treatment arms
0.454
PFS HR
95% CI 0.346–0.596
<0.0001
PFS P-value
Two-sided log-rank
24
Outcome Measures
34 statistical analyses
FeaturePROFILE 1014
TrialPROFILE 1014
NCT identifierNCT01154140
PhasePhase 3
ConditionNon Squamous Lung Cancer
Population described in the brief titlePatients with ALK-positive non-squamous cancer of the lung
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment343
Arms2
Primary endpointProgression-Free Survival (PFS) Based on IRR
Primary endpoint typeTime-to-event
Hypothesis typeSuperiority
Results postedYes
Statistical analyses posted34
Lead sponsorPfizer

2. Clinical Question

The central comparison was whether crizotinib differed from standard chemotherapy consisting of pemetrexed plus cisplatin or carboplatin in patients with ALK-positive non-squamous lung cancer.

Population
Patients with ALK-positive non-squamous cancer of the lung, as described in the trial's brief title.
Intervention
Crizotinib.
Comparator
Standard chemotherapy with pemetrexed plus cisplatin or carboplatin.
Primary question
Compare progression-free survival from randomization using a superiority hypothesis.

Statistically, the question is expressed as a comparison of time-to-event distributions rather than simply comparing the proportion of participants who experienced progression by a fixed calendar date. That distinction is important because participants can be followed for different lengths of time and may be censored before an event occurs.

3. Trial Design

PROFILE 1014 used randomized allocation and a parallel design. The registry classifies masking as none, so the trial was not described as masked. The trial had two treatment arms and a planned treatment purpose.

Step 1
Randomization
Step 2
Two arms
Step 3
Follow-up
Step 4
Event assessment
Step 5
Statistical comparison

Crizotinib

Treatment arm
  • Intervention classified as treatment (drug).
  • Compared with the chemotherapy arm for the primary and reported secondary analyses.

Chemotherapy

Comparator arm
  • Pemetrexed plus cisplatin or carboplatin.
  • Used as the comparator for the randomized analysis.

The analysis population for the formal efficacy comparisons was the FA population, defined in the registry as all participants who were randomized with study treatment assignment designated according to the initial randomization. The patient-reported outcome analyses used a narrower PRO evaluable population when specified.

Trial timeline

2011-01-13

Trial start

The registered study start date was 2011-01-13.

2013-11-30

Primary completion

The registered primary completion date was 2013-11-30.

Up to 72 months

Overall survival endpoint window

The registered overall-survival analysis was defined from randomization to death or the last date known alive for participants not known to have died, up to 72 months.

4. Endpoints

Primary endpoint: Progression-Free Survival Based on IRR

Registered endpoint: Progression-Free Survival (PFS) Based on IRR.

Time frame: Randomization to objective progression, death or last tumor assessment without progression before any additional anti-cancer treatment.

Endpoint type: Time-to-event.

Secondary endpoints represented in the statistical record

EndpointTypeMethodEffect measure
Overall Survival (OS)Time-to-eventLog-rank; stratified Cox model in analysis textHazard ratio
Objective Response Rate (ORR), assessed by IRRBinaryPearson chi-squareRisk difference
Disease control at Week 12, based on IRRBinaryPearson chi-squareRisk difference
Time to Progression (TTP), based on IRRTime-to-eventLog-rank; Cox model in analysis textHazard ratio
Time to Intracranial Progression (IC-TTP), based on IRRTime-to-eventLog-rank; Cox model in analysis textHazard ratio
Time to Extracranial Progression (EC-TTP), based on IRRTime-to-eventLog-rank; Cox model in analysis textHazard ratio
Time to Deterioration in chest pain, dyspnea or coughTime-to-eventLog-rank; Cox model in analysis textHazard ratio
EORTC-QLQ-C30 functioning and global quality of lifeContinuousRepeated-measures mixed-effects modelMean difference
EORTC-QLQ-C30 symptomsContinuousRepeated-measures mixed-effects modelMean or median difference
EORTC-QLQ-LC13 lung-cancer symptomsContinuousRepeated-measures mixed-effects modelMean difference
EQ-5D VAS general health statusContinuousRepeated-measures mixed-effects modelMean difference

5. Statistical Methodology

The statistical record contains three principal methodological families: survival analysis using log-rank testing and Cox proportional-hazards models, categorical analysis using Pearson's chi-square test, and longitudinal analysis using repeated-measures mixed-effects models.

Time-to-event analysis

The primary PFS comparison used a log-rank test with a hazard ratio as the effect measure. The same general survival framework was used for OS, TTP, IC-TTP, EC-TTP and time to deterioration.

Categorical analysis

ORR and disease control at Week 12 were analysed with Pearson's chi-square test, with the treatment difference expressed as a risk difference.

Repeated measures

Patient-reported functioning, quality-of-life, symptom and health-status outcomes were analysed using repeated-measures mixed-effects models.

Stratification

The OS analysis text specifies a Cox proportional-hazards model stratified by ECOG performance status, race group and brain metastases.

Primary PFS analysis

The primary endpoint was analysed with a log-rank test in the FA population. The effect measure was the hazard ratio comparing crizotinib with chemotherapy. The reported hypothesis type was superiority and the confidence interval was two-sided at the 95% level.

Primary statistical comparison
Log-rank test  →  hazard ratio  →  95% two-sided confidence interval

The registry's primary analysis reports HR = 0.454, with 95% CI 0.346 to 0.596 and p < 0.0001.

Multiplicity and the ORR threshold

The ClinicalTrials.gov record gives an explicit multiplicity-related rule for objective response rate. If the PFS endpoint was significant, ORR was to be considered significant if the two-sided Pearson chi-square p-value was ≤ 0.0494. The ORR confidence interval was calculated based on the normal distribution.

This is an important distinction from treating every secondary p-value as though it had an independent, unrestricted 0.05 threshold. The registry record specifically documents a decision rule for ORR tied to the primary PFS result.

Longitudinal patient-reported outcomes

The patient-reported outcome analyses used a repeated-measures mixed-effects model. For the EORTC-QLQ-C30 analyses, the model included an intercept, treatment, treatment-by-time interaction, and baseline EORTC-QLQ-C30 subscale baseline score. Intercept and time from first dose were included as random effects.

The PRO evaluable population was narrower than the FA population: it included participants from the FA population who completed a baseline and at least one postbaseline PRO assessment before crossover to crizotinib or the end of the randomized treatment period, according to the endpoint-specific registry wording.

6. Results: Primary Endpoint

Progression-Free Survival Based on IRR

The primary endpoint was PFS based on independent radiologic review. The comparison used the FA population and a log-rank test. The reported effect measure was a hazard ratio.

Primary PFS result

HR 0.454

95% CI 0.346 to 0.596  ·  two-sided  ·  p < 0.0001

Crizotinib vs chemotherapy; superiority hypothesis.

A hazard ratio of 0.454 means that the estimated instantaneous hazard under the proportional-hazards interpretation was 45.4% of the comparator hazard, corresponding to a 54.6% lower estimated hazard. This is a relative measure of the event hazard, not a statement that 54.6% of participants avoided progression and not a direct estimate of a difference in median PFS.

Clinical Biostats interpretation

What the estimate means: HR 0.454 summarizes the relative event rate between the two randomized groups over the analysis period under the hazard-ratio framework. Values below 1 indicate a lower estimated hazard in the crizotinib group relative to chemotherapy.

What it does not mean: The HR is not a percentage of participants who benefited, is not a relative risk, and should not be read as saying that participants receiving crizotinib had exactly 54.6% fewer progression events.

Precision: The 95% confidence interval, 0.346 to 0.596, quantifies uncertainty around the estimated hazard ratio. Its bounds remain below 1, so the interval is consistent with a lower hazard under the model used for this comparison.

The p-value: p < 0.0001 addresses the statistical evidence against the null comparison under the specified testing procedure. It does not measure the magnitude or clinical importance of the effect; the HR and its confidence interval provide that effect-size information.

Important survival-analysis caution: Interpretation of a single hazard ratio as a common relative effect over time depends on the proportional-hazards framework. The registry describes Cox proportional-hazards modelling for several secondary time-to-event endpoints and explicitly states the proportional-hazards assumption for those analyses.

Censoring: PFS is a time-to-event endpoint, so participants without an observed progression or death contribute information up to their censoring or last qualifying assessment rather than simply being classified as having no event.

7. Secondary Results: Survival and Response

Overall Survival

OS was defined from randomization to death or the last date known alive for participants not known to have died, up to 72 months. The FA population was used. The analysis used a log-rank test, while the analysis text specifies a Cox proportional-hazards model stratified by ECOG performance status, race group and brain metastases.

Overall survival

HR 0.760

95% CI 0.548 to 1.053  ·  p = 0.0489

Crizotinib vs chemotherapy; two-sided superiority analysis.

Clinical Biostats interpretation

What the estimate means: HR 0.760 corresponds to an estimated hazard that is 76.0% of the chemotherapy hazard under the proportional-hazards interpretation, or a 24.0% lower estimated hazard.

What it does not mean: It does not mean that mortality was reduced by exactly 24.0% for every participant, nor does it provide a median-survival difference.

Precision: The 95% CI of 0.548 to 1.053 is wider than the primary PFS interval and extends above 1. The interval therefore includes the possibility of no hazard difference under the model, despite the reported p-value of 0.0489.

Stratification matters: The analysis text specifies stratification by ECOG performance status, race group and brain metastases. Stratification allows the Cox analysis to account for these prespecified factors when estimating the treatment hazard ratio.

P-value versus effect size: The p-value describes evidence under the specified null hypothesis and testing procedure; it does not quantify the size or precision of the observed HR. The HR and confidence interval should be considered together.

Objective Response Rate

ORR was the percentage of participants with an objective response assessed by independent radiologic review. The FA population was analysed with Pearson's chi-square test. The treatment effect was reported as a difference in percentage, normalized here as a risk difference.

MeasureEstimate95% CIP-value
ORR difference29.419.5 to 39.3<0.0001

The registry-reported analysis notes specify that if PFS was significant, ORR was considered significant when the two-sided Pearson chi-square p-value was ≤ 0.0494. The reported ORR p-value is <0.0001, satisfying that documented criterion.

Disease Control at Week 12

MeasureEstimate95% CIP-value
Difference in percentage10.0670.8 to 19.40.0381

The Week 12 disease-control endpoint was binary and was analysed using Pearson's chi-square test. The confidence interval for the difference in percentage was based on the normal distribution.

Time to Progression

EndpointHR95% CIP-value
Time to Progression, based on IRR0.4410.335 to 0.582<0.0001
Time to Intracranial Progression, based on IRR0.5950.338 to 1.0480.0347
Time to Extracranial Progression, based on IRR0.3870.286 to 0.524<0.0001
Time to Deterioration in chest pain, dyspnea or cough0.5910.452 to 0.7730.0002

These endpoints share the time-to-event framework but answer different clinical questions. TTP focuses on objective progression, whereas IC-TTP and EC-TTP isolate intracranial and extracranial progression. Time to deterioration shifts the event definition toward patient-reported deterioration in specified symptoms.

8. Secondary Results: Patient-Reported Outcomes

The patient-reported outcome results are particularly useful statistically because the trial combines repeated observations with treatment-by-time dynamics. The registry reports net mean or median differences from repeated-measures mixed-effects models rather than reducing each participant to a single final observation.

EORTC-QLQ-C30 functioning and global quality of life

QLQ-C30 domainDifference95% CIP-value
Global QoL13.830310.74 to 16.92<0.0001
Cognitive functioning3.35320.60 to 6.110.0170
Emotional functioning7.51654.57 to 10.46<0.0001
Physical functioning10.40357.48 to 13.32<0.0001
Role functioning15.551311.29 to 19.81<0.0001
Social functioning8.76414.69 to 12.84<0.0001

These are differences on the questionnaire scales, not hazard ratios. The positive estimates indicate that the reported net difference was positive for the named functioning or global-QoL measure. Statistical interpretation should therefore preserve the scale and direction rather than translate these estimates into percentages.

EORTC-QLQ-C30 symptoms

QLQ-C30 symptom domainDifference95% CIP-value
Appetite loss-13.4976-18.03 to -8.97<0.0001
Constipation-4.4336-9.00 to 0.130.0570
Diarrhea12.49068.98 to 16.00<0.0001
Dysponea-13.4622-17.20 to -9.73<0.0001
Fatigue-14.9987-18.52 to -11.48<0.0001
Financial difficulties-0.8186-4.56 to 2.920.6681
Insomnia-10.0430-14.22 to -5.87<0.0001
Nausea and vomiting-3.4446-6.84 to -0.050.0468
Pain-9.9277-13.23 to -6.62<0.0001

One useful feature of these results is that the confidence intervals make the distinction between statistical evidence and precision visible. For example, the constipation estimate is -4.4336 with a 95% CI from -9.00 to 0.13 and p = 0.0570, whereas the appetite-loss estimate is -13.4976 with a 95% CI from -18.03 to -8.97 and p < 0.0001.

EORTC-QLQ-LC13 lung-cancer symptoms

QLQ-LC13 domainDifference95% CIP-value
Alopecia-4.8149-8.52 to -1.110.0108
Coughing-8.3926-12.06 to -4.72<0.0001
Dysphagia0.6651-1.78 to 3.110.5938
Dyspnoea-9.0080-11.96 to -6.06<0.0001
Haemoptysis-0.8828-1.82 to 0.060.0656
Pain in arm or shoulder-6.0475-9.22 to -2.880.0002
Pain in chest-8.0959-11.35 to -4.84<0.0001
Pain in other parts-6.7717-10.24 to -3.310.0001
Peripheral neuropathy3.35210.11 to 6.590.0427
Sore mouth-2.1521-5.00 to 0.690.1382

EQ-5D visual analog scale

EndpointDifference95% CIP-value
General Health Status, EQ-5D VAS3.99080.81 to 7.170.0139

The EQ-5D VAS analysis also used a repeated-measures mixed-effects model, with an intercept, treatment, treatment-by-time interaction and baseline EQ-5D VAS score; intercept and time from first dose were included as random effects.

9. Safety

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk. The available record does not provide a detailed breakdown by individual adverse-event term, severity category or treatment relatedness, so the safety summary is restricted to the reported serious-adverse-event counts.

ArmParticipants with serious adverse eventsAt risk
Crizotinib71171
Chemotherapy49169

The denominators here are not the overall randomized enrollment of 343; they are the arm-specific numbers posted on ClinicalTrials.gov for the serious-adverse-event summary. Accordingly, these figures should not be silently converted into percentages or interpreted as though the denominators represented every randomized participant.

10. Statistical Methods Explained

Why was a log-rank test used for PFS?

PFS records the timing of an event rather than simply whether an event occurred. The log-rank test compares the survival experience between randomized groups across follow-up while accommodating censored observations. In PROFILE 1014, the primary PFS analysis used this test and expressed the treatment contrast with a hazard ratio.

What does an HR of 0.454 mean?

It means the estimated hazard in the crizotinib group was 0.454 times the chemotherapy hazard under the proportional-hazards interpretation. Equivalently, 1 - 0.454 = 0.546, so the estimated hazard was 54.6% lower. That is a statement about the hazard ratio, not a direct statement about the proportion of participants who avoided progression.

Why is the confidence interval important?

A point estimate alone does not show how precisely the treatment effect was estimated. For PFS, the 95% CI was 0.346 to 0.596. The interval gives a range of values compatible with the statistical analysis at the stated confidence level and is therefore more informative than the p-value alone for understanding uncertainty around the estimated effect.

Why was Pearson's chi-square test used for ORR?

ORR is a binary endpoint: participants are classified according to whether they had an objective response. Pearson's chi-square test provides a comparison of the categorical outcome distributions between the two treatment groups. The registry expresses the treatment effect as a difference in percentage rather than as a hazard ratio.

Why use a mixed-effects model for quality-of-life outcomes?

The QLQ-C30, QLQ-LC13 and EQ-5D analyses involve repeated observations over time. A repeated-measures mixed-effects model can represent the correlation among observations from the same participant instead of treating every measurement as statistically independent. In PROFILE 1014, the documented models included treatment, treatment-by-time interaction and baseline score, with random effects for intercept and time from first dose.

Why does the OS analysis mention stratification?

The OS analysis text specifies that the Cox proportional-hazards model was stratified by ECOG performance status, race group and brain metastases. Stratification allows the survival comparison to be made while accounting for the specified stratification factors rather than treating them as irrelevant to the survival-risk structure.

Why should the many secondary p-values be interpreted cautiously?

The trial record contains multiple secondary endpoints and multiple patient-reported domains. Each additional statistical test creates opportunities for apparently small p-values even when the underlying null hypotheses are true. The ClinicalTrials.gov record explicitly documents a significance threshold of ≤ 0.0494 for ORR conditional on significant PFS, demonstrating that at least part of the multiplicity structure was handled through a prespecified testing rule. The individual PRO p-values should therefore be interpreted in the context of their multiplicity rather than treated as a collection of independent confirmatory findings.

11. Interpreting Hazard Ratios Across the Trial

PROFILE 1014 provides several examples of why hazard ratios should be interpreted endpoint by endpoint rather than as interchangeable measures.

EndpointHR95% CIP-value
Primary PFS0.4540.346 to 0.596<0.0001
Overall Survival0.7600.548 to 1.0530.0489
Time to Progression0.4410.335 to 0.582<0.0001
Intracranial Progression0.5950.338 to 1.0480.0347
Extracranial Progression0.3870.286 to 0.524<0.0001
Time to Deterioration0.5910.452 to 0.7730.0002

The estimates are not identical because the endpoints define different events and therefore measure different aspects of follow-up. PFS includes objective progression or death, TTP focuses on progression, IC-TTP focuses on intracranial progression, EC-TTP focuses on extracranial progression, and time to deterioration concerns specified symptom deterioration. A hazard ratio must therefore always be read together with its endpoint definition.

The OS result also illustrates why the confidence interval deserves equal attention to the point estimate. An HR of 0.760 is below 1, but its 95% CI of 0.548 to 1.053 extends above 1. The statistical interpretation is therefore more nuanced than simply classifying the estimate as "positive" or "negative."

12. Limitations

13. Why This Trial Matters Statistically

PROFILE 1014 is a useful statistical teaching case because several major clinical-trial methods appear in one randomized comparison. The primary endpoint is a classic time-to-event outcome, requiring methods that account for event timing and censoring rather than a simple two-proportion comparison.

The trial also shows the complementary roles of different effect measures. Hazard ratios summarize relative event hazards; risk differences summarize differences in binary response outcomes; and mean or median differences describe changes on repeated patient-reported scales. These quantities cannot be substituted for one another without changing the question being answered.

The OS analysis provides a clear example of stratified Cox modelling. The registry specifies stratification by ECOG performance status, race group and brain metastases, illustrating how clinically relevant factors can be incorporated into a survival-analysis framework without turning them into the primary treatment comparison.

The PRO analyses add another important layer. Rather than analysing only one postbaseline measurement, the trial used repeated-measures mixed-effects models containing treatment, treatment-by-time interaction and baseline score, with random effects for participant-level intercept and time structure. This is a fundamentally different statistical problem from the primary PFS analysis, even though both compare the same randomized treatment groups.

Finally, the explicit ORR significance rule demonstrates why a clinical-trial statistical analysis should examine the testing hierarchy rather than simply count p-values below 0.05. The ClinicalTrials.gov record states that ORR was to be considered significant after significant PFS when its two-sided p-value was ≤ 0.0494. That conditional rule is part of the interpretation of the ORR result.

14. A Practical Reading of the Statistical Story

A useful way to read PROFILE 1014 is to move from design to estimand to analysis and then to uncertainty:

1. Start with randomization

The treatment groups were created through randomized allocation, making the between-group comparison the central structure of the trial.

2. Identify the endpoint

The primary endpoint was PFS based on independent radiologic review, a time-to-event outcome.

3. Match method to endpoint

The primary comparison used the log-rank test and a hazard ratio; binary secondary endpoints used chi-square testing; longitudinal PRO endpoints used mixed-effects models.

4. Read the interval

The confidence interval shows the precision of the estimated treatment contrast and can reveal uncertainty that is not obvious from the point estimate alone.

This sequence prevents a common statistical mistake: beginning with the p-value and working backward. The more informative approach begins by asking what was measured, how the endpoint behaves statistically, which population was analysed, what effect measure was selected, and what uncertainty surrounds the estimate.

15. Related Tutorials

Learn more about the methods used in this trial:

16. Related Calculators

17. Sources

Continue with the Clinical Biostats methods library

Explore the statistical methods that connect randomized clinical-trial design with survival analysis, categorical outcomes, longitudinal models and clinical-trial calculations.

18. Record Summary

PROFILE 1014 provides a compact example of how a randomized phase 3 trial can require several distinct statistical frameworks. Its primary PFS endpoint was analysed as a time-to-event outcome using a log-rank test and hazard ratio, with HR 0.454, 95% CI 0.346 to 0.596 and p < 0.0001. Secondary analyses extended the survival framework to OS, TTP, intracranial and extracranial progression, and symptom deterioration; categorical endpoints used Pearson's chi-square test; and patient-reported outcomes used repeated-measures mixed-effects models.

The central statistical lesson is that the treatment effect cannot be represented by one number alone. The hazard ratio describes relative event hazards, risk differences describe binary outcome contrasts, and mean or median differences describe changes on patient-reported scales. Confidence intervals provide the corresponding uncertainty, while the endpoint definition, analysis population, censoring framework, stratification and multiplicity rules determine how each estimate should be interpreted.

Clinical Biostats methodology: A trial-results page should distinguish the reported numerical evidence from statistical interpretation. PROFILE 1014 demonstrates why endpoint definition, analysis population, effect measure, confidence interval and testing framework should be read together rather than treating the p-value as the complete statistical result.