This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
PROFILE 1014 was a randomized, parallel-group, open-label phase 3 treatment trial enrolling 343 participants with non-squamous lung cancer. The statistical record compares crizotinib with chemotherapy and contains a formal time-to-event analysis for the primary endpoint of progression-free survival based on independent radiologic review.
| Feature | PROFILE 1014 |
|---|---|
| Trial | PROFILE 1014 |
| NCT identifier | NCT01154140 |
| Phase | Phase 3 |
| Condition | Non Squamous Lung Cancer |
| Population described in the brief title | Patients with ALK-positive non-squamous cancer of the lung |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 343 |
| Arms | 2 |
| Primary endpoint | Progression-Free Survival (PFS) Based on IRR |
| Primary endpoint type | Time-to-event |
| Hypothesis type | Superiority |
| Results posted | Yes |
| Statistical analyses posted | 34 |
| Lead sponsor | Pfizer |
2. Clinical Question
The central comparison was whether crizotinib differed from standard chemotherapy consisting of pemetrexed plus cisplatin or carboplatin in patients with ALK-positive non-squamous lung cancer.
Statistically, the question is expressed as a comparison of time-to-event distributions rather than simply comparing the proportion of participants who experienced progression by a fixed calendar date. That distinction is important because participants can be followed for different lengths of time and may be censored before an event occurs.
3. Trial Design
PROFILE 1014 used randomized allocation and a parallel design. The registry classifies masking as none, so the trial was not described as masked. The trial had two treatment arms and a planned treatment purpose.
Crizotinib
- Intervention classified as treatment (drug).
- Compared with the chemotherapy arm for the primary and reported secondary analyses.
Chemotherapy
- Pemetrexed plus cisplatin or carboplatin.
- Used as the comparator for the randomized analysis.
The analysis population for the formal efficacy comparisons was the FA population, defined in the registry as all participants who were randomized with study treatment assignment designated according to the initial randomization. The patient-reported outcome analyses used a narrower PRO evaluable population when specified.
Trial timeline
Trial start
The registered study start date was 2011-01-13.
Primary completion
The registered primary completion date was 2013-11-30.
Overall survival endpoint window
The registered overall-survival analysis was defined from randomization to death or the last date known alive for participants not known to have died, up to 72 months.
4. Endpoints
Primary endpoint: Progression-Free Survival Based on IRR
Registered endpoint: Progression-Free Survival (PFS) Based on IRR.
Time frame: Randomization to objective progression, death or last tumor assessment without progression before any additional anti-cancer treatment.
Endpoint type: Time-to-event.
Secondary endpoints represented in the statistical record
| Endpoint | Type | Method | Effect measure |
|---|---|---|---|
| Overall Survival (OS) | Time-to-event | Log-rank; stratified Cox model in analysis text | Hazard ratio |
| Objective Response Rate (ORR), assessed by IRR | Binary | Pearson chi-square | Risk difference |
| Disease control at Week 12, based on IRR | Binary | Pearson chi-square | Risk difference |
| Time to Progression (TTP), based on IRR | Time-to-event | Log-rank; Cox model in analysis text | Hazard ratio |
| Time to Intracranial Progression (IC-TTP), based on IRR | Time-to-event | Log-rank; Cox model in analysis text | Hazard ratio |
| Time to Extracranial Progression (EC-TTP), based on IRR | Time-to-event | Log-rank; Cox model in analysis text | Hazard ratio |
| Time to Deterioration in chest pain, dyspnea or cough | Time-to-event | Log-rank; Cox model in analysis text | Hazard ratio |
| EORTC-QLQ-C30 functioning and global quality of life | Continuous | Repeated-measures mixed-effects model | Mean difference |
| EORTC-QLQ-C30 symptoms | Continuous | Repeated-measures mixed-effects model | Mean or median difference |
| EORTC-QLQ-LC13 lung-cancer symptoms | Continuous | Repeated-measures mixed-effects model | Mean difference |
| EQ-5D VAS general health status | Continuous | Repeated-measures mixed-effects model | Mean difference |
5. Statistical Methodology
The statistical record contains three principal methodological families: survival analysis using log-rank testing and Cox proportional-hazards models, categorical analysis using Pearson's chi-square test, and longitudinal analysis using repeated-measures mixed-effects models.
Time-to-event analysis
The primary PFS comparison used a log-rank test with a hazard ratio as the effect measure. The same general survival framework was used for OS, TTP, IC-TTP, EC-TTP and time to deterioration.
Categorical analysis
ORR and disease control at Week 12 were analysed with Pearson's chi-square test, with the treatment difference expressed as a risk difference.
Repeated measures
Patient-reported functioning, quality-of-life, symptom and health-status outcomes were analysed using repeated-measures mixed-effects models.
Stratification
The OS analysis text specifies a Cox proportional-hazards model stratified by ECOG performance status, race group and brain metastases.
Primary PFS analysis
The primary endpoint was analysed with a log-rank test in the FA population. The effect measure was the hazard ratio comparing crizotinib with chemotherapy. The reported hypothesis type was superiority and the confidence interval was two-sided at the 95% level.
The registry's primary analysis reports HR = 0.454, with 95% CI 0.346 to 0.596 and p < 0.0001.
Multiplicity and the ORR threshold
The ClinicalTrials.gov record gives an explicit multiplicity-related rule for objective response rate. If the PFS endpoint was significant, ORR was to be considered significant if the two-sided Pearson chi-square p-value was ≤ 0.0494. The ORR confidence interval was calculated based on the normal distribution.
This is an important distinction from treating every secondary p-value as though it had an independent, unrestricted 0.05 threshold. The registry record specifically documents a decision rule for ORR tied to the primary PFS result.
Longitudinal patient-reported outcomes
The patient-reported outcome analyses used a repeated-measures mixed-effects model. For the EORTC-QLQ-C30 analyses, the model included an intercept, treatment, treatment-by-time interaction, and baseline EORTC-QLQ-C30 subscale baseline score. Intercept and time from first dose were included as random effects.
The PRO evaluable population was narrower than the FA population: it included participants from the FA population who completed a baseline and at least one postbaseline PRO assessment before crossover to crizotinib or the end of the randomized treatment period, according to the endpoint-specific registry wording.
6. Results: Primary Endpoint
Progression-Free Survival Based on IRR
The primary endpoint was PFS based on independent radiologic review. The comparison used the FA population and a log-rank test. The reported effect measure was a hazard ratio.
Primary PFS result
95% CI 0.346 to 0.596 · two-sided · p < 0.0001
Crizotinib vs chemotherapy; superiority hypothesis.
A hazard ratio of 0.454 means that the estimated instantaneous hazard under the proportional-hazards interpretation was 45.4% of the comparator hazard, corresponding to a 54.6% lower estimated hazard. This is a relative measure of the event hazard, not a statement that 54.6% of participants avoided progression and not a direct estimate of a difference in median PFS.
What the estimate means: HR 0.454 summarizes the relative event rate between the two randomized groups over the analysis period under the hazard-ratio framework. Values below 1 indicate a lower estimated hazard in the crizotinib group relative to chemotherapy.
What it does not mean: The HR is not a percentage of participants who benefited, is not a relative risk, and should not be read as saying that participants receiving crizotinib had exactly 54.6% fewer progression events.
Precision: The 95% confidence interval, 0.346 to 0.596, quantifies uncertainty around the estimated hazard ratio. Its bounds remain below 1, so the interval is consistent with a lower hazard under the model used for this comparison.
The p-value: p < 0.0001 addresses the statistical evidence against the null comparison under the specified testing procedure. It does not measure the magnitude or clinical importance of the effect; the HR and its confidence interval provide that effect-size information.
Important survival-analysis caution: Interpretation of a single hazard ratio as a common relative effect over time depends on the proportional-hazards framework. The registry describes Cox proportional-hazards modelling for several secondary time-to-event endpoints and explicitly states the proportional-hazards assumption for those analyses.
Censoring: PFS is a time-to-event endpoint, so participants without an observed progression or death contribute information up to their censoring or last qualifying assessment rather than simply being classified as having no event.
7. Secondary Results: Survival and Response
Overall Survival
OS was defined from randomization to death or the last date known alive for participants not known to have died, up to 72 months. The FA population was used. The analysis used a log-rank test, while the analysis text specifies a Cox proportional-hazards model stratified by ECOG performance status, race group and brain metastases.
Overall survival
95% CI 0.548 to 1.053 · p = 0.0489
Crizotinib vs chemotherapy; two-sided superiority analysis.
What the estimate means: HR 0.760 corresponds to an estimated hazard that is 76.0% of the chemotherapy hazard under the proportional-hazards interpretation, or a 24.0% lower estimated hazard.
What it does not mean: It does not mean that mortality was reduced by exactly 24.0% for every participant, nor does it provide a median-survival difference.
Precision: The 95% CI of 0.548 to 1.053 is wider than the primary PFS interval and extends above 1. The interval therefore includes the possibility of no hazard difference under the model, despite the reported p-value of 0.0489.
Stratification matters: The analysis text specifies stratification by ECOG performance status, race group and brain metastases. Stratification allows the Cox analysis to account for these prespecified factors when estimating the treatment hazard ratio.
P-value versus effect size: The p-value describes evidence under the specified null hypothesis and testing procedure; it does not quantify the size or precision of the observed HR. The HR and confidence interval should be considered together.
Objective Response Rate
ORR was the percentage of participants with an objective response assessed by independent radiologic review. The FA population was analysed with Pearson's chi-square test. The treatment effect was reported as a difference in percentage, normalized here as a risk difference.
| Measure | Estimate | 95% CI | P-value |
|---|---|---|---|
| ORR difference | 29.4 | 19.5 to 39.3 | <0.0001 |
The registry-reported analysis notes specify that if PFS was significant, ORR was considered significant when the two-sided Pearson chi-square p-value was ≤ 0.0494. The reported ORR p-value is <0.0001, satisfying that documented criterion.
Disease Control at Week 12
| Measure | Estimate | 95% CI | P-value |
|---|---|---|---|
| Difference in percentage | 10.067 | 0.8 to 19.4 | 0.0381 |
The Week 12 disease-control endpoint was binary and was analysed using Pearson's chi-square test. The confidence interval for the difference in percentage was based on the normal distribution.
Time to Progression
| Endpoint | HR | 95% CI | P-value |
|---|---|---|---|
| Time to Progression, based on IRR | 0.441 | 0.335 to 0.582 | <0.0001 |
| Time to Intracranial Progression, based on IRR | 0.595 | 0.338 to 1.048 | 0.0347 |
| Time to Extracranial Progression, based on IRR | 0.387 | 0.286 to 0.524 | <0.0001 |
| Time to Deterioration in chest pain, dyspnea or cough | 0.591 | 0.452 to 0.773 | 0.0002 |
These endpoints share the time-to-event framework but answer different clinical questions. TTP focuses on objective progression, whereas IC-TTP and EC-TTP isolate intracranial and extracranial progression. Time to deterioration shifts the event definition toward patient-reported deterioration in specified symptoms.
8. Secondary Results: Patient-Reported Outcomes
The patient-reported outcome results are particularly useful statistically because the trial combines repeated observations with treatment-by-time dynamics. The registry reports net mean or median differences from repeated-measures mixed-effects models rather than reducing each participant to a single final observation.
EORTC-QLQ-C30 functioning and global quality of life
| QLQ-C30 domain | Difference | 95% CI | P-value |
|---|---|---|---|
| Global QoL | 13.8303 | 10.74 to 16.92 | <0.0001 |
| Cognitive functioning | 3.3532 | 0.60 to 6.11 | 0.0170 |
| Emotional functioning | 7.5165 | 4.57 to 10.46 | <0.0001 |
| Physical functioning | 10.4035 | 7.48 to 13.32 | <0.0001 |
| Role functioning | 15.5513 | 11.29 to 19.81 | <0.0001 |
| Social functioning | 8.7641 | 4.69 to 12.84 | <0.0001 |
These are differences on the questionnaire scales, not hazard ratios. The positive estimates indicate that the reported net difference was positive for the named functioning or global-QoL measure. Statistical interpretation should therefore preserve the scale and direction rather than translate these estimates into percentages.
EORTC-QLQ-C30 symptoms
| QLQ-C30 symptom domain | Difference | 95% CI | P-value |
|---|---|---|---|
| Appetite loss | -13.4976 | -18.03 to -8.97 | <0.0001 |
| Constipation | -4.4336 | -9.00 to 0.13 | 0.0570 |
| Diarrhea | 12.4906 | 8.98 to 16.00 | <0.0001 |
| Dysponea | -13.4622 | -17.20 to -9.73 | <0.0001 |
| Fatigue | -14.9987 | -18.52 to -11.48 | <0.0001 |
| Financial difficulties | -0.8186 | -4.56 to 2.92 | 0.6681 |
| Insomnia | -10.0430 | -14.22 to -5.87 | <0.0001 |
| Nausea and vomiting | -3.4446 | -6.84 to -0.05 | 0.0468 |
| Pain | -9.9277 | -13.23 to -6.62 | <0.0001 |
One useful feature of these results is that the confidence intervals make the distinction between statistical evidence and precision visible. For example, the constipation estimate is -4.4336 with a 95% CI from -9.00 to 0.13 and p = 0.0570, whereas the appetite-loss estimate is -13.4976 with a 95% CI from -18.03 to -8.97 and p < 0.0001.
EORTC-QLQ-LC13 lung-cancer symptoms
| QLQ-LC13 domain | Difference | 95% CI | P-value |
|---|---|---|---|
| Alopecia | -4.8149 | -8.52 to -1.11 | 0.0108 |
| Coughing | -8.3926 | -12.06 to -4.72 | <0.0001 |
| Dysphagia | 0.6651 | -1.78 to 3.11 | 0.5938 |
| Dyspnoea | -9.0080 | -11.96 to -6.06 | <0.0001 |
| Haemoptysis | -0.8828 | -1.82 to 0.06 | 0.0656 |
| Pain in arm or shoulder | -6.0475 | -9.22 to -2.88 | 0.0002 |
| Pain in chest | -8.0959 | -11.35 to -4.84 | <0.0001 |
| Pain in other parts | -6.7717 | -10.24 to -3.31 | 0.0001 |
| Peripheral neuropathy | 3.3521 | 0.11 to 6.59 | 0.0427 |
| Sore mouth | -2.1521 | -5.00 to 0.69 | 0.1382 |
EQ-5D visual analog scale
| Endpoint | Difference | 95% CI | P-value |
|---|---|---|---|
| General Health Status, EQ-5D VAS | 3.9908 | 0.81 to 7.17 | 0.0139 |
The EQ-5D VAS analysis also used a repeated-measures mixed-effects model, with an intercept, treatment, treatment-by-time interaction and baseline EQ-5D VAS score; intercept and time from first dose were included as random effects.
9. Safety
The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk. The available record does not provide a detailed breakdown by individual adverse-event term, severity category or treatment relatedness, so the safety summary is restricted to the reported serious-adverse-event counts.
| Arm | Participants with serious adverse events | At risk |
|---|---|---|
| Crizotinib | 71 | 171 |
| Chemotherapy | 49 | 169 |
The denominators here are not the overall randomized enrollment of 343; they are the arm-specific numbers posted on ClinicalTrials.gov for the serious-adverse-event summary. Accordingly, these figures should not be silently converted into percentages or interpreted as though the denominators represented every randomized participant.
10. Statistical Methods Explained
Why was a log-rank test used for PFS?
PFS records the timing of an event rather than simply whether an event occurred. The log-rank test compares the survival experience between randomized groups across follow-up while accommodating censored observations. In PROFILE 1014, the primary PFS analysis used this test and expressed the treatment contrast with a hazard ratio.
What does an HR of 0.454 mean?
It means the estimated hazard in the crizotinib group was 0.454 times the chemotherapy hazard under the proportional-hazards interpretation. Equivalently, 1 - 0.454 = 0.546, so the estimated hazard was 54.6% lower. That is a statement about the hazard ratio, not a direct statement about the proportion of participants who avoided progression.
Why is the confidence interval important?
A point estimate alone does not show how precisely the treatment effect was estimated. For PFS, the 95% CI was 0.346 to 0.596. The interval gives a range of values compatible with the statistical analysis at the stated confidence level and is therefore more informative than the p-value alone for understanding uncertainty around the estimated effect.
Why was Pearson's chi-square test used for ORR?
ORR is a binary endpoint: participants are classified according to whether they had an objective response. Pearson's chi-square test provides a comparison of the categorical outcome distributions between the two treatment groups. The registry expresses the treatment effect as a difference in percentage rather than as a hazard ratio.
Why use a mixed-effects model for quality-of-life outcomes?
The QLQ-C30, QLQ-LC13 and EQ-5D analyses involve repeated observations over time. A repeated-measures mixed-effects model can represent the correlation among observations from the same participant instead of treating every measurement as statistically independent. In PROFILE 1014, the documented models included treatment, treatment-by-time interaction and baseline score, with random effects for intercept and time from first dose.
Why does the OS analysis mention stratification?
The OS analysis text specifies that the Cox proportional-hazards model was stratified by ECOG performance status, race group and brain metastases. Stratification allows the survival comparison to be made while accounting for the specified stratification factors rather than treating them as irrelevant to the survival-risk structure.
Why should the many secondary p-values be interpreted cautiously?
The trial record contains multiple secondary endpoints and multiple patient-reported domains. Each additional statistical test creates opportunities for apparently small p-values even when the underlying null hypotheses are true. The ClinicalTrials.gov record explicitly documents a significance threshold of ≤ 0.0494 for ORR conditional on significant PFS, demonstrating that at least part of the multiplicity structure was handled through a prespecified testing rule. The individual PRO p-values should therefore be interpreted in the context of their multiplicity rather than treated as a collection of independent confirmatory findings.
11. Interpreting Hazard Ratios Across the Trial
PROFILE 1014 provides several examples of why hazard ratios should be interpreted endpoint by endpoint rather than as interchangeable measures.
| Endpoint | HR | 95% CI | P-value |
|---|---|---|---|
| Primary PFS | 0.454 | 0.346 to 0.596 | <0.0001 |
| Overall Survival | 0.760 | 0.548 to 1.053 | 0.0489 |
| Time to Progression | 0.441 | 0.335 to 0.582 | <0.0001 |
| Intracranial Progression | 0.595 | 0.338 to 1.048 | 0.0347 |
| Extracranial Progression | 0.387 | 0.286 to 0.524 | <0.0001 |
| Time to Deterioration | 0.591 | 0.452 to 0.773 | 0.0002 |
The estimates are not identical because the endpoints define different events and therefore measure different aspects of follow-up. PFS includes objective progression or death, TTP focuses on progression, IC-TTP focuses on intracranial progression, EC-TTP focuses on extracranial progression, and time to deterioration concerns specified symptom deterioration. A hazard ratio must therefore always be read together with its endpoint definition.
The OS result also illustrates why the confidence interval deserves equal attention to the point estimate. An HR of 0.760 is below 1, but its 95% CI of 0.548 to 1.053 extends above 1. The statistical interpretation is therefore more nuanced than simply classifying the estimate as "positive" or "negative."
12. Limitations
- Registry-based scope: This analysis is limited to the ClinicalTrials.gov record and does not incorporate additional numerical results from the linked publications.
- Multiple secondary analyses: The record contains numerous secondary endpoints and patient-reported domains. Their p-values should not be interpreted as though every test were an isolated confirmatory hypothesis.
- Time-to-event assumptions: Hazard-ratio interpretation relies on the survival-analysis framework, including the proportional-hazards assumption where Cox proportional-hazards modelling is specified.
- Censoring: Time-to-event analyses depend on how follow-up and censoring are handled. The registry-reported endpoint definitions specify last qualifying assessments for participants without observed events.
- Different analysis populations: Efficacy analyses use the FA population, while PRO analyses use a PRO evaluable population requiring baseline and postbaseline assessments before crossover or the end of the randomized treatment period.
- Safety detail: Only arm-level serious-adverse-event counts are reported. No detailed adverse-event table is included in the available trial data.
- No unreported results are inferred: Median survival, subgroup estimates, additional baseline characteristics and other quantities not present in the ClinicalTrials.gov record is not reconstructed from outside sources.
13. Why This Trial Matters Statistically
PROFILE 1014 is a useful statistical teaching case because several major clinical-trial methods appear in one randomized comparison. The primary endpoint is a classic time-to-event outcome, requiring methods that account for event timing and censoring rather than a simple two-proportion comparison.
The trial also shows the complementary roles of different effect measures. Hazard ratios summarize relative event hazards; risk differences summarize differences in binary response outcomes; and mean or median differences describe changes on repeated patient-reported scales. These quantities cannot be substituted for one another without changing the question being answered.
The OS analysis provides a clear example of stratified Cox modelling. The registry specifies stratification by ECOG performance status, race group and brain metastases, illustrating how clinically relevant factors can be incorporated into a survival-analysis framework without turning them into the primary treatment comparison.
The PRO analyses add another important layer. Rather than analysing only one postbaseline measurement, the trial used repeated-measures mixed-effects models containing treatment, treatment-by-time interaction and baseline score, with random effects for participant-level intercept and time structure. This is a fundamentally different statistical problem from the primary PFS analysis, even though both compare the same randomized treatment groups.
Finally, the explicit ORR significance rule demonstrates why a clinical-trial statistical analysis should examine the testing hierarchy rather than simply count p-values below 0.05. The ClinicalTrials.gov record states that ORR was to be considered significant after significant PFS when its two-sided p-value was ≤ 0.0494. That conditional rule is part of the interpretation of the ORR result.
14. A Practical Reading of the Statistical Story
A useful way to read PROFILE 1014 is to move from design to estimand to analysis and then to uncertainty:
1. Start with randomization
The treatment groups were created through randomized allocation, making the between-group comparison the central structure of the trial.
2. Identify the endpoint
The primary endpoint was PFS based on independent radiologic review, a time-to-event outcome.
3. Match method to endpoint
The primary comparison used the log-rank test and a hazard ratio; binary secondary endpoints used chi-square testing; longitudinal PRO endpoints used mixed-effects models.
4. Read the interval
The confidence interval shows the precision of the estimated treatment contrast and can reveal uncertainty that is not obvious from the point estimate alone.
This sequence prevents a common statistical mistake: beginning with the p-value and working backward. The more informative approach begins by asking what was measured, how the endpoint behaves statistically, which population was analysed, what effect measure was selected, and what uncertainty surrounds the estimate.
15. Related Tutorials
Learn more about the methods used in this trial:
16. Related Calculators
17. Sources
- ClinicalTrials.gov: PROFILE 1014, NCT01154140.
- PubMed: PMID 30822515.
- PubMed: PMID 30652510.
- PubMed: PMID 29768118.
- PubMed: PMID 28373069.
- PubMed: PMID 27022118.
Continue with the Clinical Biostats methods library
Explore the statistical methods that connect randomized clinical-trial design with survival analysis, categorical outcomes, longitudinal models and clinical-trial calculations.
18. Record Summary
PROFILE 1014 provides a compact example of how a randomized phase 3 trial can require several distinct statistical frameworks. Its primary PFS endpoint was analysed as a time-to-event outcome using a log-rank test and hazard ratio, with HR 0.454, 95% CI 0.346 to 0.596 and p < 0.0001. Secondary analyses extended the survival framework to OS, TTP, intracranial and extracranial progression, and symptom deterioration; categorical endpoints used Pearson's chi-square test; and patient-reported outcomes used repeated-measures mixed-effects models.
The central statistical lesson is that the treatment effect cannot be represented by one number alone. The hazard ratio describes relative event hazards, risk differences describe binary outcome contrasts, and mean or median differences describe changes on patient-reported scales. Confidence intervals provide the corresponding uncertainty, while the endpoint definition, analysis population, censoring framework, stratification and multiplicity rules determine how each estimate should be interpreted.