This page separates reported trial results from statistical interpretation. The numerical results on this page are restricted to the ClinicalTrials.gov record for NCT00805194.
1. Trial at a Glance
LUME-Lung 1 was a completed, randomized, double-blind, parallel phase 3 trial evaluating BIBF 1120 plus docetaxel versus placebo plus docetaxel in second-line non-small cell lung cancer. The trial enrolled 1314 participants and registered one primary time-to-event endpoint: progression-free survival as assessed by central independent review.
| Feature | LUME-Lung 1 |
|---|---|
| Phase | Phase 3 |
| Condition | Carcinoma, Non-Small-Cell Lung |
| Clinical setting | Second-line non-small cell lung cancer |
| Design | Randomized, double-blind, parallel |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 1314 |
| Arms | 2 |
| Primary endpoint | Progression Free Survival (PFS) as Assessed by Central Independent Review |
| Primary endpoint type | Time-to-event |
| Primary analysis population | Randomised Set |
| Lead sponsor | Boehringer Ingelheim |
| Sponsor type | Industry |
| Trial dates | Start: 2008-12-03 · Primary completion: 2010-11-02 |
2. Clinical Question
The central statistical question was whether adding BIBF 1120 to docetaxel changed progression-free survival compared with placebo plus docetaxel in participants with second-line non-small cell lung cancer.
Population
Participants with carcinoma, non-small-cell lung, in the second-line treatment setting.
Intervention
BIBF 1120 plus docetaxel.
Comparator
Placebo plus docetaxel.
Primary question
Does BIBF 1120 plus docetaxel improve centrally assessed progression-free survival relative to placebo plus docetaxel?
3. Trial Design
BIBF 1120 plus docetaxel
- BIBF 1120
- Docetaxel
- Randomized assignment
Placebo plus docetaxel
- Placebo
- Docetaxel
- Randomized assignment
The combination of randomization and double masking is important statistically. Randomization establishes the treatment comparison, while masking can reduce the potential for knowledge of treatment assignment to influence assessment or other trial conduct. The registry does not provide additional allocation-ratio information in the ClinicalTrials.gov record.
4. Endpoints
The registry lists one primary endpoint and multiple secondary outcomes. The primary endpoint is a time-to-event outcome; several secondary outcomes use either Cox proportional-hazards models, logistic regression, or ANOVA.
| Endpoint | Registry definition / time frame | Analysis |
|---|---|---|
| Primary: Progression Free Survival (PFS) as Assessed by Central Independent Review | From randomisation until cut-off date 2 November 2010, when 713 PFS events were observed. PFS is defined as the duration of time from date of randomisation to date of progression or death, whichever occurs earlier, according to modified RECIST v1.0. | Stratified Cox proportional-hazards model |
| Overall Survival (Key Secondary Endpoint) | From randomisation until cut-off date 15 February 2013, approximately 48 months or 1151 deaths among all patients. The overall alpha level followed a Lan-DeMets spending function with O'Brien-Fleming shape parameter to preserve an overall 2-sided alpha level of 0.05. HR below 1 favors nintedanib | Cox proportional-hazards model with hierarchical testing |
| Follow-up Analysis of PFS by Central Independent Review | From randomisation until cut-off date 15 February 2013. | Cox proportional-hazards model |
| Follow-up Analysis of PFS by Investigator | From randomisation until cut-off date 15 February 2013. | Cox proportional-hazards model |
| Objective Tumour Response | From randomisation until cut-off date 15 February 2013. | Logistic regression |
| Disease Control | From randomisation until cut-off date 15 February 2013. | Logistic regression |
| Clinical Improvement | From randomisation until cut-off date 15 February 2013. | Cox proportional-hazards model |
| Quality of Life (QoL) | From randomisation until cut-off date 15 February 2013. | Cox proportional-hazards model |
| Change From Baseline in Tumour Size | From randomisation until cut-off date 15 February 2013; outcome unit is percentage of change in tumor size in mm. | ANOVA |
5. Statistical Methodology
Primary time-to-event analysis
The primary PFS analysis used a Cox proportional-hazards model in the Randomised Set. The analysis was stratified by baseline ECOG performance status, tumour histology, brain metastases at baseline, and prior treatment with bevacizumab.
The Cox model estimates a relative hazard through the regression coefficient. In this trial, the reported effect measure was the hazard ratio comparing BIBF 1120 plus docetaxel with placebo plus docetaxel.
Hazard ratio
The registry reports the hazard ratio as the primary effect measure for PFS and for several secondary time-to-event analyses. An HR below 1 favors the BIBF 1120-containing treatment according to the registry analysis notes.
The hazard ratio is a relative time-to-event measure. It should not be interpreted as a percentage of participants who progressed, nor as an absolute difference in PFS probability.
Logistic regression
Objective tumour response and disease control were analyzed using logistic regression, with odds ratios as the effect measure. The registry identifies the central independent review and investigator assessment separately for these endpoints.
ANOVA
Change from baseline in tumour size was analyzed using ANOVA. The registry reports two analyses of this endpoint, one based on central independent review and one based on investigator assessment.
Stratified analysis
The primary PFS Cox model incorporated stratification by baseline ECOG performance status (0 vs 1), tumour histology (squamous vs non-squamous), brain metastases at baseline (yes vs no), and prior treatment with bevacizumab (yes vs no). This is important because the hazard ratio was not simply an unstratified comparison of the two randomized groups.
6. Results: Primary Endpoint
Progression-Free Survival as Assessed by Central Independent Review
Primary PFS hazard ratio
95% CI: 0.68–0.92 · P = 0.0019
713 PFS events were observed at the 2 November 2010 cutoff.
| Primary endpoint | Result |
|---|---|
| Analysis population | Randomised Set |
| Comparison | Nintedanib Plus Docetaxel vs Placebo Plus Docetaxel |
| Method | Stratified Cox proportional-hazards model |
| Hazard ratio | 0.79 |
| 95% CI | 0.68–0.92 |
| P-value | 0.0019 |
| Hypothesis type | Superiority |
The reported HR of 0.79 means that, under the fitted stratified Cox model, the estimated instantaneous rate of progression or death was 21% lower in the BIBF 1120-containing group than in the placebo-containing group. The 21% figure is a direct interpretation of 1 − 0.79; it is a relative hazard interpretation, not an absolute reduction in the probability of progression or death.
The HR does not mean that 21% fewer participants progressed, that every participant experienced a 21% reduction in risk, or that PFS duration increased by 21%.
The 95% CI of 0.68–0.92 describes statistical uncertainty around the estimated hazard ratio under the model and analysis framework. It does not describe the range of treatment effects experienced by individual patients.
The p-value of 0.0019 addresses the statistical evidence against the null hypothesis under the prespecified superiority analysis. A p-value is not a measure of effect size and does not tell us how clinically important the estimated HR is.
Because the estimate comes from a Cox proportional-hazards model, its interpretation depends on the model framework, censoring, and the proportional-hazards assumption. The registry does not provide enough information in the ClinicalTrials.gov record to independently assess the proportional-hazards assumption.
What the primary PFS result establishes statistically
The registry reports a formal superiority analysis with an HR of 0.79, a two-sided 95% CI of 0.68–0.92, and a p-value of 0.0019. The confidence interval lies below 1, consistent with the direction of the registry's stated interpretation that an HR below 1 favors nintedanib.
Importantly, the analysis is based on the Randomised Set. This anchors the treatment comparison to randomized assignment rather than selectively comparing only participants who remained on treatment or who had complete follow-up.
7. Secondary Endpoint Results
Overall Survival: Hierarchical Key Secondary Analysis
Overall survival was analyzed from randomisation through the 15 February 2013 cutoff, described in the registry as approximately 48 months or 1151 deaths among all patients. The registry reports three Cox-model estimates corresponding to a fixed sequence of hypotheses.
| Hierarchical hypothesis | Hazard ratio | 95% CI | P-value |
|---|---|---|---|
| Patients with adenocarcinoma and <9 months since start of first-line therapy | 0.75 | 0.60–0.92 | 0.0073 |
| Patients with adenocarcinoma | 0.83 | 0.70–0.99 | 0.0359 |
| All patients | 0.94 | 0.83–1.05 | 0.2720 |
For the first hierarchical OS comparison, the HR of 0.75 corresponds to a 25% lower estimated instantaneous hazard of death under the Cox model. Its 95% CI of 0.60–0.92 quantifies uncertainty around that estimate, while the p-value of 0.0073 addresses evidence against the corresponding null hypothesis.
For the adenocarcinoma population in the second step, the HR of 0.83 corresponds to a 17% lower estimated instantaneous hazard of death. The 95% CI is 0.70–0.99, with p = 0.0359.
For all patients, the reported HR is 0.94, with a 95% CI of 0.83–1.05 and p = 0.2720. This estimate is closer to 1 than the two preceding estimates. The confidence interval spans 1, so the ClinicalTrials.gov record does not provide evidence of a statistically significant superiority comparison for this final hypothesis under the reported two-sided test.
These three results should not be read as three independent hypothesis tests. The fixed-sequence design makes the order of testing part of the statistical interpretation.
Follow-up Analysis of Progression-Free Survival: Central Independent Review
Follow-up PFS hazard ratio
95% CI: 0.75–0.96 · P = 0.0070
Cutoff: 15 February 2013
The follow-up HR of 0.85 corresponds to a 15% lower estimated instantaneous hazard of progression or death under the fitted Cox model. The 95% CI of 0.75–0.96 indicates uncertainty around the estimate, while p = 0.0070 describes the evidence against the superiority null hypothesis under the reported analysis.
This is a follow-up PFS analysis, not the original primary PFS analysis. It therefore should not be silently substituted for the primary endpoint result of HR 0.79.
Follow-up Analysis of PFS by Investigator Assessment
Investigator-assessed PFS hazard ratio
95% CI: 0.73–0.93 · P = 0.0012
Cutoff: 15 February 2013
The investigator-assessed HR of 0.82 corresponds to an 18% lower estimated instantaneous hazard of progression or death. Its 95% CI of 0.73–0.93 provides the uncertainty interval around the estimate, and p = 0.0012 is the reported hypothesis-test result.
The comparison is not identical to the centrally reviewed PFS analysis because the endpoint assessment source differs. The two estimates therefore provide related but distinct statistical summaries.
Objective Tumour Response
| Assessment | Odds ratio | 95% CI | P-value |
|---|---|---|---|
| Central independent review | 1.34 | 0.76–2.39 | 0.3067 |
| Investigator assessment | 1.41 | 0.96–2.08 | 0.0761 |
Both analyses used logistic regression in the Randomised Set. The registry states that an odds ratio greater than 1 indicates a benefit to nintedanib.
The central-review OR of 1.34 indicates that the estimated odds of objective tumour response were higher in the BIBF 1120-containing group, but the 95% CI of 0.76–2.39 includes 1 and the reported p-value is 0.3067. The investigator-assessed OR of 1.41 has a 95% CI of 0.96–2.08 and p = 0.0761.
An odds ratio is not a risk ratio or a percentage-point difference in response rates. It also does not describe the duration of response or time until progression. Those are different clinical quantities.
Disease Control
| Assessment | Odds ratio | 95% CI | P-value |
|---|---|---|---|
| Central independent review | 1.68 | 1.35–2.09 | <0.0001 |
| Investigator assessment | 1.64 | 1.31–2.05 | <0.0001 |
The registry reports odds ratios above 1 as favoring nintedanib for these disease-control analyses.
The central-review OR of 1.68 indicates higher estimated odds of disease control in the BIBF 1120-containing group relative to the control group. The 95% CI of 1.35–2.09 does not include 1. The investigator-assessed OR is 1.64, with a 95% CI of 1.31–2.05. Both p-values are reported as <0.0001.
These results describe a binary disease-control outcome. They do not establish the magnitude of the difference in absolute response probability because the underlying response proportions are not reported in the ClinicalTrials.gov record.
Clinical Improvement
Clinical improvement hazard ratio
95% CI: 0.87–1.21 · P = 0.7282
The registry classifies clinical improvement as a time-to-event endpoint and analyzes it using a Cox proportional-hazards model.
Quality of Life
| Quality-of-life time-to-deterioration measure | HR | 95% CI | P-value |
|---|---|---|---|
| Time to deterioration of cough | 0.90 | 0.77–1.05 | 0.1858 |
| Time to deterioration of dyspnoea | 1.05 | 0.91–1.20 | 0.5203 |
| Time to deterioration of pain | 0.95 | 0.82–1.09 | 0.4373 |
The registry identifies these as Cox-model analyses of time to deterioration. For cough and pain, the registry states that HR below 1 favors nintedanib; for dyspnoea it likewise defines an HR below 1 as favoring nintedanib.
Change From Baseline in Tumour Size
| Assessment | Method | P-value |
|---|---|---|
| Central independent review | ANOVA | <0.0001 |
| Investigator assessment | ANOVA | <0.0001 |
The registry reports the outcome unit as percentage of change in tumor size in mm. It does not provide an effect estimate or confidence interval for these ANOVA analyses in the ClinicalTrials.gov record.
8. Statistical Methods Explained
Why was a Cox proportional-hazards model used?
PFS, overall survival, clinical improvement, and the reported quality-of-life deterioration outcomes are time-to-event endpoints. A Cox model is designed to compare event hazards while accommodating right-censored observations, such as participants who have not experienced the event by the end of their available follow-up.
What does an HR of 0.79 mean?
An HR of 0.79 means that the fitted model estimates the instantaneous event hazard in the BIBF 1120-containing group to be 0.79 times that in the control group, under the model and analysis framework. Equivalently, 1 − 0.79 = 0.21, or a 21% lower estimated hazard. It does not mean that 21% of patients avoided progression.
Why was the primary PFS model stratified?
The registry reports stratification by baseline ECOG performance status, tumour histology, brain metastases at baseline, and prior treatment with bevacizumab. Stratification allows the treatment comparison to account for these prespecified factors in the time-to-event analysis rather than treating all participants as belonging to one homogeneous risk stratum.
Why is the confidence interval important?
A point estimate such as HR 0.79 is only one estimate from the observed data. The 95% CI of 0.68–0.92 communicates the statistical uncertainty around that estimate. A confidence interval is therefore more informative about precision than the p-value alone.
Why does the p-value not measure effect size?
The p-value describes the compatibility of the observed data with the null hypothesis under the specified statistical test. It does not quantify the size of the treatment effect. The HR, OR, and their confidence intervals provide the effect-size information in the reported analyses.
Why is the overall-survival analysis hierarchical?
The registry describes a fixed sequence of three hypotheses: patients with adenocarcinoma and less than 9 months since start of first-line therapy, patients with adenocarcinoma, and then all patients. Each hypothesis could be tested at the prespecified alpha level only after rejection of the preceding null hypothesis. This sequencing is a multiplicity-control strategy because it prevents the three hypotheses from being treated as three unrelated opportunities for a positive finding.
Why are the central-review and investigator-assessed analyses separate?
They use different assessment sources. The registry separately reports PFS by central independent review and by investigator assessment, as well as objective tumour response and disease control based on central review versus investigator assessment. Agreement between related analyses can provide useful consistency information, but the estimates should remain labeled according to their assessment method.
9. Multiplicity and Hierarchical Testing
Multiplicity is directly relevant to the overall-survival key secondary endpoint. The registry specifies a fixed sequence of statistical hypotheses:
| Sequence | Population | Rule |
|---|---|---|
| 1 | Patients with adenocarcinoma and <9 months since start of first-line therapy | Could be tested at the prespecified alpha level. |
| 2 | Patients with adenocarcinoma | Could be tested only if the preceding null hypothesis was rejected. |
| 3 | All patients | Could be tested only if the preceding null hypothesis was rejected. |
This is different from conducting three unrelated two-sided tests and simply counting how many have p-values below a chosen threshold. In a hierarchical strategy, the order of hypotheses is part of the prespecified inferential procedure.
10. Stratification and the Cox Model
The primary PFS analysis was stratified by four baseline characteristics:
| Stratification factor | Levels reported |
|---|---|
| Baseline ECOG performance status | 0 vs 1 |
| Tumour histology | Squamous vs non-squamous |
| Brain metastases at baseline | Yes vs no |
| Prior treatment with bevacizumab | Yes vs no |
These variables are not themselves treatment effects. Their role in the reported primary analysis is to define the strata used by the proportional-hazards model. The resulting HR of 0.79 therefore represents the treatment comparison from the stratified model specified in the registry record.
A stratification variable helps structure the analysis; it does not mean that the trial was designed to estimate a separate treatment effect for every level of that variable.
11. Randomization and Analysis Population
The trial was randomized and the primary PFS analysis was performed in the Randomised Set. This is statistically important because randomized assignment defines the principal comparison between the two treatment strategies.
Analyzing the Randomised Set avoids redefining the comparison according to post-randomization treatment experience. In a randomized efficacy analysis, preserving the original treatment assignment maintains the connection between randomization and the estimated treatment contrast.
The registry does not provide additional definitions of per-protocol, as-treated, or safety populations in the ClinicalTrials.gov record beyond the reported serious-adverse-event counts by arm.
12. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk.
| Safety measure | Affected / at risk |
|---|---|
| Nintedanib Plus Docetaxel | 224 / 652 |
| Placebo Plus Docetaxel | 206 / 655 |
The reported safety data use Nintedanib Plus Docetaxel terminology for the intervention arm, whereas the ClinicalTrials.gov record identifies the investigational intervention as BIBF 1120 plus docetaxel. The ClinicalTrials.gov record does not provide additional safety estimates, confidence intervals, exposure-adjusted rates, or formal between-arm hypothesis tests.
13. Understanding the Secondary Results as a Statistical System
LUME-Lung 1 illustrates why a clinical trial should not be reduced to a single p-value. Different endpoints answer different questions and use different effect measures.
| Endpoint family | Effect measure | Question answered |
|---|---|---|
| PFS | Hazard ratio | How does the estimated time-to-progression-or-death hazard compare between treatment groups? |
| Overall survival | Hazard ratio | How does the estimated hazard of death compare between treatment groups? |
| Objective tumour response | Odds ratio | How do the odds of the binary response outcome compare? |
| Disease control | Odds ratio | How do the odds of disease control compare? |
| Clinical improvement | Hazard ratio | How does the time-to-event hazard compare? |
| Quality of life | Hazard ratio | How does time to deterioration compare? |
| Change in tumour size | ANOVA / p-value | Is there statistical evidence of a difference in change from baseline? |
This distinction prevents a common statistical mistake: treating every endpoint as if it were measuring the same outcome. An HR, OR, and ANOVA p-value have different meanings and cannot be compared numerically as though they were interchangeable effect measures.
14. Planned Analysis vs Reported Analysis
For LUME-Lung 1, results are posted and a formal statistical analysis is reported for the primary endpoint. The appropriate distinction is therefore between the registered endpoint definition and the analysis actually reported.
Registered endpoint
PFS from randomisation to progression or death, whichever occurred earlier, assessed by central independent review according to modified RECIST v1.0.
Reported analysis
Stratified Cox proportional-hazards model in the Randomised Set, with HR 0.79, 95% CI 0.68–0.92, and p = 0.0019.
This correspondence between endpoint type and analysis method is a useful example of prespecified statistical planning: a time-to-event endpoint is paired with a survival-analysis framework rather than a simple comparison of proportions.
15. Interpreting Confidence Intervals Correctly
Several LUME-Lung 1 results illustrate how confidence intervals communicate more than a binary significant/non-significant label.
| Result | 95% CI | What the interval communicates |
|---|---|---|
| Primary PFS HR 0.79 | 0.68–0.92 | Uncertainty around the estimated relative hazard from the stratified Cox model. |
| OS HR 0.75 | 0.60–0.92 | Uncertainty around the first hierarchical OS estimate. |
| OS HR 0.83 | 0.70–0.99 | Uncertainty around the second hierarchical OS estimate. |
| OS HR 0.94 | 0.83–1.05 | Uncertainty includes the null value of 1 for the all-patient hypothesis. |
| Disease-control OR 1.68 | 1.35–2.09 | Uncertainty around the estimated odds ratio from central review. |
For hazard ratios and odds ratios, 1 is the usual null value. An interval entirely below 1 for an HR is consistent with the treatment direction specified by the registry; an interval entirely above 1 for an OR is consistent with higher odds in the BIBF 1120-containing group. Neither interpretation removes the need to consider the endpoint definition, analysis population, and multiplicity structure.
16. Important Limitations and Interpretation Issues
- Hazard-ratio assumptions: the primary and several secondary time-to-event analyses use Cox proportional-hazards models. The ClinicalTrials.gov record does not provide a formal assessment of the proportional-hazards assumption.
- Relative rather than absolute effects: HRs and ORs are relative measures. The statistical analyses posted on ClinicalTrials.gov do not provide the underlying absolute event probabilities or response proportions needed to translate every result into an absolute treatment difference.
- Hierarchical testing: the three reported OS hypotheses are linked by a fixed sequence. They should not be interpreted as three independent hypothesis tests.
- Secondary endpoints: the trial contains multiple secondary analyses using different endpoints and methods. A nominal p-value for one secondary endpoint should not automatically be interpreted as independent confirmatory evidence without considering the trial's multiplicity framework.
- Assessment source: central independent review and investigator assessment are distinct analysis sources. Their estimates should not be merged.
- Analysis population: the primary PFS result is explicitly reported for the Randomised Set. Results from other populations should not be substituted for it.
- Incomplete effect-size reporting: for change from baseline in tumour size, the ClinicalTrials.gov record reports p-values but no estimated difference or confidence interval.
- Safety information: the ClinicalTrials.gov record provides serious-adverse-event counts by arm but does not provide a formal statistical comparison or additional exposure-adjusted measures.
17. Why This Trial Matters Statistically
LUME-Lung 1 is a useful teaching case because it combines several core clinical-trial statistical methods within a single randomized phase 3 study. The primary endpoint is a time-to-event outcome analyzed with a stratified Cox model, while secondary endpoints use Cox regression, logistic regression, and ANOVA.
| Concept | How it appears in LUME-Lung 1 |
|---|---|
| Randomization | 1314 participants assigned in a randomized parallel phase 3 trial. |
| Double masking | the ClinicalTrials.gov record identifies the study as double masked. |
| Time-to-event endpoint | Primary PFS endpoint defined from randomisation to progression or death. |
| Kaplan-Meier concept | The registered endpoint definition states that median, 25th and 75th percentiles are calculated from an unadjusted Kaplan-Meier analysis. |
| Cox model | Used for primary PFS and multiple secondary time-to-event analyses. |
| Hazard ratio | Primary PFS HR 0.79; additional OS and time-to-event HRs are also reported. |
| Stratified analysis | Primary PFS model stratified by four baseline factors. |
| Logistic regression | Used for objective tumour response and disease control. |
| Odds ratio | Reported for objective tumour response and disease control. |
| ANOVA | Used for change from baseline in tumour size. |
| Multiplicity adjustment | Overall survival used a fixed sequence of hypotheses. |
| Interim analysis / alpha spending | The ClinicalTrials.gov record identifies interim analysis / alpha spending as a concept in the OS analyses. |
18. Statistical Interpretation of the Primary Result
The primary PFS HR of 0.79 indicates a 21% lower estimated instantaneous hazard of progression or death for BIBF 1120 plus docetaxel relative to placebo plus docetaxel, under the reported stratified Cox model.
The 95% CI of 0.68–0.92 gives the statistical uncertainty around the estimated HR. It is narrower than a very imprecise interval would be, but it still does not identify the treatment effect for any individual participant.
The reported two-sided p-value of 0.0019 indicates strong statistical evidence against the null hypothesis under the specified superiority analysis. It is not itself an estimate of clinical magnitude.
The HR does not provide an absolute PFS probability, median PFS, number needed to treat, or percentage of participants who benefited. Those quantities require additional information not reported in the ClinicalTrials.gov record.
19. Primary Endpoint in Context
The primary PFS analysis is especially instructive because the registry supplies all of the essential components of a model-based treatment comparison: a prespecified time-to-event endpoint, a defined analysis population, a Cox regression method, stratification factors, an effect measure, a confidence interval, and a p-value.
Endpoint
PFS from randomisation to progression or death, whichever occurred earlier.
Analysis
Stratified Cox proportional-hazards model.
Effect
HR 0.79, with 95% CI 0.68–0.92.
Evidence
Two-sided p = 0.0019, based on 713 observed PFS events at the cutoff.
This structure is a useful template for reading other oncology trials: first identify exactly what constitutes the event, then identify who was analyzed, then identify the statistical model, and only after that interpret the effect estimate and p-value.
20. Related Tutorials
Learn more about the methods used in this trial:
21. Related Statistical Calculators
22. Sources
- ClinicalTrials.gov: NCT00805194 — LUME-Lung 1.
- Linked PubMed record: PubMed PMID 28702806.
- Linked PubMed record: PubMed PMID 24411639.
Continue through the Clinical Biostats statistical pathway
Explore the statistical concepts behind randomized trials, survival analysis, regression, confidence intervals, and multiple testing.
23. Record Summary
LUME-Lung 1 is a randomized, double-blind, parallel phase 3 trial with 1314 participants and two treatment arms. Its registered primary endpoint was progression-free survival assessed by central independent review, defined as time from randomisation to progression or death, whichever occurred earlier. The primary analysis used a stratified Cox proportional-hazards model in the Randomised Set and reported an HR of 0.79 (95% CI 0.68–0.92; p = 0.0019) at the 2 November 2010 cutoff, when 713 PFS events had been observed.
The secondary analyses show how a single randomized trial can require several statistical frameworks. Overall survival was analyzed with Cox regression and a fixed hierarchical testing sequence. Objective tumour response and disease control were analyzed with logistic regression and odds ratios. Change from baseline in tumour size was analyzed with ANOVA. Clinical improvement and quality-of-life deterioration outcomes were analyzed as time-to-event outcomes using Cox models.
The most important statistical lesson is that the interpretation of each result depends on its endpoint definition, analysis population, model, effect measure, confidence interval, and multiplicity structure. The primary HR of 0.79 is therefore most appropriately understood as a model-based relative treatment effect for centrally assessed PFS, not as an absolute probability difference or a universal measure of individual patient benefit.