This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
LUME-Lung 2 was a randomized, double-blind, parallel phase 3 trial evaluating nintedanib (BIBF 1120) plus pemetrexed versus placebo plus pemetrexed in second-line nonsquamous non-small-cell lung cancer. The registry reports 718 enrolled participants, a time-to-event primary endpoint, 14 posted outcome measures, and 14 posted statistical analyses.
| Feature | LUME-Lung 2 |
|---|---|
| Trial name | LUME-Lung 2 |
| Brief title | BIBF 1120 Plus Pemetrexed Compared to Placebo Plus Pemetrexed in 2nd Line Nonsquamous NSCLC |
| Phase | Phase 3 |
| Condition | Carcinoma, Non-Small-Cell Lung |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Double |
| Primary purpose | Treatment |
| Enrollment | 718 |
| Primary endpoint type | Time-to-event |
| Results posted | Yes |
| Outcome measures posted | 14 |
| Statistical analyses posted | 14 |
| Lead sponsor | Boehringer Ingelheim |
| Sponsor type | Industry |
| Trial status | Completed |
2. Clinical Question
The central statistical question was whether adding nintedanib to pemetrexed changed progression-free survival compared with placebo plus pemetrexed in the registered second-line nonsquamous NSCLC population.
Population
Participants enrolled in the phase 3 LUME-Lung 2 study with carcinoma, non-small-cell lung, in the registered second-line nonsquamous NSCLC setting.
Intervention
Nintedanib (BIBF 1120) plus pemetrexed.
Comparator
Placebo plus pemetrexed.
Primary question
Does nintedanib plus pemetrexed improve progression-free survival relative to placebo plus pemetrexed?
3. Trial Design
The registry describes LUME-Lung 2 as a randomized, double-blind, parallel-group, phase 3 treatment trial. The registry lists five intervention arms, while the posted comparative statistical analyses are framed as a two-group comparison of nintedanib plus pemetrexed versus placebo plus pemetrexed.
Nintedanib Plus Pemetrexed
- Nintedanib (BIBF 1120)
- Pemetrexed
Placebo Plus Pemetrexed
- Placebo
- Pemetrexed
The intervention listing also contains B12, dexamethasone or corticosteroid equivalent, and folic acid. The ClinicalTrials.gov record does not provide enough detail to reconstruct a complete dosing schedule or treatment sequence, so those details are not inferred here.
Study initiated
The registered trial start is December 2008.
Primary completion
The registered primary completion date is June 2011.
Central-review PFS analysis
The registered primary endpoint was analyzed from randomisation through the cutoff date of 9 July 2012.
Follow-up analyses
The posted secondary analyses use a data cutoff of 15 February 2013, with the registry describing follow-up of up to 30 months.
4. Endpoints
Primary endpoint
| Endpoint | Registered definition | Time frame |
|---|---|---|
| Progression Free Survival (PFS) as Assessed by Central Independent Review | Progression Free Survival (PFS) as assessed by central independent review according to the modified RECIST (version 1.0) criteria. Progression free survival (PFS) is defined as the duration of time from date of randomisation to date of progression or death (whatever occurs earlier). Median, 25th and 75th percentiles are calculated from an unadjusted Kaplan-Meier curve. | From randomisation until cut-off date 9 July 2012 |
Posted secondary endpoints
| Endpoint | Type | Time frame | Analysis method |
|---|---|---|---|
| Overall Survival (Key Secondary Endpoint) | Time-to-event | From randomisation until data cut-off (15 February 2013), Up to 30 months | Cox proportional-hazards model |
| Follow-up Analysis of Progression Free Survival (PFS) as Assessed by Central Independent Review | Time-to-event | From randomisation until data cut-off (15 February 2013), Up to 30 months | Cox proportional-hazards model |
| Follow-up Analysis of Progression Free Survival (PFS) as Assessed by Investigator | Time-to-event | From randomisation until data cut-off (15 February 2013), Up to 30 months | Cox proportional-hazards model |
| Objective Tumor Response | Binary | From randomisation until data cut-off (15 February 2013), Up to 30 months | Logistic regression |
| Disease Control | Binary | From randomisation until data cut-off (15 February 2013), Up to 30 months | Logistic regression |
| Clinical Improvement. | Time-to-event | From randomisation until data cut-off (15 February 2013), Up to 30 months | Cox proportional-hazards model |
| Quality of Life (QoL) | Time-to-event | From randomisation until data cut-off (15 February 2013), Up to 30 months | Cox proportional-hazards model |
| Change From Baseline in Tumour Size | Binary as registered in the posted analysis | From randomisation until data cut-off (15 February 2013), Up to 30 months | ANOVA |
The registry supplies a detailed formal definition for the primary PFS endpoint. For several secondary outcomes, the ClinicalTrials.gov record provides the endpoint name, time frame, analysis population, statistical method, and effect measure but do not provide a separate narrative endpoint definition. This page therefore does not add definitions that are not present in the ClinicalTrials.gov record.
5. Statistical Methodology
The posted analyses use three main statistical families: Cox proportional-hazards models for time-to-event outcomes, logistic regression for binary outcomes, and ANOVA for change from baseline in tumour size. The Cox analyses are described as stratified analyses, with the registry specifying four baseline stratification factors.
Stratification in the Cox analyses
The posted Cox analyses were stratified by:
- Baseline ECOG PS: 0 vs 1.
- Tumour histology: adenocarcinoma vs non-adenocarcinoma.
- Brain metastases at baseline: yes vs no.
- Prior treatment with bevacizumab: yes vs no.
Stratification allows the baseline hazard to differ across the specified strata while estimating the treatment comparison across those strata. It is particularly relevant here because the registry explicitly states that the HR, confidence interval, and p-value for the posted Cox analyses were obtained from models stratified on these factors.
For this trial, the registry states that an HR below 1 favors nintedanib. That directional statement applies to the nintedanib-plus-pemetrexed versus placebo-plus-pemetrexed comparisons reported here.
Analysis populations
The primary PFS analysis and the posted time-to-event secondary analyses use the RS analysis population. Objective tumor response and its follow-up analysis use the Randomised Set. The ClinicalTrials.gov record does not define the abbreviation “RS,” so this page does not expand it beyond the label used in the posted analysis.
| Endpoint family | Analysis population | Method | Effect measure |
|---|---|---|---|
| Primary PFS | RS | Cox proportional-hazards model | Hazard ratio |
| OS | RS | Cox proportional-hazards model | Hazard ratio |
| Follow-up PFS | RS | Cox proportional-hazards model | Hazard ratio |
| Objective tumor response | Randomised Set | Logistic regression | Odds ratio |
| Disease control | RS | Logistic regression | Odds ratio |
| Clinical improvement | RS | Cox proportional-hazards model | Hazard ratio |
| Quality of life | RS | Cox proportional-hazards model | Hazard ratio |
| Change from baseline in tumour size | RS | ANOVA | No estimate reported in registry-reported analysis |
6. Primary Result: Progression-Free Survival
The primary endpoint was progression-free survival as assessed by central independent review according to modified RECIST version 1.0 criteria. PFS was defined as the time from randomisation to progression or death, whichever occurred earlier. The primary analysis used the RS population and a stratified Cox proportional-hazards model.
Hazard ratio for progression or death
95% CI: 0.70–0.99 · P = 0.0435
Two-sided confidence interval; superiority hypothesis.
| Primary PFS result | Reported value |
|---|---|
| Comparison | Nintedanib Plus Pemetrexed vs Placebo Plus Pemetrexed |
| Analysis population | RS |
| Method | Stratified Cox proportional-hazards model |
| Hazard ratio | 0.83 |
| 95% CI | 0.70–0.99 |
| P-value | 0.0435 |
| Hypothesis | Superiority |
| Data cutoff | 9 July 2012 |
An HR of 0.83 means that, under the fitted stratified Cox model, the estimated instantaneous rate of progression or death in the nintedanib-plus-pemetrexed group was approximately 83% of the corresponding rate in the placebo-plus-pemetrexed group. Expressed as a relative model-based comparison, this corresponds to a 17% lower estimated hazard for progression or death.
The HR does not mean that 17% of participants avoided progression, that every participant experienced a 17% reduction, or that the median PFS differed by 17%. It is a relative time-to-event measure derived from a model.
The 95% CI of 0.70–0.99 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of individual patient effects. The interval is relatively close to 1 at its upper boundary, so the numerical precision of the estimated relative effect matters when interpreting the magnitude.
The p-value of 0.0435 addresses the statistical evidence against the null hypothesis under the prespecified analysis framework; it does not measure the size or clinical importance of the effect. A p-value is therefore not a substitute for the HR and its confidence interval.
Finally, the Cox interpretation depends on the proportional-hazards framework. The ClinicalTrials.gov record reports a Cox proportional-hazards model but do not provide a diagnostic assessment of the proportional-hazards assumption. The result should therefore be understood as the reported model-based comparison rather than as a guarantee that the hazard ratio was constant at every time point.
7. Secondary Result: Overall Survival
Overall survival was registered as a key secondary endpoint. The posted analysis used the RS population and a stratified Cox proportional-hazards model from randomisation through the 15 February 2013 data cutoff, with follow-up described as up to 30 months.
Hazard ratio for overall survival
95% CI: 0.85–1.21 · P = 0.8940
Two-sided confidence interval; superiority hypothesis.
| Overall survival result | Reported value |
|---|---|
| Comparison | Nintedanib Plus Pemetrexed vs Placebo Plus Pemetrexed |
| Analysis population | RS |
| Method | Stratified Cox proportional-hazards model |
| Hazard ratio | 1.01 |
| 95% CI | 0.85–1.21 |
| P-value | 0.8940 |
| Data cutoff | 15 February 2013 |
| Time frame | Up to 30 months |
An HR of 1.01 is very close to 1.00. Under the fitted stratified Cox model, the estimated instantaneous rate of death was therefore approximately the same between the two randomized comparison groups at the level summarized by the model.
The 95% CI of 0.85–1.21 spans 1.00. This means the reported estimate is compatible with a range of relative hazard differences in either direction under the stated model and sampling framework. The interval is more informative about uncertainty than the point estimate alone.
The p-value of 0.8940 is not an effect-size measure. It indicates little statistical evidence against the null hypothesis for this particular superiority analysis; it does not establish that the treatments are identical or that clinically meaningful differences are impossible.
As with the primary PFS analysis, the result is a model-based hazard ratio and should not be interpreted as a percentage of participants who benefited or as a statement about an individual patient's survival. The analysis also uses the registry-specified stratification factors and the RS population.
8. Secondary Result: Follow-up Central-Review PFS
A later analysis evaluated progression-free survival by central independent review using the 15 February 2013 cutoff. The same comparison and stratification factors were used in the posted Cox analysis.
Follow-up PFS hazard ratio
95% CI: 0.70–1.00 · P = 0.0506
Two-sided confidence interval; superiority hypothesis.
The estimated HR of 0.84 corresponds to an approximately 16% lower estimated hazard of progression or death under the fitted Cox model for nintedanib plus pemetrexed relative to placebo plus pemetrexed.
The 95% CI of 0.70–1.00 reaches 1.00 at its upper boundary. The estimate therefore carries more uncertainty than the primary PFS estimate, and the interval illustrates why a point estimate should not be interpreted without its precision.
The p-value of 0.0506 is close to 0.05, but the p-value itself does not quantify the magnitude of the treatment effect. The appropriate reading is the combination of the HR, confidence interval, analysis population, data cutoff, and prespecified analysis framework.
9. Secondary Result: Follow-up Investigator-Assessed PFS
The registry also reports a follow-up PFS analysis based on investigator assessment.
Investigator-assessed PFS hazard ratio
95% CI: 0.73–1.02 · P = 0.0865
Two-sided confidence interval; superiority hypothesis.
An HR of 0.86 corresponds to an approximately 14% lower estimated hazard of progression or death under the fitted model for the nintedanib-plus-pemetrexed group.
The 95% CI of 0.73–1.02 includes 1.00. Thus the confidence interval does not exclude a null relative hazard at the conventional level associated with a two-sided 95% interval. The result also differs somewhat from the central independent review estimate of 0.84, illustrating that the assessment source is part of the statistical definition of a time-to-event endpoint.
The p-value of 0.0865 should not be interpreted as a probability that the treatment effect is absent, nor as a measure of clinical importance. It is a test statistic summary under the specified statistical model.
10. Secondary Result: Objective Tumor Response
Objective tumor response was analyzed as a binary outcome using logistic regression. The registry provides two posted analyses: one based on central independent review and one based on investigator assessment.
| Assessment | Analysis population | Odds ratio | 95% CI | P-value |
|---|---|---|---|---|
| Central independent review | Randomised Set | 1.10 | 0.65–1.85 | 0.7279 |
| Investigator assessment | Randomised Set | 1.15 | 0.75–1.76 | 0.5180 |
Central independent review
The OR of 1.10 means the estimated odds of objective tumor response were 1.10 times the corresponding odds in the comparator group under the logistic regression analysis. The 95% CI was 0.65–1.85 and the p-value was 0.7279.
Investigator assessment
The OR of 1.15 means the estimated odds of objective tumor response were 1.15 times the corresponding odds in the comparator group. The 95% CI was 0.75–1.76 and the p-value was 0.5180.
An odds ratio is not the same as a risk ratio or a difference in response percentages. For example, an OR of 1.10 does not mean that the response probability increased by 10 percentage points or even necessarily by 10%.
Both reported confidence intervals include 1.00. The central-review interval of 0.65–1.85 is especially useful for showing how much uncertainty surrounds the estimate. The investigator-assessed interval of 0.75–1.76 likewise permits a range of possible relative odds in either direction.
The registry states that an odds ratio greater than 1 indicates a benefit to nintedanib for these response analyses. That directional convention is important because the interpretation of an odds ratio depends on which group is placed in the numerator.
11. Secondary Result: Disease Control
Disease control was also analyzed with logistic regression. Two analyses were posted: one based on central independent review and one based on investigator assessment.
| Assessment | Analysis population | Odds ratio | 95% CI | P-value |
|---|---|---|---|---|
| Central independent review | RS | 1.37 | 1.02–1.85 | 0.0387 |
| Investigator assessment | RS | 1.29 | 0.95–1.75 | 0.1071 |
For the central-review analysis, an OR of 1.37 indicates that the estimated odds of disease control were 1.37 times those in the placebo-plus-pemetrexed group under the logistic regression model. The 95% CI was 1.02–1.85 and the p-value was 0.0387.
The investigator-assessed analysis produced an OR of 1.29, with a 95% CI of 0.95–1.75 and a p-value of 0.1071. The difference between the two estimates illustrates why assessment method and analysis definition matter.
Neither OR should be translated directly into an absolute probability without the underlying response rates. An odds ratio is a relative comparison of odds, not a percentage-point difference.
12. Secondary Result: Clinical Improvement
Clinical Improvement. was analyzed as a time-to-event outcome in the RS population using a stratified Cox proportional-hazards model.
Clinical Improvement. hazard ratio
95% CI: 0.74–1.16 · P = 0.5068
Two-sided confidence interval; superiority hypothesis.
The HR of 0.93 corresponds to an approximately 7% lower estimated hazard under the fitted model for the nintedanib-plus-pemetrexed group. The 95% CI of 0.74–1.16 spans 1.00, so the point estimate should not be read as establishing a directional treatment effect on its own.
The p-value of 0.5068 is a hypothesis-test result, not a measure of effect magnitude. Without the underlying event curves or absolute time-to-event summaries, the HR and CI are the principal numerical description reported by the registry for this endpoint.
13. Secondary Results: Quality of Life
The posted Quality of Life (QoL) analyses are time-to-event analyses evaluating time to deterioration of specific symptoms. The registry identifies cough, dyspnoea, and pain as the specific deterioration outcomes in the analysis notes.
| QoL outcome | Analysis | 95% CI | P-value |
|---|---|---|---|
| Time to deterioration of cough | HR 0.83 | 0.66–1.05 | 0.1181 |
| Time to deterioration of dyspnoea | HR 0.93 | 0.77–1.12 | 0.4264 |
| Time to deterioration of pain | HR 1.01 | 0.84–1.23 | 0.8929 |
The cough analysis produced an HR of 0.83, corresponding to an approximately 17% lower estimated hazard of deterioration under the fitted model. Its 95% CI was 0.66–1.05 and its p-value was 0.1181.
The dyspnoea analysis produced an HR of 0.93, with a 95% CI of 0.77–1.12 and p-value of 0.4264. The pain analysis produced an HR of 1.01, with a 95% CI of 0.84–1.23 and p-value of 0.8929.
These are separate time-to-event outcomes. Their p-values should not be treated as interchangeable measures of overall quality-of-life benefit, and the ClinicalTrials.gov record does not provide a multiplicity procedure for combining these outcomes into a single confirmatory claim.
14. Secondary Result: Change From Baseline in Tumour Size
Change from baseline in tumour size was analyzed using ANOVA. The ClinicalTrials.gov record describes the outcome unit as percentage of change in tumor size in mm. No treatment-effect estimate or confidence interval is provided in the ClinicalTrials.gov record; the registry reports p-values for the two assessment approaches.
| Assessment | Method | P-value |
|---|---|---|
| Central independent review | ANOVA | 0.1558 |
| Investigator assessment | ANOVA | 0.0565 |
ANOVA tests differences in the modeled outcome between groups; in this registry record, the registry-reported numerical result is the p-value rather than a treatment-effect estimate with a confidence interval.
The central-review p-value was 0.1558, while the investigator-assessment p-value was 0.0565. These p-values should not be converted into effect sizes. Without a reported difference estimate and confidence interval, the ClinicalTrials.gov record does not quantify the magnitude or precision of the between-group change in tumour size.
15. Serious Adverse Events by Arm
The ClinicalTrials.gov record reports serious adverse events by treatment group as affected participants divided by participants at risk.
| Treatment group | Affected | At risk |
|---|---|---|
| Nintedanib Plus Pemetrexed | 104 | 347 |
| Placebo Plus Pemetrexed | 117 | 357 |
The affected-to-at-risk values are approximately 30.0% and 32.8% when expressed as simple proportions, but those percentages are derived from the registry-reported counts and are not additional registry-reported estimates. This page therefore keeps the primary safety presentation in the exact affected/at-risk format reported.
16. Statistical Methods Explained
Why was a Cox proportional-hazards model used for PFS?
PFS is a time-to-event endpoint: each participant is followed until progression or death, or until censoring if an event has not occurred by the relevant observation endpoint. A Cox model uses the ordering and timing of events while accommodating censored observations. The resulting hazard ratio summarizes the relative event rate under the fitted model.
What does a hazard ratio of 0.83 mean?
An HR of 0.83 means that the modeled instantaneous rate of progression or death was estimated at approximately 83% of the comparator rate, conditional on the model. It can be described as an approximately 17% lower estimated hazard. It does not mean a 17% absolute improvement in PFS probability, nor does it mean that every participant experienced the same reduction.
Why was stratification used?
The registry reports that the Cox analyses were stratified by baseline ECOG PS, tumour histology, brain metastases at baseline, and prior bevacizumab treatment. Stratification permits the underlying baseline hazard to differ across those categories rather than forcing a single baseline hazard across all strata. The treatment effect is then estimated within the stratified modeling framework.
What is the difference between an odds ratio and a hazard ratio?
An odds ratio compares the odds of a binary outcome, such as objective tumor response or disease control. A hazard ratio compares instantaneous event rates over time for a time-to-event endpoint. An OR of 1.37 for disease control therefore cannot be interpreted in the same way as an HR of 0.83 for PFS.
Why does the p-value not tell us the size of the treatment effect?
A p-value summarizes the compatibility of the observed data with a null hypothesis under the specified statistical model. It is affected by sample size, event information, variability, and the magnitude of the observed difference. Effect size is better represented by the HR or OR together with its confidence interval, while absolute event probabilities or time summaries provide additional clinical context.
Why are central-review and investigator-assessed PFS results listed separately?
The registry explicitly distinguishes PFS assessed by central independent review from PFS assessed by the investigator. These are different assessment sources and therefore different operational measurements of the endpoint. Comparing their numerical results can be informative, but they should not be silently combined into one estimate.
Why is the primary PFS cutoff different from the follow-up cutoff?
The primary PFS endpoint was defined through the cutoff date of 9 July 2012. Several secondary analyses use a later data cutoff of 15 February 2013, with follow-up described as up to 30 months. A later cutoff incorporates additional follow-up information and therefore represents a different analysis dataset.
17. Early Stopping for Futility
The registry contains an important design caveat: recruitment for the study was stopped early based on the results of a pre defined futility analysis.
What futility means
A futility analysis is designed to assess whether continuing a study is unlikely to produce the prespecified objective under the assumptions of the trial design. It is conceptually different from an efficacy analysis, where the question is whether evidence is sufficiently strong to support a treatment effect.
Why early stopping matters
Stopping recruitment early changes the amount of information accumulated and can affect the precision and operating characteristics of the final evidence. The ClinicalTrials.gov record does not provide the futility boundary or its numerical operating characteristics.
The important statistical point is that the registry explicitly identifies the futility analysis as the reason recruitment was stopped early. This should be considered when interpreting the later posted estimates and the amount of information available for each endpoint.
18. Multiplicity and Multiple Posted Analyses
The registry reports 14 statistical analyses covering the primary PFS endpoint and multiple secondary outcomes. The ClinicalTrials.gov record identifies all analyses as superiority hypotheses and provide two-sided confidence intervals where an estimate and interval are reported.
| Analysis family | Number of posted analyses represented here | Primary statistical role |
|---|---|---|
| Primary PFS | 1 | Primary time-to-event comparison |
| Overall survival | 1 | Key secondary time-to-event comparison |
| Follow-up PFS | 2 | Central-review and investigator-assessed follow-up analyses |
| Objective tumor response | 2 | Central-review and investigator-assessed binary analyses |
| Disease control | 2 | Central-review and investigator-assessed binary analyses |
| Clinical Improvement. | 1 | Time-to-event analysis |
| Quality of Life (QoL) | 3 | Cough, dyspnoea, and pain deterioration analyses |
| Change From Baseline in Tumour Size | 2 | Central-review and investigator-assessed ANOVA analyses |
Because many endpoints and assessment approaches were analyzed, a nominal p-value should not automatically be treated as evidence that a result is confirmatory in the same sense as a prespecified primary endpoint. The ClinicalTrials.gov record does not provide a multiplicity adjustment procedure or an alpha-allocation hierarchy for all 14 posted analyses, so this page does not invent one.
19. Randomization and Blinding
The trial is registered as randomized and double masked, with a parallel design. Randomization is central to the causal interpretation of a treatment comparison because assignment, rather than observed outcome, determines which comparison group a participant belongs to.
Randomization
Randomized allocation is intended to balance known and unknown prognostic factors on average across treatment groups, subject to chance variation.
Double masking
Double masking can reduce the opportunity for knowledge of treatment assignment to influence participant or investigator behavior and assessment. The registry does not provide additional masking-role details in the ClinicalTrials.gov record.
The distinction between randomized assignment and the later analysis population is important. A randomized trial's treatment comparison is anchored in the assigned groups, while the registry's individual statistical analyses specify either RS or the Randomised Set as their analysis population.
20. Understanding the Primary PFS Estimate
The primary HR of 0.83 is a relative model-based measure. It indicates a lower estimated hazard of progression or death for nintedanib plus pemetrexed under the reported stratified Cox model.
The ClinicalTrials.gov record does not provide median PFS values, PFS rates at a specified time point, or absolute event probabilities. Therefore, the primary HR should not be converted into an invented absolute benefit.
The 95% CI of 0.70–0.99 is essential because it shows the uncertainty surrounding the point estimate. The upper boundary is close to 1.00, so the magnitude of the estimated effect should be interpreted with attention to that uncertainty.
The two-sided p-value of 0.0435 provides the hypothesis-test component of the result. It should be interpreted as evidence relative to the specified null hypothesis, not as a measure of the size, importance, or probability of the treatment effect.
21. Comparing the Posted Time-to-Event Analyses
| Endpoint | HR | 95% CI | P-value | Cutoff |
|---|---|---|---|---|
| Primary central-review PFS | 0.83 | 0.70–0.99 | 0.0435 | 9 July 2012 |
| Overall survival | 1.01 | 0.85–1.21 | 0.8940 | 15 February 2013 |
| Follow-up central-review PFS | 0.84 | 0.70–1.00 | 0.0506 | 15 February 2013 |
| Follow-up investigator-assessed PFS | 0.86 | 0.73–1.02 | 0.0865 | 15 February 2013 |
| Clinical Improvement. | 0.93 | 0.74–1.16 | 0.5068 | 15 February 2013 |
| QoL: cough deterioration | 0.83 | 0.66–1.05 | 0.1181 | 15 February 2013 |
| QoL: dyspnoea deterioration | 0.93 | 0.77–1.12 | 0.4264 | 15 February 2013 |
| QoL: pain deterioration | 1.01 | 0.84–1.23 | 0.8929 | 15 February 2013 |
The table demonstrates why a trial cannot be summarized adequately by a single p-value. The primary PFS analysis has an HR below 1 with a 95% CI that remains below 1, while the later central-review PFS estimate is slightly closer to the null and has a confidence interval reaching 1.00. Overall survival is centered very close to 1.00. The other time-to-event endpoints have their own estimates and uncertainty.
These differences do not require a contradiction: the endpoints measure different events, use different cutoffs, and in some cases use different assessment definitions. A statistical analysis should preserve those distinctions rather than collapsing them into one overall treatment-effect statement.
22. Limitations
- Early stopping for futility: the registry states that recruitment was stopped early based on a pre defined futility analysis. The ClinicalTrials.gov record does not provide the futility boundary or its numerical decision rule.
- No median PFS or OS reported: although the primary endpoint definition states that median, 25th, and 75th percentiles are calculated from an unadjusted Kaplan-Meier curve, the ClinicalTrials.gov record does not provide those numerical summaries.
- Multiple analyses: 14 statistical analyses were posted. The ClinicalTrials.gov record does not provide a complete multiplicity-adjustment strategy across all of them.
- Model assumptions: Cox proportional-hazards models rely on the proportional-hazards framework. The ClinicalTrials.gov record does not report a diagnostic assessment of that assumption.
- Assessment source: central independent review and investigator assessment produce separate analyses for some endpoints. Their results should not be conflated.
- Analysis populations: the registry specifies RS for most posted analyses and Randomised Set for objective tumor response. The ClinicalTrials.gov record does not define the RS abbreviation.
- Safety analysis: serious adverse event counts are available by arm, but the ClinicalTrials.gov record does not provide a formal statistical comparison.
- Missing data: the ClinicalTrials.gov record does not specify an imputation strategy for missing observations. No imputation method is therefore inferred.
- Stratification: the Cox models are stratified, but the ClinicalTrials.gov record does not provide stratum-specific estimates or an assessment of treatment-by-stratum interaction.
- Absolute effects: the ClinicalTrials.gov record contains hazard ratios and odds ratios but do not provide enough information to calculate validated absolute survival or response differences without adding external data.
23. Why This Trial Matters Statistically
LUME-Lung 2 is a useful statistical teaching case because the registry combines a randomized, double-blind phase 3 design with a time-to-event primary endpoint, stratified Cox modeling, binary response analyses, ANOVA, multiple assessment sources, and a prespecified futility-based early stopping decision.
| Concept | How it appears in LUME-Lung 2 |
|---|---|
| Randomization | The trial is registered as randomized with a parallel design. |
| Blinding | The trial is registered as double masked. |
| Time-to-event endpoint | Primary PFS is defined from randomisation to progression or death. |
| Kaplan-Meier estimation | The registered PFS definition states that median, 25th and 75th percentiles are calculated from an unadjusted Kaplan-Meier curve. |
| Cox model | Primary PFS and multiple secondary time-to-event outcomes use Cox proportional-hazards models. |
| Stratified analysis | Cox analyses are stratified by ECOG PS, histology, brain metastases, and prior bevacizumab treatment. |
| Hazard ratio | The primary PFS effect measure is HR 0.83 with a 95% CI of 0.70–0.99. |
| Logistic regression | Objective tumor response and disease control are analyzed with logistic regression. |
| Odds ratio | Binary analyses report odds ratios such as 1.10, 1.15, 1.37, and 1.29. |
| ANOVA | Change from baseline in tumour size is analyzed using ANOVA. |
| Multiple endpoints | The registry reports 14 statistical analyses across primary and secondary outcomes. |
| Early stopping | Recruitment was stopped early based on a pre defined futility analysis. |
24. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
25. Related Statistical Calculators
Use these calculator pathways to explore the statistical quantities that appear in this trial:
26. Sources
- ClinicalTrials.gov: LUME-Lung 2, NCT00806819. Trial design, endpoint definitions, posted results, statistical analyses, analysis populations, and registry caveats used on this page.
- PubMed: PMID 28184431. Linked publication record associated with the trial.
Continue through the Clinical Biostats statistical pathway
Use the trial as a practical example of time-to-event analysis, hazard ratios, logistic regression, odds ratios, ANOVA, confidence intervals, and statistical interpretation.
27. Record Summary
LUME-Lung 2 provides a compact example of how a randomized phase 3 trial can generate several distinct statistical questions from the same treatment comparison. The registered primary endpoint was progression-free survival assessed by central independent review, defined as time from randomisation to progression or death. Its posted analysis used a stratified Cox proportional-hazards model and reported HR 0.83, 95% CI 0.70–0.99, P = 0.0435.
The secondary analyses show why endpoint-specific interpretation matters. Overall survival had an HR of 1.01 with a 95% CI of 0.85–1.21 and P = 0.8940. Follow-up central-review PFS had HR 0.84, 95% CI 0.70–1.00, and P = 0.0506, while investigator-assessed PFS had HR 0.86, 95% CI 0.73–1.02, and P = 0.0865. Binary outcomes were evaluated using logistic regression, and change from baseline in tumour size was analyzed with ANOVA.
The statistical story also includes an important design feature: recruitment was stopped early based on a pre defined futility analysis. Because the registry reports 14 statistical analyses across several endpoints and assessment approaches, individual p-values should be interpreted in their proper endpoint and multiplicity context rather than treated as interchangeable evidence.