← Clinical Trials
Nonsquamous NSCLC Phase 3 Completed NCT00806819

LUME-Lung 2: Complete Statistical Analysis of Nintedanib in Nonsquamous NSCLC

An independent statistical review of the randomized phase 3 LUME-Lung 2 trial comparing nintedanib plus pemetrexed with placebo plus pemetrexed in second-line nonsquamous non-small-cell lung cancer, with emphasis on progression-free survival, overall survival, response, disease control, quality of life, and the statistical models used to analyze them.

Trial period: 2008-12 to 2011-06  ·  Enrollment: 718  ·  Lead sponsor: Boehringer Ingelheim
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

LUME-Lung 2 was a randomized, double-blind, parallel phase 3 trial evaluating nintedanib (BIBF 1120) plus pemetrexed versus placebo plus pemetrexed in second-line nonsquamous non-small-cell lung cancer. The registry reports 718 enrolled participants, a time-to-event primary endpoint, 14 posted outcome measures, and 14 posted statistical analyses.

718
Enrollment
ClinicalTrials.gov
3
Phase
Phase 3
0.83
Primary PFS HR
95% CI 0.70–0.99
0.0435
Primary PFS P-value
Two-sided
FeatureLUME-Lung 2
Trial nameLUME-Lung 2
Brief titleBIBF 1120 Plus Pemetrexed Compared to Placebo Plus Pemetrexed in 2nd Line Nonsquamous NSCLC
PhasePhase 3
ConditionCarcinoma, Non-Small-Cell Lung
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
Enrollment718
Primary endpoint typeTime-to-event
Results postedYes
Outcome measures posted14
Statistical analyses posted14
Lead sponsorBoehringer Ingelheim
Sponsor typeIndustry
Trial statusCompleted

2. Clinical Question

The central statistical question was whether adding nintedanib to pemetrexed changed progression-free survival compared with placebo plus pemetrexed in the registered second-line nonsquamous NSCLC population.

Population

Participants enrolled in the phase 3 LUME-Lung 2 study with carcinoma, non-small-cell lung, in the registered second-line nonsquamous NSCLC setting.

Intervention

Nintedanib (BIBF 1120) plus pemetrexed.

Comparator

Placebo plus pemetrexed.

Primary question

Does nintedanib plus pemetrexed improve progression-free survival relative to placebo plus pemetrexed?

3. Trial Design

The registry describes LUME-Lung 2 as a randomized, double-blind, parallel-group, phase 3 treatment trial. The registry lists five intervention arms, while the posted comparative statistical analyses are framed as a two-group comparison of nintedanib plus pemetrexed versus placebo plus pemetrexed.

01
Randomize718 enrolled
02
Double blindRandomized treatment comparison
03
TreatmentNintedanib + pemetrexed or placebo + pemetrexed
04
AssessPFS, survival, response and other outcomes
05
AnalyzeCox, logistic regression and ANOVA
COMPARISON GROUP 1

Nintedanib Plus Pemetrexed

  • Nintedanib (BIBF 1120)
  • Pemetrexed
COMPARISON GROUP 2

Placebo Plus Pemetrexed

  • Placebo
  • Pemetrexed

The intervention listing also contains B12, dexamethasone or corticosteroid equivalent, and folic acid. The ClinicalTrials.gov record does not provide enough detail to reconstruct a complete dosing schedule or treatment sequence, so those details are not inferred here.

2008-12 · Trial start

Study initiated

The registered trial start is December 2008.

2011-06 · Primary completion

Primary completion

The registered primary completion date is June 2011.

9 July 2012 · Primary PFS cutoff

Central-review PFS analysis

The registered primary endpoint was analyzed from randomisation through the cutoff date of 9 July 2012.

15 February 2013 · Later cutoff

Follow-up analyses

The posted secondary analyses use a data cutoff of 15 February 2013, with the registry describing follow-up of up to 30 months.

4. Endpoints

Primary endpoint

EndpointRegistered definitionTime frame
Progression Free Survival (PFS) as Assessed by Central Independent Review Progression Free Survival (PFS) as assessed by central independent review according to the modified RECIST (version 1.0) criteria. Progression free survival (PFS) is defined as the duration of time from date of randomisation to date of progression or death (whatever occurs earlier). Median, 25th and 75th percentiles are calculated from an unadjusted Kaplan-Meier curve. From randomisation until cut-off date 9 July 2012

Posted secondary endpoints

EndpointTypeTime frameAnalysis method
Overall Survival (Key Secondary Endpoint)Time-to-eventFrom randomisation until data cut-off (15 February 2013), Up to 30 monthsCox proportional-hazards model
Follow-up Analysis of Progression Free Survival (PFS) as Assessed by Central Independent ReviewTime-to-eventFrom randomisation until data cut-off (15 February 2013), Up to 30 monthsCox proportional-hazards model
Follow-up Analysis of Progression Free Survival (PFS) as Assessed by InvestigatorTime-to-eventFrom randomisation until data cut-off (15 February 2013), Up to 30 monthsCox proportional-hazards model
Objective Tumor ResponseBinaryFrom randomisation until data cut-off (15 February 2013), Up to 30 monthsLogistic regression
Disease ControlBinaryFrom randomisation until data cut-off (15 February 2013), Up to 30 monthsLogistic regression
Clinical Improvement.Time-to-eventFrom randomisation until data cut-off (15 February 2013), Up to 30 monthsCox proportional-hazards model
Quality of Life (QoL)Time-to-eventFrom randomisation until data cut-off (15 February 2013), Up to 30 monthsCox proportional-hazards model
Change From Baseline in Tumour SizeBinary as registered in the posted analysisFrom randomisation until data cut-off (15 February 2013), Up to 30 monthsANOVA

The registry supplies a detailed formal definition for the primary PFS endpoint. For several secondary outcomes, the ClinicalTrials.gov record provides the endpoint name, time frame, analysis population, statistical method, and effect measure but do not provide a separate narrative endpoint definition. This page therefore does not add definitions that are not present in the ClinicalTrials.gov record.

5. Statistical Methodology

The posted analyses use three main statistical families: Cox proportional-hazards models for time-to-event outcomes, logistic regression for binary outcomes, and ANOVA for change from baseline in tumour size. The Cox analyses are described as stratified analyses, with the registry specifying four baseline stratification factors.

Time-to-event
Cox proportional-hazards model for PFS, overall survival, clinical improvement, and the listed quality-of-life time-to-event outcomes.
Binary outcomes
Logistic regression for objective tumor response and disease control.
Continuous change
ANOVA for change from baseline in tumour size.
Effect measures
Hazard ratios for time-to-event analyses and odds ratios for binary outcomes.

Stratification in the Cox analyses

The posted Cox analyses were stratified by:

Stratification allows the baseline hazard to differ across the specified strata while estimating the treatment comparison across those strata. It is particularly relevant here because the registry explicitly states that the HR, confidence interval, and p-value for the posted Cox analyses were obtained from models stratified on these factors.

Cox model interpretation
HR = relative instantaneous event rate under the fitted proportional-hazards model

For this trial, the registry states that an HR below 1 favors nintedanib. That directional statement applies to the nintedanib-plus-pemetrexed versus placebo-plus-pemetrexed comparisons reported here.

Analysis populations

The primary PFS analysis and the posted time-to-event secondary analyses use the RS analysis population. Objective tumor response and its follow-up analysis use the Randomised Set. The ClinicalTrials.gov record does not define the abbreviation “RS,” so this page does not expand it beyond the label used in the posted analysis.

Endpoint familyAnalysis populationMethodEffect measure
Primary PFSRSCox proportional-hazards modelHazard ratio
OSRSCox proportional-hazards modelHazard ratio
Follow-up PFSRSCox proportional-hazards modelHazard ratio
Objective tumor responseRandomised SetLogistic regressionOdds ratio
Disease controlRSLogistic regressionOdds ratio
Clinical improvementRSCox proportional-hazards modelHazard ratio
Quality of lifeRSCox proportional-hazards modelHazard ratio
Change from baseline in tumour sizeRSANOVANo estimate reported in registry-reported analysis

6. Primary Result: Progression-Free Survival

The primary endpoint was progression-free survival as assessed by central independent review according to modified RECIST version 1.0 criteria. PFS was defined as the time from randomisation to progression or death, whichever occurred earlier. The primary analysis used the RS population and a stratified Cox proportional-hazards model.

Hazard ratio for progression or death

0.83

95% CI: 0.70–0.99   ·   P = 0.0435

Two-sided confidence interval; superiority hypothesis.

Primary PFS resultReported value
ComparisonNintedanib Plus Pemetrexed vs Placebo Plus Pemetrexed
Analysis populationRS
MethodStratified Cox proportional-hazards model
Hazard ratio0.83
95% CI0.70–0.99
P-value0.0435
HypothesisSuperiority
Data cutoff9 July 2012
Clinical Biostats interpretation

An HR of 0.83 means that, under the fitted stratified Cox model, the estimated instantaneous rate of progression or death in the nintedanib-plus-pemetrexed group was approximately 83% of the corresponding rate in the placebo-plus-pemetrexed group. Expressed as a relative model-based comparison, this corresponds to a 17% lower estimated hazard for progression or death.

The HR does not mean that 17% of participants avoided progression, that every participant experienced a 17% reduction, or that the median PFS differed by 17%. It is a relative time-to-event measure derived from a model.

The 95% CI of 0.70–0.99 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of individual patient effects. The interval is relatively close to 1 at its upper boundary, so the numerical precision of the estimated relative effect matters when interpreting the magnitude.

The p-value of 0.0435 addresses the statistical evidence against the null hypothesis under the prespecified analysis framework; it does not measure the size or clinical importance of the effect. A p-value is therefore not a substitute for the HR and its confidence interval.

Finally, the Cox interpretation depends on the proportional-hazards framework. The ClinicalTrials.gov record reports a Cox proportional-hazards model but do not provide a diagnostic assessment of the proportional-hazards assumption. The result should therefore be understood as the reported model-based comparison rather than as a guarantee that the hazard ratio was constant at every time point.

7. Secondary Result: Overall Survival

Overall survival was registered as a key secondary endpoint. The posted analysis used the RS population and a stratified Cox proportional-hazards model from randomisation through the 15 February 2013 data cutoff, with follow-up described as up to 30 months.

Hazard ratio for overall survival

1.01

95% CI: 0.85–1.21   ·   P = 0.8940

Two-sided confidence interval; superiority hypothesis.

Overall survival resultReported value
ComparisonNintedanib Plus Pemetrexed vs Placebo Plus Pemetrexed
Analysis populationRS
MethodStratified Cox proportional-hazards model
Hazard ratio1.01
95% CI0.85–1.21
P-value0.8940
Data cutoff15 February 2013
Time frameUp to 30 months
Clinical Biostats interpretation

An HR of 1.01 is very close to 1.00. Under the fitted stratified Cox model, the estimated instantaneous rate of death was therefore approximately the same between the two randomized comparison groups at the level summarized by the model.

The 95% CI of 0.85–1.21 spans 1.00. This means the reported estimate is compatible with a range of relative hazard differences in either direction under the stated model and sampling framework. The interval is more informative about uncertainty than the point estimate alone.

The p-value of 0.8940 is not an effect-size measure. It indicates little statistical evidence against the null hypothesis for this particular superiority analysis; it does not establish that the treatments are identical or that clinically meaningful differences are impossible.

As with the primary PFS analysis, the result is a model-based hazard ratio and should not be interpreted as a percentage of participants who benefited or as a statement about an individual patient's survival. The analysis also uses the registry-specified stratification factors and the RS population.

8. Secondary Result: Follow-up Central-Review PFS

A later analysis evaluated progression-free survival by central independent review using the 15 February 2013 cutoff. The same comparison and stratification factors were used in the posted Cox analysis.

Follow-up PFS hazard ratio

0.84

95% CI: 0.70–1.00   ·   P = 0.0506

Two-sided confidence interval; superiority hypothesis.

Clinical Biostats interpretation

The estimated HR of 0.84 corresponds to an approximately 16% lower estimated hazard of progression or death under the fitted Cox model for nintedanib plus pemetrexed relative to placebo plus pemetrexed.

The 95% CI of 0.70–1.00 reaches 1.00 at its upper boundary. The estimate therefore carries more uncertainty than the primary PFS estimate, and the interval illustrates why a point estimate should not be interpreted without its precision.

The p-value of 0.0506 is close to 0.05, but the p-value itself does not quantify the magnitude of the treatment effect. The appropriate reading is the combination of the HR, confidence interval, analysis population, data cutoff, and prespecified analysis framework.

9. Secondary Result: Follow-up Investigator-Assessed PFS

The registry also reports a follow-up PFS analysis based on investigator assessment.

Investigator-assessed PFS hazard ratio

0.86

95% CI: 0.73–1.02   ·   P = 0.0865

Two-sided confidence interval; superiority hypothesis.

Clinical Biostats interpretation

An HR of 0.86 corresponds to an approximately 14% lower estimated hazard of progression or death under the fitted model for the nintedanib-plus-pemetrexed group.

The 95% CI of 0.73–1.02 includes 1.00. Thus the confidence interval does not exclude a null relative hazard at the conventional level associated with a two-sided 95% interval. The result also differs somewhat from the central independent review estimate of 0.84, illustrating that the assessment source is part of the statistical definition of a time-to-event endpoint.

The p-value of 0.0865 should not be interpreted as a probability that the treatment effect is absent, nor as a measure of clinical importance. It is a test statistic summary under the specified statistical model.

10. Secondary Result: Objective Tumor Response

Objective tumor response was analyzed as a binary outcome using logistic regression. The registry provides two posted analyses: one based on central independent review and one based on investigator assessment.

AssessmentAnalysis populationOdds ratio95% CIP-value
Central independent reviewRandomised Set1.100.65–1.850.7279
Investigator assessmentRandomised Set1.150.75–1.760.5180

Central independent review

The OR of 1.10 means the estimated odds of objective tumor response were 1.10 times the corresponding odds in the comparator group under the logistic regression analysis. The 95% CI was 0.65–1.85 and the p-value was 0.7279.

Investigator assessment

The OR of 1.15 means the estimated odds of objective tumor response were 1.15 times the corresponding odds in the comparator group. The 95% CI was 0.75–1.76 and the p-value was 0.5180.

Clinical Biostats interpretation

An odds ratio is not the same as a risk ratio or a difference in response percentages. For example, an OR of 1.10 does not mean that the response probability increased by 10 percentage points or even necessarily by 10%.

Both reported confidence intervals include 1.00. The central-review interval of 0.65–1.85 is especially useful for showing how much uncertainty surrounds the estimate. The investigator-assessed interval of 0.75–1.76 likewise permits a range of possible relative odds in either direction.

The registry states that an odds ratio greater than 1 indicates a benefit to nintedanib for these response analyses. That directional convention is important because the interpretation of an odds ratio depends on which group is placed in the numerator.

11. Secondary Result: Disease Control

Disease control was also analyzed with logistic regression. Two analyses were posted: one based on central independent review and one based on investigator assessment.

AssessmentAnalysis populationOdds ratio95% CIP-value
Central independent reviewRS1.371.02–1.850.0387
Investigator assessmentRS1.290.95–1.750.1071
Clinical Biostats interpretation

For the central-review analysis, an OR of 1.37 indicates that the estimated odds of disease control were 1.37 times those in the placebo-plus-pemetrexed group under the logistic regression model. The 95% CI was 1.02–1.85 and the p-value was 0.0387.

The investigator-assessed analysis produced an OR of 1.29, with a 95% CI of 0.95–1.75 and a p-value of 0.1071. The difference between the two estimates illustrates why assessment method and analysis definition matter.

Neither OR should be translated directly into an absolute probability without the underlying response rates. An odds ratio is a relative comparison of odds, not a percentage-point difference.

12. Secondary Result: Clinical Improvement

Clinical Improvement. was analyzed as a time-to-event outcome in the RS population using a stratified Cox proportional-hazards model.

Clinical Improvement. hazard ratio

0.93

95% CI: 0.74–1.16   ·   P = 0.5068

Two-sided confidence interval; superiority hypothesis.

Clinical Biostats interpretation

The HR of 0.93 corresponds to an approximately 7% lower estimated hazard under the fitted model for the nintedanib-plus-pemetrexed group. The 95% CI of 0.74–1.16 spans 1.00, so the point estimate should not be read as establishing a directional treatment effect on its own.

The p-value of 0.5068 is a hypothesis-test result, not a measure of effect magnitude. Without the underlying event curves or absolute time-to-event summaries, the HR and CI are the principal numerical description reported by the registry for this endpoint.

13. Secondary Results: Quality of Life

The posted Quality of Life (QoL) analyses are time-to-event analyses evaluating time to deterioration of specific symptoms. The registry identifies cough, dyspnoea, and pain as the specific deterioration outcomes in the analysis notes.

QoL outcomeAnalysis95% CIP-value
Time to deterioration of coughHR 0.830.66–1.050.1181
Time to deterioration of dyspnoeaHR 0.930.77–1.120.4264
Time to deterioration of painHR 1.010.84–1.230.8929
Clinical Biostats interpretation

The cough analysis produced an HR of 0.83, corresponding to an approximately 17% lower estimated hazard of deterioration under the fitted model. Its 95% CI was 0.66–1.05 and its p-value was 0.1181.

The dyspnoea analysis produced an HR of 0.93, with a 95% CI of 0.77–1.12 and p-value of 0.4264. The pain analysis produced an HR of 1.01, with a 95% CI of 0.84–1.23 and p-value of 0.8929.

These are separate time-to-event outcomes. Their p-values should not be treated as interchangeable measures of overall quality-of-life benefit, and the ClinicalTrials.gov record does not provide a multiplicity procedure for combining these outcomes into a single confirmatory claim.

14. Secondary Result: Change From Baseline in Tumour Size

Change from baseline in tumour size was analyzed using ANOVA. The ClinicalTrials.gov record describes the outcome unit as percentage of change in tumor size in mm. No treatment-effect estimate or confidence interval is provided in the ClinicalTrials.gov record; the registry reports p-values for the two assessment approaches.

AssessmentMethodP-value
Central independent reviewANOVA0.1558
Investigator assessmentANOVA0.0565
Clinical Biostats interpretation

ANOVA tests differences in the modeled outcome between groups; in this registry record, the registry-reported numerical result is the p-value rather than a treatment-effect estimate with a confidence interval.

The central-review p-value was 0.1558, while the investigator-assessment p-value was 0.0565. These p-values should not be converted into effect sizes. Without a reported difference estimate and confidence interval, the ClinicalTrials.gov record does not quantify the magnitude or precision of the between-group change in tumour size.

15. Serious Adverse Events by Arm

The ClinicalTrials.gov record reports serious adverse events by treatment group as affected participants divided by participants at risk.

Treatment groupAffectedAt risk
Nintedanib Plus Pemetrexed104347
Placebo Plus Pemetrexed117357
Serious adverse events: affected participants / participants at risk
Nintedanib + pemetrexed
104 / 347
Placebo + pemetrexed
117 / 357

The affected-to-at-risk values are approximately 30.0% and 32.8% when expressed as simple proportions, but those percentages are derived from the registry-reported counts and are not additional registry-reported estimates. This page therefore keeps the primary safety presentation in the exact affected/at-risk format reported.

Safety interpretation: the safety counts are descriptive. They should not be treated as a formal hypothesis test or as an adjusted comparison of treatment risk because the ClinicalTrials.gov record does not provide a statistical analysis for serious adverse events.

16. Statistical Methods Explained

Why was a Cox proportional-hazards model used for PFS?

PFS is a time-to-event endpoint: each participant is followed until progression or death, or until censoring if an event has not occurred by the relevant observation endpoint. A Cox model uses the ordering and timing of events while accommodating censored observations. The resulting hazard ratio summarizes the relative event rate under the fitted model.

What does a hazard ratio of 0.83 mean?

An HR of 0.83 means that the modeled instantaneous rate of progression or death was estimated at approximately 83% of the comparator rate, conditional on the model. It can be described as an approximately 17% lower estimated hazard. It does not mean a 17% absolute improvement in PFS probability, nor does it mean that every participant experienced the same reduction.

Why was stratification used?

The registry reports that the Cox analyses were stratified by baseline ECOG PS, tumour histology, brain metastases at baseline, and prior bevacizumab treatment. Stratification permits the underlying baseline hazard to differ across those categories rather than forcing a single baseline hazard across all strata. The treatment effect is then estimated within the stratified modeling framework.

What is the difference between an odds ratio and a hazard ratio?

An odds ratio compares the odds of a binary outcome, such as objective tumor response or disease control. A hazard ratio compares instantaneous event rates over time for a time-to-event endpoint. An OR of 1.37 for disease control therefore cannot be interpreted in the same way as an HR of 0.83 for PFS.

Why does the p-value not tell us the size of the treatment effect?

A p-value summarizes the compatibility of the observed data with a null hypothesis under the specified statistical model. It is affected by sample size, event information, variability, and the magnitude of the observed difference. Effect size is better represented by the HR or OR together with its confidence interval, while absolute event probabilities or time summaries provide additional clinical context.

Why are central-review and investigator-assessed PFS results listed separately?

The registry explicitly distinguishes PFS assessed by central independent review from PFS assessed by the investigator. These are different assessment sources and therefore different operational measurements of the endpoint. Comparing their numerical results can be informative, but they should not be silently combined into one estimate.

Why is the primary PFS cutoff different from the follow-up cutoff?

The primary PFS endpoint was defined through the cutoff date of 9 July 2012. Several secondary analyses use a later data cutoff of 15 February 2013, with follow-up described as up to 30 months. A later cutoff incorporates additional follow-up information and therefore represents a different analysis dataset.

17. Early Stopping for Futility

The registry contains an important design caveat: recruitment for the study was stopped early based on the results of a pre defined futility analysis.

What futility means

A futility analysis is designed to assess whether continuing a study is unlikely to produce the prespecified objective under the assumptions of the trial design. It is conceptually different from an efficacy analysis, where the question is whether evidence is sufficiently strong to support a treatment effect.

Why early stopping matters

Stopping recruitment early changes the amount of information accumulated and can affect the precision and operating characteristics of the final evidence. The ClinicalTrials.gov record does not provide the futility boundary or its numerical operating characteristics.

The important statistical point is that the registry explicitly identifies the futility analysis as the reason recruitment was stopped early. This should be considered when interpreting the later posted estimates and the amount of information available for each endpoint.

18. Multiplicity and Multiple Posted Analyses

The registry reports 14 statistical analyses covering the primary PFS endpoint and multiple secondary outcomes. The ClinicalTrials.gov record identifies all analyses as superiority hypotheses and provide two-sided confidence intervals where an estimate and interval are reported.

Analysis familyNumber of posted analyses represented herePrimary statistical role
Primary PFS1Primary time-to-event comparison
Overall survival1Key secondary time-to-event comparison
Follow-up PFS2Central-review and investigator-assessed follow-up analyses
Objective tumor response2Central-review and investigator-assessed binary analyses
Disease control2Central-review and investigator-assessed binary analyses
Clinical Improvement.1Time-to-event analysis
Quality of Life (QoL)3Cough, dyspnoea, and pain deterioration analyses
Change From Baseline in Tumour Size2Central-review and investigator-assessed ANOVA analyses

Because many endpoints and assessment approaches were analyzed, a nominal p-value should not automatically be treated as evidence that a result is confirmatory in the same sense as a prespecified primary endpoint. The ClinicalTrials.gov record does not provide a multiplicity adjustment procedure or an alpha-allocation hierarchy for all 14 posted analyses, so this page does not invent one.

Interpretation principle: the primary PFS result should be read as the registered primary endpoint analysis. Secondary p-values are best interpreted together with their endpoint role, analysis population, confidence interval, assessment method, and the fact that multiple analyses were posted.

19. Randomization and Blinding

The trial is registered as randomized and double masked, with a parallel design. Randomization is central to the causal interpretation of a treatment comparison because assignment, rather than observed outcome, determines which comparison group a participant belongs to.

Randomization

Randomized allocation is intended to balance known and unknown prognostic factors on average across treatment groups, subject to chance variation.

Double masking

Double masking can reduce the opportunity for knowledge of treatment assignment to influence participant or investigator behavior and assessment. The registry does not provide additional masking-role details in the ClinicalTrials.gov record.

The distinction between randomized assignment and the later analysis population is important. A randomized trial's treatment comparison is anchored in the assigned groups, while the registry's individual statistical analyses specify either RS or the Randomised Set as their analysis population.

20. Understanding the Primary PFS Estimate

Relative effect

The primary HR of 0.83 is a relative model-based measure. It indicates a lower estimated hazard of progression or death for nintedanib plus pemetrexed under the reported stratified Cox model.

Absolute effect

The ClinicalTrials.gov record does not provide median PFS values, PFS rates at a specified time point, or absolute event probabilities. Therefore, the primary HR should not be converted into an invented absolute benefit.

Precision

The 95% CI of 0.70–0.99 is essential because it shows the uncertainty surrounding the point estimate. The upper boundary is close to 1.00, so the magnitude of the estimated effect should be interpreted with attention to that uncertainty.

Statistical evidence

The two-sided p-value of 0.0435 provides the hypothesis-test component of the result. It should be interpreted as evidence relative to the specified null hypothesis, not as a measure of the size, importance, or probability of the treatment effect.

21. Comparing the Posted Time-to-Event Analyses

EndpointHR95% CIP-valueCutoff
Primary central-review PFS0.830.70–0.990.04359 July 2012
Overall survival1.010.85–1.210.894015 February 2013
Follow-up central-review PFS0.840.70–1.000.050615 February 2013
Follow-up investigator-assessed PFS0.860.73–1.020.086515 February 2013
Clinical Improvement.0.930.74–1.160.506815 February 2013
QoL: cough deterioration0.830.66–1.050.118115 February 2013
QoL: dyspnoea deterioration0.930.77–1.120.426415 February 2013
QoL: pain deterioration1.010.84–1.230.892915 February 2013

The table demonstrates why a trial cannot be summarized adequately by a single p-value. The primary PFS analysis has an HR below 1 with a 95% CI that remains below 1, while the later central-review PFS estimate is slightly closer to the null and has a confidence interval reaching 1.00. Overall survival is centered very close to 1.00. The other time-to-event endpoints have their own estimates and uncertainty.

These differences do not require a contradiction: the endpoints measure different events, use different cutoffs, and in some cases use different assessment definitions. A statistical analysis should preserve those distinctions rather than collapsing them into one overall treatment-effect statement.

22. Limitations

23. Why This Trial Matters Statistically

LUME-Lung 2 is a useful statistical teaching case because the registry combines a randomized, double-blind phase 3 design with a time-to-event primary endpoint, stratified Cox modeling, binary response analyses, ANOVA, multiple assessment sources, and a prespecified futility-based early stopping decision.

ConceptHow it appears in LUME-Lung 2
RandomizationThe trial is registered as randomized with a parallel design.
BlindingThe trial is registered as double masked.
Time-to-event endpointPrimary PFS is defined from randomisation to progression or death.
Kaplan-Meier estimationThe registered PFS definition states that median, 25th and 75th percentiles are calculated from an unadjusted Kaplan-Meier curve.
Cox modelPrimary PFS and multiple secondary time-to-event outcomes use Cox proportional-hazards models.
Stratified analysisCox analyses are stratified by ECOG PS, histology, brain metastases, and prior bevacizumab treatment.
Hazard ratioThe primary PFS effect measure is HR 0.83 with a 95% CI of 0.70–0.99.
Logistic regressionObjective tumor response and disease control are analyzed with logistic regression.
Odds ratioBinary analyses report odds ratios such as 1.10, 1.15, 1.37, and 1.29.
ANOVAChange from baseline in tumour size is analyzed using ANOVA.
Multiple endpointsThe registry reports 14 statistical analyses across primary and secondary outcomes.
Early stoppingRecruitment was stopped early based on a pre defined futility analysis.

24. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

25. Related Statistical Calculators

Use these calculator pathways to explore the statistical quantities that appear in this trial:

26. Sources

Continue through the Clinical Biostats statistical pathway

Use the trial as a practical example of time-to-event analysis, hazard ratios, logistic regression, odds ratios, ANOVA, confidence intervals, and statistical interpretation.

27. Record Summary

LUME-Lung 2 provides a compact example of how a randomized phase 3 trial can generate several distinct statistical questions from the same treatment comparison. The registered primary endpoint was progression-free survival assessed by central independent review, defined as time from randomisation to progression or death. Its posted analysis used a stratified Cox proportional-hazards model and reported HR 0.83, 95% CI 0.70–0.99, P = 0.0435.

The secondary analyses show why endpoint-specific interpretation matters. Overall survival had an HR of 1.01 with a 95% CI of 0.85–1.21 and P = 0.8940. Follow-up central-review PFS had HR 0.84, 95% CI 0.70–1.00, and P = 0.0506, while investigator-assessed PFS had HR 0.86, 95% CI 0.73–1.02, and P = 0.0865. Binary outcomes were evaluated using logistic regression, and change from baseline in tumour size was analyzed with ANOVA.

The statistical story also includes an important design feature: recruitment was stopped early based on a pre defined futility analysis. Because the registry reports 14 statistical analyses across several endpoints and assessment approaches, individual p-values should be interpreted in their proper endpoint and multiplicity context rather than treated as interchangeable evidence.

Clinical Biostats methodology: A trial-results page should distinguish the registered endpoint, analysis population, statistical model, effect measure, uncertainty interval, and p-value. The goal is not simply to reproduce numerical results, but to explain what each statistical quantity means and what it does not establish.