← Clinical Trials
Non-Small Cell Lung Cancer Phase 3 Time-to-Event Analysis NCT00805194

LUME-Lung 1: Complete Statistical Analysis of BIBF 1120 in Second-Line Non-Small Cell Lung Cancer

An independent statistical analysis of the randomized phase 3 LUME-Lung 1 trial comparing BIBF 1120 plus docetaxel with placebo plus docetaxel in second-line non-small cell lung cancer, with emphasis on progression-free survival, overall survival, response, disease control, quality of life, and the statistical methods used.

Trial status: Completed  ·  Enrollment: 1314  ·  Primary completion: 2 November 2010
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results on this page are restricted to the ClinicalTrials.gov record for NCT00805194.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

LUME-Lung 1 was a completed, randomized, double-blind, parallel phase 3 trial evaluating BIBF 1120 plus docetaxel versus placebo plus docetaxel in second-line non-small cell lung cancer. The trial enrolled 1314 participants and registered one primary time-to-event endpoint: progression-free survival as assessed by central independent review.

1314
Enrollment
Randomized trial
2
Treatment arms
Parallel design
0.79
Primary PFS HR
95% CI 0.68–0.92
0.0019
Primary PFS p-value
Two-sided
FeatureLUME-Lung 1
PhasePhase 3
ConditionCarcinoma, Non-Small-Cell Lung
Clinical settingSecond-line non-small cell lung cancer
DesignRandomized, double-blind, parallel
AllocationRandomized
Primary purposeTreatment
Enrollment1314
Arms2
Primary endpointProgression Free Survival (PFS) as Assessed by Central Independent Review
Primary endpoint typeTime-to-event
Primary analysis populationRandomised Set
Lead sponsorBoehringer Ingelheim
Sponsor typeIndustry
Trial datesStart: 2008-12-03  ·  Primary completion: 2010-11-02

2. Clinical Question

The central statistical question was whether adding BIBF 1120 to docetaxel changed progression-free survival compared with placebo plus docetaxel in participants with second-line non-small cell lung cancer.

Population

Participants with carcinoma, non-small-cell lung, in the second-line treatment setting.

Intervention

BIBF 1120 plus docetaxel.

Comparator

Placebo plus docetaxel.

Primary question

Does BIBF 1120 plus docetaxel improve centrally assessed progression-free survival relative to placebo plus docetaxel?

3. Trial Design

01
Randomize1314 participants
02
Two armsBIBF 1120 or placebo
03
DocetaxelCombination treatment
04
AssessPFS and secondary outcomes
05
AnalyzeTime-to-event and binary endpoints
INTERVENTION ARM

BIBF 1120 plus docetaxel

  • BIBF 1120
  • Docetaxel
  • Randomized assignment
CONTROL ARM

Placebo plus docetaxel

  • Placebo
  • Docetaxel
  • Randomized assignment
Allocation
Randomized.
Masking
Double.
Design model
Parallel.
Primary purpose
Treatment.

The combination of randomization and double masking is important statistically. Randomization establishes the treatment comparison, while masking can reduce the potential for knowledge of treatment assignment to influence assessment or other trial conduct. The registry does not provide additional allocation-ratio information in the ClinicalTrials.gov record.

4. Endpoints

The registry lists one primary endpoint and multiple secondary outcomes. The primary endpoint is a time-to-event outcome; several secondary outcomes use either Cox proportional-hazards models, logistic regression, or ANOVA.

EndpointRegistry definition / time frameAnalysis
Primary: Progression Free Survival (PFS) as Assessed by Central Independent Review From randomisation until cut-off date 2 November 2010, when 713 PFS events were observed. PFS is defined as the duration of time from date of randomisation to date of progression or death, whichever occurs earlier, according to modified RECIST v1.0. Stratified Cox proportional-hazards model
Overall Survival (Key Secondary Endpoint) From randomisation until cut-off date 15 February 2013, approximately 48 months or 1151 deaths among all patients. The overall alpha level followed a Lan-DeMets spending function with O'Brien-Fleming shape parameter to preserve an overall 2-sided alpha level of 0.05. HR below 1 favors nintedanib Cox proportional-hazards model with hierarchical testing
Follow-up Analysis of PFS by Central Independent Review From randomisation until cut-off date 15 February 2013. Cox proportional-hazards model
Follow-up Analysis of PFS by Investigator From randomisation until cut-off date 15 February 2013. Cox proportional-hazards model
Objective Tumour Response From randomisation until cut-off date 15 February 2013. Logistic regression
Disease Control From randomisation until cut-off date 15 February 2013. Logistic regression
Clinical Improvement From randomisation until cut-off date 15 February 2013. Cox proportional-hazards model
Quality of Life (QoL) From randomisation until cut-off date 15 February 2013. Cox proportional-hazards model
Change From Baseline in Tumour Size From randomisation until cut-off date 15 February 2013; outcome unit is percentage of change in tumor size in mm. ANOVA

5. Statistical Methodology

Primary time-to-event analysis

The primary PFS analysis used a Cox proportional-hazards model in the Randomised Set. The analysis was stratified by baseline ECOG performance status, tumour histology, brain metastases at baseline, and prior treatment with bevacizumab.

Primary model
h(t|X) = h0(t) exp(βX)

The Cox model estimates a relative hazard through the regression coefficient. In this trial, the reported effect measure was the hazard ratio comparing BIBF 1120 plus docetaxel with placebo plus docetaxel.

Hazard ratio

The registry reports the hazard ratio as the primary effect measure for PFS and for several secondary time-to-event analyses. An HR below 1 favors the BIBF 1120-containing treatment according to the registry analysis notes.

Primary effect measure
HR = estimated hazard in BIBF 1120 + docetaxel ÷ estimated hazard in placebo + docetaxel

The hazard ratio is a relative time-to-event measure. It should not be interpreted as a percentage of participants who progressed, nor as an absolute difference in PFS probability.

Logistic regression

Objective tumour response and disease control were analyzed using logistic regression, with odds ratios as the effect measure. The registry identifies the central independent review and investigator assessment separately for these endpoints.

ANOVA

Change from baseline in tumour size was analyzed using ANOVA. The registry reports two analyses of this endpoint, one based on central independent review and one based on investigator assessment.

Stratified analysis

The primary PFS Cox model incorporated stratification by baseline ECOG performance status (0 vs 1), tumour histology (squamous vs non-squamous), brain metastases at baseline (yes vs no), and prior treatment with bevacizumab (yes vs no). This is important because the hazard ratio was not simply an unstratified comparison of the two randomized groups.

6. Results: Primary Endpoint

Progression-Free Survival as Assessed by Central Independent Review

Primary PFS hazard ratio

0.79

95% CI: 0.68–0.92   ·   P = 0.0019

713 PFS events were observed at the 2 November 2010 cutoff.

Primary endpointResult
Analysis populationRandomised Set
ComparisonNintedanib Plus Docetaxel vs Placebo Plus Docetaxel
MethodStratified Cox proportional-hazards model
Hazard ratio0.79
95% CI0.68–0.92
P-value0.0019
Hypothesis typeSuperiority
Clinical Biostats interpretation

The reported HR of 0.79 means that, under the fitted stratified Cox model, the estimated instantaneous rate of progression or death was 21% lower in the BIBF 1120-containing group than in the placebo-containing group. The 21% figure is a direct interpretation of 1 − 0.79; it is a relative hazard interpretation, not an absolute reduction in the probability of progression or death.

The HR does not mean that 21% fewer participants progressed, that every participant experienced a 21% reduction in risk, or that PFS duration increased by 21%.

The 95% CI of 0.68–0.92 describes statistical uncertainty around the estimated hazard ratio under the model and analysis framework. It does not describe the range of treatment effects experienced by individual patients.

The p-value of 0.0019 addresses the statistical evidence against the null hypothesis under the prespecified superiority analysis. A p-value is not a measure of effect size and does not tell us how clinically important the estimated HR is.

Because the estimate comes from a Cox proportional-hazards model, its interpretation depends on the model framework, censoring, and the proportional-hazards assumption. The registry does not provide enough information in the ClinicalTrials.gov record to independently assess the proportional-hazards assumption.

What the primary PFS result establishes statistically

The registry reports a formal superiority analysis with an HR of 0.79, a two-sided 95% CI of 0.68–0.92, and a p-value of 0.0019. The confidence interval lies below 1, consistent with the direction of the registry's stated interpretation that an HR below 1 favors nintedanib.

Importantly, the analysis is based on the Randomised Set. This anchors the treatment comparison to randomized assignment rather than selectively comparing only participants who remained on treatment or who had complete follow-up.

7. Secondary Endpoint Results

Overall Survival: Hierarchical Key Secondary Analysis

Overall survival was analyzed from randomisation through the 15 February 2013 cutoff, described in the registry as approximately 48 months or 1151 deaths among all patients. The registry reports three Cox-model estimates corresponding to a fixed sequence of hypotheses.

Hierarchical hypothesisHazard ratio95% CIP-value
Patients with adenocarcinoma and <9 months since start of first-line therapy0.750.60–0.920.0073
Patients with adenocarcinoma0.830.70–0.990.0359
All patients0.940.83–1.050.2720
Hierarchy matters. The registry states that the hypotheses were tested in a fixed sequence: first patients with adenocarcinoma and <9 months since start of first-line therapy, then patients with adenocarcinoma, and then all patients. Each hypothesis could be tested at the prespecified alpha level only if the preceding null hypothesis had been rejected. The ClinicalTrials.gov record does not specify the numerical alpha level.
Clinical Biostats interpretation

For the first hierarchical OS comparison, the HR of 0.75 corresponds to a 25% lower estimated instantaneous hazard of death under the Cox model. Its 95% CI of 0.60–0.92 quantifies uncertainty around that estimate, while the p-value of 0.0073 addresses evidence against the corresponding null hypothesis.

For the adenocarcinoma population in the second step, the HR of 0.83 corresponds to a 17% lower estimated instantaneous hazard of death. The 95% CI is 0.70–0.99, with p = 0.0359.

For all patients, the reported HR is 0.94, with a 95% CI of 0.83–1.05 and p = 0.2720. This estimate is closer to 1 than the two preceding estimates. The confidence interval spans 1, so the ClinicalTrials.gov record does not provide evidence of a statistically significant superiority comparison for this final hypothesis under the reported two-sided test.

These three results should not be read as three independent hypothesis tests. The fixed-sequence design makes the order of testing part of the statistical interpretation.

Follow-up Analysis of Progression-Free Survival: Central Independent Review

Follow-up PFS hazard ratio

0.85

95% CI: 0.75–0.96   ·   P = 0.0070

Cutoff: 15 February 2013

Clinical Biostats interpretation

The follow-up HR of 0.85 corresponds to a 15% lower estimated instantaneous hazard of progression or death under the fitted Cox model. The 95% CI of 0.75–0.96 indicates uncertainty around the estimate, while p = 0.0070 describes the evidence against the superiority null hypothesis under the reported analysis.

This is a follow-up PFS analysis, not the original primary PFS analysis. It therefore should not be silently substituted for the primary endpoint result of HR 0.79.

Follow-up Analysis of PFS by Investigator Assessment

Investigator-assessed PFS hazard ratio

0.82

95% CI: 0.73–0.93   ·   P = 0.0012

Cutoff: 15 February 2013

Clinical Biostats interpretation

The investigator-assessed HR of 0.82 corresponds to an 18% lower estimated instantaneous hazard of progression or death. Its 95% CI of 0.73–0.93 provides the uncertainty interval around the estimate, and p = 0.0012 is the reported hypothesis-test result.

The comparison is not identical to the centrally reviewed PFS analysis because the endpoint assessment source differs. The two estimates therefore provide related but distinct statistical summaries.

Objective Tumour Response

AssessmentOdds ratio95% CIP-value
Central independent review1.340.76–2.390.3067
Investigator assessment1.410.96–2.080.0761

Both analyses used logistic regression in the Randomised Set. The registry states that an odds ratio greater than 1 indicates a benefit to nintedanib.

Clinical Biostats interpretation

The central-review OR of 1.34 indicates that the estimated odds of objective tumour response were higher in the BIBF 1120-containing group, but the 95% CI of 0.76–2.39 includes 1 and the reported p-value is 0.3067. The investigator-assessed OR of 1.41 has a 95% CI of 0.96–2.08 and p = 0.0761.

An odds ratio is not a risk ratio or a percentage-point difference in response rates. It also does not describe the duration of response or time until progression. Those are different clinical quantities.

Disease Control

AssessmentOdds ratio95% CIP-value
Central independent review1.681.35–2.09<0.0001
Investigator assessment1.641.31–2.05<0.0001

The registry reports odds ratios above 1 as favoring nintedanib for these disease-control analyses.

Clinical Biostats interpretation

The central-review OR of 1.68 indicates higher estimated odds of disease control in the BIBF 1120-containing group relative to the control group. The 95% CI of 1.35–2.09 does not include 1. The investigator-assessed OR is 1.64, with a 95% CI of 1.31–2.05. Both p-values are reported as <0.0001.

These results describe a binary disease-control outcome. They do not establish the magnitude of the difference in absolute response probability because the underlying response proportions are not reported in the ClinicalTrials.gov record.

Clinical Improvement

Clinical improvement hazard ratio

1.03

95% CI: 0.87–1.21   ·   P = 0.7282

The registry classifies clinical improvement as a time-to-event endpoint and analyzes it using a Cox proportional-hazards model.

Quality of Life

Quality-of-life time-to-deterioration measureHR95% CIP-value
Time to deterioration of cough0.900.77–1.050.1858
Time to deterioration of dyspnoea1.050.91–1.200.5203
Time to deterioration of pain0.950.82–1.090.4373

The registry identifies these as Cox-model analyses of time to deterioration. For cough and pain, the registry states that HR below 1 favors nintedanib; for dyspnoea it likewise defines an HR below 1 as favoring nintedanib.

Change From Baseline in Tumour Size

AssessmentMethodP-value
Central independent reviewANOVA<0.0001
Investigator assessmentANOVA<0.0001

The registry reports the outcome unit as percentage of change in tumor size in mm. It does not provide an effect estimate or confidence interval for these ANOVA analyses in the ClinicalTrials.gov record.

Interpretation boundary: because the registry supplies p-values but no corresponding group means, mean difference, or confidence interval for change from baseline in tumour size, the statistical evidence can be described but the magnitude of the estimated difference cannot be reconstructed from the ClinicalTrials.gov record.

8. Statistical Methods Explained

Why was a Cox proportional-hazards model used?

PFS, overall survival, clinical improvement, and the reported quality-of-life deterioration outcomes are time-to-event endpoints. A Cox model is designed to compare event hazards while accommodating right-censored observations, such as participants who have not experienced the event by the end of their available follow-up.

What does an HR of 0.79 mean?

An HR of 0.79 means that the fitted model estimates the instantaneous event hazard in the BIBF 1120-containing group to be 0.79 times that in the control group, under the model and analysis framework. Equivalently, 1 − 0.79 = 0.21, or a 21% lower estimated hazard. It does not mean that 21% of patients avoided progression.

Why was the primary PFS model stratified?

The registry reports stratification by baseline ECOG performance status, tumour histology, brain metastases at baseline, and prior treatment with bevacizumab. Stratification allows the treatment comparison to account for these prespecified factors in the time-to-event analysis rather than treating all participants as belonging to one homogeneous risk stratum.

Why is the confidence interval important?

A point estimate such as HR 0.79 is only one estimate from the observed data. The 95% CI of 0.68–0.92 communicates the statistical uncertainty around that estimate. A confidence interval is therefore more informative about precision than the p-value alone.

Why does the p-value not measure effect size?

The p-value describes the compatibility of the observed data with the null hypothesis under the specified statistical test. It does not quantify the size of the treatment effect. The HR, OR, and their confidence intervals provide the effect-size information in the reported analyses.

Why is the overall-survival analysis hierarchical?

The registry describes a fixed sequence of three hypotheses: patients with adenocarcinoma and less than 9 months since start of first-line therapy, patients with adenocarcinoma, and then all patients. Each hypothesis could be tested at the prespecified alpha level only after rejection of the preceding null hypothesis. This sequencing is a multiplicity-control strategy because it prevents the three hypotheses from being treated as three unrelated opportunities for a positive finding.

Why are the central-review and investigator-assessed analyses separate?

They use different assessment sources. The registry separately reports PFS by central independent review and by investigator assessment, as well as objective tumour response and disease control based on central review versus investigator assessment. Agreement between related analyses can provide useful consistency information, but the estimates should remain labeled according to their assessment method.

9. Multiplicity and Hierarchical Testing

Multiplicity is directly relevant to the overall-survival key secondary endpoint. The registry specifies a fixed sequence of statistical hypotheses:

SequencePopulationRule
1Patients with adenocarcinoma and <9 months since start of first-line therapyCould be tested at the prespecified alpha level.
2Patients with adenocarcinomaCould be tested only if the preceding null hypothesis was rejected.
3All patientsCould be tested only if the preceding null hypothesis was rejected.

This is different from conducting three unrelated two-sided tests and simply counting how many have p-values below a chosen threshold. In a hierarchical strategy, the order of hypotheses is part of the prespecified inferential procedure.

Important: the ClinicalTrials.gov record identifies the hierarchical testing structure and refer to a prespecified alpha level, but they do not give the numerical alpha value. No numerical alpha should therefore be inferred from the reported p-values.

10. Stratification and the Cox Model

The primary PFS analysis was stratified by four baseline characteristics:

Stratification factorLevels reported
Baseline ECOG performance status0 vs 1
Tumour histologySquamous vs non-squamous
Brain metastases at baselineYes vs no
Prior treatment with bevacizumabYes vs no

These variables are not themselves treatment effects. Their role in the reported primary analysis is to define the strata used by the proportional-hazards model. The resulting HR of 0.79 therefore represents the treatment comparison from the stratified model specified in the registry record.

A useful distinction
Treatment effect ≠ stratification factor

A stratification variable helps structure the analysis; it does not mean that the trial was designed to estimate a separate treatment effect for every level of that variable.

11. Randomization and Analysis Population

The trial was randomized and the primary PFS analysis was performed in the Randomised Set. This is statistically important because randomized assignment defines the principal comparison between the two treatment strategies.

Analyzing the Randomised Set avoids redefining the comparison according to post-randomization treatment experience. In a randomized efficacy analysis, preserving the original treatment assignment maintains the connection between randomization and the estimated treatment contrast.

The registry does not provide additional definitions of per-protocol, as-treated, or safety populations in the ClinicalTrials.gov record beyond the reported serious-adverse-event counts by arm.

12. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk.

Safety measureAffected / at risk
Nintedanib Plus Docetaxel224 / 652
Placebo Plus Docetaxel206 / 655
Serious adverse events: affected participants
Nintedanib + docetaxel
224
Placebo + docetaxel
206

The reported safety data use Nintedanib Plus Docetaxel terminology for the intervention arm, whereas the ClinicalTrials.gov record identifies the investigational intervention as BIBF 1120 plus docetaxel. The ClinicalTrials.gov record does not provide additional safety estimates, confidence intervals, exposure-adjusted rates, or formal between-arm hypothesis tests.

Do not infer a p-value from these counts. Serious adverse events are reported here as affected participants divided by those at risk. A formal statistical comparison would require an explicitly specified analysis framework, and the ClinicalTrials.gov record does not report one for these safety counts.

13. Understanding the Secondary Results as a Statistical System

LUME-Lung 1 illustrates why a clinical trial should not be reduced to a single p-value. Different endpoints answer different questions and use different effect measures.

Endpoint familyEffect measureQuestion answered
PFSHazard ratioHow does the estimated time-to-progression-or-death hazard compare between treatment groups?
Overall survivalHazard ratioHow does the estimated hazard of death compare between treatment groups?
Objective tumour responseOdds ratioHow do the odds of the binary response outcome compare?
Disease controlOdds ratioHow do the odds of disease control compare?
Clinical improvementHazard ratioHow does the time-to-event hazard compare?
Quality of lifeHazard ratioHow does time to deterioration compare?
Change in tumour sizeANOVA / p-valueIs there statistical evidence of a difference in change from baseline?

This distinction prevents a common statistical mistake: treating every endpoint as if it were measuring the same outcome. An HR, OR, and ANOVA p-value have different meanings and cannot be compared numerically as though they were interchangeable effect measures.

14. Planned Analysis vs Reported Analysis

For LUME-Lung 1, results are posted and a formal statistical analysis is reported for the primary endpoint. The appropriate distinction is therefore between the registered endpoint definition and the analysis actually reported.

Registered endpoint

PFS from randomisation to progression or death, whichever occurred earlier, assessed by central independent review according to modified RECIST v1.0.

Reported analysis

Stratified Cox proportional-hazards model in the Randomised Set, with HR 0.79, 95% CI 0.68–0.92, and p = 0.0019.

This correspondence between endpoint type and analysis method is a useful example of prespecified statistical planning: a time-to-event endpoint is paired with a survival-analysis framework rather than a simple comparison of proportions.

15. Interpreting Confidence Intervals Correctly

Several LUME-Lung 1 results illustrate how confidence intervals communicate more than a binary significant/non-significant label.

Result95% CIWhat the interval communicates
Primary PFS HR 0.790.68–0.92Uncertainty around the estimated relative hazard from the stratified Cox model.
OS HR 0.750.60–0.92Uncertainty around the first hierarchical OS estimate.
OS HR 0.830.70–0.99Uncertainty around the second hierarchical OS estimate.
OS HR 0.940.83–1.05Uncertainty includes the null value of 1 for the all-patient hypothesis.
Disease-control OR 1.681.35–2.09Uncertainty around the estimated odds ratio from central review.

For hazard ratios and odds ratios, 1 is the usual null value. An interval entirely below 1 for an HR is consistent with the treatment direction specified by the registry; an interval entirely above 1 for an OR is consistent with higher odds in the BIBF 1120-containing group. Neither interpretation removes the need to consider the endpoint definition, analysis population, and multiplicity structure.

16. Important Limitations and Interpretation Issues

17. Why This Trial Matters Statistically

LUME-Lung 1 is a useful teaching case because it combines several core clinical-trial statistical methods within a single randomized phase 3 study. The primary endpoint is a time-to-event outcome analyzed with a stratified Cox model, while secondary endpoints use Cox regression, logistic regression, and ANOVA.

ConceptHow it appears in LUME-Lung 1
Randomization1314 participants assigned in a randomized parallel phase 3 trial.
Double maskingthe ClinicalTrials.gov record identifies the study as double masked.
Time-to-event endpointPrimary PFS endpoint defined from randomisation to progression or death.
Kaplan-Meier conceptThe registered endpoint definition states that median, 25th and 75th percentiles are calculated from an unadjusted Kaplan-Meier analysis.
Cox modelUsed for primary PFS and multiple secondary time-to-event analyses.
Hazard ratioPrimary PFS HR 0.79; additional OS and time-to-event HRs are also reported.
Stratified analysisPrimary PFS model stratified by four baseline factors.
Logistic regressionUsed for objective tumour response and disease control.
Odds ratioReported for objective tumour response and disease control.
ANOVAUsed for change from baseline in tumour size.
Multiplicity adjustmentOverall survival used a fixed sequence of hypotheses.
Interim analysis / alpha spendingThe ClinicalTrials.gov record identifies interim analysis / alpha spending as a concept in the OS analyses.

18. Statistical Interpretation of the Primary Result

Relative effect

The primary PFS HR of 0.79 indicates a 21% lower estimated instantaneous hazard of progression or death for BIBF 1120 plus docetaxel relative to placebo plus docetaxel, under the reported stratified Cox model.

Precision

The 95% CI of 0.68–0.92 gives the statistical uncertainty around the estimated HR. It is narrower than a very imprecise interval would be, but it still does not identify the treatment effect for any individual participant.

Evidence against the null

The reported two-sided p-value of 0.0019 indicates strong statistical evidence against the null hypothesis under the specified superiority analysis. It is not itself an estimate of clinical magnitude.

What cannot be concluded from the HR alone

The HR does not provide an absolute PFS probability, median PFS, number needed to treat, or percentage of participants who benefited. Those quantities require additional information not reported in the ClinicalTrials.gov record.

19. Primary Endpoint in Context

The primary PFS analysis is especially instructive because the registry supplies all of the essential components of a model-based treatment comparison: a prespecified time-to-event endpoint, a defined analysis population, a Cox regression method, stratification factors, an effect measure, a confidence interval, and a p-value.

Endpoint

PFS from randomisation to progression or death, whichever occurred earlier.

Analysis

Stratified Cox proportional-hazards model.

Effect

HR 0.79, with 95% CI 0.68–0.92.

Evidence

Two-sided p = 0.0019, based on 713 observed PFS events at the cutoff.

This structure is a useful template for reading other oncology trials: first identify exactly what constitutes the event, then identify who was analyzed, then identify the statistical model, and only after that interpret the effect estimate and p-value.

20. Related Tutorials

Learn more about the methods used in this trial:

21. Related Statistical Calculators

22. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical concepts behind randomized trials, survival analysis, regression, confidence intervals, and multiple testing.

23. Record Summary

LUME-Lung 1 is a randomized, double-blind, parallel phase 3 trial with 1314 participants and two treatment arms. Its registered primary endpoint was progression-free survival assessed by central independent review, defined as time from randomisation to progression or death, whichever occurred earlier. The primary analysis used a stratified Cox proportional-hazards model in the Randomised Set and reported an HR of 0.79 (95% CI 0.68–0.92; p = 0.0019) at the 2 November 2010 cutoff, when 713 PFS events had been observed.

The secondary analyses show how a single randomized trial can require several statistical frameworks. Overall survival was analyzed with Cox regression and a fixed hierarchical testing sequence. Objective tumour response and disease control were analyzed with logistic regression and odds ratios. Change from baseline in tumour size was analyzed with ANOVA. Clinical improvement and quality-of-life deterioration outcomes were analyzed as time-to-event outcomes using Cox models.

The most important statistical lesson is that the interpretation of each result depends on its endpoint definition, analysis population, model, effect measure, confidence interval, and multiplicity structure. The primary HR of 0.79 is therefore most appropriately understood as a model-based relative treatment effect for centrally assessed PFS, not as an absolute probability difference or a universal measure of individual patient benefit.

Clinical Biostats methodology: A trial-results page should not merely repeat reported numbers. The goal is to reconstruct the statistical structure of the trial—what was measured, how it was analyzed, what the effect estimate means, and what it does not mean—while keeping reported evidence separate from educational interpretation.