← Clinical Trials
Non-Small Cell Lung Cancer Phase 3 Time-to-Event Analysis NCT01121393

LUX-Lung 6: Complete Statistical Analysis of Afatinib in Non-Small Cell Lung Cancer

An independent statistical review of the randomized phase 3 LUX-Lung 6 trial comparing afatinib 40 mg with gemcitabine-cisplatin chemotherapy in first-line non-small cell lung cancer, with progression-free survival as the registered primary endpoint.

Trial status: Completed  ·  Enrollment: 364  ·  Sponsor: Boehringer Ingelheim
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

LUX-Lung 6 was a randomized, open-label, parallel-group phase 3 treatment trial with 364 participants. It compared BIBW 2992 (afatinib) with gemcitabine-cisplatin chemotherapy in first-line non-small cell lung cancer, with progression-free survival designated as the registered primary endpoint.

364
Enrollment
Total participants
2
Treatment Arms
Parallel design
0.281
Primary PFS HR
95% CI 0.203–0.389
<0.0001
Primary PFS P-value
Stratified log-rank
FeatureLUX-Lung 6
Trial nameLUX-Lung 6
ClinicalTrials.gov identifierNCT01121393
Brief titleBIBW 2992 (Afatinib) vs Gemcitabine-cisplatin in 1st Line Non-Small Cell Lung Cancer (NSCLC)
PhasePhase 3
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment364
Arms2
InterventionsGemcitabine + Cisplatin; BIBW 2992
Primary endpointProgression-free Survival
Primary endpoint typeTime-to-event
Hypothesis typeSuperiority
Trial statusCompleted
Start2010-04-19
Primary completion2017-11-23
Lead sponsorBoehringer Ingelheim
Sponsor typeIndustry

2. Clinical Question

The statistical question was whether treatment with BIBW 2992 (afatinib) differed from gemcitabine-cisplatin chemotherapy with respect to the registered primary endpoint, progression-free survival, in patients with non-small cell lung cancer and adenocarcinoma represented in the trial record.

Population

Participants enrolled in the first-line non-small cell lung cancer trial described by the registry, including the listed conditions of carcinoma, non-small-cell lung and adenocarcinoma.

Intervention

BIBW 2992 (afatinib), with the reported primary efficacy comparison identified as afatinib 40 mg.

Comparator

Gemcitabine / cisplatin chemotherapy.

Primary question

Does afatinib produce a different progression-free survival experience from gemcitabine-cisplatin chemotherapy under the prespecified superiority framework?

Registry scope: The ClinicalTrials.gov record does not provide a complete baseline-characteristics table or treatment-specific enrollment counts for the two randomized arms. Those details are therefore not reconstructed here.

3. Trial Design

01
Randomize 364 participants
02
Two arms Afatinib vs chemotherapy
03
Follow-up Tumour assessments
04
Time-to-event PFS and OS
05
Additional outcomes Response, shrinkage, HRQOL
Allocation
Randomized.
Model
Parallel-group design.
Masking
None.
Primary purpose
Treatment.
Phase
Phase 3.
Hypothesis
Superiority.
ARM A

Afatinib

  • BIBW 2992 (afatinib)
  • Primary analysis identifies the treatment group as afatinib 40 mg
ARM B

Gemcitabine / Cisplatin Chemotherapy

  • Gemcitabine + cisplatin
  • Used as the comparator treatment group in the posted analyses

The absence of masking is statistically relevant because treatment assignment was not concealed from participants or investigators after randomization. Randomization nevertheless provides the structural basis for comparing outcomes between the assigned treatment groups. For time-to-event outcomes, the analysis must additionally account for censoring and differing follow-up times.

4. Endpoints

The registry identifies one primary endpoint: progression-free survival. The ClinicalTrials.gov record also contain secondary analyses of objective response, disease control, overall survival, tumour shrinkage, and health-related quality-of-life deterioration endpoints.

EndpointTypeRegistered / reported time frameAnalysis
Progression-free Survival Time-to-event Tumour assessment were performed at screening, week 6, 12, 18, 24, 30, 36, 42, 48 and then every 12 weeks until progression or death whichever occurs first up to week 374 Stratified Cox proportional-hazards model; stratified log-rank test
Objective Response (OR) Binary Tumour assessment were performed at screening, week 6, 12, 18, 24, 30, 36, 42, 48 and then every 12 weeks until progression or death whichever occurs first up to week 374 Stratified logistic regression
Disease Control (DC) Binary Tumour assessment were performed at screening, week 6, 12, 18, 24, 30, 36, 42, 48 and then every 12 weeks until progression or death whichever occurs first up to week 374 Stratified logistic regression
Overall Survival (OS) Time-to-event From randomisation up to 374 weeks Stratified Cox proportional-hazards model; stratified log-rank test
Tumour Shrinkage Continuous Tumour assessment were performed at screening, week 6, 12, 18, 24, 30, 36, 42, 48 and then every 12 weeks until progression or death whichever occurs first up to week 374 ANCOVA
HRQOL: Time of Deterioration in Coughing Time-to-event Baseline and throughout the study (every 3 weeks) until progression or death (whichever occurs first) up to 374 weeks. Stratified Cox model; stratified log-rank test
HRQOL: Time of Deterioration in Dyspnoea Time-to-event Baseline and throughout the study (every 3 weeks) until progression or death (whichever occurs first) up to 374 weeks. Stratified Cox model; stratified log-rank test
HRQOL: Time of Deterioration in Pain Time-to-event Baseline and throughout the study (every 3 weeks) until progression or death (whichever occurs first) up to 374 weeks. Stratified Cox model; stratified log-rank test

Primary endpoint definition

Progression-free survival was defined as the time from randomisation to disease progression or death, whichever occurs earlier. The registry states that PFS was assessed by central independent review according to Response Evaluation Criteria in Solid Tumours (RECIST) version 1.1 and that a pre-defined set of censoring rules was used for patients who did not have a PFS.

Why the definition matters: PFS is not simply a measurement of tumour size at one visit. It is a time-to-event endpoint in which the event is progression or death, whichever occurs first. Patients who have not experienced the event at their last usable assessment contribute censored follow-up according to the prespecified censoring rules.

5. Statistical Methodology

The registry analyses show four principal statistical methods: Cox proportional-hazards models, log-rank tests, logistic regression, and ANCOVA. The methods correspond naturally to the endpoint structures: time-to-event outcomes, binary outcomes, and a continuous tumour-shrinkage outcome.

Stratified Cox proportional-hazards model

The primary PFS hazard ratio was estimated with a Cox proportional-hazards model stratified by EGFR mutation category. The same stratification concept was used for the posted OS and HRQOL time-to-event analyses.

Hazard-ratio framework
HR = hAfatinib(t) / hGemcitabine/Cisplatin(t)

An HR below 1 indicates a lower estimated instantaneous event rate in the afatinib group than in the comparator group under the fitted model. It is a relative time-to-event measure, not an absolute probability of an event.

Stratified log-rank test

The primary PFS comparison used a stratified log-rank test based on the EGFR mutation category used as a stratification factor at randomisation. The same framework was used for the posted OS and HRQOL log-rank analyses.

Logistic regression

Objective response and disease control were binary outcomes. The registry reports stratified logistic regression for both, with the treatment effect expressed as an odds ratio. Logistic regression estimates the association between randomized treatment group and the odds of the binary outcome while incorporating the specified stratification structure.

ANCOVA

Tumour shrinkage was analyzed using ANCOVA. The registry states that the analysis was adjusted for baseline sum of diameters and EGFR mutation group. This is important because the comparison concerns final tumour measurements while accounting for baseline tumour burden and the specified mutation-group factor.

ANCOVA concept
Adjusted outcome = treatment effect + baseline covariate + other prespecified adjustment

The purpose of covariate adjustment is to compare treatment groups after accounting for relevant baseline information. It does not turn an observational comparison into a randomized trial; here, the randomized design provides the principal basis for the treatment comparison.

Analysis populations

The primary PFS analysis used the randomised set (RS), defined in the registry as all patients randomized to receive treatment, whether treated or not. The objective response, disease control, overall survival, and HRQOL analyses reported here also identify the randomized set as the analysis population.

EndpointAnalysis populationKey method
Progression-free SurvivalRandomised setStratified Cox; stratified log-rank
Objective ResponseRandomised setStratified logistic regression
Disease ControlRandomised setStratified logistic regression
Overall SurvivalRandomised setStratified Cox; stratified log-rank
Tumour ShrinkageRandomised set with baseline and post-baseline target-lesion measurementsANCOVA
HRQOL deterioration endpointsRandomised setStratified Cox; stratified log-rank

6. Primary Result: Progression-Free Survival

The registered primary endpoint was progression-free survival. The ClinicalTrials.gov record contains both a model-based effect estimate and a hypothesis test.

Hazard ratio for progression or death

0.281

95% CI: 0.203–0.389   ·   Two-sided

Stratified log-rank P < 0.0001

Primary PFS analysisReported result
ComparisonAfatinib 40 mg vs Gemcitabine / Cisplatin Chemotherapy
Analysis populationRandomised set
Effect measureHazard ratio
HR0.281
95% CI0.203–0.389
CI typeTwo-sided
ModelCox proportional-hazards model stratified by EGFR mutation category
Hypothesis testStratified log-rank test
P-value<0.0001
Hypothesis typeSuperiority
Clinical Biostats interpretation

An HR of 0.281 means that the estimated instantaneous rate of progression or death in the afatinib group was approximately 28.1% of the corresponding estimated rate in the gemcitabine-cisplatin group under the fitted Cox model. Equivalently, 0.281 corresponds to an estimated 71.9% lower hazard relative to the comparator.

The HR does not mean that 71.9% of participants avoided progression, nor does it mean that each participant experienced exactly a 71.9% reduction in individual risk. It is a relative model-based measure of the event rate over time.

The 95% CI of 0.203–0.389 describes statistical uncertainty around the estimated hazard ratio. It is not a prediction interval for individual patients and does not mean that individual treatment effects must fall inside this range.

The P-value of <0.0001 addresses evidence against the relevant null hypothesis under the specified stratified log-rank testing framework. It does not measure the size or clinical importance of the treatment effect. Effect size is described by the hazard ratio and its confidence interval.

Because this is a Cox analysis, interpretation also depends on the proportional-hazards framework underlying a single HR summary. The ClinicalTrials.gov record does not report a formal assessment of that assumption, so the HR should be understood as the reported model-based summary rather than as a guarantee that the relative hazards were identical at every point in follow-up.

Why both Cox and log-rank analyses were useful

The two primary analyses answer related but different statistical questions. The stratified log-rank test provides the formal hypothesis test comparing the time-to-event distributions while respecting the EGFR mutation-category stratification. The stratified Cox model supplies an interpretable relative effect estimate—the hazard ratio—together with its confidence interval.

Educational note: the registry supplies the hazard ratio, confidence interval, and log-rank P-value, but not the individual event and censoring times needed to reconstruct a Kaplan-Meier curve. This page therefore does not fabricate a Kaplan-Meier curve from summary statistics.

7. Secondary Efficacy Results

Objective Response

Odds ratio for objective response

7.572

95% CI: 4.522–12.679   ·   P < 0.0001

Stratified logistic regression

Clinical Biostats interpretation

The reported odds ratio of 7.572 means that the estimated odds of objective response were 7.572 times as high in the afatinib group as in the gemcitabine-cisplatin group under the specified stratified logistic regression model.

An odds ratio is not the same as a risk ratio or a difference in response percentages. Without the underlying response counts in the ClinicalTrials.gov record, an absolute response-rate difference cannot be calculated without introducing information that is not provided here.

The 95% CI of 4.522–12.679 quantifies uncertainty around the odds-ratio estimate. The P-value of <0.0001 addresses the statistical test; it does not say that the odds ratio is clinically important merely because the P-value is small.

Disease Control

Odds ratio for disease control

3.843

95% CI: 2.039–7.240   ·   P < 0.0001

Stratified logistic regression

Clinical Biostats interpretation

The disease-control odds ratio of 3.843 indicates that the estimated odds of disease control were 3.843 times as high in the afatinib group as in the comparator group under the reported stratified logistic model.

The confidence interval, 2.039–7.240, expresses uncertainty around that relative odds estimate. As with objective response, the odds ratio should not be read as a percentage-point improvement in disease control.

The P-value of <0.0001 provides evidence against the null hypothesis used for the reported comparison, but it does not quantify the magnitude of benefit. The magnitude is represented by the odds ratio and its confidence interval.

Tumour Shrinkage

Adjusted mean difference in final tumour measurements

−13.64 mm

95% CI: −17.10 to −10.19 mm   ·   P < 0.0001

ANCOVA adjusted for baseline sum of diameters and EGFR mutation group

Clinical Biostats interpretation

The reported mean difference of −13.64 mm indicates a lower adjusted final value in the afatinib group relative to the gemcitabine-cisplatin group, using the ANCOVA model specified in the registry.

The confidence interval of −17.10 to −10.19 mm describes uncertainty around the adjusted mean difference. Because the analysis is adjusted for baseline sum of diameters and EGFR mutation group, this is not simply the raw difference between two unadjusted final means.

The tumour-shrinkage analysis had a more restricted measurement population: the registry states that there were only 220 patients in the afatinib arm and 101 in the gemcitabine-cisplatin arm with both baseline and post-baseline target-lesion measurements. That restriction is important when interpreting this endpoint because it differs from the broader randomized-set framework used for the primary PFS analysis.

8. Overall Survival

Overall survival was a secondary time-to-event endpoint measured from randomisation up to 374 weeks. The registry provides both a stratified Cox hazard ratio and a stratified log-rank P-value.

Hazard ratio for overall survival

0.904

95% CI: 0.715–1.144   ·   Stratified log-rank P = 0.4013

Stratified by EGFR mutation category

OS analysisReported result
Time frameFrom randomisation up to 374 weeks
Analysis populationRandomised set
Effect measureHazard ratio
HR0.904
95% CI0.715–1.144
CI typeTwo-sided
Cox modelStratified by EGFR mutation category
Log-rank P-value0.4013
Hypothesis typeSuperiority
Clinical Biostats interpretation

An OS HR of 0.904 corresponds to an estimated instantaneous rate of death about 9.6% lower in the afatinib group relative to the comparator under the reported Cox model.

The 95% CI of 0.715–1.144 spans 1.00. Thus, the interval includes both values corresponding to a lower estimated hazard and values corresponding to a higher estimated hazard. The registry-reported two-sided log-rank P-value of 0.4013 is not small under conventional hypothesis-testing interpretation.

The P-value does not establish the size of any possible treatment effect. The hazard ratio and confidence interval are the appropriate reported quantities for describing the estimated relative effect and its uncertainty.

No median overall survival or absolute survival probabilities are included in the ClinicalTrials.gov record. They are therefore not presented here.

9. Health-Related Quality of Life

The registry contains three HRQOL time-to-deterioration endpoints: coughing, dyspnoea, and pain. Each was assessed from baseline throughout the study every 3 weeks until progression or death, whichever occurred first, up to 374 weeks. Both Cox and log-rank analyses were posted for each endpoint.

HRQOL endpointHR95% CILog-rank P-value
Time of Deterioration in Coughing 0.458 0.303–0.692 0.0001
Time of Deterioration in Dyspnoea 0.534 0.394–0.724 <0.0001
Time of Deterioration in Pain 0.699 0.511–0.956 0.0220
Clinical Biostats interpretation

For all three reported HRQOL deterioration endpoints, the estimated hazard ratios are below 1. Under the Cox framework, this corresponds to lower estimated instantaneous rates of deterioration in the afatinib group relative to the comparator.

For coughing, the HR was 0.458, corresponding to an estimated 54.2% lower hazard of deterioration. The 95% CI was 0.303–0.692, with a log-rank P-value of 0.0001.

For dyspnoea, the HR was 0.534, corresponding to an estimated 46.6% lower hazard of deterioration. The 95% CI was 0.394–0.724, with a log-rank P-value of <0.0001.

For pain, the HR was 0.699, corresponding to an estimated 30.1% lower hazard of deterioration. The 95% CI was 0.511–0.956, with a log-rank P-value of 0.0220.

These endpoints should be distinguished from PFS and OS. They describe time to deterioration in specific quality-of-life domains rather than time to tumour progression or death.

10. Secondary Time-to-Event Analysis: What the HRQOL Results Teach

The HRQOL results illustrate an important statistical principle: a clinical outcome can be converted into a time-to-event endpoint when the analysis focuses on time until deterioration rather than simply whether deterioration occurred at any point.

Censoring

A participant who has not experienced the specified deterioration by the end of usable follow-up does not necessarily contribute a complete event time. The survival-analysis framework allows such observations to contribute censored follow-up.

Relative effect

The HR summarizes the relative instantaneous event rate under the Cox model. It does not directly state how many months a participant's deterioration was delayed.

Endpoint-specific meaning

An HR for coughing should not be interpreted as an HR for overall survival or progression-free survival. Each endpoint has its own event definition.

Multiple outcomes

When several secondary endpoints are tested, the collection of P-values requires careful interpretation. The ClinicalTrials.gov record does not provide a multiplicity-adjustment strategy for these analyses.

11. Statistical Methods Explained

Why was a Cox proportional-hazards model used for PFS?

PFS is a time-to-event endpoint. Participants can experience progression or death at different times, while others may remain event-free at their last assessment and therefore be censored. The Cox model is designed for this structure and produces a hazard ratio that summarizes the relative event rate between treatment groups. In LUX-Lung 6, the model was stratified by EGFR mutation category.

What does an HR of 0.281 mean?

An HR of 0.281 means that the fitted model estimates the instantaneous rate of progression or death in the afatinib group to be 0.281 times the comparator rate. Expressed as a relative reduction, this corresponds to approximately 71.9% lower estimated hazard. It does not mean that 71.9% of patients avoided progression.

Why was a stratified log-rank test used?

The log-rank test compares time-to-event experience between treatment groups while accounting for the survival-analysis structure. Here, the registry specifically reports stratification by EGFR mutation category, the factor used at randomisation. Stratification prevents the comparison from ignoring a prespecified factor incorporated into the trial design.

Why was logistic regression used for objective response?

Objective response is recorded as a binary outcome: a participant either satisfies the response definition or does not. Logistic regression is therefore appropriate for modeling the odds of response. The registry reports a stratified logistic regression model, with the treatment effect expressed as an odds ratio of 7.572.

What does an odds ratio of 7.572 mean?

An odds ratio of 7.572 means that the estimated odds of objective response were 7.572 times as high in the afatinib group as in the comparator group under the specified model. Odds are not probabilities, so the odds ratio cannot be converted into a percentage-point response difference without the underlying response probabilities.

Why was ANCOVA used for tumour shrinkage?

Tumour shrinkage was a continuous measurement in millimetres. ANCOVA allows the treatment comparison to be adjusted for baseline sum of diameters and EGFR mutation group. This can improve the precision and interpretability of the treatment comparison by accounting for relevant baseline information.

Why does the P-value not tell us the size of the treatment effect?

A P-value measures the compatibility of the observed data with a specified null hypothesis under the statistical model. It is affected by both the magnitude of an observed effect and the amount of information in the analysis. Effect magnitude is better described by measures such as the HR or odds ratio together with a confidence interval.

12. Confidence Intervals and Statistical Precision

Confidence intervals are particularly important in a trial with several types of effect measures because they distinguish the point estimate from the uncertainty surrounding it.

EndpointPoint estimate95% CIEffect measure
Primary PFS0.2810.203–0.389Hazard ratio
Objective Response7.5724.522–12.679Odds ratio
Disease Control3.8432.039–7.240Odds ratio
Tumour Shrinkage−13.64 mm−17.10 to −10.19 mmMean difference
Overall Survival0.9040.715–1.144Hazard ratio
HRQOL: Coughing0.4580.303–0.692Hazard ratio
HRQOL: Dyspnoea0.5340.394–0.724Hazard ratio
HRQOL: Pain0.6990.511–0.956Hazard ratio

The confidence intervals also illustrate why a point estimate should never be interpreted in isolation. The PFS HR is estimated relatively far below 1, while the OS HR is closer to 1 and its interval includes 1. The response and disease-control analyses use odds ratios rather than hazard ratios, while tumour shrinkage uses a mean difference. These measures are not interchangeable.

A useful distinction
Effect size ≠ P-value ≠ confidence interval

The effect measure describes the estimated magnitude and direction; the confidence interval describes uncertainty around that estimate; the P-value addresses a hypothesis test. Keeping these three concepts separate prevents many common interpretation errors.

13. Stratified Analysis

EGFR mutation category appears repeatedly in the statistical methods reported by the registry. It was used as a stratification factor in the primary PFS analysis and in the secondary OS and HRQOL time-to-event analyses. It was also included in the ANCOVA adjustment for tumour shrinkage and in the logistic-regression analyses for objective response and disease control.

EndpointRole of EGFR mutation category
Progression-free SurvivalStratification factor in Cox model and stratified log-rank test
Objective ResponseStratification factor in logistic regression
Disease ControlStratified logistic regression
Overall SurvivalStratification factor in Cox model and stratified log-rank test
Tumour ShrinkageAdjustment variable in ANCOVA
HRQOL deteriorationStratification factor in Cox models and stratified log-rank tests

Stratification is not the same as claiming that EGFR mutation category modifies the treatment effect. A stratification factor is incorporated to account for an important design characteristic. Demonstrating effect modification would require an appropriate interaction or heterogeneity analysis, which is not reported in the ClinicalTrials.gov record.

14. Safety Results

The ClinicalTrials.gov record provides serious adverse events by treatment arm as affected participants over participants at risk.

Safety measureAfatinib 40 mgGemcitabine / Cisplatin Chemotherapy
Serious adverse events 40 / 239 12 / 113
Serious adverse events: affected / at risk
Afatinib
40 / 239
Gemcitabine / Cisplatin
12 / 113

The affected and at-risk counts are reported exactly as reported in the registry. They should not be treated as equivalent to a formal comparative risk analysis unless the corresponding safety analysis specification and statistical comparison are available.

Safety interpretation: the ClinicalTrials.gov record does not provide a formal confidence interval, P-value, relative risk, odds ratio, or time-to-event analysis for serious adverse events. The page therefore reports the arm-specific affected/at-risk counts without constructing an unreported inferential comparison.

15. Missing Data, Censoring, and Measurement Populations

Time-to-event endpoints such as PFS, OS, and HRQOL deterioration require special treatment of incomplete follow-up. The PFS definition explicitly states that predefined censoring rules were used for patients who did not have a PFS. This is different from simply treating an incomplete observation as if no event had occurred.

PFS censoring

The registry explicitly states that predefined censoring rules were used for patients who did not have a PFS.

Randomised-set principle

The primary PFS analysis used the randomized set, including patients randomized to receive treatment whether treated or not.

Tumour shrinkage

This endpoint had a restricted measurement population because baseline and post-baseline target-lesion measurements were required.

Registry limitation

The ClinicalTrials.gov record does not describe a general missing-data imputation method for the reported endpoints.

This distinction matters because the primary PFS estimate is based on a time-to-event framework that can incorporate censored observations, whereas tumour shrinkage was analyzed only among participants with the specified baseline and post-baseline measurements. The two analyses therefore answer different statistical questions and rely on different available measurement sets.

16. Multiplicity and Multiple Endpoints

The registry lists 18 posted outcome measures and 13 posted statistical analyses, including one primary endpoint with two posted primary analyses and multiple secondary endpoints. This creates an important interpretive distinction between the prespecified primary endpoint and the collection of secondary findings.

Analysis familyRole in the ClinicalTrials.gov recordInterpretive point
Progression-free SurvivalRegistered primary endpointPrimary confirmatory treatment comparison under the reported superiority framework
Objective ResponseSecondary endpointBinary efficacy outcome analyzed using logistic regression
Disease ControlSecondary endpointBinary efficacy outcome analyzed using logistic regression
Overall SurvivalSecondary endpointTime-to-event outcome analyzed using Cox and log-rank methods
Tumour ShrinkageSecondary endpointContinuous outcome analyzed using ANCOVA
HRQOL deterioration endpointsSecondary endpointsMultiple time-to-event outcomes

Multiple statistical tests create the possibility of false-positive findings across the collection of analyses even when each individual test is conducted correctly. The ClinicalTrials.gov record does not provide a multiplicity-adjustment procedure or endpoint hierarchy for the secondary analyses. Consequently, the individual secondary P-values should be interpreted as the reported results of those analyses rather than as evidence that every secondary endpoint was independently powered and error-controlled at the same level as the primary endpoint.

17. Primary Endpoint vs Secondary Endpoints

The statistical structure of LUX-Lung 6 becomes clearer when the endpoints are grouped by what they measure.

QuestionEndpointEffect measureMethod
How long until progression or death? Progression-free Survival Hazard ratio Stratified Cox + log-rank
How often was objective response achieved? Objective Response Odds ratio Stratified logistic regression
How often was disease controlled? Disease Control Odds ratio Stratified logistic regression
How long until death? Overall Survival Hazard ratio Stratified Cox + log-rank
How did final tumour measurements differ after adjustment? Tumour Shrinkage Mean difference ANCOVA
How long until specific HRQOL deterioration? Coughing, dyspnoea, pain Hazard ratio Stratified Cox + log-rank

No single endpoint captures the entire statistical evidence. PFS measures disease progression or death, OS measures death, response measures a binary tumour outcome, tumour shrinkage measures a continuous radiologic quantity, and HRQOL deterioration measures time until deterioration in specific patient-reported domains.

18. A Worked Interpretation of the Primary Analysis

Suppose a reader sees only the headline result: HR 0.281, 95% CI 0.203–0.389, P < 0.0001. A statistically literate interpretation proceeds in several steps.

Step 1 · Identify the endpoint

PFS is not OS

The event is disease progression or death, whichever occurs earlier. The estimate therefore describes PFS, not mortality alone.

Step 2 · Identify the effect measure

HR is a relative time-to-event measure

The value 0.281 compares estimated instantaneous event rates under the Cox model. It is not a percentage of patients and is not a median difference.

Step 3 · Examine precision

Read the confidence interval

The 95% CI is 0.203–0.389. This gives the statistical uncertainty around the estimated HR and remains below 1 throughout the interval.

Step 4 · Read the hypothesis test

Separate P-value from effect size

The stratified log-rank P-value is <0.0001. This is evidence against the null hypothesis under the reported test, but it does not quantify the magnitude of the effect.

Step 5 · Check the analysis structure

Stratification and population matter

The analysis used the randomized set and stratified the Cox model and log-rank test by EGFR mutation category.

19. Important Limitations and Interpretation Issues

20. Why This Trial Matters Statistically

LUX-Lung 6 is a useful teaching example because a single randomized trial incorporates several major branches of clinical-trial statistics. The primary endpoint is a censored time-to-event outcome; secondary outcomes include binary response measures, a continuous tumour measurement, overall survival, and several patient-reported time-to-deterioration endpoints.

ConceptHow it appears in LUX-Lung 6
RandomizationRandomized phase 3 parallel-group design
Primary endpointProgression-free survival
Time-to-event analysisPFS, OS, and HRQOL deterioration
Kaplan-Meier frameworkThe natural descriptive framework for the reported time-to-event endpoints
Hazard ratioPrimary PFS and secondary OS/HRQOL effect measure
Log-rank testFormal time-to-event comparison for PFS, OS, and HRQOL endpoints
Cox modelEstimation of relative time-to-event effects
Stratified analysisEGFR mutation category used in multiple analyses
Logistic regressionObjective response and disease control
Odds ratioEffect measure for binary response outcomes
ANCOVATumour shrinkage adjusted for baseline sum of diameters and EGFR mutation group
Covariate adjustmentBaseline tumour measurement and EGFR mutation group in tumour-shrinkage analysis
Multiple endpointsOne primary endpoint plus multiple secondary outcomes
Analysis populationsRandomised set for the reported efficacy analyses, with a restricted measurement population for tumour shrinkage

21. What the Primary PFS Result Does — and Does Not — Establish

What it establishes statistically

The registry reports a stratified Cox HR of 0.281 with a two-sided 95% CI of 0.203–0.389, together with a stratified log-rank P-value of <0.0001. These are the reported statistical results for the registered primary PFS endpoint under a superiority hypothesis.

What it does not establish by itself

The HR does not provide a median PFS, an absolute difference in survival probability at a particular time, the proportion of patients who benefit, or an individual patient's probability of remaining progression-free. Those quantities require corresponding data that are not reported in the ClinicalTrials.gov record used for this page.

Why OS must be considered separately

The secondary OS analysis reports HR 0.904 with 95% CI 0.715–1.144 and stratified log-rank P = 0.4013. PFS and OS are different endpoints, so the primary PFS result should not be presented as if it were an overall-survival result.

22. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

23. Related Statistical Calculators

24. Sources

Continue with the statistical methods

Explore the underlying survival-analysis, regression, confidence-interval, and clinical-trial concepts used to interpret LUX-Lung 6.

25. Record Summary

LUX-Lung 6 provides a compact example of how several statistical methods work together in a randomized phase 3 clinical trial. The registered primary endpoint, progression-free survival, was analyzed in the randomized set using a Cox proportional-hazards model stratified by EGFR mutation category and a stratified log-rank test. The reported HR was 0.281 with a 95% CI of 0.203–0.389, and the stratified log-rank P-value was <0.0001.

The secondary analyses illustrate different statistical data structures. Objective response and disease control were analyzed with stratified logistic regression and reported as odds ratios. Tumour shrinkage was analyzed using ANCOVA adjusted for baseline sum of diameters and EGFR mutation group. Overall survival and three HRQOL time-of-deterioration endpoints used stratified time-to-event methods. Serious adverse events were reported as affected participants over participants at risk by treatment arm.

The most important statistical lesson is that these results should be interpreted according to their endpoint and effect measure. A hazard ratio describes a relative time-to-event quantity, an odds ratio describes relative odds for a binary outcome, and a mean difference describes an adjusted difference in a continuous outcome. Confidence intervals describe uncertainty around these estimates, while P-values address the corresponding hypothesis tests. None of these quantities, considered alone, describes every dimension of clinical outcome.

Clinical Biostats methodology: This page separates the numerical evidence reported by the trial registry from statistical interpretation. Where the ClinicalTrials.gov record does not contain a number, subgroup result, baseline characteristic, median survival, or formal multiplicity procedure, that information has not been reconstructed or inferred.