← Clinical Trials
Lung Adenocarcinoma Phase 2 Randomized NCT01466660

LUX-Lung 7: Complete Statistical Analysis of Afatinib in EGFR Mutation-Positive Lung Adenocarcinoma

An independent statistical analysis of the randomized phase 2 LUX-Lung 7 trial comparing afatinib with gefitinib as first-line treatment for EGFR mutation-positive adenocarcinoma of the lung.

Trial status: Completed  ·  Enrollment: 319  ·  Primary completion: 08 April 2016
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the information reported in the ClinicalTrials.gov record.

1. Trial at a Glance

LUX-Lung 7 was a randomized, parallel, open-label phase 2 trial comparing afatinib with gefitinib in first-line EGFR mutation-positive adenocarcinoma of the lung. The registry reports 319 randomized participants, two treatment arms, three registered primary time-to-event endpoints, and formal analyses for each primary endpoint.

319
Enrolled
Randomized set
2
Arms
Afatinib vs gefitinib
0.750
TTF HR
95% CI 0.595–0.944
0.862
OS HR
95% CI 0.674–1.101
FeatureLUX-Lung 7
Trial nameLUX-Lung 7
Brief titleA Phase IIb Trial of Afatinib(BIBW2992) Versus Gefitinib for the Treatment of 1st Line EGFR Mutation Positive Adenocarcinoma of the Lung
PhasePhase 2
StatusCompleted
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment319
InterventionsAfatinib (drug); gefitinib (drug)
Registered primary endpoints3
Primary endpoint typeTime-to-event
Results postedYes
Statistical analyses posted9
Lead sponsorBoehringer Ingelheim
Sponsor typeIndustry

2. Clinical Question

The central statistical question was whether first-line afatinib produced different time-to-event outcomes from gefitinib in participants with EGFR mutation-positive adenocarcinoma of the lung.

Population

Participants with 1st line EGFR mutation positive adenocarcinoma of the lung, as described by the registered trial title.

Intervention

Afatinib.

Comparator

Gefitinib.

Primary question

How do afatinib and gefitinib compare for progression-free survival, time to treatment failure, and overall survival?

3. Trial Design

01
Randomize 319 participants
02
Treat Afatinib or gefitinib
03
Assess Time-to-event outcomes
04
Follow Events and censoring
05
Analyze Stratified survival methods
ARM A · AFATINIB

Afatinib

  • Afatinib (drug)
  • Compared directly with gefitinib
  • Included in the randomized set used for primary efficacy analyses
ARM B · GEFITINIB

Gefitinib

  • Gefitinib (drug)
  • Comparator treatment
  • Included in the randomized set used for primary efficacy analyses
Allocation
Randomized.
Design
Parallel-group design.
Masking
None; the registry describes the trial as unmasked.
Hypothesis type
Superiority is listed for the posted statistical analyses, while the analysis notes state that the trial was exploratory and no formal hypotheses were tested.

4. Endpoints

The registry identifies three primary endpoints, all of which are time-to-event outcomes. Their definitions and time frames are important because the event definition determines what the hazard ratio represents.

EndpointRegistered definitionTime frame
Progression-free Survival Progression-free survival (PFS) defined as the time from date of randomisation to date of disease progression, or date of death if a patient died earlier. Participants with no event (Disease progression (PD) or death) were censored. PD was primarily evaluated for the primary analysis by an independent central imaging review according to Response Evaluation Criteria in Solid Tumours (RECIST) version 1.1. Per RECIST version 1.1. for target lesions and assessed by Computed Tomography (CT)-scan or Magnetic Resonance Imaging (MRI): PD, At least a 20% increase in the sum of the longest diameter (SoD) of target lesions taking as reference the smallest SoD of target lesions recorded since the treatment started, together with an absolute increase in the SoD of target lesions of at least 5 millimetre (mm) or the appearance of one or more new lesions. For the final analysis (analysis cut-off date 12 April 2019) status and date of PD were determined by investigator assessment. From first drug administration until 28 days after last drug administration + Follow-Up period for collecting information on disease progression or death, up to 2465 days.
Time to Treatment Failure (TTF) (Main Overall Survival Analysis Cut-off Date, 08 April 2016) Time to Treatment Failure (TTF) which was the time from the date of randomisation to the date of i.e. permanent treatment discontinuation for any reason. From first drug administration until last drug administration, up to 1482 days
Overall Survival Overall survival (OS) which was defined as the time from the date of randomisation to the date of death. Participants for whom there is no evidence of death at the time of the analysis will be censored at the date that they were last known to be alive. From first drug administration until 28 days after last drug administration + Follow-Up period for collecting information on disease progression or death, up to 2465 days.

Secondary endpoints with posted analyses

EndpointTypeAnalysis methodEffect measure
Objective Response RateCount / rateLogistic regressionOdds ratio
Disease ControlBinaryLogistic regressionOdds ratio
Tumour Shrinkage (Main Overall Survival Analysis Cut-off Date, 08 April 2016)Time-to-eventANCOVAMean difference
Health-related Quality of Life (Primary Analysis Cut-off Date, 21 August 2015)ContinuousMixed-effects modelMean difference
Health-related Quality of Life (Primary Analysis Cut-off Date, 21 August 2015)ContinuousMixed-effects modelMean difference
Health-related Quality of Life (Primary Analysis Cut-off Date, 21 August 2015)ContinuousMixed-effects modelMean difference

5. Analysis Populations and Stratification

The three primary endpoint analyses used the randomised set, defined in the registry as all patients randomized to receive treatment, whether treated or not. This is an important feature of the primary efficacy analysis because treatment assignment, rather than subsequent treatment exposure, defines membership in the analysis population.

AnalysisPopulationKey analytical feature
Primary endpoints Randomised set All patients randomized to receive treatment, whether treated or not
Objective Response Rate Randomised set Logistic regression, stratified analysis
Disease Control Randomised set Logistic regression, stratified analysis
Tumour Shrinkage Participants in the randomised set with tumour assessments ANCOVA with baseline and stratification-related covariate adjustment
Health-related Quality of Life All randomised subjects with health-related quality of life data Mixed-effects model with covariate adjustment

The primary survival analyses were stratified by EGFR mutation group and presence of brain metastases at baseline. This same stratification structure appears in the primary Cox proportional-hazards analyses and the secondary logistic-regression analyses.

6. Statistical Methodology

Kaplan-Meier estimation and time-to-event analysis

The primary endpoints are time-to-event outcomes. For PFS, an event is disease progression or death, while the registered OS definition uses death as the event and censors participants with no evidence of death at the analysis time. TTF instead uses permanent treatment discontinuation for any reason.

These definitions illustrate why the endpoint itself matters. Three endpoints can all be expressed in months and analyzed using survival methods while answering different clinical questions:

EndpointEvent representedInterpretive question
PFSDisease progression or deathHow long until progression or death?
TTFPermanent treatment discontinuation for any reasonHow long until treatment failure?
OSDeathHow long until death?
Conceptual Kaplan-Meier estimator
S(t) = ∏ti ≤ t (1 − di/ni)

The Kaplan-Meier framework estimates the probability of remaining event-free over time while accommodating right-censored observations.

Log-rank testing

The registry reports the log-rank test for all three primary endpoints. The log-rank approach compares the observed and expected numbers of events between treatment groups over the follow-up period. In this trial, the analysis text additionally identifies stratification by EGFR mutation group and presence of brain metastases at baseline.

Cox proportional-hazards modeling

The registry analysis notes state that a Cox proportional-hazards model, stratified by EGFR mutation group and presence of baseline brain metastases, was used to estimate the hazard ratio for PFS and OS. For these endpoints, the reported ratio was calculated as Afatinib divided by Gefitinib.

Hazard-ratio interpretation
HR = hazard in Afatinib group / hazard in Gefitinib group

An HR below 1 indicates a lower estimated instantaneous event rate in the afatinib group under the fitted model. It is a relative time-to-event measure, not a probability and not an absolute difference in survival time.

Logistic regression

Objective Response Rate and Disease Control were analyzed using logistic regression. Both analyses used the randomized set and incorporated stratification for EGFR mutation group and presence of brain metastases at baseline. The effect measure was an odds ratio calculated as Afatinib divided by Gefitinib.

ANCOVA

Tumour Shrinkage was analyzed using ANCOVA among participants in the randomized set with tumour assessments. The analysis adjusted for baseline sum of diameters, EGFR mutation group, and presence of brain metastases at baseline. The reported mean difference was calculated as Afatinib minus Gefitinib.

Mixed-effects models

Health-related quality of life was measured every 8 weeks, up to 56 weeks, and analyzed with mixed-effects models. The reported analyses adjusted for EGFR mutation group and presence of brain metastases at baseline. This is a natural setting for a longitudinal model because the same participants can contribute repeated quality-of-life observations over time.

7. Primary Results: Progression-free Survival

The registry reports a formal primary analysis of PFS using a log-rank test and a hazard ratio from a stratified Cox proportional-hazards model. The analysis population was the randomized set.

Hazard ratio for progression or death

0.822

95% CI: 0.655–1.032   ·   P = 0.0891

Hazard ratio calculated as Afatinib divided by Gefitinib.

Clinical Biostats interpretation

The estimated HR of 0.822 means that the fitted model estimated the instantaneous rate of progression or death in the afatinib group to be about 17.8% lower than in the gefitinib group, under the proportional-hazards model and the specified stratification.

The HR does not mean that 17.8% fewer participants progressed, nor does it mean that every participant experienced a 17.8% reduction in risk. It is a relative time-to-event estimate.

The 95% CI of 0.655–1.032 describes uncertainty around the estimated hazard ratio. Its width indicates that the point estimate should not be interpreted as an exact treatment effect. Because the interval includes 1, the data are compatible with both a lower and a higher hazard under the model.

The p-value of 0.0891 measures evidence against the specified null comparison under the analysis framework; it does not measure the size or clinical importance of the effect. Most importantly, the registry analysis notes state that this was an exploratory trial and no formal hypotheses were tested, so the p-value should not be converted into a claim of formal statistical success or failure.

The interpretation also depends on censoring and on the proportional-hazards modeling framework. A single HR summarizes a relative hazard over the analyzed follow-up; it is not a direct description of absolute survival probabilities at a particular time.

8. Primary Results: Time to Treatment Failure

TTF was a registered primary endpoint with a time frame extending from first drug administration until last drug administration, up to 1482 days. The registered definition describes TTF as the time from randomization to permanent treatment discontinuation for any reason.

Hazard ratio for treatment failure

0.750

95% CI: 0.595–0.944   ·   P = 0.0136

Ratio calculated as Afatinib divided by Gefitinib.

Clinical Biostats interpretation

The estimated HR of 0.750 corresponds to approximately a 25% lower estimated instantaneous rate of treatment failure in the afatinib group relative to the gefitinib group, using the reported hazard-ratio definition.

This does not mean that 25% of participants avoided treatment failure, nor does it establish a 25% increase in treatment duration for an individual participant. The hazard ratio describes a relative event rate over time.

The 95% CI of 0.595–0.944 quantifies uncertainty around the estimated relative hazard. Unlike the PFS interval, the reported interval lies below 1, indicating that the estimated treatment-failure hazard was lower for afatinib across the interval represented by this analysis.

The p-value of 0.0136 quantifies evidence under the specified statistical comparison; it is not an effect-size measure. The registry explicitly describes the trial as exploratory and states that no formal hypotheses were tested. Therefore, the numerical p-value should be interpreted as part of the reported exploratory analysis rather than as a standalone confirmatory claim.

9. Primary Results: Overall Survival

OS was defined as the time from randomization to death. Participants with no evidence of death at analysis were censored at the date they were last known to be alive.

Hazard ratio for death

0.862

95% CI: 0.674–1.101   ·   P = 0.2343

Hazard ratio calculated as Afatinib divided by Gefitinib.

Clinical Biostats interpretation

The HR of 0.862 corresponds to an estimated instantaneous rate of death approximately 13.8% lower in the afatinib group than in the gefitinib group under the fitted stratified Cox model.

Again, the HR does not represent an absolute survival difference, a difference in median survival, or the percentage of participants who benefited. Those quantities require different information and different estimands.

The 95% CI of 0.674–1.101 spans 1. The interval therefore reflects substantial uncertainty about the direction and magnitude of the relative hazard, even though the point estimate is below 1.

The p-value of 0.2343 is evidence summarized under the specified statistical test; it does not quantify the magnitude of the observed HR. As with the other primary endpoints, the registry states that the trial was exploratory and that no formal hypotheses were tested.

OS is also particularly sensitive to events and censoring during long-term follow-up. The registered definition makes clear that participants without evidence of death were censored at the date they were last known to be alive.

Educational note: the available trial data provide hazard-ratio estimates and confidence intervals but do not provide the underlying event-by-event data needed to reconstruct a valid Kaplan-Meier curve. This page therefore does not fabricate a survival curve from summary statistics.

10. Comparison of the Three Primary Endpoints

Primary endpointMethodEffect95% CIP-value
Progression-free SurvivalLog-rank; stratified Cox modelHR 0.8220.655–1.0320.0891
Time to Treatment FailureLog-rank; stratified analysisHR 0.7500.595–0.9440.0136
Overall SurvivalLog-rank; stratified Cox modelHR 0.8620.674–1.1010.2343

The three estimates should not be collapsed into a single measure of "benefit." They represent different event definitions. PFS captures progression or death, TTF captures permanent treatment discontinuation for any reason, and OS captures death. Consequently, the fact that their hazard ratios differ is not itself surprising or evidence of inconsistency.

PFS

The point estimate was below 1, with HR 0.822, but the 95% CI extended above 1.

TTF

The point estimate was 0.750, with a 95% CI of 0.595–0.944.

OS

The point estimate was below 1 at 0.862, while the 95% CI extended from 0.674 to 1.101.

Exploratory status

The registry analysis notes explicitly state that no formal hypotheses were tested.

11. Secondary Results: Objective Response Rate

Objective Response Rate was analyzed as a percentage of participants using logistic regression. The analysis used the randomized set and was stratified for EGFR mutation group and presence of brain metastases at baseline.

Odds ratio for objective response

1.307

95% CI: 0.768–2.223   ·   P = 0.3235

Odds ratio calculated as Afatinib divided by Gefitinib.

Clinical Biostats interpretation

An OR of 1.307 means that the estimated odds of objective response were 1.307 times as high with afatinib as with gefitinib under the reported logistic-regression model.

Odds are not probabilities. An odds ratio of 1.307 therefore cannot be interpreted as a 30.7% higher response rate without knowing the underlying response probabilities.

The 95% CI of 0.768–2.223 is relatively broad and includes 1. The p-value of 0.3235 describes the evidence under the model and does not measure the size of the estimated odds ratio.

The registry again describes this as an exploratory trial with no formal hypotheses tested. The reported stratification also matters: the OR is not simply an unadjusted ratio of two raw response proportions.

12. Secondary Results: Disease Control

Disease Control was a binary secondary endpoint analyzed using logistic regression in the randomized set, with stratification for EGFR mutation group and presence of brain metastases at baseline.

Odds ratio for disease control

1.138

95% CI: 0.447–2.896   ·   P = 0.7856

Odds ratio calculated as Afatinib divided by Gefitinib.

Clinical Biostats interpretation

The OR of 1.138 indicates an estimated odds of disease control 1.138 times the odds under gefitinib, according to the reported logistic-regression analysis.

The confidence interval, 0.447–2.896, is wide and spans 1 by a substantial amount. This indicates that the point estimate alone provides an incomplete picture of the uncertainty in the treatment comparison.

The p-value of 0.7856 is a measure associated with the statistical comparison, not a probability that the treatment effect is zero and not a measure of clinical relevance.

13. Secondary Results: Tumour Shrinkage

Tumour Shrinkage was reported in millimetres and analyzed using ANCOVA among participants in the randomized set with tumour assessments. The analysis adjusted for baseline sum of diameters, EGFR mutation group, and presence of brain metastases at baseline.

Adjusted mean difference in final values

-3.45 mm

95% CI: -7.13 to 0.23   ·   P = 0.0657

Difference calculated as Afatinib minus Gefitinib.

Clinical Biostats interpretation

The negative mean difference of -3.45 means that the adjusted final tumour-shrinkage measure was estimated to be 3.45 mm lower with afatinib than with gefitinib, using the registry's stated direction of subtraction.

The estimate is adjusted rather than simply comparing two unadjusted means. The covariates were baseline sum of diameters, EGFR mutation group, and presence of brain metastases at baseline.

The 95% CI of -7.13 to 0.23 includes zero, which means the uncertainty interval contains both a negative and a small positive mean difference. The p-value of 0.0657 does not quantify the magnitude of the 3.45 mm estimate.

14. Secondary Results: Health-Related Quality of Life

Health-related quality of life was assessed every 8 weeks, up to 56 weeks. Three posted analyses are reported in the ClinicalTrials.gov record: EQ-5D UK utility score, EQ-5D Belgium utility score, and EQ-VAS utility score. All used mixed-effects models adjusted for EGFR mutation group and presence of brain metastases at baseline.

Quality-of-life measureMean difference95% CIP-value
EQ-5D UK utility score-0.02-0.06 to 0.010.1422
EQ-5D Belgium utility score-0.03-0.06 to 0.000.0540
EQ-VAS utility score-1.5-3.9 to 0.80.2032
Clinical Biostats interpretation

Each mean difference is calculated with afatinib as the treatment group and gefitinib as the comparator. A negative value therefore indicates a lower estimated final value in the afatinib group under the reported model.

The mixed-effects approach is important because quality-of-life data were collected repeatedly from participants over time. Rather than treating every measurement as an independent observation, a longitudinal mixed model can account for the repeated-measures structure while incorporating the stated covariate adjustments.

The confidence intervals should be read independently for each quality-of-life measure. For example, the EQ-5D UK interval of -0.06 to 0.01 includes zero, as does the EQ-VAS interval of -3.9 to 0.8. The EQ-5D Belgium interval extends to 0.00. None of these p-values should be interpreted as a measure of the size of the mean difference.

15. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected participants divided by participants at risk.

Safety measureAfatinibGefitinib
Serious adverse events75/16064/159
Serious adverse events · affected / at risk
Afatinib
75/160
Gefitinib
64/159

These are counts of affected participants over the reported at-risk denominators. They are not a hazard ratio, odds ratio, or risk difference, and the ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse events by arm.

16. Statistical Methods Explained

Why was a stratified Cox model used?

The registry states that PFS and OS were analyzed with a Cox proportional-hazards model stratified by EGFR mutation group and presence of baseline brain metastases. Stratification allows the baseline hazard to differ across these strata while estimating a treatment hazard ratio within the specified model structure.

What does an HR of 0.750 mean for TTF?

Because the registry defines the ratio as afatinib divided by gefitinib, 0.750 indicates a lower estimated instantaneous rate of treatment failure for afatinib. The corresponding descriptive transformation is a 25% lower estimated hazard. It does not mean that treatment failure was prevented in exactly 25% of participants.

Why is PFS different from TTF?

PFS is defined around disease progression or death. TTF is defined around permanent treatment discontinuation for any reason. Discontinuation and disease progression are therefore different events, even though both can be analyzed using time-to-event methods.

What does an odds ratio of 1.307 mean?

The OR of 1.307 for objective response means that the modeled odds of response were 1.307 times those under gefitinib. An odds ratio is not a risk ratio or percentage-point difference, so the underlying response probabilities would be required to translate it into an absolute probability comparison.

Why was ANCOVA used for tumour shrinkage?

The registry reports ANCOVA with adjustment for baseline sum of diameters, EGFR mutation group, and presence of brain metastases at baseline. Baseline tumour size is especially relevant because a final tumour measurement can be influenced by the starting measurement. Covariate adjustment estimates the treatment-group difference after accounting for the specified baseline information.

Why use a mixed-effects model for quality of life?

Quality of life was measured every 8 weeks up to 56 weeks, so a participant could contribute multiple observations. A mixed-effects model is designed for this repeated-measures structure and can estimate treatment differences while accounting for the longitudinal nature of the data and the stated covariate adjustments.

Does P = 0.0136 prove that afatinib is superior?

No. The p-value is part of the reported exploratory analysis and does not measure effect size. The registry analysis notes explicitly state that the trial was exploratory and that no formal hypotheses were tested. The appropriate interpretation is therefore to consider the estimated effect, confidence interval, endpoint definition, analysis model, and exploratory status together.

17. Understanding Confidence Intervals in LUX-Lung 7

Confidence intervals are particularly informative in this trial because the same type of effect measure can have very different levels of precision across endpoints.

EndpointEstimate95% CIWhat the interval tells us
PFS HR0.8220.655–1.032The interval includes 1, so uncertainty includes both a lower and higher hazard relative to the comparator.
TTF HR0.7500.595–0.944The interval is below 1 and therefore excludes an equal hazard within this reported interval.
OS HR0.8620.674–1.101The interval includes 1 and extends above it.
ORR OR1.3070.768–2.223The interval includes 1 and is relatively broad.
Disease Control OR1.1380.447–2.896The interval includes 1 and is broad relative to the point estimate.
Tumour Shrinkage mean difference-3.45-7.13 to 0.23The interval includes zero, the value representing no mean difference.

The null reference depends on the effect measure. For hazard ratios and odds ratios, the reference value is 1. For a mean difference, the reference value is 0. This distinction is fundamental when reading confidence intervals.

18. Multiplicity, Hypothesis Testing, and the Exploratory Design

The registry lists three primary endpoints, all with a superiority hypothesis type in the posted statistical analyses. However, the analysis notes for the primary endpoints explicitly state: "Exploratory trial, no formal hypotheses were tested."

Interpretation caution: the presence of reported p-values does not by itself turn an exploratory analysis into a confirmatory hypothesis-testing framework. The trial's stated exploratory status should remain part of the interpretation of all reported p-values.
FeatureWhat the ClinicalTrials.gov record reportsStatistical implication
Primary endpoints3PFS, TTF, and OS each have their own treatment comparison.
Hypothesis typeSuperiorityThe posted analysis fields identify superiority as the hypothesis type.
Trial characterizationExploratoryThe analysis notes state that no formal hypotheses were tested.
Statistical analyses9 postedPrimary and secondary endpoints use several different model families.

Because the ClinicalTrials.gov record does not report a multiplicity-adjustment procedure, alpha allocation, or a formal hierarchical testing strategy, this page does not infer one. The safest statistical reading is to preserve the registry's explicit distinction between reported p-values and the absence of formal hypothesis testing.

19. Censoring and the Meaning of Time-to-Event Results

Time-to-event analysis differs from a simple comparison of event proportions because participants may have different lengths of observation. The registry explicitly describes censoring for PFS and OS.

PFS censoring

Participants with no event, defined as disease progression or death, were censored.

OS censoring

Participants with no evidence of death at analysis were censored at the date they were last known to be alive.

TTF event

The event was permanent treatment discontinuation for any reason.

Why this matters

The definition of the event and the rules for censoring determine the estimand being analyzed.

A hazard ratio should therefore be interpreted in the context of the event definition and censoring rules. The OS HR and PFS HR are not interchangeable simply because both are produced by Cox models.

20. Covariate Adjustment and Stratified Analysis

LUX-Lung 7 provides several examples of how baseline information can enter an analysis. EGFR mutation group and baseline brain metastases appear repeatedly as stratification factors. Baseline sum of diameters additionally enters the ANCOVA for tumour shrinkage.

MethodAdjustment / stratification reported
PFS Cox modelEGFR mutation group; presence of baseline brain metastases
OS Cox modelEGFR mutation group; presence of baseline brain metastases
TTF analysisEGFR mutation group; presence of baseline brain metastases
Objective Response RateEGFR mutation group; presence of baseline brain metastases
Disease ControlEGFR mutation group; presence of baseline brain metastases
Tumour Shrinkage ANCOVABaseline sum of diameters; EGFR mutation group; presence of baseline brain metastases
Quality of Life mixed-effects modelsEGFR mutation group; presence of baseline brain metastases

Stratification and covariate adjustment are not interchangeable concepts. A stratified survival model allows the baseline hazard to vary by stratum, whereas ANCOVA directly adjusts a continuous outcome for specified covariates. The choice should therefore be understood as part of the endpoint-specific statistical model.

21. Primary and Secondary Analysis Map

RoleEndpointModel familyEffect measure
PrimaryProgression-free SurvivalSurvival analysisHazard ratio
PrimaryTime to Treatment FailureSurvival analysisHazard ratio
PrimaryOverall SurvivalSurvival analysisHazard ratio
SecondaryObjective Response RateCategorical dataOdds ratio
SecondaryDisease ControlCategorical dataOdds ratio
SecondaryTumour ShrinkageLinear modelsMean difference
SecondaryHealth-related Quality of LifeLongitudinal / mixed modelsMean difference

This endpoint map is statistically useful because it demonstrates that a randomized clinical trial can contain several distinct estimands and model families. Survival outcomes require handling of event times and censoring; binary outcomes use odds-based regression; continuous tumour measurements use covariate-adjusted linear modeling; and repeated quality-of-life measurements use longitudinal models.

22. Limitations

23. Why This Trial Matters Statistically

LUX-Lung 7 is a useful teaching example because the same randomized comparison generates several different statistical questions. The primary endpoints all use time-to-event methods, but their event definitions differ. Secondary endpoints then move into logistic regression, ANCOVA, and mixed-effects modeling.

ConceptHow it appears in LUX-Lung 7
Randomization319 participants were randomized in a parallel two-arm design.
Time-to-event endpointsPFS, TTF, and OS were registered as primary endpoints.
Log-rank testReported for each primary endpoint.
Hazard ratioUsed for PFS, TTF, and OS treatment comparisons.
Stratified analysisEGFR mutation group and presence of baseline brain metastases were used in the reported analyses.
Logistic regressionUsed for Objective Response Rate and Disease Control.
Odds ratioReported for Objective Response Rate and Disease Control.
ANCOVAUsed for tumour shrinkage with baseline covariate adjustment.
Mixed-effects modelUsed for repeated health-related quality-of-life assessments.
Confidence intervalsReported for all posted statistical effect estimates in the ClinicalTrials.gov record.
Exploratory analysisThe registry notes state that no formal hypotheses were tested.

The trial therefore illustrates an important principle in clinical biostatistics: the statistical method should follow the endpoint. A hazard ratio is appropriate to describe a modeled relative event rate for a time-to-event endpoint, whereas an odds ratio describes a binary outcome and a mean difference describes a continuous outcome on its measurement scale.

24. What the Primary Hazard Ratios Do — and Do Not — Mean

PFS HR 0.822

The model estimated a lower instantaneous rate of progression or death with afatinib relative to gefitinib. The descriptive 17.8% lower hazard interpretation is a transformation of the reported HR, not a statement about the percentage of patients who avoided progression.

TTF HR 0.750

The model estimated a lower instantaneous rate of treatment failure with afatinib. The descriptive 25% lower hazard interpretation does not mean that treatment duration increased by 25% for every participant.

OS HR 0.862

The model estimated a lower instantaneous rate of death with afatinib. The descriptive 13.8% lower hazard interpretation does not represent an absolute survival difference or the percentage of participants whose lives were extended.

Across all three endpoints, the most important distinction is between a relative hazard and an absolute clinical quantity. A hazard ratio summarizes relative event rates under a model. It does not directly provide the probability of an event by a specified time, the number needed to treat, or the difference in median event time.

25. Interpreting the P-values Correctly

EndpointP-valueWhat it should not be interpreted as
PFS0.0891Not the probability that the null hypothesis is true; not the size of the treatment effect.
TTF0.0136Not the probability that afatinib is superior; not the percentage reduction in treatment failure.
OS0.2343Not the probability that there is no treatment effect; not a measure of clinical importance.
Objective Response Rate0.3235Not the probability of no response benefit; not a measure of the odds ratio's magnitude.
Disease Control0.7856Not the probability that the two treatments are identical.
Tumour Shrinkage0.0657Not a measure of the size of the -3.45 mm estimate.
EQ-5D UK0.1422Not a measure of the size of the -0.02 mean difference.
EQ-5D Belgium0.0540Not a probability that the treatments are equal.
EQ-VAS0.2032Not a measure of the magnitude of the -1.5 mean difference.

Because the trial is explicitly characterized as exploratory, the p-values are most appropriately read alongside the point estimates, confidence intervals, endpoint definitions, and model specifications. A p-value alone cannot answer whether an effect is clinically meaningful.

26. Trial Timeline

13 December 2011

Trial start

The registry lists 2011-12-13 as the study start date.

08 April 2016

Primary completion

The registry lists 2016-04-08 as the primary completion date. This date is also identified in the TTF endpoint title as the main overall survival analysis cut-off date.

21 August 2015

Quality-of-life analysis cut-off

The quality-of-life analyses use the registered primary analysis cut-off date of 21 August 2015 and assessments every 8 weeks up to 56 weeks.

27. Record-Level Statistical Summary

DimensionSummary
Trial phasePhase 2
Enrollment319
Arms2
AllocationRandomized
DesignParallel
MaskingNone
Primary endpointsProgression-free Survival; Time to Treatment Failure; Overall Survival
Primary endpoint typeTime-to-event
Primary methodsLog-rank test with stratified Cox modeling for reported PFS and OS hazard ratios
Secondary methodsLogistic regression; ANCOVA; mixed-effects model
Effect measuresHazard ratio; odds ratio; mean difference
Results postedYes
Statistical analyses posted9
Exploratory statusAnalysis notes state that no formal hypotheses were tested

28. Related Tutorials

Learn more about the methods used in this trial:

29. Related Statistical Calculators

30. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical methods behind randomized trials, survival analysis, regression, covariate adjustment, and longitudinal modeling.

31. Why the Statistical Story Matters

LUX-Lung 7 demonstrates why clinical-trial interpretation should begin with the estimand rather than the p-value. PFS, TTF, and OS were all primary time-to-event endpoints, but they represent different events. Their hazard ratios therefore answer different questions. The secondary analyses extend the statistical framework further: objective response and disease control use odds ratios, tumour shrinkage uses an adjusted mean difference, and repeated quality-of-life measurements use mixed-effects models.

The primary PFS estimate of 0.822, TTF estimate of 0.750, and OS estimate of 0.862 should consequently be read as three distinct model-based summaries rather than as interchangeable measures of treatment effect. The corresponding confidence intervals show how uncertainty differs among the endpoints. The registry's explicit characterization of the trial as exploratory is equally important: reported p-values are evidence summaries from the posted analyses, not standalone measures of effect magnitude or clinical importance.

Clinical Biostats methodology: A useful trial-analysis page separates the reported numerical result from its statistical interpretation. For LUX-Lung 7, that means preserving the endpoint definitions, analysis populations, stratification factors, model families, effect measures, confidence intervals, p-values, and the registry's explicit exploratory designation rather than reducing the trial to a single statistical conclusion.