This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the information reported in the ClinicalTrials.gov record.
1. Trial at a Glance
LUX-Lung 7 was a randomized, parallel, open-label phase 2 trial comparing afatinib with gefitinib in first-line EGFR mutation-positive adenocarcinoma of the lung. The registry reports 319 randomized participants, two treatment arms, three registered primary time-to-event endpoints, and formal analyses for each primary endpoint.
| Feature | LUX-Lung 7 |
|---|---|
| Trial name | LUX-Lung 7 |
| Brief title | A Phase IIb Trial of Afatinib(BIBW2992) Versus Gefitinib for the Treatment of 1st Line EGFR Mutation Positive Adenocarcinoma of the Lung |
| Phase | Phase 2 |
| Status | Completed |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 319 |
| Interventions | Afatinib (drug); gefitinib (drug) |
| Registered primary endpoints | 3 |
| Primary endpoint type | Time-to-event |
| Results posted | Yes |
| Statistical analyses posted | 9 |
| Lead sponsor | Boehringer Ingelheim |
| Sponsor type | Industry |
2. Clinical Question
The central statistical question was whether first-line afatinib produced different time-to-event outcomes from gefitinib in participants with EGFR mutation-positive adenocarcinoma of the lung.
Population
Participants with 1st line EGFR mutation positive adenocarcinoma of the lung, as described by the registered trial title.
Intervention
Afatinib.
Comparator
Gefitinib.
Primary question
How do afatinib and gefitinib compare for progression-free survival, time to treatment failure, and overall survival?
3. Trial Design
Afatinib
- Afatinib (drug)
- Compared directly with gefitinib
- Included in the randomized set used for primary efficacy analyses
Gefitinib
- Gefitinib (drug)
- Comparator treatment
- Included in the randomized set used for primary efficacy analyses
4. Endpoints
The registry identifies three primary endpoints, all of which are time-to-event outcomes. Their definitions and time frames are important because the event definition determines what the hazard ratio represents.
| Endpoint | Registered definition | Time frame |
|---|---|---|
| Progression-free Survival | Progression-free survival (PFS) defined as the time from date of randomisation to date of disease progression, or date of death if a patient died earlier. Participants with no event (Disease progression (PD) or death) were censored. PD was primarily evaluated for the primary analysis by an independent central imaging review according to Response Evaluation Criteria in Solid Tumours (RECIST) version 1.1. Per RECIST version 1.1. for target lesions and assessed by Computed Tomography (CT)-scan or Magnetic Resonance Imaging (MRI): PD, At least a 20% increase in the sum of the longest diameter (SoD) of target lesions taking as reference the smallest SoD of target lesions recorded since the treatment started, together with an absolute increase in the SoD of target lesions of at least 5 millimetre (mm) or the appearance of one or more new lesions. For the final analysis (analysis cut-off date 12 April 2019) status and date of PD were determined by investigator assessment. | From first drug administration until 28 days after last drug administration + Follow-Up period for collecting information on disease progression or death, up to 2465 days. |
| Time to Treatment Failure (TTF) (Main Overall Survival Analysis Cut-off Date, 08 April 2016) | Time to Treatment Failure (TTF) which was the time from the date of randomisation to the date of i.e. permanent treatment discontinuation for any reason. | From first drug administration until last drug administration, up to 1482 days |
| Overall Survival | Overall survival (OS) which was defined as the time from the date of randomisation to the date of death. Participants for whom there is no evidence of death at the time of the analysis will be censored at the date that they were last known to be alive. | From first drug administration until 28 days after last drug administration + Follow-Up period for collecting information on disease progression or death, up to 2465 days. |
Secondary endpoints with posted analyses
| Endpoint | Type | Analysis method | Effect measure |
|---|---|---|---|
| Objective Response Rate | Count / rate | Logistic regression | Odds ratio |
| Disease Control | Binary | Logistic regression | Odds ratio |
| Tumour Shrinkage (Main Overall Survival Analysis Cut-off Date, 08 April 2016) | Time-to-event | ANCOVA | Mean difference |
| Health-related Quality of Life (Primary Analysis Cut-off Date, 21 August 2015) | Continuous | Mixed-effects model | Mean difference |
| Health-related Quality of Life (Primary Analysis Cut-off Date, 21 August 2015) | Continuous | Mixed-effects model | Mean difference |
| Health-related Quality of Life (Primary Analysis Cut-off Date, 21 August 2015) | Continuous | Mixed-effects model | Mean difference |
5. Analysis Populations and Stratification
The three primary endpoint analyses used the randomised set, defined in the registry as all patients randomized to receive treatment, whether treated or not. This is an important feature of the primary efficacy analysis because treatment assignment, rather than subsequent treatment exposure, defines membership in the analysis population.
| Analysis | Population | Key analytical feature |
|---|---|---|
| Primary endpoints | Randomised set | All patients randomized to receive treatment, whether treated or not |
| Objective Response Rate | Randomised set | Logistic regression, stratified analysis |
| Disease Control | Randomised set | Logistic regression, stratified analysis |
| Tumour Shrinkage | Participants in the randomised set with tumour assessments | ANCOVA with baseline and stratification-related covariate adjustment |
| Health-related Quality of Life | All randomised subjects with health-related quality of life data | Mixed-effects model with covariate adjustment |
The primary survival analyses were stratified by EGFR mutation group and presence of brain metastases at baseline. This same stratification structure appears in the primary Cox proportional-hazards analyses and the secondary logistic-regression analyses.
6. Statistical Methodology
Kaplan-Meier estimation and time-to-event analysis
The primary endpoints are time-to-event outcomes. For PFS, an event is disease progression or death, while the registered OS definition uses death as the event and censors participants with no evidence of death at the analysis time. TTF instead uses permanent treatment discontinuation for any reason.
These definitions illustrate why the endpoint itself matters. Three endpoints can all be expressed in months and analyzed using survival methods while answering different clinical questions:
| Endpoint | Event represented | Interpretive question |
|---|---|---|
| PFS | Disease progression or death | How long until progression or death? |
| TTF | Permanent treatment discontinuation for any reason | How long until treatment failure? |
| OS | Death | How long until death? |
The Kaplan-Meier framework estimates the probability of remaining event-free over time while accommodating right-censored observations.
Log-rank testing
The registry reports the log-rank test for all three primary endpoints. The log-rank approach compares the observed and expected numbers of events between treatment groups over the follow-up period. In this trial, the analysis text additionally identifies stratification by EGFR mutation group and presence of brain metastases at baseline.
Cox proportional-hazards modeling
The registry analysis notes state that a Cox proportional-hazards model, stratified by EGFR mutation group and presence of baseline brain metastases, was used to estimate the hazard ratio for PFS and OS. For these endpoints, the reported ratio was calculated as Afatinib divided by Gefitinib.
An HR below 1 indicates a lower estimated instantaneous event rate in the afatinib group under the fitted model. It is a relative time-to-event measure, not a probability and not an absolute difference in survival time.
Logistic regression
Objective Response Rate and Disease Control were analyzed using logistic regression. Both analyses used the randomized set and incorporated stratification for EGFR mutation group and presence of brain metastases at baseline. The effect measure was an odds ratio calculated as Afatinib divided by Gefitinib.
ANCOVA
Tumour Shrinkage was analyzed using ANCOVA among participants in the randomized set with tumour assessments. The analysis adjusted for baseline sum of diameters, EGFR mutation group, and presence of brain metastases at baseline. The reported mean difference was calculated as Afatinib minus Gefitinib.
Mixed-effects models
Health-related quality of life was measured every 8 weeks, up to 56 weeks, and analyzed with mixed-effects models. The reported analyses adjusted for EGFR mutation group and presence of brain metastases at baseline. This is a natural setting for a longitudinal model because the same participants can contribute repeated quality-of-life observations over time.
7. Primary Results: Progression-free Survival
The registry reports a formal primary analysis of PFS using a log-rank test and a hazard ratio from a stratified Cox proportional-hazards model. The analysis population was the randomized set.
Hazard ratio for progression or death
95% CI: 0.655–1.032 · P = 0.0891
Hazard ratio calculated as Afatinib divided by Gefitinib.
The estimated HR of 0.822 means that the fitted model estimated the instantaneous rate of progression or death in the afatinib group to be about 17.8% lower than in the gefitinib group, under the proportional-hazards model and the specified stratification.
The HR does not mean that 17.8% fewer participants progressed, nor does it mean that every participant experienced a 17.8% reduction in risk. It is a relative time-to-event estimate.
The 95% CI of 0.655–1.032 describes uncertainty around the estimated hazard ratio. Its width indicates that the point estimate should not be interpreted as an exact treatment effect. Because the interval includes 1, the data are compatible with both a lower and a higher hazard under the model.
The p-value of 0.0891 measures evidence against the specified null comparison under the analysis framework; it does not measure the size or clinical importance of the effect. Most importantly, the registry analysis notes state that this was an exploratory trial and no formal hypotheses were tested, so the p-value should not be converted into a claim of formal statistical success or failure.
The interpretation also depends on censoring and on the proportional-hazards modeling framework. A single HR summarizes a relative hazard over the analyzed follow-up; it is not a direct description of absolute survival probabilities at a particular time.
8. Primary Results: Time to Treatment Failure
TTF was a registered primary endpoint with a time frame extending from first drug administration until last drug administration, up to 1482 days. The registered definition describes TTF as the time from randomization to permanent treatment discontinuation for any reason.
Hazard ratio for treatment failure
95% CI: 0.595–0.944 · P = 0.0136
Ratio calculated as Afatinib divided by Gefitinib.
The estimated HR of 0.750 corresponds to approximately a 25% lower estimated instantaneous rate of treatment failure in the afatinib group relative to the gefitinib group, using the reported hazard-ratio definition.
This does not mean that 25% of participants avoided treatment failure, nor does it establish a 25% increase in treatment duration for an individual participant. The hazard ratio describes a relative event rate over time.
The 95% CI of 0.595–0.944 quantifies uncertainty around the estimated relative hazard. Unlike the PFS interval, the reported interval lies below 1, indicating that the estimated treatment-failure hazard was lower for afatinib across the interval represented by this analysis.
The p-value of 0.0136 quantifies evidence under the specified statistical comparison; it is not an effect-size measure. The registry explicitly describes the trial as exploratory and states that no formal hypotheses were tested. Therefore, the numerical p-value should be interpreted as part of the reported exploratory analysis rather than as a standalone confirmatory claim.
9. Primary Results: Overall Survival
OS was defined as the time from randomization to death. Participants with no evidence of death at analysis were censored at the date they were last known to be alive.
Hazard ratio for death
95% CI: 0.674–1.101 · P = 0.2343
Hazard ratio calculated as Afatinib divided by Gefitinib.
The HR of 0.862 corresponds to an estimated instantaneous rate of death approximately 13.8% lower in the afatinib group than in the gefitinib group under the fitted stratified Cox model.
Again, the HR does not represent an absolute survival difference, a difference in median survival, or the percentage of participants who benefited. Those quantities require different information and different estimands.
The 95% CI of 0.674–1.101 spans 1. The interval therefore reflects substantial uncertainty about the direction and magnitude of the relative hazard, even though the point estimate is below 1.
The p-value of 0.2343 is evidence summarized under the specified statistical test; it does not quantify the magnitude of the observed HR. As with the other primary endpoints, the registry states that the trial was exploratory and that no formal hypotheses were tested.
OS is also particularly sensitive to events and censoring during long-term follow-up. The registered definition makes clear that participants without evidence of death were censored at the date they were last known to be alive.
10. Comparison of the Three Primary Endpoints
| Primary endpoint | Method | Effect | 95% CI | P-value |
|---|---|---|---|---|
| Progression-free Survival | Log-rank; stratified Cox model | HR 0.822 | 0.655–1.032 | 0.0891 |
| Time to Treatment Failure | Log-rank; stratified analysis | HR 0.750 | 0.595–0.944 | 0.0136 |
| Overall Survival | Log-rank; stratified Cox model | HR 0.862 | 0.674–1.101 | 0.2343 |
The three estimates should not be collapsed into a single measure of "benefit." They represent different event definitions. PFS captures progression or death, TTF captures permanent treatment discontinuation for any reason, and OS captures death. Consequently, the fact that their hazard ratios differ is not itself surprising or evidence of inconsistency.
PFS
The point estimate was below 1, with HR 0.822, but the 95% CI extended above 1.
TTF
The point estimate was 0.750, with a 95% CI of 0.595–0.944.
OS
The point estimate was below 1 at 0.862, while the 95% CI extended from 0.674 to 1.101.
Exploratory status
The registry analysis notes explicitly state that no formal hypotheses were tested.
11. Secondary Results: Objective Response Rate
Objective Response Rate was analyzed as a percentage of participants using logistic regression. The analysis used the randomized set and was stratified for EGFR mutation group and presence of brain metastases at baseline.
Odds ratio for objective response
95% CI: 0.768–2.223 · P = 0.3235
Odds ratio calculated as Afatinib divided by Gefitinib.
An OR of 1.307 means that the estimated odds of objective response were 1.307 times as high with afatinib as with gefitinib under the reported logistic-regression model.
Odds are not probabilities. An odds ratio of 1.307 therefore cannot be interpreted as a 30.7% higher response rate without knowing the underlying response probabilities.
The 95% CI of 0.768–2.223 is relatively broad and includes 1. The p-value of 0.3235 describes the evidence under the model and does not measure the size of the estimated odds ratio.
The registry again describes this as an exploratory trial with no formal hypotheses tested. The reported stratification also matters: the OR is not simply an unadjusted ratio of two raw response proportions.
12. Secondary Results: Disease Control
Disease Control was a binary secondary endpoint analyzed using logistic regression in the randomized set, with stratification for EGFR mutation group and presence of brain metastases at baseline.
Odds ratio for disease control
95% CI: 0.447–2.896 · P = 0.7856
Odds ratio calculated as Afatinib divided by Gefitinib.
The OR of 1.138 indicates an estimated odds of disease control 1.138 times the odds under gefitinib, according to the reported logistic-regression analysis.
The confidence interval, 0.447–2.896, is wide and spans 1 by a substantial amount. This indicates that the point estimate alone provides an incomplete picture of the uncertainty in the treatment comparison.
The p-value of 0.7856 is a measure associated with the statistical comparison, not a probability that the treatment effect is zero and not a measure of clinical relevance.
13. Secondary Results: Tumour Shrinkage
Tumour Shrinkage was reported in millimetres and analyzed using ANCOVA among participants in the randomized set with tumour assessments. The analysis adjusted for baseline sum of diameters, EGFR mutation group, and presence of brain metastases at baseline.
Adjusted mean difference in final values
95% CI: -7.13 to 0.23 · P = 0.0657
Difference calculated as Afatinib minus Gefitinib.
The negative mean difference of -3.45 means that the adjusted final tumour-shrinkage measure was estimated to be 3.45 mm lower with afatinib than with gefitinib, using the registry's stated direction of subtraction.
The estimate is adjusted rather than simply comparing two unadjusted means. The covariates were baseline sum of diameters, EGFR mutation group, and presence of brain metastases at baseline.
The 95% CI of -7.13 to 0.23 includes zero, which means the uncertainty interval contains both a negative and a small positive mean difference. The p-value of 0.0657 does not quantify the magnitude of the 3.45 mm estimate.
14. Secondary Results: Health-Related Quality of Life
Health-related quality of life was assessed every 8 weeks, up to 56 weeks. Three posted analyses are reported in the ClinicalTrials.gov record: EQ-5D UK utility score, EQ-5D Belgium utility score, and EQ-VAS utility score. All used mixed-effects models adjusted for EGFR mutation group and presence of brain metastases at baseline.
| Quality-of-life measure | Mean difference | 95% CI | P-value |
|---|---|---|---|
| EQ-5D UK utility score | -0.02 | -0.06 to 0.01 | 0.1422 |
| EQ-5D Belgium utility score | -0.03 | -0.06 to 0.00 | 0.0540 |
| EQ-VAS utility score | -1.5 | -3.9 to 0.8 | 0.2032 |
Each mean difference is calculated with afatinib as the treatment group and gefitinib as the comparator. A negative value therefore indicates a lower estimated final value in the afatinib group under the reported model.
The mixed-effects approach is important because quality-of-life data were collected repeatedly from participants over time. Rather than treating every measurement as an independent observation, a longitudinal mixed model can account for the repeated-measures structure while incorporating the stated covariate adjustments.
The confidence intervals should be read independently for each quality-of-life measure. For example, the EQ-5D UK interval of -0.06 to 0.01 includes zero, as does the EQ-VAS interval of -3.9 to 0.8. The EQ-5D Belgium interval extends to 0.00. None of these p-values should be interpreted as a measure of the size of the mean difference.
15. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected participants divided by participants at risk.
| Safety measure | Afatinib | Gefitinib |
|---|---|---|
| Serious adverse events | 75/160 | 64/159 |
These are counts of affected participants over the reported at-risk denominators. They are not a hazard ratio, odds ratio, or risk difference, and the ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse events by arm.
16. Statistical Methods Explained
Why was a stratified Cox model used?
The registry states that PFS and OS were analyzed with a Cox proportional-hazards model stratified by EGFR mutation group and presence of baseline brain metastases. Stratification allows the baseline hazard to differ across these strata while estimating a treatment hazard ratio within the specified model structure.
What does an HR of 0.750 mean for TTF?
Because the registry defines the ratio as afatinib divided by gefitinib, 0.750 indicates a lower estimated instantaneous rate of treatment failure for afatinib. The corresponding descriptive transformation is a 25% lower estimated hazard. It does not mean that treatment failure was prevented in exactly 25% of participants.
Why is PFS different from TTF?
PFS is defined around disease progression or death. TTF is defined around permanent treatment discontinuation for any reason. Discontinuation and disease progression are therefore different events, even though both can be analyzed using time-to-event methods.
What does an odds ratio of 1.307 mean?
The OR of 1.307 for objective response means that the modeled odds of response were 1.307 times those under gefitinib. An odds ratio is not a risk ratio or percentage-point difference, so the underlying response probabilities would be required to translate it into an absolute probability comparison.
Why was ANCOVA used for tumour shrinkage?
The registry reports ANCOVA with adjustment for baseline sum of diameters, EGFR mutation group, and presence of brain metastases at baseline. Baseline tumour size is especially relevant because a final tumour measurement can be influenced by the starting measurement. Covariate adjustment estimates the treatment-group difference after accounting for the specified baseline information.
Why use a mixed-effects model for quality of life?
Quality of life was measured every 8 weeks up to 56 weeks, so a participant could contribute multiple observations. A mixed-effects model is designed for this repeated-measures structure and can estimate treatment differences while accounting for the longitudinal nature of the data and the stated covariate adjustments.
Does P = 0.0136 prove that afatinib is superior?
No. The p-value is part of the reported exploratory analysis and does not measure effect size. The registry analysis notes explicitly state that the trial was exploratory and that no formal hypotheses were tested. The appropriate interpretation is therefore to consider the estimated effect, confidence interval, endpoint definition, analysis model, and exploratory status together.
17. Understanding Confidence Intervals in LUX-Lung 7
Confidence intervals are particularly informative in this trial because the same type of effect measure can have very different levels of precision across endpoints.
| Endpoint | Estimate | 95% CI | What the interval tells us |
|---|---|---|---|
| PFS HR | 0.822 | 0.655–1.032 | The interval includes 1, so uncertainty includes both a lower and higher hazard relative to the comparator. |
| TTF HR | 0.750 | 0.595–0.944 | The interval is below 1 and therefore excludes an equal hazard within this reported interval. |
| OS HR | 0.862 | 0.674–1.101 | The interval includes 1 and extends above it. |
| ORR OR | 1.307 | 0.768–2.223 | The interval includes 1 and is relatively broad. |
| Disease Control OR | 1.138 | 0.447–2.896 | The interval includes 1 and is broad relative to the point estimate. |
| Tumour Shrinkage mean difference | -3.45 | -7.13 to 0.23 | The interval includes zero, the value representing no mean difference. |
The null reference depends on the effect measure. For hazard ratios and odds ratios, the reference value is 1. For a mean difference, the reference value is 0. This distinction is fundamental when reading confidence intervals.
18. Multiplicity, Hypothesis Testing, and the Exploratory Design
The registry lists three primary endpoints, all with a superiority hypothesis type in the posted statistical analyses. However, the analysis notes for the primary endpoints explicitly state: "Exploratory trial, no formal hypotheses were tested."
| Feature | What the ClinicalTrials.gov record reports | Statistical implication |
|---|---|---|
| Primary endpoints | 3 | PFS, TTF, and OS each have their own treatment comparison. |
| Hypothesis type | Superiority | The posted analysis fields identify superiority as the hypothesis type. |
| Trial characterization | Exploratory | The analysis notes state that no formal hypotheses were tested. |
| Statistical analyses | 9 posted | Primary and secondary endpoints use several different model families. |
Because the ClinicalTrials.gov record does not report a multiplicity-adjustment procedure, alpha allocation, or a formal hierarchical testing strategy, this page does not infer one. The safest statistical reading is to preserve the registry's explicit distinction between reported p-values and the absence of formal hypothesis testing.
19. Censoring and the Meaning of Time-to-Event Results
Time-to-event analysis differs from a simple comparison of event proportions because participants may have different lengths of observation. The registry explicitly describes censoring for PFS and OS.
PFS censoring
Participants with no event, defined as disease progression or death, were censored.
OS censoring
Participants with no evidence of death at analysis were censored at the date they were last known to be alive.
TTF event
The event was permanent treatment discontinuation for any reason.
Why this matters
The definition of the event and the rules for censoring determine the estimand being analyzed.
A hazard ratio should therefore be interpreted in the context of the event definition and censoring rules. The OS HR and PFS HR are not interchangeable simply because both are produced by Cox models.
20. Covariate Adjustment and Stratified Analysis
LUX-Lung 7 provides several examples of how baseline information can enter an analysis. EGFR mutation group and baseline brain metastases appear repeatedly as stratification factors. Baseline sum of diameters additionally enters the ANCOVA for tumour shrinkage.
| Method | Adjustment / stratification reported |
|---|---|
| PFS Cox model | EGFR mutation group; presence of baseline brain metastases |
| OS Cox model | EGFR mutation group; presence of baseline brain metastases |
| TTF analysis | EGFR mutation group; presence of baseline brain metastases |
| Objective Response Rate | EGFR mutation group; presence of baseline brain metastases |
| Disease Control | EGFR mutation group; presence of baseline brain metastases |
| Tumour Shrinkage ANCOVA | Baseline sum of diameters; EGFR mutation group; presence of baseline brain metastases |
| Quality of Life mixed-effects models | EGFR mutation group; presence of baseline brain metastases |
Stratification and covariate adjustment are not interchangeable concepts. A stratified survival model allows the baseline hazard to vary by stratum, whereas ANCOVA directly adjusts a continuous outcome for specified covariates. The choice should therefore be understood as part of the endpoint-specific statistical model.
21. Primary and Secondary Analysis Map
| Role | Endpoint | Model family | Effect measure |
|---|---|---|---|
| Primary | Progression-free Survival | Survival analysis | Hazard ratio |
| Primary | Time to Treatment Failure | Survival analysis | Hazard ratio |
| Primary | Overall Survival | Survival analysis | Hazard ratio |
| Secondary | Objective Response Rate | Categorical data | Odds ratio |
| Secondary | Disease Control | Categorical data | Odds ratio |
| Secondary | Tumour Shrinkage | Linear models | Mean difference |
| Secondary | Health-related Quality of Life | Longitudinal / mixed models | Mean difference |
This endpoint map is statistically useful because it demonstrates that a randomized clinical trial can contain several distinct estimands and model families. Survival outcomes require handling of event times and censoring; binary outcomes use odds-based regression; continuous tumour measurements use covariate-adjusted linear modeling; and repeated quality-of-life measurements use longitudinal models.
22. Limitations
- Exploratory status: the registry analysis notes state that the trial was exploratory and that no formal hypotheses were tested. Reported p-values therefore should not be treated as if they came from a conventional confirmatory testing hierarchy unless additional statistical documentation establishes that framework.
- Multiple primary endpoints: three primary endpoints were analyzed. The ClinicalTrials.gov record does not specify a multiplicity-adjustment strategy, alpha allocation, or hierarchical testing procedure.
- Hazard-ratio interpretation: the primary effect measures are relative hazards. They do not directly describe absolute differences in survival probability or treatment duration.
- Proportional-hazards assumption: Cox hazard ratios rely on a proportional-hazards model. The ClinicalTrials.gov record does not report a formal assessment of that assumption.
- Censoring: PFS and OS analyses depend on the stated censoring rules. The ClinicalTrials.gov record does not provide individual-level event and censoring records with which to independently verify the survival estimates.
- Analysis populations differ: primary survival analyses use the randomized set, while tumour shrinkage requires participants with tumour assessments and quality-of-life analyses require randomized subjects with quality-of-life data.
- Secondary endpoints: several secondary analyses were conducted, and the ClinicalTrials.gov record does not specify a multiplicity-control strategy for these comparisons.
- Safety comparison: the ClinicalTrials.gov record provides serious adverse-event counts by arm but do not provide a formal statistical comparison of those safety data.
23. Why This Trial Matters Statistically
LUX-Lung 7 is a useful teaching example because the same randomized comparison generates several different statistical questions. The primary endpoints all use time-to-event methods, but their event definitions differ. Secondary endpoints then move into logistic regression, ANCOVA, and mixed-effects modeling.
| Concept | How it appears in LUX-Lung 7 |
|---|---|
| Randomization | 319 participants were randomized in a parallel two-arm design. |
| Time-to-event endpoints | PFS, TTF, and OS were registered as primary endpoints. |
| Log-rank test | Reported for each primary endpoint. |
| Hazard ratio | Used for PFS, TTF, and OS treatment comparisons. |
| Stratified analysis | EGFR mutation group and presence of baseline brain metastases were used in the reported analyses. |
| Logistic regression | Used for Objective Response Rate and Disease Control. |
| Odds ratio | Reported for Objective Response Rate and Disease Control. |
| ANCOVA | Used for tumour shrinkage with baseline covariate adjustment. |
| Mixed-effects model | Used for repeated health-related quality-of-life assessments. |
| Confidence intervals | Reported for all posted statistical effect estimates in the ClinicalTrials.gov record. |
| Exploratory analysis | The registry notes state that no formal hypotheses were tested. |
The trial therefore illustrates an important principle in clinical biostatistics: the statistical method should follow the endpoint. A hazard ratio is appropriate to describe a modeled relative event rate for a time-to-event endpoint, whereas an odds ratio describes a binary outcome and a mean difference describes a continuous outcome on its measurement scale.
24. What the Primary Hazard Ratios Do — and Do Not — Mean
The model estimated a lower instantaneous rate of progression or death with afatinib relative to gefitinib. The descriptive 17.8% lower hazard interpretation is a transformation of the reported HR, not a statement about the percentage of patients who avoided progression.
The model estimated a lower instantaneous rate of treatment failure with afatinib. The descriptive 25% lower hazard interpretation does not mean that treatment duration increased by 25% for every participant.
The model estimated a lower instantaneous rate of death with afatinib. The descriptive 13.8% lower hazard interpretation does not represent an absolute survival difference or the percentage of participants whose lives were extended.
Across all three endpoints, the most important distinction is between a relative hazard and an absolute clinical quantity. A hazard ratio summarizes relative event rates under a model. It does not directly provide the probability of an event by a specified time, the number needed to treat, or the difference in median event time.
25. Interpreting the P-values Correctly
| Endpoint | P-value | What it should not be interpreted as |
|---|---|---|
| PFS | 0.0891 | Not the probability that the null hypothesis is true; not the size of the treatment effect. |
| TTF | 0.0136 | Not the probability that afatinib is superior; not the percentage reduction in treatment failure. |
| OS | 0.2343 | Not the probability that there is no treatment effect; not a measure of clinical importance. |
| Objective Response Rate | 0.3235 | Not the probability of no response benefit; not a measure of the odds ratio's magnitude. |
| Disease Control | 0.7856 | Not the probability that the two treatments are identical. |
| Tumour Shrinkage | 0.0657 | Not a measure of the size of the -3.45 mm estimate. |
| EQ-5D UK | 0.1422 | Not a measure of the size of the -0.02 mean difference. |
| EQ-5D Belgium | 0.0540 | Not a probability that the treatments are equal. |
| EQ-VAS | 0.2032 | Not a measure of the magnitude of the -1.5 mean difference. |
Because the trial is explicitly characterized as exploratory, the p-values are most appropriately read alongside the point estimates, confidence intervals, endpoint definitions, and model specifications. A p-value alone cannot answer whether an effect is clinically meaningful.
26. Trial Timeline
Trial start
The registry lists 2011-12-13 as the study start date.
Primary completion
The registry lists 2016-04-08 as the primary completion date. This date is also identified in the TTF endpoint title as the main overall survival analysis cut-off date.
Quality-of-life analysis cut-off
The quality-of-life analyses use the registered primary analysis cut-off date of 21 August 2015 and assessments every 8 weeks up to 56 weeks.
27. Record-Level Statistical Summary
| Dimension | Summary |
|---|---|
| Trial phase | Phase 2 |
| Enrollment | 319 |
| Arms | 2 |
| Allocation | Randomized |
| Design | Parallel |
| Masking | None |
| Primary endpoints | Progression-free Survival; Time to Treatment Failure; Overall Survival |
| Primary endpoint type | Time-to-event |
| Primary methods | Log-rank test with stratified Cox modeling for reported PFS and OS hazard ratios |
| Secondary methods | Logistic regression; ANCOVA; mixed-effects model |
| Effect measures | Hazard ratio; odds ratio; mean difference |
| Results posted | Yes |
| Statistical analyses posted | 9 |
| Exploratory status | Analysis notes state that no formal hypotheses were tested |
28. Related Tutorials
Learn more about the methods used in this trial:
29. Related Statistical Calculators
30. Sources
- ClinicalTrials.gov: LUX-Lung 7, NCT01466660.
- PubMed: Publication associated with LUX-Lung 7, PMID 29653820.
- PubMed: Publication associated with LUX-Lung 7, PMID 28426106.
- PubMed: Publication associated with LUX-Lung 7, PMID 27083334.
Continue through the Clinical Biostats statistical pathway
Explore the statistical methods behind randomized trials, survival analysis, regression, covariate adjustment, and longitudinal modeling.
31. Why the Statistical Story Matters
LUX-Lung 7 demonstrates why clinical-trial interpretation should begin with the estimand rather than the p-value. PFS, TTF, and OS were all primary time-to-event endpoints, but they represent different events. Their hazard ratios therefore answer different questions. The secondary analyses extend the statistical framework further: objective response and disease control use odds ratios, tumour shrinkage uses an adjusted mean difference, and repeated quality-of-life measurements use mixed-effects models.
The primary PFS estimate of 0.822, TTF estimate of 0.750, and OS estimate of 0.862 should consequently be read as three distinct model-based summaries rather than as interchangeable measures of treatment effect. The corresponding confidence intervals show how uncertainty differs among the endpoints. The registry's explicit characterization of the trial as exploratory is equally important: reported p-values are evidence summaries from the posted analyses, not standalone measures of effect magnitude or clinical importance.