This page separates reported trial results from statistical interpretation. This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
LUX-Lung 8 was a randomized, parallel-group, phase 3 trial comparing afatinib with erlotinib in patients with carcinoma of the non-small-cell lung after at least one prior platinum-based chemotherapy. The registry identifies progression-free survival based on central independent review according to RECIST 1.1 as the single registered primary endpoint.
| Feature | LUX-Lung 8 |
|---|---|
| Phase | Phase 3 |
| Condition | Carcinoma, Non-Small-Cell Lung |
| Trial population described in the brief title | Squamous cell lung cancer after at least one prior platinum-based chemotherapy |
| Design | Randomized, parallel-group, open-label |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 795 |
| Interventions | Afatinib and erlotinib |
| Primary endpoint type | Time-to-event |
| Results posted | Yes |
| Outcome measures posted | 8 |
| Statistical analyses posted | 11 |
| Hypothesis type | Superiority |
| Lead sponsor | Boehringer Ingelheim |
| Trial status | Completed |
| ClinicalTrials.gov | NCT01523587 |
2. Clinical Question
The trial addresses whether afatinib provides a different time-to-event outcome from erlotinib in patients with squamous cell lung cancer after at least one prior platinum-based chemotherapy. The registered hypothesis type is superiority, so the statistical question is framed as whether the treatment groups differ in the prespecified direction rather than whether afatinib merely achieves a predefined non-inferiority standard.
Population
Patients with squamous cell lung cancer, as described in the trial brief title, after at least one prior platinum-based chemotherapy.
Intervention
Afatinib.
Comparator
Erlotinib.
Primary question
Does afatinib improve progression-free survival compared with erlotinib when progression is determined by central independent review according to RECIST 1.1?
3. Trial Design
Afatinib
- Afatinib was the investigational treatment.
- The registry compares this group directly with erlotinib.
- Primary efficacy analysis used the randomized set.
Erlotinib
- Erlotinib was the comparator treatment.
- The registry compares this group directly with afatinib.
- Primary efficacy analysis used the randomized set.
The trial began on 05 March 2012 and had a primary completion date of 21 October 2013. The primary progression-free survival analysis used a later cutoff date of 02 March 2015, while several secondary outcomes used a study-closure period extending to 27 December 2017.
4. Endpoints
Primary endpoint
| Endpoint | Registry definition / time frame | Type |
|---|---|---|
| Progression-free Survival, Based on Central Independent Review as Determined by Response Evaluation Criteria in Solid Tumours 1.1 | First treatment administration up until cut off date of 02 March 2015 (up to 1058 days). Progression Free Survival was defined as the time from randomization to disease progression (or death if the patient died before progression) by central independent review according to Response Evaluation Criteria in Solid Tumours (RECIST) version 1.1. | Time-to-event |
Secondary endpoints with posted statistical analyses
| Endpoint | Time frame | Type |
|---|---|---|
| Overall Survival | From first drug administration from 9 April 2012 until study closure on 27 Dec 2017 (approximately 2089 days). | Time-to-event |
| Number of Participants With Objective Response According to RECIST 1.1 | First treatment administration up until cut off date of 02 March 2015 (up to 1058 days). | Binary |
| Number of Participants With Disease Control According to RECIST 1.1 | First treatment administration up until cut off date of 02 March 2015 (up to 1058 days). | Binary |
| Tumour Shrinkage | First treatment administration up until cut off date of 02 March 2015 (up to 1058 days). | Continuous |
| Summary of Time to Deterioration in Coughing, Dyspnoea and Pain | From first drug administration from 9 April 2012 until study closure on 27 Dec 2017 (approximately 2089 days). | Time-to-event |
| Change in Score Over Time in Coughing,Dyspnoea and Pain | From first drug administration from 9 April 2012 until study closure on 27 Dec 2017 (approximately 2089 days). | Time-to-event in the registry analysis record |
5. Statistical Methodology
The posted analyses use four main statistical method families: log-rank testing and Cox proportional-hazards models for time-to-event outcomes, logistic regression for binary outcomes, and ANCOVA for tumour shrinkage. The registry also identifies stratified analysis and covariate adjustment in specific analyses.
| Method | Used for | Effect measure | Statistical role |
|---|---|---|---|
| Log-rank test | Progression-free survival; overall survival | Hazard ratio reported alongside the comparison | Comparison of time-to-event distributions |
| Cox proportional-hazards model | Overall survival; symptom deterioration; registry-reported symptom score analyses | Hazard ratio or, for the score analyses, mean difference in final values | Model-based time-to-event comparison |
| Logistic regression | Objective response; disease control | Odds ratio | Comparison of binary outcomes |
| ANCOVA | Tumour shrinkage | Adjusted mean / mean difference | Covariate-adjusted continuous-outcome comparison |
Analysis populations
The primary progression-free survival analysis was conducted in the Randomized Set (RS), defined in the registry as all patients who were randomized, regardless of whether they received investigational treatment. The overall survival, response, disease-control, and symptom analyses also identify the RS as their analysis population.
Tumour shrinkage is different: the registry states that patients from the randomized set with tumour assessments were considered for that endpoint. That distinction matters because availability of a tumour assessment can differ from simply being randomized.
Stratified analysis
Several analyses are identified as stratified. The overall survival analysis used a Cox proportional-hazards model stratified by race. The disease-control analysis used logistic regression stratified by race. The symptom deterioration analyses similarly identify Cox models stratified by race.
Stratification allows the analysis to compare treatment groups within levels of a factor rather than treating the factor as irrelevant to the comparison. In this registry record, race is explicitly identified as a stratification variable for several secondary analyses.
6. Primary Result: Progression-Free Survival
The primary endpoint was progression-free survival based on central independent review according to RECIST 1.1. The analysis compared afatinib with erlotinib in the randomized set using a log-rank test. The registry also notes that a Cox proportional-hazards model without the randomization stratification variable was used for each subgroup category, along with the corresponding log-rank test.
Hazard ratio for progression or death
95% CI: 0.693–0.956 · P = 0.0103
Two-sided superiority analysis in the Randomized Set.
| Primary PFS result | Reported value |
|---|---|
| Groups compared | Afatinib vs Erlotinib |
| Analysis population | Randomized Set (RS) |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.814 |
| 95% CI | 0.693–0.956 |
| P-value | 0.0103 |
| Hypothesis type | Superiority |
What the estimate means: a hazard ratio of 0.814 means that, under the time-to-event model used for the comparison, the estimated instantaneous rate of progression or death for afatinib relative to erlotinib was 0.814. Expressed as a simple relative interpretation, this corresponds to an estimated hazard that was about 18.6% lower for afatinib than erlotinib.
What it does not mean: it does not mean that 18.6% fewer patients necessarily experienced progression or death, and it does not mean that every patient had an 18.6% reduction in risk. A hazard ratio is a relative time-to-event measure, not an absolute risk difference.
What the confidence interval says: the 95% confidence interval of 0.693–0.956 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It gives a range of values compatible with the observed data and model assumptions; it is not a range of individual patient effects.
What the p-value says: the two-sided p-value of 0.0103 quantifies the evidence against the null hypothesis specified for the statistical comparison. It does not measure the size or clinical importance of the treatment effect. Effect size and uncertainty are conveyed by the hazard ratio and confidence interval.
Caution: interpretation of a single Cox hazard ratio relies on the model's proportional-hazards framework. If the relative hazard changes substantially over time, one summary hazard ratio may not fully describe the treatment difference. The ClinicalTrials.gov record does not provide a time-varying hazard assessment, so the HR should be interpreted within the stated model rather than as a universal constant risk reduction at every time point.
7. Secondary Result: Overall Survival
Overall survival was evaluated from first drug administration from 9 April 2012 until study closure on 27 Dec 2017, approximately 2089 days. The analysis population was the randomized set. The registry reports a log-rank analysis and states that a Cox proportional-hazards model stratified by race was used to estimate the hazard ratio and 95% confidence interval between the two treatment groups.
Hazard ratio for overall survival
95% CI: 0.727–0.973 · P = 0.0193
Two-sided superiority analysis in the Randomized Set.
| Overall survival result | Reported value |
|---|---|
| Groups compared | Afatinib vs Erlotinib |
| Analysis population | RS |
| Method | Log-rank test |
| Cox model | Stratified by race |
| Effect measure | Hazard ratio |
| Estimate | 0.841 |
| 95% CI | 0.727–0.973 |
| P-value | 0.0193 |
What the estimate means: the overall-survival hazard ratio of 0.841 indicates a lower estimated instantaneous hazard of death for afatinib relative to erlotinib under the reported Cox model. As a direct arithmetic interpretation of the HR, 1 − 0.841 = 0.159, so the estimated hazard is approximately 15.9% lower for afatinib under that model.
What it does not mean: it is not a 15.9% absolute increase in survival, nor does it say that 15.9% of patients benefited. It also does not provide a median survival time; no median survival value is reported in the ClinicalTrials.gov record.
Precision: the 95% CI of 0.727–0.973 indicates uncertainty around the HR estimate. Because the interval is relatively close to 1 at its upper boundary, the numerical size of the treatment effect should not be inferred to be exactly 0.841.
P-value: the p-value of 0.0193 is evidence against the null hypothesis under the specified two-sided analysis. It is not a measure of how large or clinically important the observed hazard ratio is.
Analysis caution: overall survival can reflect the complete treatment pathway rather than only the randomized therapy. The ClinicalTrials.gov record does not provide a crossover analysis or a detailed accounting of subsequent therapies, so this page does not infer such effects.
8. Secondary Result: Objective Response
Objective response according to RECIST 1.1 was analyzed as a binary endpoint in the randomized set. Logistic regression was used to compare afatinib with erlotinib, with the odds ratio as the effect measure.
Odds ratio for objective response
95% CI: 0.98–4.32 · P = 0.0551
Two-sided superiority analysis in the Randomized Set.
| Objective-response result | Reported value |
|---|---|
| Endpoint | Number of Participants With Objective Response According to RECIST 1.1 |
| Analysis population | RS |
| Method | Logistic regression |
| Effect measure | Odds ratio |
| Estimate | 2.06 |
| 95% CI | 0.98–4.32 |
| P-value | 0.0551 |
What the estimate means: an odds ratio of 2.06 means that the estimated odds of objective response were 2.06 times as high in the afatinib group as in the erlotinib group under the logistic regression analysis.
Odds are not probabilities: an odds ratio of 2.06 does not mean that the probability of response was 2.06 times as high. Converting odds ratios into probability differences requires the underlying event probabilities or an equivalent model specification, which are not provided here.
Precision: the 95% CI of 0.98–4.32 is relatively wide and includes 1.00. That indicates substantial statistical uncertainty about the magnitude of the odds-ratio estimate.
P-value: the two-sided p-value of 0.0551 is close to, but above, the conventional 0.05 reference point. More importantly, the p-value should not be used as a substitute for the effect estimate and confidence interval. The appropriate statistical description is that the registry reports an OR of 2.06 with a 95% CI of 0.98–4.32 and a p-value of 0.0551.
9. Secondary Result: Disease Control
Disease control according to RECIST 1.1 was also analyzed as a binary endpoint in the randomized set. The registry reports logistic regression stratified by race.
Odds ratio for disease control
95% CI: 1.18–2.06 · P = 0.0020
Two-sided superiority analysis in the Randomized Set; logistic regression stratified by race.
| Disease-control result | Reported value |
|---|---|
| Endpoint | Number of Participants With Disease Control According to RECIST 1.1 |
| Analysis population | RS |
| Method | Logistic regression stratified by race |
| Effect measure | Odds ratio |
| Estimate | 1.56 |
| 95% CI | 1.18–2.06 |
| P-value | 0.0020 |
An odds ratio of 1.56 indicates that the estimated odds of disease control were 1.56 times as high for afatinib as for erlotinib in the reported stratified logistic regression. This is an odds comparison, not a direct comparison of response probabilities.
The 95% CI of 1.18–2.06 quantifies uncertainty around the estimate and remains above 1.00. The p-value of 0.0020 provides evidence against the null hypothesis under the two-sided superiority analysis.
The disease-control endpoint should also be distinguished from objective response. A patient can contribute to disease control without necessarily meeting the criteria for objective response. Therefore, the two odds ratios answer related but different binary-outcome questions.
10. Secondary Result: Tumour Shrinkage
Tumour shrinkage was analyzed using ANCOVA. The registry states that the analysis compared treatments using ANCOVA for minimum sum of diameters, using baseline sum of diameters as a covariate. The randomization strata were included as classification factors, and the mean was adjusted for baseline sum of diameters and race.
Adjusted mean difference in tumour shrinkage
95% CI: -4.67 to 2.28 · P = 0.500
Two-sided superiority analysis.
| Tumour-shrinkage result | Reported value |
|---|---|
| Analysis population | Patients from the randomized set with tumour assessments |
| Method | ANCOVA |
| Effect measure | Adjusted mean / mean difference |
| Estimate | -1.2 |
| 95% CI | -4.67 to 2.28 |
| P-value | 0.500 |
| Covariate adjustment | Baseline sum of diameters; mean adjusted for baseline sum of diameters and race |
What ANCOVA contributes: rather than simply comparing raw post-treatment tumour measurements, ANCOVA accounts for baseline sum of diameters as a covariate. This can improve the precision of the treatment comparison when baseline measurements explain part of the variation in the outcome.
What the estimate means: the reported mean difference is -1.2, with afatinib compared with erlotinib. The sign therefore describes the direction of the adjusted mean difference under the registry's treatment-group ordering.
Precision: the 95% CI of -4.67 to 2.28 spans zero. The ClinicalTrials.gov record therefore does not establish a clear adjusted mean difference in tumour shrinkage between the groups.
P-value: the p-value of 0.500 does not measure the magnitude of tumour shrinkage. It describes the evidence against the null hypothesis for the ANCOVA comparison. The estimated difference and its confidence interval remain the more informative description of the size and uncertainty of the effect.
11. Secondary Results: Time to Deterioration
The registry reports three separate analyses under the outcome measure Summary of Time to Deterioration in Coughing, Dyspnoea and Pain. Each uses a Cox proportional-hazards model, with the analysis notes specifying stratification by race.
| Symptom | HR | 95% CI | P-value | Interpretation of HR |
|---|---|---|---|---|
| Coughing | 0.89 | 0.72–1.09 | 0.2562 | Estimated hazard of deterioration lower under afatinib in the reported model, but with substantial uncertainty. |
| Dyspnoea | 0.79 | 0.66–0.94 | 0.0078 | Estimated hazard of deterioration lower under afatinib in the reported model. |
| Pain | 0.99 | 0.82–1.18 | 0.8690 | Estimated hazards were very close to one in the reported model. |
The three estimates should not be collapsed into a single overall symptom effect. They represent distinct symptom-specific analyses under the same broad registry outcome-measure label.
Coughing
The HR of 0.89 has a 95% CI of 0.72–1.09 and a p-value of 0.2562. The confidence interval crosses 1.00, so the estimate is uncertain.
Dyspnoea
The HR of 0.79 has a 95% CI of 0.66–0.94 and a p-value of 0.0078. The estimated hazard ratio is below 1 under the reported Cox model.
Pain
The HR of 0.99 has a 95% CI of 0.82–1.18 and a p-value of 0.8690. The estimate is close to 1, with uncertainty extending on both sides.
Modeling point
All three analyses are Cox models and therefore require careful interpretation of the hazard-ratio scale and the proportional-hazards framework.
12. Secondary Results: Change in Symptom Scores Over Time
The registry separately reports analyses for Change in Score Over Time in Coughing,Dyspnoea and Pain. Although the registry analysis record identifies the endpoint type as time-to-event and the method as Cox regression, the reported effect measure is Mean Difference (Final Values). The three results below are therefore presented exactly as reported rather than reclassifying the endpoint.
| Symptom | Mean difference | 95% CI | P-value |
|---|---|---|---|
| Coughing | -3.5 | -6.15 to -0.88 | 0.0091 |
| Dyspnoea | -3.5 | -5.75 to -1.25 | 0.0024 |
| Pain | -2.7 | -5.33 to -0.15 | 0.0384 |
For these analyses, the reported effect measure is a mean difference in final values. A negative estimate indicates a lower final value for afatinib relative to erlotinib under the registry's treatment-group ordering. The estimates are -3.5 for coughing, -3.5 for dyspnoea, and -2.7 for pain.
The corresponding confidence intervals are -6.15 to -0.88, -5.75 to -1.25, and -5.33 to -0.15. Each interval is entirely below zero. The corresponding p-values are 0.0091, 0.0024, and 0.0384.
These results should not be translated into a percentage improvement or a clinical-importance statement without knowing the underlying scale, its direction, and the prespecified minimally important difference. Those quantities are not provided in the ClinicalTrials.gov record.
13. Statistical Methods Explained
Why was a log-rank test used for progression-free survival?
Progression-free survival is a time-to-event endpoint. Some participants experience progression or death during follow-up, while others may be censored because the event has not occurred by their last available assessment. The log-rank test is designed to compare the event-time distributions of randomized groups while accounting for this censoring structure. In LUX-Lung 8, the primary PFS comparison was reported using the log-rank method.
What does a hazard ratio of 0.814 mean?
The hazard ratio compares the modeled instantaneous event rate between treatment groups. An HR of 0.814 indicates a lower estimated hazard for afatinib than erlotinib under the reported PFS analysis. It does not mean that the probability of progression or death is exactly 0.814 times as large at every individual time point, and it is not an absolute risk difference.
Why does the confidence interval matter as much as the p-value?
The p-value addresses evidence against a null hypothesis, whereas the confidence interval communicates the statistical uncertainty around the estimated effect. For the primary PFS result, the HR is 0.814, while the 95% CI is 0.693–0.956. Reporting both prevents a statistically significant result from being mistaken for a precise or necessarily large treatment effect.
Why was logistic regression used for objective response and disease control?
Both objective response and disease control are binary outcomes: each participant is classified according to whether the specified event occurred. Logistic regression models the odds of that binary outcome and naturally produces an odds ratio for the treatment comparison. The disease-control analysis additionally specifies stratification by race.
What does an odds ratio of 1.56 mean?
An odds ratio of 1.56 means that the estimated odds of disease control were 1.56 times the odds in the comparator group under the reported logistic model. It does not mean that the disease-control probability increased by 56 percentage points or that the probability itself was 1.56 times larger.
Why was ANCOVA used for tumour shrinkage?
ANCOVA allows the treatment comparison to account for a baseline continuous measurement. In this trial's registry analysis, baseline sum of diameters was used as a covariate, while randomization strata were included as classification factors and the mean was adjusted for baseline sum of diameters and race. This is a different statistical problem from comparing time-to-event or binary outcomes.
Why is the analysis population important?
The primary PFS analysis used the Randomized Set, defined as all randomized patients regardless of whether they received investigational treatment. This preserves the treatment assignment created by randomization. Tumour shrinkage used the subset of randomized patients with tumour assessments, which means its analysis population is not identical to the primary PFS population.
14. Understanding Hazard Ratios in LUX-Lung 8
For LUX-Lung 8, the direction of the HR is especially useful when comparing progression-free survival, overall survival, and symptom deterioration. The HR must still be interpreted together with its confidence interval, analysis population, time frame, and model specification.
| Endpoint | HR | 95% CI | What the estimate represents |
|---|---|---|---|
| Primary PFS | 0.814 | 0.693–0.956 | Relative time-to-event effect for progression or death |
| Overall survival | 0.841 | 0.727–0.973 | Relative time-to-event effect for death |
| Time to deterioration in coughing | 0.89 | 0.72–1.09 | Relative time-to-event effect for deterioration in coughing |
| Time to deterioration in dyspnoea | 0.79 | 0.66–0.94 | Relative time-to-event effect for deterioration in dyspnoea |
| Time to deterioration in pain | 0.99 | 0.82–1.18 | Relative time-to-event effect for deterioration in pain |
The HRs also demonstrate why a trial should not be reduced to a single p-value. The estimates range from 0.79 to 0.99 for the three symptom-deterioration analyses, and their confidence intervals differ substantially. The uncertainty around each estimate is part of the result.
15. Interpreting the Odds Ratios
| Binary endpoint | OR | 95% CI | P-value |
|---|---|---|---|
| Objective response according to RECIST 1.1 | 2.06 | 0.98–4.32 | 0.0551 |
| Disease control according to RECIST 1.1 | 1.56 | 1.18–2.06 | 0.0020 |
The two odds ratios illustrate an important statistical distinction. The objective-response estimate is 2.06, but its confidence interval extends from 0.98 to 4.32. The disease-control estimate is 1.56, with a confidence interval of 1.18–2.06. The statistical evidence and precision are therefore not identical for the two binary endpoints.
16. Covariate Adjustment and ANCOVA
The tumour-shrinkage analysis is the clearest example in the ClinicalTrials.gov record of covariate adjustment. The registry states that ANCOVA used baseline sum of diameters as a covariate, included randomization strata as classification factors, and adjusted the mean for baseline sum of diameters and race.
Baseline adjustment
Using baseline sum of diameters accounts for an important continuous measurement before treatment when estimating the treatment comparison.
Classification factors
The randomization strata were included as classification factors in the ANCOVA model.
Adjusted mean
The registry reports an adjusted mean difference rather than an unadjusted difference between raw measurements.
Interpretation
Adjustment changes the model used to estimate the comparison; it does not turn a continuous outcome into a binary response endpoint.
The reported estimate of -1.2 with a 95% CI of -4.67 to 2.28 demonstrates why the adjusted estimate should be interpreted alongside its uncertainty. The interval spans zero, and the p-value is 0.500.
17. Safety
The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm. These figures are presented as affected participants divided by participants at risk.
| Safety measure | Afatinib | Erlotinib |
|---|---|---|
| Serious adverse events, affected / at risk | 174/392 | 175/395 |
The serious-adverse-event figures show the number of affected participants and the corresponding number at risk in each arm exactly as reported: 174/392 for afatinib and 175/395 for erlotinib.
These counts should not be transformed into a comparative safety conclusion without specifying the safety analysis population, follow-up exposure, adverse-event definitions, and the prespecified statistical approach. The ClinicalTrials.gov record does not provide a formal between-group statistical analysis for serious adverse events, so this page does not manufacture one.
Safety also answers a different question from efficacy. The PFS and OS hazard ratios describe time-to-event efficacy outcomes, whereas serious adverse events describe treatment-associated safety experience. They should therefore be considered as separate evidence streams rather than combined into a single numerical treatment effect.
18. What the P-Values Do — and Do Not — Tell Us
LUX-Lung 8 provides several useful examples of why p-values should not be interpreted in isolation.
| Analysis | Estimate | P-value | Interpretive point |
|---|---|---|---|
| Primary PFS | HR 0.814 | 0.0103 | The p-value describes evidence against the null; the HR and CI describe magnitude and uncertainty. |
| Overall survival | HR 0.841 | 0.0193 | The p-value does not mean the treatment effect is 1.93% or any other percentage. |
| Objective response | OR 2.06 | 0.0551 | The effect estimate remains important even though the p-value is above 0.05. |
| Disease control | OR 1.56 | 0.0020 | The p-value is not a measure of the 56% magnitude implied by the odds ratio. |
| Tumour shrinkage | Mean difference -1.2 | 0.500 | The p-value does not imply that the true mean difference is exactly zero. |
A p-value is conditional on a statistical model, null hypothesis, analysis population, and other design choices. It is not a universal measure of clinical importance. The most informative reading of a result combines the effect estimate, confidence interval, analysis population, endpoint definition, and statistical method.
19. Multiplicity and Multiple Secondary Endpoints
The registry reports one primary endpoint and multiple secondary endpoint analyses. These include overall survival, objective response, disease control, tumour shrinkage, symptom deterioration, and change in symptom scores. The ClinicalTrials.gov record identifies the hypothesis type as superiority but do not provide a multiplicity-adjustment procedure, alpha allocation scheme, or hierarchical testing strategy for the collection of secondary analyses.
| Analysis family | Number / structure in the ClinicalTrials.gov record | Interpretive implication |
|---|---|---|
| Primary endpoint | 1 registered primary endpoint | Primary confirmatory question defined around PFS. |
| Overall survival | 1 secondary analysis | Important secondary time-to-event result, but secondary in the registry hierarchy. |
| Binary outcomes | Objective response and disease control | Separate binary questions with separate odds ratios. |
| Tumour shrinkage | 1 ANCOVA analysis | Continuous outcome requiring covariate adjustment. |
| Symptom deterioration | 3 symptom-specific analyses | Separate Cox analyses for coughing, dyspnoea, and pain. |
| Change in symptom scores | 3 symptom-specific analyses | Separate reported mean differences for coughing, dyspnoea, and pain. |
20. Censoring and Time-to-Event Analysis
Progression-free survival and overall survival are time-to-event endpoints. Such analyses have an important feature that ordinary two-group comparisons do not: not every participant necessarily contributes a fully observed event time. Participants can be censored when the event has not been observed during their available follow-up.
where di represents events at time ti and ni represents participants at risk immediately before that time.
The ClinicalTrials.gov record does not provide Kaplan-Meier median estimates, survival probabilities at specified time points, event counts for the PFS analysis, or the detailed censoring rules. Those quantities are therefore not added here. The reported hazard ratios are sufficient to describe the posted formal treatment comparisons without inventing unreported survival summaries.
21. Analysis Population and Randomization
Randomization is central to the interpretation of the primary comparison. The registry defines the Randomized Set as all patients who were randomized, regardless of whether they received investigational treatment. This population was used for the primary PFS analysis and for the listed overall-survival and binary-outcome analyses.
Why randomization matters
Randomization creates the treatment groups used for the causal comparison. An analysis based on randomized assignment preserves that design principle.
Why the RS definition matters
The Randomized Set includes randomized participants regardless of whether they received investigational treatment, according to the registry definition.
Tumour assessments
Tumour shrinkage uses randomized patients with tumour assessments, a more restricted analysis population than the general RS.
Safety
The ClinicalTrials.gov record is reported as affected participants divided by participants at risk; the ClinicalTrials.gov record does not specify an additional formal safety comparison.
22. Design Features Not Reported in the Supplied Data
Several methodological topics commonly important in phase 3 trials are not supported by the registry-reported LUX-Lung 8 data. They are therefore not assigned a value or interpretation here.
| Topic | What can be concluded from the ClinicalTrials.gov record |
|---|---|
| Non-inferiority margin | Not applicable to the reported hypothesis description; the registry identifies the hypothesis type as superiority and does not provide a non-inferiority margin. |
| Crossover | No crossover analysis or crossover rate is provided. |
| Factorial design | The design is identified as parallel, not factorial. |
| Interim analysis | No interim-analysis procedure is provided in the ClinicalTrials.gov record. |
| Missing-data imputation | No formal missing-data or imputation strategy is provided in the ClinicalTrials.gov record. |
| Bayesian methods | No Bayesian method is identified. The normalized methods are ANCOVA, Cox proportional-hazards model, log-rank test, and logistic regression. |
| Multiplicity adjustment | No specific adjustment procedure is posted on ClinicalTrials.gov for the secondary analyses. |
This distinction is important for statistical interpretation: an absent methodological detail should not be filled in merely because a particular approach is common in other clinical trials.
23. Limitations
- Registry-level detail: this analysis is constrained to the ClinicalTrials.gov record. It does not add estimates or design details from publications beyond the links registry-reported as allowed sources.
- No median survival values: the ClinicalTrials.gov record contains hazard ratios for PFS and OS but do not provide median PFS or median OS, so no median survival values are presented.
- No baseline table: the ClinicalTrials.gov record does not provide baseline demographic or disease-characteristic values, so treatment-group baseline balance cannot be numerically assessed here.
- No subgroup estimates: although several analyses mention stratification by race, no treatment-effect subgroup estimates are provided in the ClinicalTrials.gov record.
- No crossover information: the ClinicalTrials.gov record does not identify crossover, so no adjustment or sensitivity analysis for crossover is presented.
- Hazard-ratio assumptions: Cox-model estimates depend on the proportional-hazards framework. The ClinicalTrials.gov record does not provide an assessment of that assumption.
- Secondary-endpoint multiplicity: multiple secondary analyses are posted, but the ClinicalTrials.gov record does not specify a multiplicity-control strategy for interpreting the complete family of secondary p-values.
- Different analysis populations: most analyses use the randomized set, whereas tumour shrinkage is restricted to randomized patients with tumour assessments.
- Registry endpoint labeling: the change-in-score analyses are labeled as time-to-event in the registry while reporting a mean difference in final values. This page preserves the registry description rather than reconstructing an alternative model.
- Safety interpretation: serious adverse-event counts are reported by arm, but the ClinicalTrials.gov record does not include a formal comparative statistical analysis for that safety outcome.
24. Why This Trial Matters Statistically
LUX-Lung 8 is a useful teaching case because the registry results connect several major clinical-trial methods within one randomized comparison. The primary endpoint is a time-to-event outcome analyzed with a log-rank test and hazard ratio, while secondary outcomes demonstrate logistic regression, stratified Cox modeling, and ANCOVA with baseline covariate adjustment.
| Concept | How it appears in LUX-Lung 8 |
|---|---|
| Randomization | 795 participants were enrolled in a randomized two-arm phase 3 parallel-group trial. |
| Time-to-event analysis | Progression-free survival and overall survival are analyzed as time-to-event outcomes. |
| Log-rank test | Used for the posted PFS and overall-survival comparisons. |
| Hazard ratio | Used for PFS, OS, and symptom-deterioration analyses. |
| Confidence interval | Reported alongside the major effect estimates. |
| Cox model | Used for OS and symptom deterioration, with race identified as a stratification factor in several analyses. |
| Logistic regression | Used for objective response and disease control. |
| Odds ratio | Quantifies the binary-outcome comparisons for objective response and disease control. |
| ANCOVA | Used for tumour shrinkage with baseline sum of diameters as a covariate. |
| Covariate adjustment | Tumour-shrinkage means were adjusted for baseline sum of diameters and race. |
| Analysis populations | The Randomized Set is used for most efficacy analyses; tumour shrinkage is restricted to patients with tumour assessments. |
| Multiple secondary analyses | Separate efficacy and symptom endpoints illustrate why each estimate must be interpreted within its endpoint definition and analysis method. |
25. A Statistical Reading of the Complete Results
The most important result is the primary progression-free survival comparison: HR 0.814, 95% CI 0.693–0.956, P = 0.0103. This is a time-to-event treatment comparison in the randomized set, with a hazard ratio below 1 and a confidence interval that remains below 1.
The overall-survival analysis reports a similar direction but a somewhat different estimate: HR 0.841, 95% CI 0.727–0.973, P = 0.0193. The two estimates should not be treated as interchangeable. PFS measures progression or death, whereas OS measures death, and the two endpoints have different event definitions and follow-up periods.
The binary outcomes add another dimension. Objective response has an odds ratio of 2.06 with a 95% CI of 0.98–4.32 and a p-value of 0.0551. Disease control has an odds ratio of 1.56 with a 95% CI of 1.18–2.06 and a p-value of 0.0020. These are odds comparisons, not direct probability differences.
The tumour-shrinkage analysis demonstrates a different modeling strategy. ANCOVA produced an adjusted mean difference of -1.2 with a 95% CI of -4.67 to 2.28 and a p-value of 0.500. That result should be interpreted on the continuous-outcome scale and should not be conflated with the binary response analysis.
The symptom analyses further illustrate endpoint-specific interpretation. The time-to-deterioration HRs are 0.89 for coughing, 0.79 for dyspnoea, and 0.99 for pain. The separate change-in-score analyses report mean differences of -3.5, -3.5, and -2.7, respectively. These results should remain separated because they represent different statistical quantities even though they concern related symptoms.
The statistical story is therefore broader than the primary p-value. The trial contains a primary time-to-event analysis, a secondary survival endpoint, binary response and disease-control analyses, a covariate-adjusted continuous endpoint, symptom-deterioration survival analyses, and reported symptom-score differences. The appropriate interpretation is to examine each result on its own statistical scale while retaining the randomized design and the hierarchy between the primary and secondary endpoints.
26. Related Tutorials
Learn more about the methods used in this trial:
27. Related Statistical Calculators
28. Sources
- ClinicalTrials.gov: LUX-Lung 8, NCT01523587.
- PubMed record: PMID 30573970.
- PubMed record: PMID 29902295.
- PubMed record: PMID 28577938.
- PubMed record: PMID 26156651.
Continue through Clinical Biostats
Explore the statistical methods behind randomized clinical trials, time-to-event endpoints, regression models, and clinical-trial analysis workflows.
29. Record Summary
LUX-Lung 8 provides a compact but methodologically diverse example of phase 3 clinical-trial statistics. The randomized trial enrolled 795 participants and compared afatinib with erlotinib. Its registered primary endpoint was progression-free survival, analyzed using a log-rank test with a reported hazard ratio of 0.814, 95% CI 0.693–0.956, and p-value 0.0103. Secondary analyses extend the statistical framework to overall survival, objective response, disease control, tumour shrinkage, symptom deterioration, and symptom-score changes.
The most useful statistical reading combines the effect measure, confidence interval, p-value, endpoint definition, analysis population, and model specification. Hazard ratios, odds ratios, and adjusted mean differences are not interchangeable quantities. Each describes a different aspect of the treatment comparison and must be interpreted on its own statistical scale.