← Clinical Trials
Squamous Cell Lung Cancer Phase 3 Completed NCT01523587

LUX-Lung 8: Complete Statistical Analysis of Afatinib in Squamous Cell Lung Cancer

An independent statistical analysis of the randomized phase 3 LUX-Lung 8 trial comparing afatinib with erlotinib for the treatment of squamous cell lung cancer after at least one prior platinum-based chemotherapy.

Randomized phase 3  ·  795 participants  ·  Primary endpoint: progression-free survival
Scope of this record

This page separates reported trial results from statistical interpretation. This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

LUX-Lung 8 was a randomized, parallel-group, phase 3 trial comparing afatinib with erlotinib in patients with carcinoma of the non-small-cell lung after at least one prior platinum-based chemotherapy. The registry identifies progression-free survival based on central independent review according to RECIST 1.1 as the single registered primary endpoint.

795
Enrolled
2 treatment arms
2
Arms
Afatinib vs erlotinib
0.814
Primary PFS HR
95% CI 0.693–0.956
0.0103
Primary PFS P-value
Two-sided
FeatureLUX-Lung 8
PhasePhase 3
ConditionCarcinoma, Non-Small-Cell Lung
Trial population described in the brief titleSquamous cell lung cancer after at least one prior platinum-based chemotherapy
DesignRandomized, parallel-group, open-label
AllocationRandomized
Primary purposeTreatment
Enrollment795
InterventionsAfatinib and erlotinib
Primary endpoint typeTime-to-event
Results postedYes
Outcome measures posted8
Statistical analyses posted11
Hypothesis typeSuperiority
Lead sponsorBoehringer Ingelheim
Trial statusCompleted
ClinicalTrials.govNCT01523587

2. Clinical Question

The trial addresses whether afatinib provides a different time-to-event outcome from erlotinib in patients with squamous cell lung cancer after at least one prior platinum-based chemotherapy. The registered hypothesis type is superiority, so the statistical question is framed as whether the treatment groups differ in the prespecified direction rather than whether afatinib merely achieves a predefined non-inferiority standard.

Population

Patients with squamous cell lung cancer, as described in the trial brief title, after at least one prior platinum-based chemotherapy.

Intervention

Afatinib.

Comparator

Erlotinib.

Primary question

Does afatinib improve progression-free survival compared with erlotinib when progression is determined by central independent review according to RECIST 1.1?

3. Trial Design

01
Randomize795 participants
02
Two armsAfatinib vs erlotinib
03
Open labelNo masking
04
AssessTime-to-event and other outcomes
05
CompareRegistry statistical analyses
ARM A

Afatinib

  • Afatinib was the investigational treatment.
  • The registry compares this group directly with erlotinib.
  • Primary efficacy analysis used the randomized set.
ARM B

Erlotinib

  • Erlotinib was the comparator treatment.
  • The registry compares this group directly with afatinib.
  • Primary efficacy analysis used the randomized set.
Allocation
Randomized allocation to the two treatment groups.
Model
Parallel-group design.
Masking
None. The trial was open label.
Primary purpose
Treatment.

The trial began on 05 March 2012 and had a primary completion date of 21 October 2013. The primary progression-free survival analysis used a later cutoff date of 02 March 2015, while several secondary outcomes used a study-closure period extending to 27 December 2017.

4. Endpoints

Primary endpoint

EndpointRegistry definition / time frameType
Progression-free Survival, Based on Central Independent Review as Determined by Response Evaluation Criteria in Solid Tumours 1.1 First treatment administration up until cut off date of 02 March 2015 (up to 1058 days). Progression Free Survival was defined as the time from randomization to disease progression (or death if the patient died before progression) by central independent review according to Response Evaluation Criteria in Solid Tumours (RECIST) version 1.1. Time-to-event

Secondary endpoints with posted statistical analyses

EndpointTime frameType
Overall SurvivalFrom first drug administration from 9 April 2012 until study closure on 27 Dec 2017 (approximately 2089 days).Time-to-event
Number of Participants With Objective Response According to RECIST 1.1First treatment administration up until cut off date of 02 March 2015 (up to 1058 days).Binary
Number of Participants With Disease Control According to RECIST 1.1First treatment administration up until cut off date of 02 March 2015 (up to 1058 days).Binary
Tumour ShrinkageFirst treatment administration up until cut off date of 02 March 2015 (up to 1058 days).Continuous
Summary of Time to Deterioration in Coughing, Dyspnoea and PainFrom first drug administration from 9 April 2012 until study closure on 27 Dec 2017 (approximately 2089 days).Time-to-event
Change in Score Over Time in Coughing,Dyspnoea and PainFrom first drug administration from 9 April 2012 until study closure on 27 Dec 2017 (approximately 2089 days).Time-to-event in the registry analysis record
Endpoint-label caution: the registry contains multiple analyses under the same broad outcome-measure names. For the symptom deterioration endpoint, the individual analyses identify coughing, dyspnoea, and pain separately. For the change-in-score endpoint, the registry labels the endpoint type as time-to-event but reports a final-value mean difference as the effect measure. This page preserves that registry characterization rather than silently substituting a different endpoint definition.

5. Statistical Methodology

The posted analyses use four main statistical method families: log-rank testing and Cox proportional-hazards models for time-to-event outcomes, logistic regression for binary outcomes, and ANCOVA for tumour shrinkage. The registry also identifies stratified analysis and covariate adjustment in specific analyses.

MethodUsed forEffect measureStatistical role
Log-rank testProgression-free survival; overall survivalHazard ratio reported alongside the comparisonComparison of time-to-event distributions
Cox proportional-hazards modelOverall survival; symptom deterioration; registry-reported symptom score analysesHazard ratio or, for the score analyses, mean difference in final valuesModel-based time-to-event comparison
Logistic regressionObjective response; disease controlOdds ratioComparison of binary outcomes
ANCOVATumour shrinkageAdjusted mean / mean differenceCovariate-adjusted continuous-outcome comparison

Analysis populations

The primary progression-free survival analysis was conducted in the Randomized Set (RS), defined in the registry as all patients who were randomized, regardless of whether they received investigational treatment. The overall survival, response, disease-control, and symptom analyses also identify the RS as their analysis population.

Tumour shrinkage is different: the registry states that patients from the randomized set with tumour assessments were considered for that endpoint. That distinction matters because availability of a tumour assessment can differ from simply being randomized.

Stratified analysis

Several analyses are identified as stratified. The overall survival analysis used a Cox proportional-hazards model stratified by race. The disease-control analysis used logistic regression stratified by race. The symptom deterioration analyses similarly identify Cox models stratified by race.

Why stratification matters
Treatment effect  →  estimated while accounting for the specified stratification factor

Stratification allows the analysis to compare treatment groups within levels of a factor rather than treating the factor as irrelevant to the comparison. In this registry record, race is explicitly identified as a stratification variable for several secondary analyses.

6. Primary Result: Progression-Free Survival

The primary endpoint was progression-free survival based on central independent review according to RECIST 1.1. The analysis compared afatinib with erlotinib in the randomized set using a log-rank test. The registry also notes that a Cox proportional-hazards model without the randomization stratification variable was used for each subgroup category, along with the corresponding log-rank test.

Hazard ratio for progression or death

0.814

95% CI: 0.693–0.956   ·   P = 0.0103

Two-sided superiority analysis in the Randomized Set.

Primary PFS resultReported value
Groups comparedAfatinib vs Erlotinib
Analysis populationRandomized Set (RS)
MethodLog-rank test
Effect measureHazard ratio
Estimate0.814
95% CI0.693–0.956
P-value0.0103
Hypothesis typeSuperiority
Clinical Biostats interpretation

What the estimate means: a hazard ratio of 0.814 means that, under the time-to-event model used for the comparison, the estimated instantaneous rate of progression or death for afatinib relative to erlotinib was 0.814. Expressed as a simple relative interpretation, this corresponds to an estimated hazard that was about 18.6% lower for afatinib than erlotinib.

What it does not mean: it does not mean that 18.6% fewer patients necessarily experienced progression or death, and it does not mean that every patient had an 18.6% reduction in risk. A hazard ratio is a relative time-to-event measure, not an absolute risk difference.

What the confidence interval says: the 95% confidence interval of 0.693–0.956 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It gives a range of values compatible with the observed data and model assumptions; it is not a range of individual patient effects.

What the p-value says: the two-sided p-value of 0.0103 quantifies the evidence against the null hypothesis specified for the statistical comparison. It does not measure the size or clinical importance of the treatment effect. Effect size and uncertainty are conveyed by the hazard ratio and confidence interval.

Caution: interpretation of a single Cox hazard ratio relies on the model's proportional-hazards framework. If the relative hazard changes substantially over time, one summary hazard ratio may not fully describe the treatment difference. The ClinicalTrials.gov record does not provide a time-varying hazard assessment, so the HR should be interpreted within the stated model rather than as a universal constant risk reduction at every time point.

7. Secondary Result: Overall Survival

Overall survival was evaluated from first drug administration from 9 April 2012 until study closure on 27 Dec 2017, approximately 2089 days. The analysis population was the randomized set. The registry reports a log-rank analysis and states that a Cox proportional-hazards model stratified by race was used to estimate the hazard ratio and 95% confidence interval between the two treatment groups.

Hazard ratio for overall survival

0.841

95% CI: 0.727–0.973   ·   P = 0.0193

Two-sided superiority analysis in the Randomized Set.

Overall survival resultReported value
Groups comparedAfatinib vs Erlotinib
Analysis populationRS
MethodLog-rank test
Cox modelStratified by race
Effect measureHazard ratio
Estimate0.841
95% CI0.727–0.973
P-value0.0193
Clinical Biostats interpretation

What the estimate means: the overall-survival hazard ratio of 0.841 indicates a lower estimated instantaneous hazard of death for afatinib relative to erlotinib under the reported Cox model. As a direct arithmetic interpretation of the HR, 1 − 0.841 = 0.159, so the estimated hazard is approximately 15.9% lower for afatinib under that model.

What it does not mean: it is not a 15.9% absolute increase in survival, nor does it say that 15.9% of patients benefited. It also does not provide a median survival time; no median survival value is reported in the ClinicalTrials.gov record.

Precision: the 95% CI of 0.727–0.973 indicates uncertainty around the HR estimate. Because the interval is relatively close to 1 at its upper boundary, the numerical size of the treatment effect should not be inferred to be exactly 0.841.

P-value: the p-value of 0.0193 is evidence against the null hypothesis under the specified two-sided analysis. It is not a measure of how large or clinically important the observed hazard ratio is.

Analysis caution: overall survival can reflect the complete treatment pathway rather than only the randomized therapy. The ClinicalTrials.gov record does not provide a crossover analysis or a detailed accounting of subsequent therapies, so this page does not infer such effects.

8. Secondary Result: Objective Response

Objective response according to RECIST 1.1 was analyzed as a binary endpoint in the randomized set. Logistic regression was used to compare afatinib with erlotinib, with the odds ratio as the effect measure.

Odds ratio for objective response

2.06

95% CI: 0.98–4.32   ·   P = 0.0551

Two-sided superiority analysis in the Randomized Set.

Objective-response resultReported value
EndpointNumber of Participants With Objective Response According to RECIST 1.1
Analysis populationRS
MethodLogistic regression
Effect measureOdds ratio
Estimate2.06
95% CI0.98–4.32
P-value0.0551
Clinical Biostats interpretation

What the estimate means: an odds ratio of 2.06 means that the estimated odds of objective response were 2.06 times as high in the afatinib group as in the erlotinib group under the logistic regression analysis.

Odds are not probabilities: an odds ratio of 2.06 does not mean that the probability of response was 2.06 times as high. Converting odds ratios into probability differences requires the underlying event probabilities or an equivalent model specification, which are not provided here.

Precision: the 95% CI of 0.98–4.32 is relatively wide and includes 1.00. That indicates substantial statistical uncertainty about the magnitude of the odds-ratio estimate.

P-value: the two-sided p-value of 0.0551 is close to, but above, the conventional 0.05 reference point. More importantly, the p-value should not be used as a substitute for the effect estimate and confidence interval. The appropriate statistical description is that the registry reports an OR of 2.06 with a 95% CI of 0.98–4.32 and a p-value of 0.0551.

9. Secondary Result: Disease Control

Disease control according to RECIST 1.1 was also analyzed as a binary endpoint in the randomized set. The registry reports logistic regression stratified by race.

Odds ratio for disease control

1.56

95% CI: 1.18–2.06   ·   P = 0.0020

Two-sided superiority analysis in the Randomized Set; logistic regression stratified by race.

Disease-control resultReported value
EndpointNumber of Participants With Disease Control According to RECIST 1.1
Analysis populationRS
MethodLogistic regression stratified by race
Effect measureOdds ratio
Estimate1.56
95% CI1.18–2.06
P-value0.0020
Clinical Biostats interpretation

An odds ratio of 1.56 indicates that the estimated odds of disease control were 1.56 times as high for afatinib as for erlotinib in the reported stratified logistic regression. This is an odds comparison, not a direct comparison of response probabilities.

The 95% CI of 1.18–2.06 quantifies uncertainty around the estimate and remains above 1.00. The p-value of 0.0020 provides evidence against the null hypothesis under the two-sided superiority analysis.

The disease-control endpoint should also be distinguished from objective response. A patient can contribute to disease control without necessarily meeting the criteria for objective response. Therefore, the two odds ratios answer related but different binary-outcome questions.

10. Secondary Result: Tumour Shrinkage

Tumour shrinkage was analyzed using ANCOVA. The registry states that the analysis compared treatments using ANCOVA for minimum sum of diameters, using baseline sum of diameters as a covariate. The randomization strata were included as classification factors, and the mean was adjusted for baseline sum of diameters and race.

Adjusted mean difference in tumour shrinkage

-1.2 mm

95% CI: -4.67 to 2.28   ·   P = 0.500

Two-sided superiority analysis.

Tumour-shrinkage resultReported value
Analysis populationPatients from the randomized set with tumour assessments
MethodANCOVA
Effect measureAdjusted mean / mean difference
Estimate-1.2
95% CI-4.67 to 2.28
P-value0.500
Covariate adjustmentBaseline sum of diameters; mean adjusted for baseline sum of diameters and race
Clinical Biostats interpretation

What ANCOVA contributes: rather than simply comparing raw post-treatment tumour measurements, ANCOVA accounts for baseline sum of diameters as a covariate. This can improve the precision of the treatment comparison when baseline measurements explain part of the variation in the outcome.

What the estimate means: the reported mean difference is -1.2, with afatinib compared with erlotinib. The sign therefore describes the direction of the adjusted mean difference under the registry's treatment-group ordering.

Precision: the 95% CI of -4.67 to 2.28 spans zero. The ClinicalTrials.gov record therefore does not establish a clear adjusted mean difference in tumour shrinkage between the groups.

P-value: the p-value of 0.500 does not measure the magnitude of tumour shrinkage. It describes the evidence against the null hypothesis for the ANCOVA comparison. The estimated difference and its confidence interval remain the more informative description of the size and uncertainty of the effect.

11. Secondary Results: Time to Deterioration

The registry reports three separate analyses under the outcome measure Summary of Time to Deterioration in Coughing, Dyspnoea and Pain. Each uses a Cox proportional-hazards model, with the analysis notes specifying stratification by race.

SymptomHR95% CIP-valueInterpretation of HR
Coughing0.890.72–1.090.2562Estimated hazard of deterioration lower under afatinib in the reported model, but with substantial uncertainty.
Dyspnoea0.790.66–0.940.0078Estimated hazard of deterioration lower under afatinib in the reported model.
Pain0.990.82–1.180.8690Estimated hazards were very close to one in the reported model.

The three estimates should not be collapsed into a single overall symptom effect. They represent distinct symptom-specific analyses under the same broad registry outcome-measure label.

Coughing

The HR of 0.89 has a 95% CI of 0.72–1.09 and a p-value of 0.2562. The confidence interval crosses 1.00, so the estimate is uncertain.

Dyspnoea

The HR of 0.79 has a 95% CI of 0.66–0.94 and a p-value of 0.0078. The estimated hazard ratio is below 1 under the reported Cox model.

Pain

The HR of 0.99 has a 95% CI of 0.82–1.18 and a p-value of 0.8690. The estimate is close to 1, with uncertainty extending on both sides.

Modeling point

All three analyses are Cox models and therefore require careful interpretation of the hazard-ratio scale and the proportional-hazards framework.

12. Secondary Results: Change in Symptom Scores Over Time

The registry separately reports analyses for Change in Score Over Time in Coughing,Dyspnoea and Pain. Although the registry analysis record identifies the endpoint type as time-to-event and the method as Cox regression, the reported effect measure is Mean Difference (Final Values). The three results below are therefore presented exactly as reported rather than reclassifying the endpoint.

SymptomMean difference95% CIP-value
Coughing-3.5-6.15 to -0.880.0091
Dyspnoea-3.5-5.75 to -1.250.0024
Pain-2.7-5.33 to -0.150.0384
Clinical Biostats interpretation

For these analyses, the reported effect measure is a mean difference in final values. A negative estimate indicates a lower final value for afatinib relative to erlotinib under the registry's treatment-group ordering. The estimates are -3.5 for coughing, -3.5 for dyspnoea, and -2.7 for pain.

The corresponding confidence intervals are -6.15 to -0.88, -5.75 to -1.25, and -5.33 to -0.15. Each interval is entirely below zero. The corresponding p-values are 0.0091, 0.0024, and 0.0384.

These results should not be translated into a percentage improvement or a clinical-importance statement without knowing the underlying scale, its direction, and the prespecified minimally important difference. Those quantities are not provided in the ClinicalTrials.gov record.

13. Statistical Methods Explained

Why was a log-rank test used for progression-free survival?

Progression-free survival is a time-to-event endpoint. Some participants experience progression or death during follow-up, while others may be censored because the event has not occurred by their last available assessment. The log-rank test is designed to compare the event-time distributions of randomized groups while accounting for this censoring structure. In LUX-Lung 8, the primary PFS comparison was reported using the log-rank method.

What does a hazard ratio of 0.814 mean?

The hazard ratio compares the modeled instantaneous event rate between treatment groups. An HR of 0.814 indicates a lower estimated hazard for afatinib than erlotinib under the reported PFS analysis. It does not mean that the probability of progression or death is exactly 0.814 times as large at every individual time point, and it is not an absolute risk difference.

Why does the confidence interval matter as much as the p-value?

The p-value addresses evidence against a null hypothesis, whereas the confidence interval communicates the statistical uncertainty around the estimated effect. For the primary PFS result, the HR is 0.814, while the 95% CI is 0.693–0.956. Reporting both prevents a statistically significant result from being mistaken for a precise or necessarily large treatment effect.

Why was logistic regression used for objective response and disease control?

Both objective response and disease control are binary outcomes: each participant is classified according to whether the specified event occurred. Logistic regression models the odds of that binary outcome and naturally produces an odds ratio for the treatment comparison. The disease-control analysis additionally specifies stratification by race.

What does an odds ratio of 1.56 mean?

An odds ratio of 1.56 means that the estimated odds of disease control were 1.56 times the odds in the comparator group under the reported logistic model. It does not mean that the disease-control probability increased by 56 percentage points or that the probability itself was 1.56 times larger.

Why was ANCOVA used for tumour shrinkage?

ANCOVA allows the treatment comparison to account for a baseline continuous measurement. In this trial's registry analysis, baseline sum of diameters was used as a covariate, while randomization strata were included as classification factors and the mean was adjusted for baseline sum of diameters and race. This is a different statistical problem from comparing time-to-event or binary outcomes.

Why is the analysis population important?

The primary PFS analysis used the Randomized Set, defined as all randomized patients regardless of whether they received investigational treatment. This preserves the treatment assignment created by randomization. Tumour shrinkage used the subset of randomized patients with tumour assessments, which means its analysis population is not identical to the primary PFS population.

14. Understanding Hazard Ratios in LUX-Lung 8

Conceptual interpretation
HR < 1  →  lower estimated instantaneous event rate in the afatinib group

For LUX-Lung 8, the direction of the HR is especially useful when comparing progression-free survival, overall survival, and symptom deterioration. The HR must still be interpreted together with its confidence interval, analysis population, time frame, and model specification.

EndpointHR95% CIWhat the estimate represents
Primary PFS0.8140.693–0.956Relative time-to-event effect for progression or death
Overall survival0.8410.727–0.973Relative time-to-event effect for death
Time to deterioration in coughing0.890.72–1.09Relative time-to-event effect for deterioration in coughing
Time to deterioration in dyspnoea0.790.66–0.94Relative time-to-event effect for deterioration in dyspnoea
Time to deterioration in pain0.990.82–1.18Relative time-to-event effect for deterioration in pain

The HRs also demonstrate why a trial should not be reduced to a single p-value. The estimates range from 0.79 to 0.99 for the three symptom-deterioration analyses, and their confidence intervals differ substantially. The uncertainty around each estimate is part of the result.

15. Interpreting the Odds Ratios

Binary endpointOR95% CIP-value
Objective response according to RECIST 1.12.060.98–4.320.0551
Disease control according to RECIST 1.11.561.18–2.060.0020

The two odds ratios illustrate an important statistical distinction. The objective-response estimate is 2.06, but its confidence interval extends from 0.98 to 4.32. The disease-control estimate is 1.56, with a confidence interval of 1.18–2.06. The statistical evidence and precision are therefore not identical for the two binary endpoints.

Do not convert odds ratios into risk ratios without additional information. The numerical relationship between an odds ratio and a probability ratio depends on the underlying event rate. Because the ClinicalTrials.gov record does not provide the response and disease-control counts by treatment group, this page does not derive probabilities or risk ratios from the reported odds ratios.

16. Covariate Adjustment and ANCOVA

The tumour-shrinkage analysis is the clearest example in the ClinicalTrials.gov record of covariate adjustment. The registry states that ANCOVA used baseline sum of diameters as a covariate, included randomization strata as classification factors, and adjusted the mean for baseline sum of diameters and race.

Baseline adjustment

Using baseline sum of diameters accounts for an important continuous measurement before treatment when estimating the treatment comparison.

Classification factors

The randomization strata were included as classification factors in the ANCOVA model.

Adjusted mean

The registry reports an adjusted mean difference rather than an unadjusted difference between raw measurements.

Interpretation

Adjustment changes the model used to estimate the comparison; it does not turn a continuous outcome into a binary response endpoint.

The reported estimate of -1.2 with a 95% CI of -4.67 to 2.28 demonstrates why the adjusted estimate should be interpreted alongside its uncertainty. The interval spans zero, and the p-value is 0.500.

17. Safety

The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm. These figures are presented as affected participants divided by participants at risk.

Safety measureAfatinibErlotinib
Serious adverse events, affected / at risk174/392175/395
Clinical Biostats interpretation

The serious-adverse-event figures show the number of affected participants and the corresponding number at risk in each arm exactly as reported: 174/392 for afatinib and 175/395 for erlotinib.

These counts should not be transformed into a comparative safety conclusion without specifying the safety analysis population, follow-up exposure, adverse-event definitions, and the prespecified statistical approach. The ClinicalTrials.gov record does not provide a formal between-group statistical analysis for serious adverse events, so this page does not manufacture one.

Safety also answers a different question from efficacy. The PFS and OS hazard ratios describe time-to-event efficacy outcomes, whereas serious adverse events describe treatment-associated safety experience. They should therefore be considered as separate evidence streams rather than combined into a single numerical treatment effect.

18. What the P-Values Do — and Do Not — Tell Us

LUX-Lung 8 provides several useful examples of why p-values should not be interpreted in isolation.

AnalysisEstimateP-valueInterpretive point
Primary PFSHR 0.8140.0103The p-value describes evidence against the null; the HR and CI describe magnitude and uncertainty.
Overall survivalHR 0.8410.0193The p-value does not mean the treatment effect is 1.93% or any other percentage.
Objective responseOR 2.060.0551The effect estimate remains important even though the p-value is above 0.05.
Disease controlOR 1.560.0020The p-value is not a measure of the 56% magnitude implied by the odds ratio.
Tumour shrinkageMean difference -1.20.500The p-value does not imply that the true mean difference is exactly zero.

A p-value is conditional on a statistical model, null hypothesis, analysis population, and other design choices. It is not a universal measure of clinical importance. The most informative reading of a result combines the effect estimate, confidence interval, analysis population, endpoint definition, and statistical method.

19. Multiplicity and Multiple Secondary Endpoints

The registry reports one primary endpoint and multiple secondary endpoint analyses. These include overall survival, objective response, disease control, tumour shrinkage, symptom deterioration, and change in symptom scores. The ClinicalTrials.gov record identifies the hypothesis type as superiority but do not provide a multiplicity-adjustment procedure, alpha allocation scheme, or hierarchical testing strategy for the collection of secondary analyses.

Analysis familyNumber / structure in the ClinicalTrials.gov recordInterpretive implication
Primary endpoint1 registered primary endpointPrimary confirmatory question defined around PFS.
Overall survival1 secondary analysisImportant secondary time-to-event result, but secondary in the registry hierarchy.
Binary outcomesObjective response and disease controlSeparate binary questions with separate odds ratios.
Tumour shrinkage1 ANCOVA analysisContinuous outcome requiring covariate adjustment.
Symptom deterioration3 symptom-specific analysesSeparate Cox analyses for coughing, dyspnoea, and pain.
Change in symptom scores3 symptom-specific analysesSeparate reported mean differences for coughing, dyspnoea, and pain.
Multiplicity caution: because the ClinicalTrials.gov record does not specify a formal multiplicity-adjustment strategy for these secondary analyses, this page does not reinterpret their individual p-values as if they represented independently protected confirmatory hypotheses. The reported estimates and p-values are presented as registry results, while their inferential scope should be kept distinct from the single registered primary endpoint.

20. Censoring and Time-to-Event Analysis

Progression-free survival and overall survival are time-to-event endpoints. Such analyses have an important feature that ordinary two-group comparisons do not: not every participant necessarily contributes a fully observed event time. Participants can be censored when the event has not been observed during their available follow-up.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

where di represents events at time ti and ni represents participants at risk immediately before that time.

The ClinicalTrials.gov record does not provide Kaplan-Meier median estimates, survival probabilities at specified time points, event counts for the PFS analysis, or the detailed censoring rules. Those quantities are therefore not added here. The reported hazard ratios are sufficient to describe the posted formal treatment comparisons without inventing unreported survival summaries.

21. Analysis Population and Randomization

Randomization is central to the interpretation of the primary comparison. The registry defines the Randomized Set as all patients who were randomized, regardless of whether they received investigational treatment. This population was used for the primary PFS analysis and for the listed overall-survival and binary-outcome analyses.

Why randomization matters

Randomization creates the treatment groups used for the causal comparison. An analysis based on randomized assignment preserves that design principle.

Why the RS definition matters

The Randomized Set includes randomized participants regardless of whether they received investigational treatment, according to the registry definition.

Tumour assessments

Tumour shrinkage uses randomized patients with tumour assessments, a more restricted analysis population than the general RS.

Safety

The ClinicalTrials.gov record is reported as affected participants divided by participants at risk; the ClinicalTrials.gov record does not specify an additional formal safety comparison.

22. Design Features Not Reported in the Supplied Data

Several methodological topics commonly important in phase 3 trials are not supported by the registry-reported LUX-Lung 8 data. They are therefore not assigned a value or interpretation here.

TopicWhat can be concluded from the ClinicalTrials.gov record
Non-inferiority marginNot applicable to the reported hypothesis description; the registry identifies the hypothesis type as superiority and does not provide a non-inferiority margin.
CrossoverNo crossover analysis or crossover rate is provided.
Factorial designThe design is identified as parallel, not factorial.
Interim analysisNo interim-analysis procedure is provided in the ClinicalTrials.gov record.
Missing-data imputationNo formal missing-data or imputation strategy is provided in the ClinicalTrials.gov record.
Bayesian methodsNo Bayesian method is identified. The normalized methods are ANCOVA, Cox proportional-hazards model, log-rank test, and logistic regression.
Multiplicity adjustmentNo specific adjustment procedure is posted on ClinicalTrials.gov for the secondary analyses.

This distinction is important for statistical interpretation: an absent methodological detail should not be filled in merely because a particular approach is common in other clinical trials.

23. Limitations

24. Why This Trial Matters Statistically

LUX-Lung 8 is a useful teaching case because the registry results connect several major clinical-trial methods within one randomized comparison. The primary endpoint is a time-to-event outcome analyzed with a log-rank test and hazard ratio, while secondary outcomes demonstrate logistic regression, stratified Cox modeling, and ANCOVA with baseline covariate adjustment.

ConceptHow it appears in LUX-Lung 8
Randomization795 participants were enrolled in a randomized two-arm phase 3 parallel-group trial.
Time-to-event analysisProgression-free survival and overall survival are analyzed as time-to-event outcomes.
Log-rank testUsed for the posted PFS and overall-survival comparisons.
Hazard ratioUsed for PFS, OS, and symptom-deterioration analyses.
Confidence intervalReported alongside the major effect estimates.
Cox modelUsed for OS and symptom deterioration, with race identified as a stratification factor in several analyses.
Logistic regressionUsed for objective response and disease control.
Odds ratioQuantifies the binary-outcome comparisons for objective response and disease control.
ANCOVAUsed for tumour shrinkage with baseline sum of diameters as a covariate.
Covariate adjustmentTumour-shrinkage means were adjusted for baseline sum of diameters and race.
Analysis populationsThe Randomized Set is used for most efficacy analyses; tumour shrinkage is restricted to patients with tumour assessments.
Multiple secondary analysesSeparate efficacy and symptom endpoints illustrate why each estimate must be interpreted within its endpoint definition and analysis method.

25. A Statistical Reading of the Complete Results

The most important result is the primary progression-free survival comparison: HR 0.814, 95% CI 0.693–0.956, P = 0.0103. This is a time-to-event treatment comparison in the randomized set, with a hazard ratio below 1 and a confidence interval that remains below 1.

The overall-survival analysis reports a similar direction but a somewhat different estimate: HR 0.841, 95% CI 0.727–0.973, P = 0.0193. The two estimates should not be treated as interchangeable. PFS measures progression or death, whereas OS measures death, and the two endpoints have different event definitions and follow-up periods.

The binary outcomes add another dimension. Objective response has an odds ratio of 2.06 with a 95% CI of 0.98–4.32 and a p-value of 0.0551. Disease control has an odds ratio of 1.56 with a 95% CI of 1.18–2.06 and a p-value of 0.0020. These are odds comparisons, not direct probability differences.

The tumour-shrinkage analysis demonstrates a different modeling strategy. ANCOVA produced an adjusted mean difference of -1.2 with a 95% CI of -4.67 to 2.28 and a p-value of 0.500. That result should be interpreted on the continuous-outcome scale and should not be conflated with the binary response analysis.

The symptom analyses further illustrate endpoint-specific interpretation. The time-to-deterioration HRs are 0.89 for coughing, 0.79 for dyspnoea, and 0.99 for pain. The separate change-in-score analyses report mean differences of -3.5, -3.5, and -2.7, respectively. These results should remain separated because they represent different statistical quantities even though they concern related symptoms.

Clinical Biostats interpretation

The statistical story is therefore broader than the primary p-value. The trial contains a primary time-to-event analysis, a secondary survival endpoint, binary response and disease-control analyses, a covariate-adjusted continuous endpoint, symptom-deterioration survival analyses, and reported symptom-score differences. The appropriate interpretation is to examine each result on its own statistical scale while retaining the randomized design and the hierarchy between the primary and secondary endpoints.

26. Related Tutorials

Learn more about the methods used in this trial:

27. Related Statistical Calculators

28. Sources

Continue through Clinical Biostats

Explore the statistical methods behind randomized clinical trials, time-to-event endpoints, regression models, and clinical-trial analysis workflows.

29. Record Summary

LUX-Lung 8 provides a compact but methodologically diverse example of phase 3 clinical-trial statistics. The randomized trial enrolled 795 participants and compared afatinib with erlotinib. Its registered primary endpoint was progression-free survival, analyzed using a log-rank test with a reported hazard ratio of 0.814, 95% CI 0.693–0.956, and p-value 0.0103. Secondary analyses extend the statistical framework to overall survival, objective response, disease control, tumour shrinkage, symptom deterioration, and symptom-score changes.

The most useful statistical reading combines the effect measure, confidence interval, p-value, endpoint definition, analysis population, and model specification. Hazard ratios, odds ratios, and adjusted mean differences are not interchangeable quantities. Each describes a different aspect of the treatment comparison and must be interpreted on its own statistical scale.

Clinical Biostats methodology: A trial-results page should not merely repeat isolated numerical results. The goal is to reconstruct the statistical story of the trial while clearly separating reported evidence from educational interpretation and avoiding unsupported assumptions about analyses that are not documented in the ClinicalTrials.gov record.