This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
KEYNOTE-671 was a randomized, double-blind, parallel phase 3 trial evaluating pembrolizumab with platinum doublet chemotherapy as neoadjuvant/adjuvant therapy compared with placebo with the same chemotherapy framework in participants with resectable stage II, IIIA, and resectable IIIB (T3-4N2) non-small cell lung cancer.
| Feature | KEYNOTE-671 |
|---|---|
| Phase | Phase 3 |
| Condition | Non-small cell lung cancer |
| Population | Participants with resectable stage II, IIIA, and resectable IIIB (T3-4N2) non-small cell lung cancer |
| Design | Randomized, double-blind, parallel |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Primary endpoints | Event Free Survival (EFS) and Overall Survival (OS) |
| Primary endpoint type | Time-to-event |
| Enrollment | 797 |
| Trial start | April 24, 2018 |
| Primary completion | July 10, 2023 |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT03425643 |
2. Clinical Question
The central statistical question was whether neoadjuvant/adjuvant pembrolizumab combined with platinum doublet chemotherapy improves the two registered primary time-to-event endpoints, EFS and OS, compared with neoadjuvant/adjuvant placebo combined with platinum doublet chemotherapy.
Population
Participants with resectable stage II, IIIA, and resectable IIIB (T3-4N2) non-small cell lung cancer.
Intervention
Neoadjuvant/adjuvant pembrolizumab with platinum doublet chemotherapy.
Comparator
Neoadjuvant/adjuvant placebo with platinum doublet chemotherapy.
Primary question
Does the pembrolizumab-containing strategy improve EFS and OS relative to the placebo-containing strategy?
3. Trial Design
The design is statistically important because randomization establishes the principal comparison between the two treatment strategies, while double masking reduces the opportunity for knowledge of treatment assignment to influence trial conduct and assessment.
4. Primary Endpoints
| Endpoint | Registry definition | Time frame | Primary analysis |
|---|---|---|---|
| Event Free Survival (EFS) | EFS is defined as the time from randomization until radiographic disease progression, local progression precluding surgery, inability to resect the tumor, local or distant recurrence, or death due to any cause. EFS is determined either by biopsy assessed by local pathologist or by investigator-assessed imaging using Response Evaluation Criteria in Solid Tumors Version 1.1 (RECIST 1.1). | Up to approximately 5 years | Log-rank test; stratified Cox model for hazard ratio |
| Overall Survival (OS) | OS is defined as the time from randomization until death from any cause. The OS for all participants is presented through the database cut-off date of 10-Jul-2023. | Up to approximately 5 years | Log-rank test; stratified Cox model for hazard ratio |
Both primary endpoints are time-to-event outcomes. That matters statistically because the analysis must account not only for whether an event occurred, but also for the time until the event and for participants whose event time is not observed during follow-up.
5. Secondary Endpoints and Outcome Measures
| Endpoint | Time frame | Type | Analysis method reported |
|---|---|---|---|
| Major Pathological Response (mPR) Rate | Up to approximately 8 weeks following completion of neoadjuvant treatment (up to Study Week 20) | Count / rate | Stratified Miettinen and Nurminen |
| Pathological Complete Response (pCR) Rate | Up to approximately 8 weeks following completion of neoadjuvant treatment (up to Study Week 20) | Count / rate | Stratified Miettinen and Nurminen |
| Change From Baseline in Neoadjuvant Phase in EORTC QLQ-C30 Global Health Status (Item 29) Score | Baseline (cycle 1 in neoadjuvant phase) and neoadjuvant week 11 | Continuous | Two-sided t-test |
| Change From Baseline in Adjuvant Phase in EORTC QLQ-C30 Global Health Status (Item 29) Score | Baseline (cycle 1 in neoadjuvant phase) and adjuvant week 10 (up to Study Week 30) | Continuous | Two-sided t-test |
6. Statistical Methodology
Time-to-event analysis
The two primary endpoints are analyzed as time-to-event outcomes. The registry reports a log-rank method for both EFS and OS and describes a stratified Cox model with Efron's tie-handling method for the hazard-ratio analysis.
A hazard ratio summarizes a relative difference in the event hazard under the fitted survival model. It is not an absolute risk difference and does not state how many individual participants benefit.
Stratified Cox model
The EFS and OS analyses used a stratified Cox model with Efron's tie-handling method. Treatment was included as a covariate, while the model was stratified by Stage (II versus III), TPS (≥50% versus <50%), Histology (Squamous versus Non-squamous), and Region (East-Asia versus non-East Asia).
Stratification is useful when prespecified baseline factors are expected to influence the underlying event hazard. Instead of assuming a common baseline hazard across all strata, the stratified Cox framework permits separate baseline hazard functions while estimating the treatment effect across the strata.
Log-rank testing
The registry identifies the log-rank test as the reported method for both primary endpoints. Conceptually, the log-rank test compares the observed and expected numbers of events between treatment groups over the course of follow-up.
The test is a hypothesis test rather than an effect-size measure. The hazard ratio provides an estimate of relative treatment effect, while the confidence interval describes statistical uncertainty around that estimate.
Score-based confidence intervals for proportions
The major pathological response and pathological complete response analyses used a stratified Miettinen and Nurminen method. The ClinicalTrials.gov record normalize this as a score-based confidence-interval approach for proportions.
For these endpoints, the reported effect is a difference in percentage between the two treatment strategies. Stratification was by Stage, TPS, Histology, and Region, matching the factors used in the primary time-to-event model.
Continuous outcomes and t-tests
The two quality-of-life change endpoints were analyzed with two-sided t-tests. The reported effect measure was the difference in least-square means, with estimates based on a constrained longitudinal data analysis model incorporating treatment-by-visit interaction and the stratification factors.
This distinction is important: although the registry method field identifies a t-test, the analysis description also identifies a longitudinal modeling framework. The numerical effect measure is therefore a difference in least-square means rather than simply a raw difference between two independently calculated arithmetic means.
7. Primary Results: Event Free Survival
The EFS analysis included all randomized participants. The registry compares neoadjuvant/adjuvant pembrolizumab plus chemotherapy with neoadjuvant/adjuvant placebo plus chemotherapy using a log-rank test and a stratified Cox model.
Event Free Survival
95% CI: 0.48–0.72 · P < 0.00001
Two-sided confidence interval · Superiority hypothesis
| Primary endpoint | Analysis population | Method | Effect | 95% CI | P-value |
|---|---|---|---|---|---|
| Event Free Survival | All randomized participants | Log-rank; stratified Cox model | HR 0.59 | 0.48–0.72 | <0.00001 |
An EFS hazard ratio of 0.59 means that, under the reported stratified Cox model, the estimated instantaneous rate of an EFS event was 41% lower in the pembrolizumab-containing strategy than in the placebo-containing strategy. The 41% figure is a direct interpretation of 1 − 0.59; it is a relative hazard interpretation, not a statement that 41% of participants avoided an event.
The 95% CI of 0.48–0.72 describes uncertainty around the estimated hazard ratio. It does not describe the range of individual patient outcomes. The entire interval is below 1, so the reported estimate and its confidence interval are consistent with a lower estimated event hazard for the pembrolizumab-containing strategy.
The P-value of <0.00001 addresses evidence against the null hypothesis within the reported testing framework. It does not measure the size of the treatment effect. The magnitude of the effect is conveyed by the hazard ratio and its confidence interval.
Because the analysis is based on a Cox model, the hazard ratio is a model-based summary of the relative event hazard. Interpretation of a single HR is most straightforward when the proportional-hazards assumption is reasonably appropriate over follow-up. The ClinicalTrials.gov record does not provide a formal assessment of that assumption.
8. Primary Results: Overall Survival
The OS analysis also included all randomized participants. OS was defined from randomization until death from any cause, with the registry specifying presentation through the database cut-off date of 10-Jul-2023.
Overall Survival
95% CI: 0.56–0.93 · P = 0.00517
Two-sided confidence interval · Superiority hypothesis
| Primary endpoint | Analysis population | Method | Effect | 95% CI | P-value |
|---|---|---|---|---|---|
| Overall Survival | All randomized participants | Log-rank; stratified Cox model | HR 0.72 | 0.56–0.93 | 0.00517 |
An OS hazard ratio of 0.72 means that, under the reported stratified Cox model, the estimated instantaneous rate of death was 28% lower in the pembrolizumab-containing strategy than in the placebo-containing strategy.
The 95% CI of 0.56–0.93 expresses uncertainty around that estimate. Because the interval remains below 1, the reported confidence interval is consistent with a lower estimated hazard of death in the pembrolizumab-containing group.
The P-value of 0.00517 measures the strength of evidence against the null hypothesis under the stated statistical framework. It is not a probability that the treatment works, nor does it quantify the clinical magnitude of the difference. The HR and confidence interval provide the effect-size information.
As with EFS, the OS estimate comes from a stratified Cox model. Censoring, the time-to-event structure, the specified stratification factors, and the proportional-hazards assumption all matter to interpretation. The ClinicalTrials.gov record does not report a separate diagnostic assessment of proportional hazards.
9. Secondary Results: Major Pathological Response
The major pathological response rate was evaluated up to approximately 8 weeks following completion of neoadjuvant treatment, up to Study Week 20. The analysis population consisted of all randomized participants.
Difference in Major Pathological Response Rate
95% CI: 13.9–24.7 · P < 0.00001
Two-sided confidence interval · Stratified Miettinen and Nurminen method
| Endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Major Pathological Response Rate | Difference in percentage | 19.2 | 13.9–24.7 | <0.00001 |
The reported estimate of 19.2 percentage points is an absolute difference in the pathological response rate between the two randomized strategies. Unlike a hazard ratio, it is expressed directly on the percentage-point scale.
The 95% CI of 13.9–24.7 describes uncertainty around that difference. It does not mean that the treatment effect for an individual participant lies somewhere between 13.9 and 24.7 percentage points; it describes uncertainty in the estimated population-level difference.
The P-value of <0.00001 addresses statistical evidence under the reported analysis. It does not indicate that the treatment effect is "19.2% significant" or provide a measure of effect size.
The analysis was stratified by Stage, TPS, Histology, and Region. That means the reported difference incorporates the prespecified stratification structure rather than simply comparing two unadjusted overall percentages.
10. Secondary Results: Pathological Complete Response
The pathological complete response rate was evaluated over the same neoadjuvant time frame: up to approximately 8 weeks following completion of neoadjuvant treatment, up to Study Week 20.
Difference in Pathological Complete Response Rate
95% CI: 10.1–18.7 · P < 0.00001
Two-sided confidence interval · Stratified Miettinen and Nurminen method
| Endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Pathological Complete Response Rate | Difference in percentage | 14.2 | 10.1–18.7 | <0.00001 |
The estimate of 14.2 percentage points represents the reported absolute difference in pathological complete response rates between the randomized strategies.
The 95% CI of 10.1–18.7 indicates the precision of the estimated difference under the reported score-based method. The interval remains positive throughout, so the estimated difference is consistently above zero within the confidence interval.
The P-value of <0.00001 is evidence against the null hypothesis under the reported testing framework. It should not be confused with the magnitude of the response difference or with the probability that the observed difference will be reproduced in every future population.
The use of a stratified Miettinen and Nurminen method is particularly relevant because the response comparison is not being presented as a simple unstratified difference alone. The registry specifies stratification by Stage, TPS, Histology, and Region.
11. Secondary Results: Quality of Life — Neoadjuvant Phase
The neoadjuvant quality-of-life endpoint assessed change from baseline in the EORTC QLQ-C30 Global Health Status (Item 29) score between baseline, defined as cycle 1 in the neoadjuvant phase, and neoadjuvant week 11.
| Endpoint | Analysis population | Effect | 95% CI | P-value |
|---|---|---|---|---|
| Change From Baseline in Neoadjuvant Phase in EORTC QLQ-C30 Global Health Status (Item 29) Score | All randomized participants who received at least one dose of study treatment and had at least one EORTC-QLQ-C30 Item 30 assessment data available | Difference in least-square means: 1.43 | -1.64–4.49 | 0.3611 |
The estimated difference in least-square means was 1.43. The corresponding 95% CI of -1.64–4.49 spans zero, indicating substantial uncertainty about the direction of the between-group difference under the reported model.
The P-value of 0.3611 does not provide evidence against the null hypothesis under the reported two-sided test. It also does not establish equivalence or prove that the two strategies have identical quality-of-life effects.
The analysis population differs from the primary efficacy population: it includes randomized participants who received at least one dose and had at least one relevant assessment available. This distinction is important when comparing the quality-of-life analysis with the primary EFS and OS analyses, which used all randomized participants.
12. Secondary Results: Quality of Life — Adjuvant Phase
The adjuvant quality-of-life endpoint assessed change from baseline in the EORTC QLQ-C30 Global Health Status (Item 29) score between baseline, defined as cycle 1 in the neoadjuvant phase, and adjuvant week 10, up to Study Week 30.
| Endpoint | Analysis population | Effect | 95% CI | P-value |
|---|---|---|---|---|
| Change From Baseline in Adjuvant Phase in EORTC QLQ-C30 Global Health Status (Item 29) Score | All randomized participants who received at least one dose of study treatment and had at least one EORTC-QLQ-C30 Item 30 assessment data available | Difference in least-square means: 2.22 | -0.58–5.02 | 0.1197 |
The estimated difference in least-square means was 2.22. The 95% CI of -0.58–5.02 includes zero, so the estimate is compatible with both a small negative and a positive between-group difference under the reported analysis.
The P-value of 0.1197 does not provide evidence against the null hypothesis under the reported two-sided test. As with the neoadjuvant analysis, a nonsignificant P-value should not be translated into proof of no difference or statistical equivalence.
The registry describes a constrained longitudinal data analysis model with treatment-by-visit interaction and the trial's stratification factors. Thus, the reported least-square-mean difference reflects a model-based comparison rather than an unadjusted difference between two simple change scores.
13. Serious Adverse Events
The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected participants over participants at risk.
| Randomized strategy | Serious adverse events | Risk among those at risk |
|---|---|---|
| Pembro + Chemo/Pembro | 165 / 396 | 165 affected among 396 at risk |
| Placebo + Chemo/Placebo | 133 / 399 | 133 affected among 399 at risk |
The ClinicalTrials.gov record does not provide a formal comparative statistical analysis for serious adverse events. Accordingly, the counts should be treated as descriptive safety information rather than as evidence of a statistically tested difference between the randomized strategies.
The serious-adverse-event data answer a different question from EFS and OS. They describe the occurrence of a safety category among participants at risk; they are not time-to-event hazard ratios and should not be interpreted using the same framework as the primary survival endpoints.
The denominators are also important. The reported figures are 165/396 and 133/399. A comparison should preserve those denominators rather than assuming that the two safety populations have identical sizes.
14. Statistical Methods Explained
Why was a stratified Cox model used for EFS and OS?
The registry specifies a stratified Cox model with treatment as a covariate and stratification by Stage, TPS, Histology, and Region. Stratification allows those prespecified factors to influence the baseline hazard separately while estimating a common treatment hazard ratio across the strata. This is different from simply inserting every factor as an ordinary covariate into an unstratified Cox model.
What does an EFS hazard ratio of 0.59 mean?
Under the reported model, an HR of 0.59 corresponds to an estimated instantaneous event hazard that is 41% lower for the pembrolizumab-containing strategy relative to the placebo-containing strategy. It does not mean that 41% of participants were spared an EFS event, nor does it mean that every participant experienced the same relative reduction.
Why is the confidence interval important?
The point estimate alone does not communicate precision. For EFS, the estimate is 0.59 and the 95% CI is 0.48–0.72. For OS, the estimate is 0.72 and the 95% CI is 0.56–0.93. The intervals show the statistical uncertainty around the estimated relative effects under the stated analysis framework.
Why does the log-rank P-value not measure effect size?
The log-rank P-value evaluates evidence against a null hypothesis concerning the survival distributions. A very small P-value can occur with a modest effect in a sufficiently informative study, while a larger P-value can occur with a potentially important effect when precision is limited. The effect size is better represented by the hazard ratio and its confidence interval.
Why use Miettinen and Nurminen for pathological response rates?
Major pathological response and pathological complete response are binary response outcomes summarized as rates. The registry specifies a stratified Miettinen and Nurminen method, which provides a score-based approach for the confidence interval around the between-group difference in proportions while incorporating the specified stratification factors.
Why are the quality-of-life results analyzed differently from EFS and OS?
EFS and OS are time-to-event outcomes, whereas the EORTC QLQ-C30 endpoints are continuous score changes. The registry therefore reports a t-test framework for the quality-of-life comparisons and describes a constrained longitudinal data analysis model with treatment-by-visit interaction and the trial's stratification factors. The resulting effect measure is a difference in least-square means rather than a hazard ratio.
Why does the analysis population matter?
The primary EFS and OS analyses used all randomized participants. The quality-of-life analyses used randomized participants who received at least one dose and had at least one relevant assessment available. Those populations are not identical, so the resulting estimates answer somewhat different statistical questions.
15. Reading the Two Primary Hazard Ratios Together
| Primary endpoint | HR | 95% CI | P-value | Direct relative-hazard interpretation |
|---|---|---|---|---|
| EFS | 0.59 | 0.48–0.72 | <0.00001 | Approximately 41% lower estimated event hazard |
| OS | 0.72 | 0.56–0.93 | 0.00517 | Approximately 28% lower estimated hazard of death |
The two estimates should not be collapsed into a single "overall treatment effect." EFS and OS are different endpoints with different event definitions. EFS includes several events specified by the registry, including disease progression, local progression precluding surgery, inability to resect the tumor, local or distant recurrence, and death from any cause. OS is defined solely by death from any cause.
The EFS HR of 0.59 and OS HR of 0.72 both fall below 1, but they measure different processes. The fact that the numerical HRs are different is not itself evidence of inconsistency. Treatment effects can differ across endpoints because the endpoints count different types and timings of events.
16. Understanding the Response Results
Major pathological response
Reported difference: 19.2 percentage points, with a 95% CI of 13.9–24.7 and P < 0.00001.
Pathological complete response
Reported difference: 14.2 percentage points, with a 95% CI of 10.1–18.7 and P < 0.00001.
These are absolute percentage-point effects, so they should not be described as hazard reductions. A difference of 19.2 percentage points means that the estimated response rate differed by 19.2 percentage points between the randomized strategies under the reported analysis.
The two pathological response endpoints also have a defined assessment window: up to approximately 8 weeks following completion of neoadjuvant treatment, up to Study Week 20. That time frame matters because response rates are not timeless properties of a treatment; they depend on when and how the response is assessed.
17. Understanding the Quality-of-Life Results
| Phase | Estimate | 95% CI | P-value | Statistical reading |
|---|---|---|---|---|
| Neoadjuvant | 1.43 | -1.64–4.49 | 0.3611 | Confidence interval includes zero |
| Adjuvant | 2.22 | -0.58–5.02 | 0.1197 | Confidence interval includes zero |
Both reported confidence intervals include zero. For a difference in least-square means, zero corresponds to no estimated between-group difference. The results therefore do not provide evidence against zero under the reported two-sided tests.
18. Stratification and Covariate Adjustment
The primary EFS and OS analyses were stratified by four factors:
| Stratification factor | Categories specified in the analysis |
|---|---|
| Stage | II versus III |
| TPS | ≥50% versus <50% |
| Histology | Squamous versus Non-squamous |
| Region | East-Asia versus non-East Asia |
These factors appear in both the primary survival analysis and the stratified Miettinen and Nurminen response analyses. This provides a consistent statistical structure across the principal efficacy outcomes.
Stratification should not be interpreted as proof that the treatment effect is identical in every stratum. A stratified model estimates the overall treatment effect while allowing baseline hazards to differ by stratum. Establishing treatment-effect heterogeneity requires an appropriate interaction analysis rather than simply comparing the numerical estimates from individual subgroups.
19. Confidence Intervals: A Unified View
| Outcome | Estimate | 95% CI | Scale |
|---|---|---|---|
| EFS | 0.59 | 0.48–0.72 | Hazard ratio |
| OS | 0.72 | 0.56–0.93 | Hazard ratio |
| mPR rate | 19.2 | 13.9–24.7 | Percentage-point difference |
| pCR rate | 14.2 | 10.1–18.7 | Percentage-point difference |
| Neoadjuvant quality of life | 1.43 | -1.64–4.49 | Difference in least-square means |
| Adjuvant quality of life | 2.22 | -0.58–5.02 | Difference in least-square means |
This table illustrates why confidence intervals should always be interpreted together with the scale of the effect measure. A hazard-ratio interval is centered on a null value of 1, whereas a difference in percentages or means is centered on a null value of 0.
For HRs, values below 1 favor the pembrolizumab-containing strategy under the event definitions used here. For percentage-point differences, positive values indicate a higher response rate in the first listed treatment group. For mean differences, zero represents no between-group difference.
20. Missing Data, Censoring, and Analysis Populations
The ClinicalTrials.gov record explicitly define the analysis population for each reported statistical analysis, but they do not provide a detailed missing-data or imputation plan for the full trial.
| Analysis | Population specified in the registry data | Key statistical issue |
|---|---|---|
| EFS | All randomized participants | Time-to-event follow-up and censoring |
| OS | All randomized participants | Time-to-event follow-up and censoring |
| mPR | All randomized participants | Binary response classification and stratified rate comparison |
| pCR | All randomized participants | Binary response classification and stratified rate comparison |
| Neoadjuvant quality of life | Randomized participants who received at least one dose and had at least one relevant assessment | Availability of longitudinal assessment data |
| Adjuvant quality of life | Randomized participants who received at least one dose and had at least one relevant assessment | Availability of longitudinal assessment data |
For the time-to-event endpoints, censoring is a central statistical concept: participants who do not experience the event during observed follow-up can still contribute information up to their censoring time. The ClinicalTrials.gov record does not provide enough detail to specify every censoring rule or missing-data procedure used in the trial.
21. Multiplicity, Interim Analysis, and Other Design Topics
The ClinicalTrials.gov record identifies two primary endpoints, both tested for superiority, and provide formal analyses for both. They do not provide an alpha-allocation scheme, interim-analysis schedule, multiplicity-adjustment procedure, non-inferiority margin, crossover rule, or Bayesian analysis.
| Design topic | What the ClinicalTrials.gov record supports |
|---|---|
| Primary endpoints | Two: EFS and OS |
| Hypothesis type | Superiority for both primary endpoint analyses |
| Interim analysis | Not specified in the ClinicalTrials.gov record |
| Multiplicity procedure | Not specified in the ClinicalTrials.gov record |
| Non-inferiority margin | Not applicable to the reported superiority analyses |
| Crossover | Not specified in the ClinicalTrials.gov record |
| Bayesian methods | Not reported |
This distinction is important because the presence of multiple endpoints does not, by itself, tell us how the type I error was controlled. A formal statement about multiplicity would require the relevant statistical analysis plan or registry information specifying the testing hierarchy or alpha allocation.
22. What the Primary Results Do — and Do Not — Establish
What the EFS result establishes statistically
The reported stratified Cox analysis produced an HR of 0.59 with a 95% CI of 0.48–0.72 and P < 0.00001 for the randomized comparison.
What the OS result establishes statistically
The reported stratified Cox analysis produced an HR of 0.72 with a 95% CI of 0.56–0.93 and P = 0.00517 for the randomized comparison.
What the response results establish
The reported stratified analyses estimated differences of 19.2 percentage points for mPR and 14.2 percentage points for pCR.
What the quality-of-life results establish
The reported least-square-mean differences were 1.43 in the neoadjuvant phase and 2.22 in the adjuvant phase, with both confidence intervals spanning zero.
These results should not be reduced to a single P-value or single summary statistic. The trial contains time-to-event, binary response, continuous longitudinal, and safety outcomes, each with its own estimand, analysis population, effect measure, and uncertainty.
23. Limitations
- Registry-level information: the ClinicalTrials.gov record does not contain every element of the protocol or statistical analysis plan. Topics such as detailed sample-size assumptions, alpha allocation, interim monitoring, and missing-data rules cannot be reconstructed reliably from the provided record.
- Hazard-ratio interpretation: the EFS and OS effects are model-based hazard ratios. A single HR does not provide an absolute treatment effect at a particular time point and relies on the assumptions underlying the Cox analysis.
- Different endpoint scales: EFS and OS are expressed as hazard ratios, pathological response outcomes as percentage-point differences, and quality-of-life outcomes as differences in least-square means. These estimates are not directly interchangeable.
- Analysis populations differ: the primary efficacy analyses used all randomized participants, while the quality-of-life analyses required treatment exposure and available assessment data.
- Response analysis is time-specific: mPR and pCR were evaluated within the specified neoadjuvant assessment window. Their estimates should not be interpreted as long-term outcomes.
- Safety comparison: the registry-reported serious-adverse-event counts are descriptive. No formal between-arm statistical comparison is provided in the ClinicalTrials.gov record.
- Subgroup interpretation: the primary models are stratified by four baseline factors, but the ClinicalTrials.gov record does not provide subgroup-specific treatment-effect estimates or interaction tests.
- Unreported design details: the ClinicalTrials.gov record does not support conclusions about crossover, Bayesian methods, interim monitoring, or a detailed multiplicity strategy.
24. Why This Trial Matters Statistically
KEYNOTE-671 is a useful teaching case because it combines several common clinical-trial statistical frameworks within one randomized study. The same trial moves from time-to-event analysis for EFS and OS, to stratified binary-outcome analysis for pathological response, to longitudinal continuous-outcome analysis for quality of life.
| Concept | How it appears in KEYNOTE-671 |
|---|---|
| Randomization | Participants were randomized to two parallel strategies. |
| Double blinding | The trial was double-blind. |
| Time-to-event endpoints | EFS and OS were the two primary endpoints. |
| Log-rank test | Reported for both primary time-to-event comparisons. |
| Hazard ratio | Used to quantify relative EFS and OS effects. |
| Stratified Cox model | Used for EFS and OS with four specified stratification factors. |
| Confidence intervals | Reported for all six statistical analyses reported. |
| Risk difference | Used for the major pathological response comparison. |
| Score-based CI for proportions | Miettinen and Nurminen method used for mPR and pCR. |
| t-test / longitudinal modeling | Used for the two quality-of-life comparisons. |
| Analysis populations | Primary survival analyses used all randomized participants; quality-of-life analyses used a more restricted assessment population. |
From a biostatistical perspective, the most instructive feature is not any single numerical result. It is the alignment between the clinical question, endpoint type, analysis population, effect measure, and statistical method.
25. Related Tutorials
Learn more about the methods used in this trial:
26. Related Calculators
27. Sources
- ClinicalTrials.gov: KEYNOTE-671, NCT03425643.
- PubMed: PMID 39288781.
- PubMed: PMID 37272513.
- PubMed: PMID 42628840.
- PubMed: PMID 41875364.
- PubMed: PMID 38717993.
Continue through the Clinical Biostats statistical pathway
Explore the statistical concepts behind randomized trials, survival endpoints, confidence intervals, response rates, and longitudinal outcomes.
28. Record Summary
KEYNOTE-671 provides a broad example of how different clinical-trial questions require different statistical estimands and methods. The two primary endpoints, EFS and OS, were analyzed as time-to-event outcomes using log-rank testing and stratified Cox models. The reported EFS HR was 0.59 (95% CI 0.48–0.72; P < 0.00001), while the reported OS HR was 0.72 (95% CI 0.56–0.93; P = 0.00517).
The secondary analyses illustrate two additional statistical scales. Major pathological response and pathological complete response were analyzed with stratified Miettinen and Nurminen methods, yielding percentage-point differences of 19.2 and 14.2, respectively. The two quality-of-life analyses used least-square-mean differences of 1.43 and 2.22, with confidence intervals that included zero.
The most important statistical lesson is that these estimates should be interpreted on their appropriate scales. A hazard ratio describes a relative event hazard, a response-rate difference describes an absolute percentage-point difference, and a least-square-mean difference describes a model-based difference in a continuous outcome. Confidence intervals and analysis populations are essential parts of interpreting each result.