← Clinical Trials
Non-Small Cell Lung Cancer Phase 3 Randomized NCT03425643

KEYNOTE-671: Complete Statistical Analysis of Pembrolizumab in Resectable Non-Small Cell Lung Cancer

An independent statistical analysis of the randomized phase 3 KEYNOTE-671 trial evaluating pembrolizumab with platinum doublet chemotherapy as neoadjuvant/adjuvant therapy for participants with resectable stage II, IIIA, and resectable IIIB (T3-4N2) non-small cell lung cancer.

Trial status: Active, not recruiting  ·  Enrollment: 797  ·  Primary completion: July 10, 2023
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

KEYNOTE-671 was a randomized, double-blind, parallel phase 3 trial evaluating pembrolizumab with platinum doublet chemotherapy as neoadjuvant/adjuvant therapy compared with placebo with the same chemotherapy framework in participants with resectable stage II, IIIA, and resectable IIIB (T3-4N2) non-small cell lung cancer.

797
Enrolled
Randomized trial
2
Arms
Parallel design
0.59
EFS HR
95% CI 0.48–0.72
0.72
OS HR
95% CI 0.56–0.93
FeatureKEYNOTE-671
PhasePhase 3
ConditionNon-small cell lung cancer
PopulationParticipants with resectable stage II, IIIA, and resectable IIIB (T3-4N2) non-small cell lung cancer
DesignRandomized, double-blind, parallel
AllocationRandomized
Primary purposeTreatment
Primary endpointsEvent Free Survival (EFS) and Overall Survival (OS)
Primary endpoint typeTime-to-event
Enrollment797
Trial startApril 24, 2018
Primary completionJuly 10, 2023
Lead sponsorMerck Sharp & Dohme LLC
Sponsor typeIndustry
ClinicalTrials.govNCT03425643

2. Clinical Question

The central statistical question was whether neoadjuvant/adjuvant pembrolizumab combined with platinum doublet chemotherapy improves the two registered primary time-to-event endpoints, EFS and OS, compared with neoadjuvant/adjuvant placebo combined with platinum doublet chemotherapy.

Population

Participants with resectable stage II, IIIA, and resectable IIIB (T3-4N2) non-small cell lung cancer.

Intervention

Neoadjuvant/adjuvant pembrolizumab with platinum doublet chemotherapy.

Comparator

Neoadjuvant/adjuvant placebo with platinum doublet chemotherapy.

Primary question

Does the pembrolizumab-containing strategy improve EFS and OS relative to the placebo-containing strategy?

3. Trial Design

01
Randomize797 participants
02
NeoadjuvantStudy treatment before surgery
03
SurgeryResectability is incorporated into EFS
04
AdjuvantStudy treatment after surgery
05
Follow-upEFS and OS
Allocation
Randomized allocation with two parallel treatment arms.
Masking
Double-blind.
Primary purpose
Treatment.
Interventions
Pembrolizumab, placebo, cisplatin, gemcitabine, and pemetrexed.

The design is statistically important because randomization establishes the principal comparison between the two treatment strategies, while double masking reduces the opportunity for knowledge of treatment assignment to influence trial conduct and assessment.

What the registry data establish: the ClinicalTrials.gov record identifies a randomized, double-blind, parallel phase 3 trial with 797 participants and two arms. It does not provide enough information in the ClinicalTrials.gov record to reconstruct treatment-arm sample sizes or a detailed treatment schedule.

4. Primary Endpoints

EndpointRegistry definitionTime framePrimary analysis
Event Free Survival (EFS) EFS is defined as the time from randomization until radiographic disease progression, local progression precluding surgery, inability to resect the tumor, local or distant recurrence, or death due to any cause. EFS is determined either by biopsy assessed by local pathologist or by investigator-assessed imaging using Response Evaluation Criteria in Solid Tumors Version 1.1 (RECIST 1.1). Up to approximately 5 years Log-rank test; stratified Cox model for hazard ratio
Overall Survival (OS) OS is defined as the time from randomization until death from any cause. The OS for all participants is presented through the database cut-off date of 10-Jul-2023. Up to approximately 5 years Log-rank test; stratified Cox model for hazard ratio

Both primary endpoints are time-to-event outcomes. That matters statistically because the analysis must account not only for whether an event occurred, but also for the time until the event and for participants whose event time is not observed during follow-up.

5. Secondary Endpoints and Outcome Measures

EndpointTime frameTypeAnalysis method reported
Major Pathological Response (mPR) Rate Up to approximately 8 weeks following completion of neoadjuvant treatment (up to Study Week 20) Count / rate Stratified Miettinen and Nurminen
Pathological Complete Response (pCR) Rate Up to approximately 8 weeks following completion of neoadjuvant treatment (up to Study Week 20) Count / rate Stratified Miettinen and Nurminen
Change From Baseline in Neoadjuvant Phase in EORTC QLQ-C30 Global Health Status (Item 29) Score Baseline (cycle 1 in neoadjuvant phase) and neoadjuvant week 11 Continuous Two-sided t-test
Change From Baseline in Adjuvant Phase in EORTC QLQ-C30 Global Health Status (Item 29) Score Baseline (cycle 1 in neoadjuvant phase) and adjuvant week 10 (up to Study Week 30) Continuous Two-sided t-test

6. Statistical Methodology

Time-to-event analysis

The two primary endpoints are analyzed as time-to-event outcomes. The registry reports a log-rank method for both EFS and OS and describes a stratified Cox model with Efron's tie-handling method for the hazard-ratio analysis.

Conceptual hazard-ratio interpretation
HR < 1  →  lower estimated instantaneous event rate in the pembrolizumab group

A hazard ratio summarizes a relative difference in the event hazard under the fitted survival model. It is not an absolute risk difference and does not state how many individual participants benefit.

Stratified Cox model

The EFS and OS analyses used a stratified Cox model with Efron's tie-handling method. Treatment was included as a covariate, while the model was stratified by Stage (II versus III), TPS (≥50% versus <50%), Histology (Squamous versus Non-squamous), and Region (East-Asia versus non-East Asia).

Stratification is useful when prespecified baseline factors are expected to influence the underlying event hazard. Instead of assuming a common baseline hazard across all strata, the stratified Cox framework permits separate baseline hazard functions while estimating the treatment effect across the strata.

Log-rank testing

The registry identifies the log-rank test as the reported method for both primary endpoints. Conceptually, the log-rank test compares the observed and expected numbers of events between treatment groups over the course of follow-up.

The test is a hypothesis test rather than an effect-size measure. The hazard ratio provides an estimate of relative treatment effect, while the confidence interval describes statistical uncertainty around that estimate.

Score-based confidence intervals for proportions

The major pathological response and pathological complete response analyses used a stratified Miettinen and Nurminen method. The ClinicalTrials.gov record normalize this as a score-based confidence-interval approach for proportions.

For these endpoints, the reported effect is a difference in percentage between the two treatment strategies. Stratification was by Stage, TPS, Histology, and Region, matching the factors used in the primary time-to-event model.

Continuous outcomes and t-tests

The two quality-of-life change endpoints were analyzed with two-sided t-tests. The reported effect measure was the difference in least-square means, with estimates based on a constrained longitudinal data analysis model incorporating treatment-by-visit interaction and the stratification factors.

This distinction is important: although the registry method field identifies a t-test, the analysis description also identifies a longitudinal modeling framework. The numerical effect measure is therefore a difference in least-square means rather than simply a raw difference between two independently calculated arithmetic means.

7. Primary Results: Event Free Survival

The EFS analysis included all randomized participants. The registry compares neoadjuvant/adjuvant pembrolizumab plus chemotherapy with neoadjuvant/adjuvant placebo plus chemotherapy using a log-rank test and a stratified Cox model.

Event Free Survival

HR 0.59

95% CI: 0.48–0.72   ·   P < 0.00001

Two-sided confidence interval  ·  Superiority hypothesis

Primary endpointAnalysis populationMethodEffect95% CIP-value
Event Free Survival All randomized participants Log-rank; stratified Cox model HR 0.59 0.48–0.72 <0.00001
Clinical Biostats interpretation

An EFS hazard ratio of 0.59 means that, under the reported stratified Cox model, the estimated instantaneous rate of an EFS event was 41% lower in the pembrolizumab-containing strategy than in the placebo-containing strategy. The 41% figure is a direct interpretation of 1 − 0.59; it is a relative hazard interpretation, not a statement that 41% of participants avoided an event.

The 95% CI of 0.48–0.72 describes uncertainty around the estimated hazard ratio. It does not describe the range of individual patient outcomes. The entire interval is below 1, so the reported estimate and its confidence interval are consistent with a lower estimated event hazard for the pembrolizumab-containing strategy.

The P-value of <0.00001 addresses evidence against the null hypothesis within the reported testing framework. It does not measure the size of the treatment effect. The magnitude of the effect is conveyed by the hazard ratio and its confidence interval.

Because the analysis is based on a Cox model, the hazard ratio is a model-based summary of the relative event hazard. Interpretation of a single HR is most straightforward when the proportional-hazards assumption is reasonably appropriate over follow-up. The ClinicalTrials.gov record does not provide a formal assessment of that assumption.

8. Primary Results: Overall Survival

The OS analysis also included all randomized participants. OS was defined from randomization until death from any cause, with the registry specifying presentation through the database cut-off date of 10-Jul-2023.

Overall Survival

HR 0.72

95% CI: 0.56–0.93   ·   P = 0.00517

Two-sided confidence interval  ·  Superiority hypothesis

Primary endpointAnalysis populationMethodEffect95% CIP-value
Overall Survival All randomized participants Log-rank; stratified Cox model HR 0.72 0.56–0.93 0.00517
Clinical Biostats interpretation

An OS hazard ratio of 0.72 means that, under the reported stratified Cox model, the estimated instantaneous rate of death was 28% lower in the pembrolizumab-containing strategy than in the placebo-containing strategy.

The 95% CI of 0.56–0.93 expresses uncertainty around that estimate. Because the interval remains below 1, the reported confidence interval is consistent with a lower estimated hazard of death in the pembrolizumab-containing group.

The P-value of 0.00517 measures the strength of evidence against the null hypothesis under the stated statistical framework. It is not a probability that the treatment works, nor does it quantify the clinical magnitude of the difference. The HR and confidence interval provide the effect-size information.

As with EFS, the OS estimate comes from a stratified Cox model. Censoring, the time-to-event structure, the specified stratification factors, and the proportional-hazards assumption all matter to interpretation. The ClinicalTrials.gov record does not report a separate diagnostic assessment of proportional hazards.

9. Secondary Results: Major Pathological Response

The major pathological response rate was evaluated up to approximately 8 weeks following completion of neoadjuvant treatment, up to Study Week 20. The analysis population consisted of all randomized participants.

Difference in Major Pathological Response Rate

19.2 percentage points

95% CI: 13.9–24.7   ·   P < 0.00001

Two-sided confidence interval  ·  Stratified Miettinen and Nurminen method

EndpointEffect measureEstimate95% CIP-value
Major Pathological Response Rate Difference in percentage 19.2 13.9–24.7 <0.00001
Clinical Biostats interpretation

The reported estimate of 19.2 percentage points is an absolute difference in the pathological response rate between the two randomized strategies. Unlike a hazard ratio, it is expressed directly on the percentage-point scale.

The 95% CI of 13.9–24.7 describes uncertainty around that difference. It does not mean that the treatment effect for an individual participant lies somewhere between 13.9 and 24.7 percentage points; it describes uncertainty in the estimated population-level difference.

The P-value of <0.00001 addresses statistical evidence under the reported analysis. It does not indicate that the treatment effect is "19.2% significant" or provide a measure of effect size.

The analysis was stratified by Stage, TPS, Histology, and Region. That means the reported difference incorporates the prespecified stratification structure rather than simply comparing two unadjusted overall percentages.

10. Secondary Results: Pathological Complete Response

The pathological complete response rate was evaluated over the same neoadjuvant time frame: up to approximately 8 weeks following completion of neoadjuvant treatment, up to Study Week 20.

Difference in Pathological Complete Response Rate

14.2 percentage points

95% CI: 10.1–18.7   ·   P < 0.00001

Two-sided confidence interval  ·  Stratified Miettinen and Nurminen method

EndpointEffect measureEstimate95% CIP-value
Pathological Complete Response Rate Difference in percentage 14.2 10.1–18.7 <0.00001
Clinical Biostats interpretation

The estimate of 14.2 percentage points represents the reported absolute difference in pathological complete response rates between the randomized strategies.

The 95% CI of 10.1–18.7 indicates the precision of the estimated difference under the reported score-based method. The interval remains positive throughout, so the estimated difference is consistently above zero within the confidence interval.

The P-value of <0.00001 is evidence against the null hypothesis under the reported testing framework. It should not be confused with the magnitude of the response difference or with the probability that the observed difference will be reproduced in every future population.

The use of a stratified Miettinen and Nurminen method is particularly relevant because the response comparison is not being presented as a simple unstratified difference alone. The registry specifies stratification by Stage, TPS, Histology, and Region.

11. Secondary Results: Quality of Life — Neoadjuvant Phase

The neoadjuvant quality-of-life endpoint assessed change from baseline in the EORTC QLQ-C30 Global Health Status (Item 29) score between baseline, defined as cycle 1 in the neoadjuvant phase, and neoadjuvant week 11.

EndpointAnalysis populationEffect95% CIP-value
Change From Baseline in Neoadjuvant Phase in EORTC QLQ-C30 Global Health Status (Item 29) Score All randomized participants who received at least one dose of study treatment and had at least one EORTC-QLQ-C30 Item 30 assessment data available Difference in least-square means: 1.43 -1.64–4.49 0.3611
Clinical Biostats interpretation

The estimated difference in least-square means was 1.43. The corresponding 95% CI of -1.64–4.49 spans zero, indicating substantial uncertainty about the direction of the between-group difference under the reported model.

The P-value of 0.3611 does not provide evidence against the null hypothesis under the reported two-sided test. It also does not establish equivalence or prove that the two strategies have identical quality-of-life effects.

The analysis population differs from the primary efficacy population: it includes randomized participants who received at least one dose and had at least one relevant assessment available. This distinction is important when comparing the quality-of-life analysis with the primary EFS and OS analyses, which used all randomized participants.

12. Secondary Results: Quality of Life — Adjuvant Phase

The adjuvant quality-of-life endpoint assessed change from baseline in the EORTC QLQ-C30 Global Health Status (Item 29) score between baseline, defined as cycle 1 in the neoadjuvant phase, and adjuvant week 10, up to Study Week 30.

EndpointAnalysis populationEffect95% CIP-value
Change From Baseline in Adjuvant Phase in EORTC QLQ-C30 Global Health Status (Item 29) Score All randomized participants who received at least one dose of study treatment and had at least one EORTC-QLQ-C30 Item 30 assessment data available Difference in least-square means: 2.22 -0.58–5.02 0.1197
Clinical Biostats interpretation

The estimated difference in least-square means was 2.22. The 95% CI of -0.58–5.02 includes zero, so the estimate is compatible with both a small negative and a positive between-group difference under the reported analysis.

The P-value of 0.1197 does not provide evidence against the null hypothesis under the reported two-sided test. As with the neoadjuvant analysis, a nonsignificant P-value should not be translated into proof of no difference or statistical equivalence.

The registry describes a constrained longitudinal data analysis model with treatment-by-visit interaction and the trial's stratification factors. Thus, the reported least-square-mean difference reflects a model-based comparison rather than an unadjusted difference between two simple change scores.

13. Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected participants over participants at risk.

Randomized strategySerious adverse eventsRisk among those at risk
Pembro + Chemo/Pembro 165 / 396 165 affected among 396 at risk
Placebo + Chemo/Placebo 133 / 399 133 affected among 399 at risk

The ClinicalTrials.gov record does not provide a formal comparative statistical analysis for serious adverse events. Accordingly, the counts should be treated as descriptive safety information rather than as evidence of a statistically tested difference between the randomized strategies.

Clinical Biostats interpretation

The serious-adverse-event data answer a different question from EFS and OS. They describe the occurrence of a safety category among participants at risk; they are not time-to-event hazard ratios and should not be interpreted using the same framework as the primary survival endpoints.

The denominators are also important. The reported figures are 165/396 and 133/399. A comparison should preserve those denominators rather than assuming that the two safety populations have identical sizes.

14. Statistical Methods Explained

Why was a stratified Cox model used for EFS and OS?

The registry specifies a stratified Cox model with treatment as a covariate and stratification by Stage, TPS, Histology, and Region. Stratification allows those prespecified factors to influence the baseline hazard separately while estimating a common treatment hazard ratio across the strata. This is different from simply inserting every factor as an ordinary covariate into an unstratified Cox model.

What does an EFS hazard ratio of 0.59 mean?

Under the reported model, an HR of 0.59 corresponds to an estimated instantaneous event hazard that is 41% lower for the pembrolizumab-containing strategy relative to the placebo-containing strategy. It does not mean that 41% of participants were spared an EFS event, nor does it mean that every participant experienced the same relative reduction.

Why is the confidence interval important?

The point estimate alone does not communicate precision. For EFS, the estimate is 0.59 and the 95% CI is 0.48–0.72. For OS, the estimate is 0.72 and the 95% CI is 0.56–0.93. The intervals show the statistical uncertainty around the estimated relative effects under the stated analysis framework.

Why does the log-rank P-value not measure effect size?

The log-rank P-value evaluates evidence against a null hypothesis concerning the survival distributions. A very small P-value can occur with a modest effect in a sufficiently informative study, while a larger P-value can occur with a potentially important effect when precision is limited. The effect size is better represented by the hazard ratio and its confidence interval.

Why use Miettinen and Nurminen for pathological response rates?

Major pathological response and pathological complete response are binary response outcomes summarized as rates. The registry specifies a stratified Miettinen and Nurminen method, which provides a score-based approach for the confidence interval around the between-group difference in proportions while incorporating the specified stratification factors.

Why are the quality-of-life results analyzed differently from EFS and OS?

EFS and OS are time-to-event outcomes, whereas the EORTC QLQ-C30 endpoints are continuous score changes. The registry therefore reports a t-test framework for the quality-of-life comparisons and describes a constrained longitudinal data analysis model with treatment-by-visit interaction and the trial's stratification factors. The resulting effect measure is a difference in least-square means rather than a hazard ratio.

Why does the analysis population matter?

The primary EFS and OS analyses used all randomized participants. The quality-of-life analyses used randomized participants who received at least one dose and had at least one relevant assessment available. Those populations are not identical, so the resulting estimates answer somewhat different statistical questions.

15. Reading the Two Primary Hazard Ratios Together

Primary endpointHR95% CIP-valueDirect relative-hazard interpretation
EFS 0.59 0.48–0.72 <0.00001 Approximately 41% lower estimated event hazard
OS 0.72 0.56–0.93 0.00517 Approximately 28% lower estimated hazard of death

The two estimates should not be collapsed into a single "overall treatment effect." EFS and OS are different endpoints with different event definitions. EFS includes several events specified by the registry, including disease progression, local progression precluding surgery, inability to resect the tumor, local or distant recurrence, and death from any cause. OS is defined solely by death from any cause.

A useful statistical distinction

The EFS HR of 0.59 and OS HR of 0.72 both fall below 1, but they measure different processes. The fact that the numerical HRs are different is not itself evidence of inconsistency. Treatment effects can differ across endpoints because the endpoints count different types and timings of events.

16. Understanding the Response Results

Major pathological response

Reported difference: 19.2 percentage points, with a 95% CI of 13.9–24.7 and P < 0.00001.

Pathological complete response

Reported difference: 14.2 percentage points, with a 95% CI of 10.1–18.7 and P < 0.00001.

These are absolute percentage-point effects, so they should not be described as hazard reductions. A difference of 19.2 percentage points means that the estimated response rate differed by 19.2 percentage points between the randomized strategies under the reported analysis.

The two pathological response endpoints also have a defined assessment window: up to approximately 8 weeks following completion of neoadjuvant treatment, up to Study Week 20. That time frame matters because response rates are not timeless properties of a treatment; they depend on when and how the response is assessed.

17. Understanding the Quality-of-Life Results

PhaseEstimate95% CIP-valueStatistical reading
Neoadjuvant 1.43 -1.64–4.49 0.3611 Confidence interval includes zero
Adjuvant 2.22 -0.58–5.02 0.1197 Confidence interval includes zero

Both reported confidence intervals include zero. For a difference in least-square means, zero corresponds to no estimated between-group difference. The results therefore do not provide evidence against zero under the reported two-sided tests.

Important interpretation: a P-value above 0.05 does not prove that the treatment strategies have identical quality-of-life effects. Likewise, a confidence interval containing zero does not establish equivalence. A formal equivalence or non-inferiority conclusion would require a prespecified equivalence or non-inferiority framework and margin; none is reported in the ClinicalTrials.gov record.

18. Stratification and Covariate Adjustment

The primary EFS and OS analyses were stratified by four factors:

Stratification factorCategories specified in the analysis
StageII versus III
TPS≥50% versus <50%
HistologySquamous versus Non-squamous
RegionEast-Asia versus non-East Asia

These factors appear in both the primary survival analysis and the stratified Miettinen and Nurminen response analyses. This provides a consistent statistical structure across the principal efficacy outcomes.

Stratification should not be interpreted as proof that the treatment effect is identical in every stratum. A stratified model estimates the overall treatment effect while allowing baseline hazards to differ by stratum. Establishing treatment-effect heterogeneity requires an appropriate interaction analysis rather than simply comparing the numerical estimates from individual subgroups.

19. Confidence Intervals: A Unified View

OutcomeEstimate95% CIScale
EFS0.590.48–0.72Hazard ratio
OS0.720.56–0.93Hazard ratio
mPR rate19.213.9–24.7Percentage-point difference
pCR rate14.210.1–18.7Percentage-point difference
Neoadjuvant quality of life1.43-1.64–4.49Difference in least-square means
Adjuvant quality of life2.22-0.58–5.02Difference in least-square means

This table illustrates why confidence intervals should always be interpreted together with the scale of the effect measure. A hazard-ratio interval is centered on a null value of 1, whereas a difference in percentages or means is centered on a null value of 0.

Statistical reading of the null value

For HRs, values below 1 favor the pembrolizumab-containing strategy under the event definitions used here. For percentage-point differences, positive values indicate a higher response rate in the first listed treatment group. For mean differences, zero represents no between-group difference.

20. Missing Data, Censoring, and Analysis Populations

The ClinicalTrials.gov record explicitly define the analysis population for each reported statistical analysis, but they do not provide a detailed missing-data or imputation plan for the full trial.

AnalysisPopulation specified in the registry dataKey statistical issue
EFS All randomized participants Time-to-event follow-up and censoring
OS All randomized participants Time-to-event follow-up and censoring
mPR All randomized participants Binary response classification and stratified rate comparison
pCR All randomized participants Binary response classification and stratified rate comparison
Neoadjuvant quality of life Randomized participants who received at least one dose and had at least one relevant assessment Availability of longitudinal assessment data
Adjuvant quality of life Randomized participants who received at least one dose and had at least one relevant assessment Availability of longitudinal assessment data

For the time-to-event endpoints, censoring is a central statistical concept: participants who do not experience the event during observed follow-up can still contribute information up to their censoring time. The ClinicalTrials.gov record does not provide enough detail to specify every censoring rule or missing-data procedure used in the trial.

21. Multiplicity, Interim Analysis, and Other Design Topics

The ClinicalTrials.gov record identifies two primary endpoints, both tested for superiority, and provide formal analyses for both. They do not provide an alpha-allocation scheme, interim-analysis schedule, multiplicity-adjustment procedure, non-inferiority margin, crossover rule, or Bayesian analysis.

Design topicWhat the ClinicalTrials.gov record supports
Primary endpointsTwo: EFS and OS
Hypothesis typeSuperiority for both primary endpoint analyses
Interim analysisNot specified in the ClinicalTrials.gov record
Multiplicity procedureNot specified in the ClinicalTrials.gov record
Non-inferiority marginNot applicable to the reported superiority analyses
CrossoverNot specified in the ClinicalTrials.gov record
Bayesian methodsNot reported

This distinction is important because the presence of multiple endpoints does not, by itself, tell us how the type I error was controlled. A formal statement about multiplicity would require the relevant statistical analysis plan or registry information specifying the testing hierarchy or alpha allocation.

22. What the Primary Results Do — and Do Not — Establish

What the EFS result establishes statistically

The reported stratified Cox analysis produced an HR of 0.59 with a 95% CI of 0.48–0.72 and P < 0.00001 for the randomized comparison.

What the OS result establishes statistically

The reported stratified Cox analysis produced an HR of 0.72 with a 95% CI of 0.56–0.93 and P = 0.00517 for the randomized comparison.

What the response results establish

The reported stratified analyses estimated differences of 19.2 percentage points for mPR and 14.2 percentage points for pCR.

What the quality-of-life results establish

The reported least-square-mean differences were 1.43 in the neoadjuvant phase and 2.22 in the adjuvant phase, with both confidence intervals spanning zero.

These results should not be reduced to a single P-value or single summary statistic. The trial contains time-to-event, binary response, continuous longitudinal, and safety outcomes, each with its own estimand, analysis population, effect measure, and uncertainty.

23. Limitations

24. Why This Trial Matters Statistically

KEYNOTE-671 is a useful teaching case because it combines several common clinical-trial statistical frameworks within one randomized study. The same trial moves from time-to-event analysis for EFS and OS, to stratified binary-outcome analysis for pathological response, to longitudinal continuous-outcome analysis for quality of life.

ConceptHow it appears in KEYNOTE-671
RandomizationParticipants were randomized to two parallel strategies.
Double blindingThe trial was double-blind.
Time-to-event endpointsEFS and OS were the two primary endpoints.
Log-rank testReported for both primary time-to-event comparisons.
Hazard ratioUsed to quantify relative EFS and OS effects.
Stratified Cox modelUsed for EFS and OS with four specified stratification factors.
Confidence intervalsReported for all six statistical analyses reported.
Risk differenceUsed for the major pathological response comparison.
Score-based CI for proportionsMiettinen and Nurminen method used for mPR and pCR.
t-test / longitudinal modelingUsed for the two quality-of-life comparisons.
Analysis populationsPrimary survival analyses used all randomized participants; quality-of-life analyses used a more restricted assessment population.

From a biostatistical perspective, the most instructive feature is not any single numerical result. It is the alignment between the clinical question, endpoint type, analysis population, effect measure, and statistical method.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Calculators

27. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical concepts behind randomized trials, survival endpoints, confidence intervals, response rates, and longitudinal outcomes.

28. Record Summary

KEYNOTE-671 provides a broad example of how different clinical-trial questions require different statistical estimands and methods. The two primary endpoints, EFS and OS, were analyzed as time-to-event outcomes using log-rank testing and stratified Cox models. The reported EFS HR was 0.59 (95% CI 0.48–0.72; P < 0.00001), while the reported OS HR was 0.72 (95% CI 0.56–0.93; P = 0.00517).

The secondary analyses illustrate two additional statistical scales. Major pathological response and pathological complete response were analyzed with stratified Miettinen and Nurminen methods, yielding percentage-point differences of 19.2 and 14.2, respectively. The two quality-of-life analyses used least-square-mean differences of 1.43 and 2.22, with confidence intervals that included zero.

The most important statistical lesson is that these estimates should be interpreted on their appropriate scales. A hazard ratio describes a relative event hazard, a response-rate difference describes an absolute percentage-point difference, and a least-square-mean difference describes a model-based difference in a continuous outcome. Confidence intervals and analysis populations are essential parts of interpreting each result.

Clinical Biostats methodology: A trial-results page should distinguish the reported numerical evidence from the statistical interpretation of that evidence. For KEYNOTE-671, that means preserving the registry definitions and reported estimates while explaining why the survival, response, quality-of-life, and safety analyses require different statistical frameworks.