This page separates reported trial results from statistical interpretation. Numerical results are taken from the ClinicalTrials.gov record. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
KEYNOTE-189 was a randomized, parallel-group, quadruple-masked phase 3 treatment trial in non-small-cell lung carcinoma. The registry reports 616 participants, two arms, two registered primary endpoints, and statistical analyses for both primary endpoints plus a secondary response endpoint and a prespecified additional PFS endpoint.
| Feature | KEYNOTE-189 |
|---|---|
| Trial name | KEYNOTE-189 |
| NCT ID | NCT02578680 |
| Phase | Phase 3 |
| Status | Completed |
| Therapeutic area | Oncology |
| Condition | Non-Small-Cell Lung Carcinoma |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 616 |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | Industry |
| Primary endpoint type | Time-to-event |
| Results posted | Yes |
2. Clinical Question
The registry describes KEYNOTE-189 as a study of pemetrexed plus platinum chemotherapy with or without pembrolizumab in participants with first-line metastatic nonsquamous non-small cell lung cancer. The primary statistical question was whether the pembrolizumab-containing regimen was superior to the control regimen with respect to progression-free survival and overall survival.
Population
Participants with first-line metastatic nonsquamous non-small cell lung cancer, represented in the registry under the condition Non-Small-Cell Lung Carcinoma.
Intervention
Pembrolizumab 200 mg with pemetrexed and platinum chemotherapy, with subsequent pembrolizumab plus pemetrexed in the reported comparison.
Comparator
The registry comparison is against the control regimen, described in the analysis as the control group receiving pemetrexed-platinum chemotherapy.
Primary question
Does the pembrolizumab-containing regimen improve the registered time-to-event endpoints PFS and OS compared with control?
3. Trial Design
The trial used randomized allocation, a parallel design, quadruple masking, and a treatment-oriented primary purpose. Enrollment was 616 participants across two arms.
Pembrolizumab-containing regimen
- Pembrolizumab 200 mg
- Pemetrexed
- Cisplatin or carboplatin
- Folic acid 350-1000 μg
- Vitamin B12 1000 μg
- Dexamethasone 4 mg
- Saline solution
Control regimen
- Pemetrexed
- Cisplatin or carboplatin
- Saline solution and associated supportive interventions as registered
The statistical comparison reported by the registry is explicitly between pembrolizumab + pemetrexed + platinum chemotherapy followed by pembrolizumab + pemetrexed and control.
4. Primary Endpoints
Both registered primary endpoints were time-to-event outcomes with a time frame of up to approximately 21 months. The registry reports a formal statistical analysis for each endpoint.
| Endpoint | Registry definition | Time frame | Analysis |
|---|---|---|---|
| Progression-Free Survival (PFS) | PFS was defined as the time from randomization to the first documented progressive disease or death due to any cause, whichever occurred first. Per RECIST 1.1, progressive disease was defined as ≥20% increase in the sum of diameters of target lesions, taking as reference the smallest sum on study, together with an absolute increase of ≥5 mm. | Up to approximately 21 months | Stratified Cox regression with log-rank testing |
| Overall Survival (OS) | OS was defined as the time from randomization to death due to any cause. Participants without documented death at the time of the interim analysis were censored at the date of the last follow-up. | Up to approximately 21 months | Stratified Cox regression with log-rank testing |
5. Statistical Methodology
Log-rank testing for time-to-event endpoints
The registry reports the log-rank test as the primary comparison method for both PFS and OS. A log-rank test evaluates whether the event-time distributions differ between randomized groups over follow-up, accounting for the timing of events rather than reducing the endpoint to a single fixed-time proportion.
Stratified Cox regression
The reported hazard ratios were based on Cox regression with treatment as a covariate and stratification by three factors: PD-L1 status (≥1% vs. <1%), platinum chemotherapy (cisplatin vs. carboplatin), and smoking status (never vs. former/current).
The registry identifies pembrolizumab as the numerator and control as the denominator for the reported hazard ratios.
Score-based comparison for response
The secondary ORR analysis used the Miettinen and Nurminen method. The registry describes the analysis as treatment adjusted and stratified by PD-L1 status, platinum chemotherapy, and smoking status.
Superiority framework
All four registry-reported formal analyses are classified as superiority analyses. This is important because the interpretation is based on whether the observed treatment comparison supports a difference in favor of the intervention, rather than whether the intervention remains within a prespecified non-inferiority margin.
Analysis populations
The registry identifies all randomized participants as the analysis population for each registry-reported formal analysis. This is consistent across the two primary endpoints, the secondary ORR analysis, and the prespecified additional PFS analysis reported in the ClinicalTrials.gov record.
6. Primary Result: Progression-Free Survival
The primary PFS analysis compared pembrolizumab + pemetrexed + platinum chemotherapy followed by pembrolizumab + pemetrexed with control. The endpoint was assessed by blinded central imaging and defined according to RECIST 1.1.
Hazard ratio for progression or death
95% CI: 0.43–0.64 · P < 0.00001
Two-sided 95% confidence interval; superiority hypothesis.
| Element | Reported result |
|---|---|
| Analysis population | All randomized participants |
| Method | Log-rank test; hazard ratio based on Cox regression |
| Effect measure | Hazard ratio |
| Estimate | 0.52 |
| 95% CI | 0.43–0.64 |
| P-value | <0.00001 |
| Hypothesis | Superiority |
An HR of 0.52 means that the fitted model estimated an instantaneous rate of progression or death approximately 48% lower in the pembrolizumab-containing group than in the control group, because 1 − 0.52 = 0.48. This is a relative time-to-event interpretation, not a statement that 48% of participants avoided progression or death.
The 95% CI of 0.43–0.64 describes uncertainty around the estimated hazard ratio under the model and analysis framework. It does not describe the range of individual patient outcomes, nor does it mean that the true treatment effect has a 95% probability of lying inside the interval.
The P < 0.00001 value addresses evidence against the null hypothesis within the specified testing framework. It does not measure the size of the treatment effect. The HR and its confidence interval describe the effect estimate and its precision; the p-value addresses statistical evidence against the null.
The analysis was stratified by PD-L1 status, platinum chemotherapy, and smoking status. Because the estimate comes from a Cox model, its interpretation also depends on the model's assumptions, including the proportional-hazards framework. The registry does not provide enough information here to independently assess that assumption from the underlying event-time data.
7. Primary Result: Overall Survival
OS was the second registered primary endpoint. It measured time from randomization to death from any cause, with participants without documented death at the interim analysis censored at their last follow-up date.
Hazard ratio for death
95% CI: 0.38–0.64 · P < 0.00001
Two-sided 95% confidence interval; superiority hypothesis.
| Element | Reported result |
|---|---|
| Analysis population | All randomized participants |
| Method | Log-rank test; hazard ratio based on Cox regression |
| Effect measure | Hazard ratio |
| Estimate | 0.49 |
| 95% CI | 0.38–0.64 |
| P-value | <0.00001 |
| Hypothesis | Superiority |
An HR of 0.49 means that the fitted model estimated an instantaneous rate of death approximately 51% lower in the pembrolizumab-containing group than in the control group. It does not mean that 51% of participants survived, that 51% of participants were cured, or that every participant experienced the same relative reduction in mortality.
The 95% CI of 0.38–0.64 provides a measure of precision around the estimated hazard ratio. The interval is narrower than a very imprecise estimate would be, but it still reflects uncertainty arising from the finite randomized sample and observed event information.
The P < 0.00001 result is evidence against the null hypothesis under the reported superiority analysis. It is not a measure of clinical magnitude. A small p-value can occur with a modest effect when information is abundant, while a large effect estimate can be imprecise in a small study.
As with the PFS analysis, the OS estimate was obtained using a stratified Cox model. The analysis incorporated PD-L1 status, platinum chemotherapy, and smoking status as stratification factors. Censoring also matters: participants without documented death at the interim analysis contributed follow-up until their last follow-up date.
8. Secondary Endpoint Result: Overall Response Rate
Overall Response Rate (ORR) was a secondary binary endpoint defined per RECIST 1.1 as assessed by blinded central imaging. The registry reports the effect as the difference in percentage versus control.
Difference in response rate versus control
95% CI: 21.1–35.4 · P < 0.0001
Two-sided 95% confidence interval; superiority hypothesis.
| Element | Reported result |
|---|---|
| Endpoint | Overall Response Rate per RECIST 1.1 as assessed by blinded central imaging |
| Analysis population | All randomized participants |
| Method | Stratified Miettinen and Nurminen |
| Effect measure | Difference in percentage vs. control |
| Risk difference | 28.5 |
| 95% CI | 21.1–35.4 |
| P-value | <0.0001 |
A reported difference of 28.5 percentage points means that the response-rate proportion in the pembrolizumab-containing group exceeded the corresponding control proportion by 28.5 percentage points under the reported analysis.
The 95% CI of 21.1–35.4 percentage points describes uncertainty around that between-group difference. It is a confidence interval for the difference in proportions, not an interval for an individual patient's probability of response.
The Miettinen-Nurminen method is a score-based approach for inference on differences between proportions. The analysis was stratified by PD-L1 status, platinum chemotherapy, and smoking status, so the reported estimate should not simply be interpreted as an unadjusted subtraction of two crude percentages.
The P < 0.0001 value addresses statistical evidence for a difference under the specified hypothesis test. It does not quantify how important a 28.5-percentage-point difference is clinically, and it should not be substituted for the confidence interval when discussing precision.
9. Other Prespecified Result: PFS by Investigator Immune-Related RECIST
The registry also reports a prespecified PFS analysis using investigator-assessed immune-related RECIST (irRECIST) response criteria. This endpoint was not identified as one of the two primary endpoints in the ClinicalTrials.gov record.
Hazard ratio for PFS by irRECIST
95% CI: 0.41–0.59 · P < 0.00001
Time frame: up to approximately 39 months.
The HR of 0.49 corresponds to an approximately 51% lower estimated instantaneous rate of the PFS event in the pembrolizumab-containing group under this analysis. This estimate uses a different endpoint assessment framework from the primary PFS endpoint, so it should not be silently substituted for the blinded-central-imaging RECIST 1.1 result.
The 95% CI of 0.41–0.59 quantifies uncertainty around the reported hazard ratio. The very small p-value indicates strong statistical evidence against the null hypothesis under the reported test, but the p-value itself does not measure effect size or clinical importance.
10. Summary of Reported Statistical Results
| Endpoint | Role | Effect | 95% CI | P-value |
|---|---|---|---|---|
| PFS by blinded central imaging, RECIST 1.1 | Primary | HR 0.52 | 0.43–0.64 | <0.00001 |
| OS | Primary | HR 0.49 | 0.38–0.64 | <0.00001 |
| ORR by blinded central imaging, RECIST 1.1 | Secondary | Risk difference 28.5 percentage points | 21.1–35.4 | <0.0001 |
| PFS by investigator irRECIST | Other prespecified | HR 0.49 | 0.41–0.59 | <0.00001 |
The ClinicalTrials.gov record therefore contain formal results for both primary endpoints, one secondary endpoint, and one additional prespecified endpoint. All registry-reported formal analyses use the all-randomized analysis population.
11. Serious Adverse Events
The ClinicalTrials.gov record reports serious adverse events by arm using affected participants over participants at risk. The ClinicalTrials.gov record contains the following counts.
| Reported group | Affected / at risk |
|---|---|
| Pembrolizumab + Pemetrexed + Platinum Chemot | 233 / 405 |
| Control | 102 / 202 |
| Control Switched Over to Pembrolizumab M | 29 / 84 |
| Control Switched Over to Pembrolizumab M | 0 / 2 |
| Pembrolizumab + Pemetrexed + Platinum Chemot | 2 / 9 |
12. Randomization and Stratified Analysis
Randomization is the design feature that creates the principal basis for comparing the treatment groups. The registry does not supply an allocation ratio in the ClinicalTrials.gov record, so no ratio is inferred here from the enrollment total.
The formal analyses nevertheless provide important detail about stratification. Both primary time-to-event analyses used a Cox regression model with treatment as a covariate stratified by:
- PD-L1 status: ≥1% vs. <1%
- Platinum chemotherapy: cisplatin vs. carboplatin
- Smoking status: never vs. former/current
The ORR analysis used the Miettinen and Nurminen method with the same treatment-covariate and stratification structure. Thus, stratification was not merely a descriptive baseline feature; it was part of the inferential analysis reported by the registry.
13. Statistical Methods Explained
Why use a hazard ratio for PFS and OS?
PFS and OS are time-to-event endpoints. A hazard ratio summarizes the relative instantaneous event rate between the randomized groups over the analyzed follow-up. Unlike a fixed-time risk difference, it uses information about when events occur and accommodates censoring.
What does an HR of 0.52 mean?
An HR of 0.52 indicates that the fitted model estimated the instantaneous rate of the event at approximately 52% of the corresponding rate in the control group. Equivalently, the estimated rate was approximately 48% lower. It does not mean that 52% of patients experienced an event, nor that each patient had exactly the same relative reduction.
Why was a stratified Cox model used?
The registry states that treatment was included as a covariate while the Cox model was stratified by PD-L1 status, platinum chemotherapy, and smoking status. Stratification allows the baseline hazard to differ across those strata while estimating the treatment effect within the specified modeling framework.
Why is the log-rank test appropriate for PFS and OS?
The log-rank test is designed for comparing survival distributions between groups when observations may be right-censored. It uses the ordering of event times rather than treating every participant as if the endpoint were observed at a common fixed time.
What does the 95% CI for the hazard ratio tell us?
The 95% CI gives a range of values reflecting statistical uncertainty around the estimated hazard ratio under the model and sampling framework. For the primary PFS analysis, the interval is 0.43–0.64; for OS, it is 0.38–0.64. Neither interval describes the distribution of treatment effects across individual patients.
Why was a different method used for ORR?
ORR is a binary endpoint rather than a time-to-event endpoint. The registry reports the stratified Miettinen and Nurminen method for comparing response proportions. Its reported effect measure is a difference in percentage versus control rather than a hazard ratio.
Why does the p-value not measure effect size?
A p-value quantifies statistical evidence against a null hypothesis under the specified test. It depends on both the observed effect and the amount of information in the study. The effect estimate and confidence interval are therefore necessary to understand magnitude and precision.
14. Kaplan-Meier Estimation and Censoring
Although the registry statistical-analyses records identify log-rank testing and Cox regression, the underlying PFS and OS endpoints are time-to-event outcomes. Kaplan-Meier estimation is the standard descriptive framework for displaying the survival function for such outcomes.
Here, di is the number of events at time ti and ni is the number at risk immediately before that time.
For OS, the registry explicitly states that participants without documented death at the time of the interim analysis were censored at the date of last follow-up. This is an important feature of survival analysis: a censored participant is not treated as having experienced the event, but their available follow-up still contributes information up to the censoring time.
15. Confidence Intervals and Effect Measures
KEYNOTE-189 illustrates two different classes of effect measure in the statistical analyses posted on ClinicalTrials.gov.
Hazard ratio
Used for PFS and OS. The reported estimates are 0.52 for primary PFS, 0.49 for OS, and 0.49 for the prespecified irRECIST PFS analysis.
Risk difference
Used for ORR. The reported difference in percentage versus control is 28.5 percentage points, with a two-sided 95% CI of 21.1–35.4.
These measures should not be compared as if they were on the same scale. A hazard ratio describes a relative time-to-event effect under a survival model; a risk difference describes an absolute difference between proportions.
16. Multiplicity and Hypothesis Testing
The trial data identify two registered primary endpoints, both analyzed under a superiority hypothesis. The ClinicalTrials.gov record does not provide an alpha-spending plan, interim-analysis boundary, formal multiplicity allocation scheme, or detailed hierarchical testing sequence.
| Feature | What the ClinicalTrials.gov record supports | What is not reported |
|---|---|---|
| Primary endpoints | 2: PFS and OS | Detailed alpha allocation between them |
| Hypothesis type | Superiority | Detailed testing hierarchy |
| Primary p-values | PFS <0.00001; OS <0.00001 | Multiplicity-adjusted alpha values |
| Interim analysis | The OS definition references an interim analysis | Boundary, alpha spending, or information fraction |
| Additional analysis | irRECIST PFS reported as Other_Pre_Specified | Formal multiplicity relationship to primary endpoints |
17. Missing Data and Imputation
The ClinicalTrials.gov record does not report a missing-data or imputation method for the primary endpoints. For OS, the registry explicitly describes censoring for participants without documented death at the interim analysis. For PFS, the endpoint is likewise a time-to-event measure, but the registry text does not provide a separate imputation algorithm.
This distinction matters statistically. Censoring is not the same as imputation. In a survival analysis, a participant whose event has not been observed by the last available follow-up is typically treated as right-censored rather than assigned an invented event time.
18. Blinding and Outcome Assessment
the ClinicalTrials.gov record identifies the study as quadruple-masked. The primary PFS endpoint was assessed by blinded central imaging, and the secondary ORR endpoint was also assessed by blinded central imaging.
Blinding
Quadruple masking reduces opportunities for knowledge of treatment assignment to influence trial conduct and assessment.
Central imaging
Blinded central imaging provides a standardized assessment framework for RECIST 1.1-based PFS and ORR endpoints.
For the additional prespecified PFS analysis, the registry identifies investigator-assessed immune-related RECIST criteria rather than blinded central imaging. That difference in assessment framework is an important reason to keep the endpoint analyses distinct.
19. Non-Inferiority, Bayesian Methods, and Other Design Topics
The ClinicalTrials.gov record classifies the formal hypotheses as superiority. No non-inferiority margin is reported, so there is no non-inferiority margin to interpret for this trial.
No Bayesian method is identified in the registry-reported statistical methodology. The normalized methods are the log-rank test and score-based confidence intervals for proportions using the Miettinen-Nurminen, Newcombe, or Wilson family of methods.
The ClinicalTrials.gov record also do not report a factorial design. The design model is parallel with two arms.
20. What the Hazard Ratios Do — and Do Not — Mean
The primary PFS HR of 0.52 represents an estimated 48% lower instantaneous rate of progression or death under the reported Cox model. It does not mean a 48% absolute improvement in PFS, and it does not imply that 48% of participants benefit.
The OS HR of 0.49 represents an estimated 51% lower instantaneous rate of death under the reported model. It is a relative time-to-event measure and should not be translated directly into an absolute survival percentage.
The PFS CI of 0.43–0.64 and OS CI of 0.38–0.64 communicate the precision of the respective estimates. They provide more information than the p-values alone because they show the range of effect estimates compatible with the statistical uncertainty represented by the analysis.
21. Important Limitations and Interpretation Issues
- Registry-level detail: The ClinicalTrials.gov record provides formal effect estimates and methods but not the complete statistical analysis plan, individual participant data, or all analysis specifications.
- Proportional-hazards interpretation: Hazard ratios are model-based. The ClinicalTrials.gov record does not provide the underlying event-time information needed to independently evaluate the proportional-hazards assumption.
- Censoring: OS participants without documented death at the interim analysis were censored at last follow-up. Interpretation therefore depends on the survival-analysis framework and censoring mechanism.
- Multiplicity: Two primary endpoints were analyzed, but the ClinicalTrials.gov record does not provide the detailed multiplicity strategy or alpha allocation needed to reconstruct the complete confirmatory testing framework.
- Analysis population: The registry-reported formal analyses use all randomized participants. This preserves the randomized comparison but does not describe treatment exposure in the same way as an as-treated safety analysis.
- Endpoint assessment: Primary PFS and secondary ORR used blinded central imaging, whereas the additional prespecified PFS analysis used investigator-assessed immune-related RECIST criteria.
- Safety denominators: The registry-reported serious-adverse-event data include multiple groups and denominators, including switched-over groups. These should not be recombined without the underlying registry result structure.
- No subgroup estimates reported: Although PD-L1 status, platinum chemotherapy, and smoking status were used for stratification, the ClinicalTrials.gov record does not provide subgroup-specific treatment estimates. Stratification should therefore not be interpreted as evidence of treatment-effect heterogeneity.
22. Why This Trial Matters Statistically
KEYNOTE-189 provides a compact teaching example of how randomized clinical-trial evidence can combine time-to-event and binary endpoints while accounting for prespecified stratification factors.
| Concept | How it appears in KEYNOTE-189 |
|---|---|
| Randomization | Randomized parallel-group phase 3 design with 616 participants |
| Blinding | Quadruple masking |
| Time-to-event endpoints | PFS and OS |
| RECIST 1.1 | Primary PFS and secondary ORR assessed by blinded central imaging |
| Log-rank test | Reported formal method for both primary endpoints |
| Hazard ratio | Effect measure for PFS and OS |
| Cox regression | Used with treatment as a covariate and prespecified stratification factors |
| Stratified analysis | PD-L1 status, platinum chemotherapy, and smoking status |
| Risk difference | Effect measure for ORR |
| Miettinen-Nurminen method | Reported method for the secondary ORR comparison |
| Confidence intervals | 95% two-sided intervals for all registry-reported formal effect estimates |
| Superiority testing | Hypothesis type for all registry-reported formal analyses |
| Censoring | Explicitly described for participants without documented death at the OS interim analysis |
23. Statistical Methods Explained: Practical Interpretation
Relative vs absolute effects
The HRs describe relative event rates, whereas the ORR risk difference describes an absolute difference in response proportions. These should not be treated as interchangeable measures.
Precision vs significance
The confidence interval addresses precision. The p-value addresses evidence against the null hypothesis. Neither quantity alone provides a complete description of the treatment effect.
Stratification vs subgroup analysis
A factor can be used for stratification without being the subject of a treatment-effect heterogeneity claim. The ClinicalTrials.gov record identifies PD-L1 status, platinum choice, and smoking status as stratification factors.
Assessment framework
Primary PFS and ORR used blinded central imaging under RECIST 1.1, while the additional PFS analysis used investigator-assessed immune-related RECIST criteria.
24. Related Tutorials
Learn more about the statistical methods used in this trial:
25. Related Calculators
26. Sources
- ClinicalTrials.gov: NCT02578680 — KEYNOTE-189.
- PubMed: PMID 40498298.
- PubMed: PMID 38642841.
- PubMed: PMID 37465924.
- PubMed: PMID 36809080.
- PubMed: PMID 36793385.
Continue with the statistical methods
Explore the survival-analysis, confidence-interval, and clinical-trial methods that appear in KEYNOTE-189.
27. Record Summary
KEYNOTE-189 illustrates a randomized phase 3 statistical framework combining two primary time-to-event endpoints with a secondary binary response endpoint. The primary PFS analysis reported an HR of 0.52 (95% CI 0.43–0.64; P < 0.00001), while the primary OS analysis reported an HR of 0.49 (95% CI 0.38–0.64; P < 0.00001). The secondary ORR analysis reported a risk difference of 28.5 percentage points (95% CI 21.1–35.4; P < 0.0001). A further prespecified irRECIST PFS analysis reported an HR of 0.49 (95% CI 0.41–0.59; P < 0.00001).
The statistical interpretation depends on understanding what each effect measure represents. Hazard ratios summarize relative time-to-event effects under a Cox model, while the ORR risk difference describes an absolute difference between binary response proportions. The trial also demonstrates the importance of prespecified stratification, blinded central assessment, censoring, confidence intervals, and the distinction between primary and additional endpoint analyses.