This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
KEYNOTE-021 was a randomized, parallel-design, open-label phase 1/2 study in non-small cell lung carcinoma. The registry reports 267 participants, 14 arms, three registered primary endpoints, six posted outcome measures, and four posted statistical analyses.
| Feature | KEYNOTE-021 |
|---|---|
| Trial name | KEYNOTE-021 |
| NCT identifier | NCT02039674 |
| Phase | Phase 1/2 |
| Status | COMPLETED |
| Therapeutic area | Oncology |
| Condition | Non-small Cell Lung Carcinoma |
| Allocation | RANDOMIZED |
| Design model | PARALLEL |
| Masking | NONE |
| Primary purpose | TREATMENT |
| Enrollment | 267.0 |
| Arms | 14 |
| Study start | 2014-02-21 |
| Primary completion | 2016-11-07 |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | INDUSTRY |
2. Clinical Question
The registry describes KEYNOTE-021 as a study of pembrolizumab in combination with chemotherapy or immunotherapy in participants with non-small cell lung carcinoma. The statistical questions represented in the posted analyses include whether objective response rates differed between randomized Cohort G treatment groups and whether the observed response rate in Cohorts D4 and H exceeded a prespecified 20% benchmark.
Population
Participants with non-small cell lung carcinoma enrolled in the phase 1/2 KEYNOTE-021 study.
Interventions
The registry lists pembrolizumab, paclitaxel, carboplatin, bevacizumab, pemetrexed, ipilimumab, erlotinib, and gefitinib among the study interventions.
Comparator
For the primary randomized Cohort G comparison, the groups were pembrolizumab 200 mg + pemetrexed + carboplatin versus placebo + pemetrexed + carboplatin.
Primary statistical questions
The posted primary analyses address objective response rate in Cohorts G+ and G- and objective response rate in Cohorts D4 and H. Dose-limiting toxicity was also registered as a primary endpoint.
3. Trial Design
The registry lists 14 arms rather than a single two-arm trial. Consequently, the overall enrollment number of 267.0 should not be interpreted as if all participants belonged to one simple two-group comparison. The posted statistical analyses identify the particular cohorts and analysis populations used for each reported comparison.
4. Registered Primary Endpoints
| Registered primary endpoint | Time frame | Type | Registry definition |
|---|---|---|---|
| Part 2 Cohorts G+ and G-: Objective Response Rate (ORR) | Up to approximately 2 years | Binary | ORR was defined as the percentage of participants in the analysis population who had a Complete Response (CR: Disappearance of all target lesions) or a Partial Response (PR: At least a 30% decrease in the sum of diameters of target lesions) per RECIST 1.1 as assessed by blinded independent central review (BICR). |
| Part 2 Cohorts D4 and H: Objective Response Rate (ORR) | Up to approximately 2 years | Binary | For participants who demonstrated a confirmed response, CR was disappearance of all target lesions and PR was at least a 30% decrease in the sum of diameters of target lesions per RECIST 1.1. Duration of response was defined as the time from first documented evidence of CR or PR until disease progression or death. DOR was assessed by BICR. |
| All Cohorts: Number of Participants Who Experienced a Dose-limiting Toxicity (DLT) | Cycle 1 (Up to 21 days) | Binary | DLTs were assessed using National Cancer Institute Common Terminology Criteria for Adverse Events version 4. The registry definition includes Grade 4 non-hematologic toxicity, Grade 4 hematologic toxicity lasting ≥7 days, and specified Grade 3 non-hematologic toxicities lasting >3 days despite optimal supportive care. |
The registry therefore contains both response-based efficacy endpoints and an early-treatment safety endpoint. These endpoints require different statistical approaches: binary response and DLT outcomes are naturally summarized as proportions, whereas progression-free survival and overall survival require time-to-event methods because not every participant necessarily experiences the event during observation.
5. Statistical Methodology
Score-based confidence interval for the randomized ORR difference
The primary Cohort G comparison used the Miettinen & Nurminen Method for the difference in percentages. This is a score-based approach for inference on the difference between two binomial proportions.
Here the registry reports the effect as a difference in percentages. A positive value means that the observed response percentage was higher in Cohort G+ than in Cohort G-.
Exact binomial testing
The D4 and H ORR analysis used the exact binomial distribution for testing. The posted hypothesis was H0: ORR ≤20% versus H1: ORR >20%.
Log-rank testing for time-to-event endpoints
The posted Cohort G analyses of progression-free survival and overall survival used the log-rank test. The registry also identifies one-sided testing in the analysis text for these comparisons.
Hazard ratios
The PFS and OS analyses report hazard ratios as the effect measure. A hazard ratio compares the modeled instantaneous event rates between groups over the analyzed follow-up; it is not a ratio of median survival times and is not an absolute difference in the percentage of participants experiencing an event.
For a superiority comparison, a hazard ratio below 1 is directionally consistent with fewer or later events in the treatment group, but the magnitude and uncertainty of the estimate must be considered together.
Kaplan-Meier estimation
Progression-free survival and overall survival are time-to-event endpoints. Kaplan-Meier estimation is the standard framework for estimating the survival function while retaining information from participants who are censored before experiencing the event.
The estimator multiplies conditional survival probabilities across observed event times, where di is the number of events and ni is the number at risk immediately before an event time.
6. Analysis Populations
| Analysis | Population reported by the registry | Database cutoff |
|---|---|---|
| Cohorts G+ and G- ORR | All randomized Cohort G participants | 08 August 2016 |
| Cohorts D4 and H ORR | All treated Cohort D4 and Cohort H participants; one Cohort H participant was excluded from the efficacy analysis population | 07 November 2016 |
| Cohorts G+ and G- PFS | All randomized Cohort G participants | 19 August 2019 |
| Cohorts G+ and G- OS | All randomized Cohort G participants | 19 August 2019 |
This distinction is important because the four posted analyses do not necessarily describe the same data snapshot. The ORR comparison and the later PFS/OS analyses use different database cutoff dates. A later cutoff can provide additional follow-up, but it should not be silently combined with an earlier analysis as if all estimates came from the same information set.
7. Results: Objective Response Rate in Cohorts G+ and G-
The primary randomized Cohort G analysis compared Part 2 Cohort G+ (pembrolizumab 200 mg + pemetrexed + carboplatin) with Part 2 Cohort G- (placebo + pemetrexed + carboplatin). The analysis population consisted of all randomized Cohort G participants, with a database cutoff date of 08 August 2016.
Difference in Objective Response Rate
95% CI: 8.9–42.1 · P = 0.0016
Miettinen & Nurminen method · Superiority hypothesis
| Feature | Reported result |
|---|---|
| Analysis population | All randomized Cohort G participants |
| Groups compared | Cohort G+ vs Cohort G- |
| Endpoint | Objective Response Rate |
| Time frame | Up to approximately 2 years |
| Effect measure | Difference in Percentages |
| Estimate | 26.3 |
| 95% CI | 8.9–42.1 |
| P-value | 0.0016 |
| Hypothesis | Superiority |
The reported estimate of 26.3 means that the percentage of participants meeting the registry definition of objective response differed by 26.3 percentage points between the two randomized Cohort G groups, with the positive direction corresponding to Cohort G+ relative to Cohort G-.
The estimate does not mean that every participant had a 26.3% increase in the probability of response, nor does it describe duration of response, progression-free survival, or overall survival. ORR is a binary tumor-response endpoint.
The 95% CI of 8.9–42.1 describes uncertainty around the estimated between-group difference under the statistical method used. It is not a range in which individual patients' treatment effects are expected to fall.
The P-value of 0.0016 addresses evidence against the null hypothesis under the specified testing framework. It is not a measure of the size or clinical importance of the response difference. The effect estimate and its confidence interval are therefore essential parts of the interpretation.
Because this is a binary endpoint, the Miettinen-Nurminen approach is appropriate for direct inference on the difference between proportions. The analysis is also tied to the stated analysis population and the 08 August 2016 data cutoff.
8. Results: Objective Response Rate in Cohorts D4 and H
The second primary efficacy analysis evaluated Part 2 Cohorts D4 and H among treated participants. The registry reports that one Cohort H participant was excluded from the efficacy analysis population. The statistical test used an exact binomial distribution.
Superiority test against a 20% ORR benchmark
H0: ORR ≤20% vs H1: ORR >20%
Exact binomial distribution for testing
| Feature | Reported result |
|---|---|
| Analysis population | All treated Cohort D4 and Cohort H participants, with one Cohort H participant excluded from the efficacy analysis population |
| Group | Part 2 Cohorts D4 & H (pembrolizumab 2 mg/kg + ipilimumab) |
| Endpoint | Objective Response Rate |
| Time frame | Up to approximately 2 years |
| Method | Exact binomial distribution for testing |
| Hypothesis | H0: ORR ≤20% vs H1: ORR >20% |
| P-value | 0.0858 |
| Hypothesis type | Superiority |
The analysis tests whether the response rate exceeded a prespecified 20% benchmark. This is different from the randomized Cohort G comparison: there is no second randomized treatment group in the posted D4/H analysis against which the response rate is directly compared.
The reported P-value of 0.0858 measures the evidence reported by the observed response data against the specified null hypothesis under the exact binomial testing framework. It does not quantify the magnitude of the response rate and does not provide a confidence interval by itself.
The exact-binomial framework is useful when the endpoint is a proportion and the analysis is framed around a fixed benchmark. The interpretation should remain tied to the stated hypothesis of ORR ≤20% versus ORR >20% rather than being converted into an unsupported between-treatment comparison.
9. Results: Progression-Free Survival in Cohorts G+ and G-
The later Cohort G progression-free survival analysis used all randomized Cohort G participants and a database cutoff date of 19 August 2019. The registry reports a log-rank analysis, a hazard ratio as the effect measure, and a one-sided P-value based on the log-rank test.
Hazard ratio for progression-free survival
95% CI: 0.35–0.83 · P = 0.00252
Log-rank test · One-sided P-value based on the log-rank test
| Feature | Reported result |
|---|---|
| Analysis population | All randomized Cohort G participants |
| Groups compared | Cohort G+ vs Cohort G- |
| Endpoint | Progression-Free Survival |
| Time frame | Up to approximately 2 years |
| Outcome unit | Months |
| Method | Log Rank |
| Effect measure | Hazard Ratio (HR) |
| Estimate | 0.54 |
| 95% CI | 0.35–0.83 |
| P-value | 0.00252 |
| Hypothesis type | Superiority |
| Analysis note | One-sided P-value based on log-rank test |
An HR of 0.54 means that the estimated instantaneous rate of progression or death was about 54% of the corresponding rate in the comparator group under the time-to-event analysis. Equivalently, expressed as a simple relative-hazard statement, this corresponds to an estimated 46% lower hazard.
That interpretation does not mean that 46% of participants avoided progression, that median PFS was 46% longer, or that each individual participant experienced a 46% reduction in risk. Hazard is an instantaneous model-based quantity rather than an absolute probability.
The 95% CI of 0.35–0.83 communicates uncertainty around the estimated hazard ratio. Its width matters: the data are compatible with a range of relative hazard reductions rather than one exact underlying effect.
The P-value of 0.00252 is explicitly described in the registry as a one-sided P-value based on the log-rank test. A P-value is evidence against the specified null hypothesis; it is not an effect-size measure and should not be used as a substitute for the HR and CI.
Because the endpoint is time-to-event, interpretation also depends on censoring and on the extent to which a single hazard ratio adequately summarizes the relative event rates over time. The ClinicalTrials.gov record does not report median PFS, Kaplan-Meier event probabilities, or a formal assessment of proportional hazards, so none of those quantities should be inferred here.
10. Results: Overall Survival in Cohorts G+ and G-
The later Cohort G overall survival analysis used the same randomized analysis population and the 19 August 2019 database cutoff. The registry again reports a log-rank test with a one-sided P-value and hazard ratio as the effect measure.
Hazard ratio for overall survival
95% CI: 0.45–1.12 · P = 0.06762
Log-rank test · One-sided P-value based on the log-rank test
| Feature | Reported result |
|---|---|
| Analysis population | All randomized Cohort G participants |
| Groups compared | Cohort G+ vs Cohort G- |
| Endpoint | Overall Survival |
| Time frame | Up to approximately 2 years |
| Outcome unit | Months |
| Method | Log Rank |
| Effect measure | Hazard Ratio (HR) |
| Estimate | 0.71 |
| 95% CI | 0.45–1.12 |
| P-value | 0.06762 |
| Hypothesis type | Superiority |
| Analysis note | One-sided P-value based on log-rank test |
An HR of 0.71 corresponds to an estimated instantaneous rate of death approximately 71% of that in the comparator group under the reported time-to-event analysis. Expressed as a simple relative-hazard interpretation, this is approximately a 29% lower estimated hazard.
The HR does not mean that 29% of participants survived, that survival time increased by 29%, or that every patient experienced the same proportional reduction in mortality hazard.
The 95% CI of 0.45–1.12 is especially important for understanding precision. It spans 1, so the interval includes a hazard ratio corresponding to no relative difference between the groups. The interval therefore conveys substantially more uncertainty about the underlying treatment effect than the point estimate alone.
The reported P-value of 0.06762 is a one-sided log-rank P-value according to the registry. It should not be interpreted as a probability that the treatment has no effect, nor as a measure of effect magnitude. The P-value and confidence interval answer different but complementary questions.
OS is also downstream of the entire treatment pathway and any subsequent care. The ClinicalTrials.gov record does not provide median OS, event counts, subsequent-treatment information, or a formal proportional-hazards assessment, so those features cannot be used to refine the numerical interpretation here.
11. Putting the Four Posted Analyses Together
The four posted statistical analyses form two distinct analytical families. The first two are response analyses: one is a randomized comparison of two Cohort G groups, while the other tests an observed response rate against a fixed 20% benchmark. The latter two are time-to-event analyses comparing the randomized Cohort G groups.
| Endpoint | Comparison | Method | Effect / test | Result |
|---|---|---|---|---|
| ORR | Cohort G+ vs G- | Miettinen & Nurminen | Risk difference | 26.3; 95% CI 8.9–42.1; P=0.0016 |
| ORR | Cohorts D4 & H vs 20% benchmark | Exact binomial | Superiority test | P=0.0858 |
| PFS | Cohort G+ vs G- | Log-rank | HR | 0.54; 95% CI 0.35–0.83; P=0.00252 |
| OS | Cohort G+ vs G- | Log-rank | HR | 0.71; 95% CI 0.45–1.12; P=0.06762 |
A useful statistical lesson is that these results cannot be reduced to a single "trial P-value." They concern different endpoints, populations or hypotheses, and analysis methods. The ORR risk difference is an absolute difference in proportions; the PFS and OS hazard ratios are relative time-to-event measures; and the D4/H analysis is a one-sample comparison with a fixed benchmark.
12. Safety: Serious Adverse Events
The ClinicalTrials.gov record includes serious adverse-event counts for several Part 1 cohorts. These are presented exactly as provided. Because the registry-reported serious-AE field ends partway through the Cohort C1 entry, it does not provide a complete arm-by-arm serious-adverse-event table for all 14 arms.
| Cohort / arm description | Affected / at risk |
|---|---|
| Part 1 Cohort A10 (Pembro 10mg/kg + Pa + C) | 5/12 |
| Part 1 Cohort A2 (Pembro 2 mg/kg + Paclita) | 7/13 |
| Cohort A Second Course (Pembro 2 mg/kg M) | 0/2 |
| Part 1 Cohort B10 (Pembro 10 mg/kg + Pa + C + ...) | 8/13 |
| Part 1 Cohort B2 (Pembro 2mg/kg + Pa + C + Bev) | 8/11 |
| Cohort B Second Course (Pembro 2 mg/kg M) | 0/1 |
| Part 1 Cohort C1 | Registry field registry-reported without a complete affected/at-risk value |
13. Statistical Methods Explained
Why was the Miettinen-Nurminen method used for Cohort G ORR?
ORR is a binary outcome: each participant either meets the response definition or does not. The Cohort G question is therefore naturally expressed as a difference between two proportions. The Miettinen-Nurminen method provides score-based inference for that difference rather than relying on a simple normal approximation that can behave poorly for some binomial settings.
What does a risk difference of 26.3 mean?
A risk difference, when expressed in percentage points, is an absolute difference between two proportions. An estimate of 26.3 means that the observed response percentage differed by 26.3 percentage points between the two Cohort G groups. It is not a relative risk, odds ratio, or hazard ratio.
Why use an exact binomial test for the D4/H analysis?
The D4/H analysis was framed as a comparison of an observed response rate with a fixed benchmark: H0: ORR ≤20% versus H1: ORR >20%. An exact binomial procedure directly evaluates that one-sample proportion hypothesis without requiring a large-sample normal approximation.
What does an HR of 0.54 mean for PFS?
An HR of 0.54 means that the estimated instantaneous rate of progression or death was 0.54 times the corresponding rate in the comparator group under the reported model-based time-to-event analysis. It does not mean that 54% of patients avoided progression or that PFS increased by 54%.
Why can an HR and a P-value tell different parts of the story?
The HR describes the estimated relative treatment effect, while the P-value describes evidence against a null hypothesis under the specified testing procedure. A very small P-value does not imply a large effect, and a point estimate alone does not communicate its statistical precision. The confidence interval connects those two ideas by showing uncertainty around the effect estimate.
Why is the one-sided P-value important?
The registry explicitly identifies the PFS and OS P-values as one-sided log-rank P-values. A one-sided test places the alternative hypothesis in a specified direction. It therefore must be interpreted according to that prespecified direction rather than being treated as interchangeable with a two-sided P-value.
Why should the ORR, PFS, and OS results not be treated as interchangeable?
ORR is a binary response endpoint assessed over a stated time frame. PFS and OS are time-to-event endpoints that incorporate event timing and censoring. Their effect measures therefore answer different questions: risk difference describes an absolute difference in proportions, whereas a hazard ratio describes a relative difference in modeled instantaneous event rates.
14. Confidence Intervals and Precision
The reported confidence intervals illustrate why point estimates should not be interpreted in isolation.
| Endpoint | Point estimate | 95% CI | What the interval describes |
|---|---|---|---|
| Cohort G ORR | 26.3 percentage points | 8.9–42.1 | Uncertainty around the difference in response percentages |
| Cohort G PFS | HR 0.54 | 0.35–0.83 | Uncertainty around the relative event-rate estimate |
| Cohort G OS | HR 0.71 | 0.45–1.12 | Uncertainty around the relative death-rate estimate |
The confidence intervals also illustrate a fundamental difference between the reported PFS and OS analyses. The PFS interval lies below 1, whereas the OS interval extends above 1. That does not turn the analyses into a simple binary "effective versus ineffective" classification; it shows that the precision and statistical evidence differ between the two time-to-event estimates.
Start with the estimand and effect measure, then read the point estimate, confidence interval, and hypothesis test together. Ask what population was analyzed, what time frame was used, what event was measured, and which data cutoff generated the estimate. Those details are part of the statistical result, not footnotes to it.
15. One-Sided Testing and Superiority
The ClinicalTrials.gov record identifies superiority as the hypothesis type for all four posted statistical analyses. The Cohort G PFS and OS analyses specifically state that the reported P-values are one-sided log-rank P-values.
Direction matters
A one-sided test evaluates evidence in a specified direction. For a hazard-ratio comparison, the relevant direction is represented by a hazard ratio below 1 when the treatment group is expected to have fewer or later events.
P-value is not effect size
The P-value answers a hypothesis-testing question. The HR or risk difference answers an effect-estimation question. Both should be reported, along with the confidence interval when available.
The D4/H analysis provides a particularly clear example: its alternative is explicitly ORR >20%. This is a benchmark-based superiority hypothesis, not a randomized comparison between two treatment groups.
16. Multiplicity and Multiple Analyses
The ClinicalTrials.gov record identifies three registered primary endpoints and four posted statistical analyses. They do not provide a complete multiplicity-adjustment strategy, alpha allocation, or endpoint hierarchy for all analyses.
| Issue | What the ClinicalTrials.gov record establishes | Interpretation |
|---|---|---|
| Primary endpoints | 3 registered primary endpoints | The study has more than one primary endpoint, so interpretation should remain endpoint-specific. |
| Statistical analyses | 4 posted analyses | The posted analyses include two primary ORR analyses and two secondary time-to-event analyses. |
| Multiplicity adjustment | Not reported in the ClinicalTrials.gov record | No specific multiplicity procedure should be attributed to the trial from the ClinicalTrials.gov record. |
| One-sided testing | Explicitly identified for PFS and OS | The stated P-values should be interpreted as one-sided log-rank P-values. |
It is therefore preferable to report each analysis with its own endpoint, population, method, estimate, confidence interval, and P-value rather than combining all four P-values into an informal overall claim.
17. Randomization and Analysis Populations
Randomization is central to the interpretation of the Cohort G comparison. The registry states that the PFS and OS analysis populations consisted of all randomized Cohort G participants. The ORR analysis used the same randomized Cohort G principle, with a database cutoff of 08 August 2016.
Using randomized participants for the efficacy comparison preserves the treatment assignment created by randomization. This is conceptually different from the D4/H ORR analysis, which used all treated participants in Cohorts D4 and H and excluded one Cohort H participant from the efficacy analysis population.
18. Time-to-Event Endpoints: What Is Being Estimated?
The registry describes PFS and OS using months as the outcome unit and identifies both as time-to-event endpoints. Time-to-event analysis has two important features that distinguish it from a simple proportion at a fixed point.
Event timing
A participant who experiences an event earlier contributes different information from a participant who remains event-free for a longer period. The analysis therefore uses the timing of events rather than only whether an event eventually occurred.
Censoring
Participants who have not experienced the event by the end of available follow-up can be censored. Kaplan-Meier and related methods are designed to retain their information up to the censoring time.
Hazard ratio
The HR summarizes the relative event-rate experience of the groups under the fitted time-to-event framework. It is not a direct measure of absolute survival probability.
Follow-up window
The registered endpoint time frame is up to approximately 2 years, while the PFS and OS analyses have a database cutoff date of 19 August 2019. The endpoint time frame and database cutoff are distinct pieces of information.
19. Limitations and Interpretation Issues
- Multiple cohorts and arms: the study contains 14 arms, so the overall enrollment of 267.0 should not be interpreted as a single simple two-arm experiment.
- Different analysis populations: the randomized Cohort G analyses and treated D4/H analysis are based on different population definitions.
- Different database cutoffs: the ORR analyses use 08 August 2016 and 07 November 2016 cutoffs, whereas the PFS and OS analyses use 19 August 2019.
- Incomplete numerical reporting: the ClinicalTrials.gov record does not provide every component that would normally accompany a full clinical-trial results table, such as median PFS, median OS, or the numerical ORR estimate for the D4/H test.
- Hazard-ratio interpretation: an HR is a relative time-to-event measure, not an absolute probability or a percentage of patients benefiting.
- Proportional-hazards considerations: a single HR is most naturally interpreted when the relative hazards can reasonably be summarized by a common ratio over time. The ClinicalTrials.gov record does not report a formal assessment of that assumption.
- One-sided testing: the registry specifically identifies the PFS and OS P-values as one-sided. They should not be silently relabeled as two-sided tests.
- Multiplicity: the ClinicalTrials.gov record identifies three primary endpoints and four posted analyses but do not specify a complete multiplicity-control strategy.
- Safety completeness: the registry-reported serious-AE field is incomplete for the 14-arm study, so a complete overall serious-AE comparison cannot be constructed from these data alone.
- Open-label design: the trial is registered with masking listed as NONE. The ClinicalTrials.gov record does not quantify the effect of this feature on any endpoint.
20. Why This Trial Matters Statistically
KEYNOTE-021 is a useful statistical teaching case because the ClinicalTrials.gov record combines randomized comparisons, benchmark-based exact testing, binary response endpoints, time-to-event outcomes, hazard ratios, score-based confidence intervals, and one-sided hypothesis testing within a multi-arm phase 1/2 design.
| Concept | How it appears in KEYNOTE-021 |
|---|---|
| Randomization | The overall study allocation is registered as RANDOMIZED. |
| Parallel design | The design model is registered as PARALLEL. |
| Binary endpoint | ORR and DLT are registered as binary endpoint types. |
| Risk difference | Cohort G ORR is reported as a difference in percentages, normalized to risk difference. |
| Score-based CI | The Miettinen & Nurminen method is used for the Cohort G ORR comparison. |
| Exact binomial testing | The D4/H ORR analysis uses an exact binomial distribution against a 20% benchmark. |
| Kaplan-Meier estimation | PFS and OS are time-to-event endpoints for which Kaplan-Meier estimation is the standard descriptive framework. |
| Log-rank testing | The posted PFS and OS analyses use Log Rank. |
| Hazard ratio | PFS and OS are reported with hazard ratios and confidence intervals. |
| One-sided testing | The registry explicitly describes the PFS and OS P-values as one-sided log-rank P-values. |
| Analysis populations | The registry distinguishes all randomized Cohort G participants from treated D4/H participants. |
| Multiple endpoints | Three primary endpoints are registered. |
21. What the Main Estimates Teach
The Cohort G ORR estimate of 26.3 percentage points is an absolute treatment-group difference. Its 95% CI of 8.9–42.1 communicates uncertainty about that difference. This is the most direct way to read the reported binary endpoint without converting it into a different effect measure.
The PFS HR of 0.54 indicates a lower estimated instantaneous rate of progression or death in the pembrolizumab group relative to the comparator under the reported analysis. The 95% CI of 0.35–0.83 quantifies uncertainty around that estimate.
The OS HR of 0.71 points in the same direction as the PFS HR, but its 95% CI of 0.45–1.12 is wider relative to the null value of 1. The reported one-sided P-value is 0.06762. These are separate statistical facts and should be reported together rather than reduced to a single label.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Statistical Calculators
24. Sources
- ClinicalTrials.gov: NCT02039674 — KEYNOTE-021.
- PubMed: PMID 27745820.
- PubMed: PMID 30138764.
- PubMed: PMID 37465924.
- PubMed: PMID 30529597.
- PubMed: PMID 30429032.
Continue through the Clinical Biostats statistical pathway
Explore the underlying clinical-trial methods through focused tutorials and statistical calculators.
25. Record Summary
KEYNOTE-021 provides a compact example of how several statistical frameworks coexist within a multi-arm clinical trial. The registered primary endpoints include two objective-response analyses and a dose-limiting-toxicity endpoint, while the posted statistical analyses also include progression-free survival and overall survival in randomized Cohort G participants. The principal methods represented in the ClinicalTrials.gov record is the Miettinen-Nurminen score-based approach for a difference in response percentages, exact binomial testing against a fixed ORR benchmark, and log-rank testing with hazard ratios for time-to-event outcomes.
The most important interpretive discipline is to preserve the distinction between effect measure, analysis population, endpoint, and hypothesis. The Cohort G ORR estimate of 26.3 is a risk difference, not a hazard ratio. The PFS HR of 0.54 and OS HR of 0.71 are time-to-event measures, not response probabilities. The D4/H P-value of 0.0858 comes from a one-sample superiority test against an ORR benchmark of 20%, not from a randomized treatment comparison.
The confidence intervals add another layer of information. The ORR difference has a 95% CI of 8.9–42.1, the PFS HR has a 95% CI of 0.35–0.83, and the OS HR has a 95% CI of 0.45–1.12. Reading these intervals alongside the corresponding estimates and P-values gives a more complete picture of statistical uncertainty than any single number can provide.