This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the values reported in the ClinicalTrials.gov record.
1. Trial at a Glance
KEYNOTE-024 was a randomized, parallel, open-label phase 3 oncology trial with 305 participants. The registered primary endpoint was progression-free survival (PFS) rate at Month 6, with the primary analysis comparing pembrolizumab with standard-of-care (SOC) chemotherapy using a stratified Cox proportional-hazards model.
| Feature | KEYNOTE-024 |
|---|---|
| Trial name | KEYNOTE-024 |
| NCT identifier | NCT02142738 |
| Phase | Phase 3 |
| Status | COMPLETED |
| Therapeutic area | Oncology |
| Condition | Non-Small Cell Lung Carcinoma |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 305 |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | Industry |
| Start | 2014-08-25 |
| Primary completion | 2016-05-09 |
| Results posted | Yes |
| Registered primary endpoint | Progression Free Survival (PFS) Rate at Month 6 |
2. Clinical Question
The central statistical question was whether participants randomized to pembrolizumab had a different progression-free survival experience than participants randomized to SOC chemotherapy, as evaluated at the registered Month 6 PFS endpoint.
Population
Participants with metastatic non-small cell lung cancer, as described in the trial's brief title and registered condition.
Intervention
Pembrolizumab.
Comparator
Standard-of-care chemotherapy. The registered interventions include paclitaxel, carboplatin, pemetrexed, cisplatin, and gemcitabine.
Primary question
Does pembrolizumab produce a different time-to-event outcome than SOC chemotherapy for the registered Month 6 PFS endpoint?
3. Trial Design
The registry describes a randomized, parallel, unmasked phase 3 treatment trial. The enrollment was 305 participants, and six trial arms were registered. The interventions listed in the registry were pembrolizumab, paclitaxel, carboplatin, pemetrexed, cisplatin, and gemcitabine.
Pembrolizumab group
- Intervention: pembrolizumab
- Primary comparison: pembrolizumab vs SOC chemotherapy
- Primary analysis population: intention-to-treat
- Time-to-event effect measure: hazard ratio
Standard-of-care chemotherapy group
- Comparator: SOC chemotherapy
- Registered chemotherapy interventions include paclitaxel, carboplatin, pemetrexed, cisplatin, and gemcitabine
- Primary comparison: pembrolizumab vs SOC chemotherapy
- Time-to-event effect measure: hazard ratio
4. Endpoints
The registry data contain three posted outcome measures and three statistical analyses. One was the registered primary endpoint; overall survival and objective response rate were secondary analyses.
| Role | Endpoint | Time frame | Type | Effect measure |
|---|---|---|---|---|
| Primary | Progression Free Survival (PFS) Rate at Month 6 | Month 6 | Time-to-event | Hazard ratio |
| Secondary | Overall Survival (OS) Rate | 12 months | Time-to-event | Hazard ratio |
| Secondary | Objective Response Rate (ORR) | Up to ~1.6 years | Binary | Risk difference |
Primary endpoint definition
PFS was defined as the time from randomization to documented disease progression per Response Evaluation Criteria in Solid Tumors version 1.1 (RECIST 1.1) or death due to any cause, whichever occurred first, and was based on blinded independent central radiologists' (BICR) review.
5. Analysis Populations and Stratification
The posted analyses use the intention-to-treat (ITT) population. The registry description states that all randomized participants were included and analyzed according to the treatment group to which they were randomized, regardless of whether or not they received study treatment.
| Analysis population | Role |
|---|---|
| Intention-to-treat | All randomized participants; participants remain in the group to which they were randomized regardless of whether or not they received study treatment. |
Stratification factors
The primary and secondary Cox analyses treated treatment as a covariate and were stratified by:
- Geographic region: East Asia vs. non-East Asia
- ECOG performance status: 0 vs. 1
- Histology: squamous vs. nonsquamous
Stratification allows the baseline hazard to vary across the specified strata while estimating the treatment effect within the Cox modeling framework.
6. Statistical Methodology
Cox proportional-hazards model
The primary PFS analysis used a Cox proportional-hazards model. Treatment was entered as a covariate, with stratification by geographic region, ECOG performance status, and histology.
The hazard ratio summarizes the relative instantaneous event rate associated with pembrolizumab compared with SOC chemotherapy under the fitted model. A value below 1 indicates a lower estimated hazard in the pembrolizumab group.
An HR of 0.50 corresponds to an estimated hazard that is 50% lower in the pembrolizumab group under the fitted model. It does not mean that exactly 50% of participants avoided progression or death.
Score-based confidence intervals for proportions
The objective response rate analysis used the Miettinen & Nurminen method. The Miettinen-Nurminen method is a score-based approach to confidence intervals for differences in proportions, related to the Newcombe and Wilson methods.
The reported effect measure was the difference in percentages, expressed here as a risk difference. The hypothesis was explicitly directional: the null hypothesis was a difference in percentages of 0, versus an alternative in which the difference was greater than 0.
Intention-to-treat analysis
The ITT principle preserves the randomized treatment comparison by keeping each participant in the treatment group assigned at randomization. This matters because treatment receipt, discontinuation, and other events occurring after randomization can otherwise create differences between groups that are no longer protected by the original randomization.
Stratified analysis
The Cox analyses were stratified by geographic region, ECOG performance status, and histology. This means the treatment comparison was estimated while allowing the underlying event pattern to differ across those predefined strata rather than assuming one common baseline hazard for every participant.
7. Primary Result: Progression-Free Survival at Month 6
The registered primary endpoint was analyzed in the ITT population using a Cox proportional-hazards model. Pembrolizumab was compared with SOC chemotherapy, with treatment as a covariate and stratification by geographic region, ECOG PS, and histology.
Primary PFS hazard ratio
95% CI: 0.37–0.68 · P < 0.001 · Two-sided CI
| Primary endpoint | Estimate | 95% CI | P-value | Hypothesis |
|---|---|---|---|---|
| PFS Rate at Month 6 | HR 0.50 | 0.37–0.68 | <0.001 | Superiority |
The estimated hazard ratio of 0.50 means that, under the fitted Cox model, the estimated instantaneous rate of the PFS event was 50% lower in the pembrolizumab group than in the SOC chemotherapy group. Because PFS events were defined as documented progression or death, the HR concerns the modeled rate of that composite time-to-event outcome.
The HR does not mean that 50% of participants were progression-free, that 50% were cured, or that every participant experienced exactly a 50% reduction in risk. A hazard ratio is a relative, model-based time-to-event measure rather than an absolute probability.
The 95% CI of 0.37–0.68 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of effects experienced by individual participants. The interval also remains below 1, which is consistent with the direction of the reported superiority result.
The P-value < 0.001 quantifies evidence against the specified null hypothesis under the statistical testing framework. It does not measure the size of the treatment effect and should not be interpreted as the probability that the null hypothesis is true.
Because the estimate comes from a Cox proportional-hazards model, interpretation of a single HR also depends on the proportional-hazards framework being a useful summary of the relative event rates over the analyzed period. The registry result itself does not provide a separate diagnostic assessment of that assumption.
What the primary result establishes statistically
The registry reports a superiority analysis with an HR of 0.50, a two-sided 95% CI of 0.37–0.68, and a P-value below 0.001. These are the quantities needed to describe the reported treatment-effect estimate, its statistical precision, and the hypothesis-test result.
What the primary result does not provide
The ClinicalTrials.gov record does not report a median PFS, a Kaplan-Meier curve, a six-month percentage for each treatment group, or subgroup-specific PFS estimates. Those quantities therefore are not reconstructed here.
8. Secondary Result: Overall Survival at 12 Months
Overall survival was a secondary time-to-event endpoint with a 12-month time frame. It was analyzed in the ITT population using a Cox proportional-hazards model with the same reported stratification factors: geographic region, ECOG PS, and histology.
Overall survival hazard ratio
95% CI: 0.47–0.86 · P = 0.002 · Two-sided CI
| Secondary endpoint | Estimate | 95% CI | P-value | Hypothesis |
|---|---|---|---|---|
| Overall Survival (OS) Rate at 12 months | HR 0.63 | 0.47–0.86 | 0.002 | Superiority |
An HR of 0.63 means that, under the fitted Cox model, the estimated instantaneous rate of death was 37% lower in the pembrolizumab group than in the SOC chemotherapy group.
This is a relative time-to-event effect. It does not mean that 37% of participants survived, that 37% of deaths were prevented, or that every individual participant experienced the same relative reduction in risk.
The 95% CI of 0.47–0.86 represents uncertainty around the estimated HR. Its endpoints describe uncertainty in the model-based treatment-effect estimate, not variability in individual patient outcomes.
The P-value of 0.002 measures the strength of evidence against the specified null hypothesis under the analysis framework. It is not a measure of clinical effect size. The effect size is better represented by the HR and its confidence interval, together with absolute survival quantities when those are available.
The analysis was secondary rather than the registered primary endpoint, so its interpretation should remain tied to its stated role in the trial's endpoint structure. The ClinicalTrials.gov record does not provide an additional multiplicity-adjustment procedure for this secondary analysis.
9. Secondary Result: Objective Response Rate
Objective response rate was a binary secondary endpoint with a time frame of up to approximately 1.6 years. The analysis used the ITT population and compared pembrolizumab with SOC chemotherapy.
Difference in objective response rates
95% CI: 6.0–27.0 · P = 0.0011 · Two-sided CI
| Secondary endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Objective Response Rate (ORR) | Risk difference | 16.6 | 6.0–27.0 | 0.0011 |
The registry reports the Miettinen & Nurminen method and states that the analysis was stratified by geographic region, ECOG PS, and histology. The hypothesis was H0: difference in percentages = 0 versus H1: difference in percentages > 0.
The estimated risk difference of 16.6 percentage points represents the reported difference in response percentages between the two randomized groups, with the direction defined by the analysis.
This is an absolute difference in proportions, not a hazard ratio. It does not describe how quickly responses occurred, how long they lasted, or whether the same participants contributed to later survival outcomes.
The 95% CI of 6.0–27.0 quantifies uncertainty around the estimated difference in response percentages. It does not indicate the range of response rates among individual participants.
The P-value of 0.0011 measures evidence against the specified null hypothesis of no difference in percentages. It does not measure the magnitude or clinical importance of the 16.6-point difference.
Because the analysis was stratified, the reported estimate should be understood in the context of the prespecified geographic-region, ECOG PS, and histology strata rather than as an unqualified unstratified difference.
10. Comparing the Three Reported Effect Measures
| Endpoint | Data type | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| PFS Rate at Month 6 | Time-to-event | Hazard ratio | 0.50 | 0.37–0.68 | <0.001 |
| OS Rate at 12 months | Time-to-event | Hazard ratio | 0.63 | 0.47–0.86 | 0.002 |
| Objective Response Rate | Binary | Risk difference | 16.6 | 6.0–27.0 | 0.0011 |
These three results should not be collapsed into a single statistic. PFS and OS are time-to-event outcomes and therefore account for follow-up and censoring through survival-analysis methods. ORR is a binary response endpoint and is summarized through a difference in percentages. The estimates answer different statistical questions.
Relative time-to-event effect
The PFS and OS hazard ratios describe relative event rates under Cox models. HR 0.50 and HR 0.63 should not be interpreted as percentages of patients with benefit.
Absolute response difference
The ORR estimate of 16.6 is expressed as a difference in percentages. It provides an absolute contrast rather than a relative hazard measure.
Precision
Each confidence interval describes uncertainty around its corresponding effect measure. The interval must be interpreted on the scale of that effect measure.
Evidence from testing
The P-values describe evidence against the respective null hypotheses. They are not interchangeable measures of effect magnitude.
11. Safety: Serious Adverse Events
The ClinicalTrials.gov record reports serious adverse events by affected participants divided by participants at risk for several treatment-course categories.
| Safety category | Serious adverse events affected / at risk |
|---|---|
| Pembrolizumab First Course | 79 / 154 |
| SOC Chemotherapy First Course | 70 / 150 |
| SOC Chemotherapy Switched Over to Pembro | 27 / 83 |
| Pembrolizumab Second Course | 2 / 12 |
| SOC Switched Over to Pembrolizumab Secon | 0 / 1 |
The safety data also illustrate why denominators matter. For example, 2 affected participants among 12 at risk represents a different evidentiary context from 0 among 1, even though both numerators are small. The ClinicalTrials.gov record does not provide confidence intervals or formal hypothesis tests for these serious-adverse-event categories.
12. Statistical Methods Explained
Why was a Cox proportional-hazards model used?
PFS and OS are time-to-event endpoints. Participants can experience the event at different times, and some observations can be censored rather than observed through the event. A Cox model provides a way to compare the instantaneous event rates between treatment groups while retaining the time-to-event structure.
What does an HR of 0.50 mean?
An HR of 0.50 means that the estimated instantaneous event rate under the fitted model is half as large in the pembrolizumab group as in the SOC chemotherapy group. It does not mean that half of participants avoided an event or that the probability of an event was reduced by exactly 50% at every time point.
Why was stratification used?
The analyses were stratified by geographic region, ECOG PS, and histology. Stratification allows the baseline event hazard to differ across those strata while focusing the treatment comparison within the Cox modeling framework. It can therefore account for important design factors without forcing their baseline hazards to be identical.
Why is a confidence interval more informative than a P-value alone?
A P-value addresses evidence against a null hypothesis, whereas a confidence interval describes uncertainty around the estimated effect. For the primary PFS analysis, the HR of 0.50 is accompanied by a 95% CI of 0.37–0.68, which provides substantially more information about the estimated effect than the P-value < 0.001 alone.
Why is ORR analyzed differently from PFS?
ORR is a binary endpoint, so each participant contributes a response or non-response classification for the relevant analysis. The registry reports the Miettinen & Nurminen method and a risk difference. PFS, by contrast, is explicitly a time-to-event endpoint and was analyzed with a Cox proportional-hazards model.
What does intention-to-treat mean here?
The ITT population included all randomized participants, with participants retained in the treatment group to which they were randomized regardless of whether they received study treatment. This preserves the randomized comparison as the basis for the efficacy analyses.
Why should the PFS and OS hazard ratios not be treated as the same quantity?
Although both are hazard ratios from Cox models, they describe different event processes. PFS concerns progression or death, whichever occurs first, while OS concerns death. Their HRs therefore summarize different endpoints and should be interpreted separately.
13. Confidence Intervals and Statistical Precision
The three posted analyses illustrate two different confidence-interval scales.
| Endpoint | Effect scale | Estimate | 95% CI |
|---|---|---|---|
| PFS | Hazard ratio | 0.50 | 0.37–0.68 |
| OS | Hazard ratio | 0.63 | 0.47–0.86 |
| ORR | Risk difference | 16.6 | 6.0–27.0 |
The PFS interval is centered on a hazard-ratio scale, whereas the ORR interval is on the percentage-point difference scale. The intervals therefore cannot be compared simply by their numerical widths. Precision must always be considered relative to the scale and meaning of the corresponding effect measure.
The interval 0.37–0.68 gives a range of plausible values for the underlying model-based hazard ratio under the stated confidence framework. The interval does not say that individual patients have hazards somewhere between 0.37 and 0.68.
The interval 6.0–27.0 is expressed in percentage points around the estimated risk difference of 16.6. It concerns the uncertainty in the population-level difference in response percentages, not the response probability of an individual participant.
14. P-values and Superiority Testing
The analyses posted on ClinicalTrials.gov identify all three treatment comparisons as superiority hypotheses.
| Endpoint | Null comparison | P-value | Hypothesis type |
|---|---|---|---|
| PFS | Hazard ratio comparison | <0.001 | Superiority |
| OS | Hazard ratio comparison | 0.002 | Superiority |
| ORR | Difference in percentages = 0 versus difference > 0 | 0.0011 | Superiority |
A superiority P-value evaluates evidence against the specified null model. It should not be converted into an effect-size ranking. The effect size comes from the estimated HR or risk difference, while the confidence interval communicates statistical precision.
15. Censoring and Time-to-Event Interpretation
The primary PFS endpoint is a time-to-event measure: time from randomization until documented progression or death, whichever occurs first. This structure means that participants who have not yet experienced the event at the relevant observation point can contribute partial follow-up information rather than simply being classified as event-free for all time.
The Cox model uses the observed event and follow-up information to estimate the relative hazard between the randomized treatment groups.
The ClinicalTrials.gov record does not report the number censored, censoring distributions, median follow-up, or a Kaplan-Meier curve. Those details are therefore not used to reconstruct additional survival statistics on this page.
16. Multiplicity, Interim Analysis, and Other Design Topics
The ClinicalTrials.gov record identifies superiority hypotheses, the primary and secondary endpoint roles, and the three posted statistical analyses. They do not provide a multiplicity-adjustment procedure, alpha-spending method, interim-analysis boundary, formal power calculation, non-inferiority margin, missing-data imputation method, or Bayesian analysis.
| Design topic | What the ClinicalTrials.gov record supports |
|---|---|
| Multiplicity | Primary and secondary endpoint roles are identified; no multiplicity-adjustment procedure is reported in the ClinicalTrials.gov record. |
| Interim analysis | No interim-analysis boundary or alpha-spending method is reported in the ClinicalTrials.gov record. |
| Non-inferiority | Not applicable to the reported superiority hypotheses; no non-inferiority margin is reported. |
| Factorial design | No factorial design is reported. |
| Bayesian methods | No Bayesian method is reported. |
| Missing-data imputation | No imputation method is reported in the registry-reported statistical-analysis fields. |
17. Why the ITT Population Matters
The ITT definition is particularly important for interpreting the efficacy results because it anchors the comparison to randomization. Participants remain associated with their assigned group even if they do not receive study treatment.
Preserves randomization
Keeping randomized participants in their assigned groups protects the treatment comparison created by random allocation.
Avoids treatment-received selection
Analyzing participants only according to treatment actually received can introduce post-randomization selection into the efficacy comparison.
Applies to posted efficacy analyses
The registry-reported PFS, OS, and ORR analyses all specify the ITT population.
Does not eliminate all limitations
ITT analysis does not by itself solve censoring, model-assumption, endpoint-definition, or multiplicity issues.
18. Understanding the Stratified Cox Analysis
The registry specifies the same three stratification factors for the PFS and OS analyses and for the ORR analysis: geographic region, ECOG PS, and histology.
| Stratification factor | Categories |
|---|---|
| Geographic region | East Asia vs. non-East Asia |
| ECOG PS | 0 vs. 1 |
| Histology | Squamous vs. nonsquamous |
In a stratified Cox analysis, the treatment coefficient is interpreted across the predefined strata while the baseline hazard is allowed to differ between strata. This is different from simply inserting the stratification variables as ordinary covariates and assuming a common baseline hazard.
The registry-reported analysis explicitly describes treatment as a covariate and geographic region, ECOG PS, and histology as stratification factors.
19. Limitations
- Registry-level reporting: the ClinicalTrials.gov record provides the principal posted estimates and methods but not the complete statistical analysis plan.
- Limited endpoint detail: the primary PFS definition is reported, but the complete RECIST 1.1 progression criteria are not reproduced in the available trial-data text.
- No median survival estimates: median PFS and median OS are not provided in the ClinicalTrials.gov record.
- No Kaplan-Meier curves: the ClinicalTrials.gov record does not contain the underlying event and censoring information needed to reconstruct survival curves.
- No subgroup estimates: although the models were stratified by geographic region, ECOG PS, and histology, subgroup-specific treatment-effect estimates are not reported.
- No baseline table: baseline demographic and clinical characteristics are not provided in the ClinicalTrials.gov record.
- No formal multiplicity procedure: the ClinicalTrials.gov record identifies primary and secondary analyses but do not report a multiplicity-adjustment strategy.
- Safety denominators differ: serious-adverse-event categories include first-course and switched-over treatment categories with different denominators, so they should not be collapsed into a single two-arm rate.
- Cox-model interpretation: a hazard ratio is a model-based relative measure and should not automatically be translated into an absolute risk reduction.
- Secondary endpoints: OS and ORR are secondary analyses and should be interpreted according to their stated endpoint roles rather than as additional primary endpoints.
20. Why This Trial Matters Statistically
KEYNOTE-024 is a useful statistical teaching example because the ClinicalTrials.gov record connect randomized treatment allocation with several different types of estimands and analysis methods.
| Concept | How it appears in KEYNOTE-024 |
|---|---|
| Randomization | The study is registered as randomized with a parallel design. |
| Intention-to-treat analysis | The PFS, OS, and ORR analyses use the ITT population. |
| Time-to-event analysis | PFS and OS are analyzed as time-to-event endpoints. |
| Cox proportional-hazards model | Used for the PFS and OS treatment comparisons. |
| Hazard ratio | Used to summarize the relative PFS and OS treatment effects. |
| Stratified analysis | Geographic region, ECOG PS, and histology are used as stratification factors. |
| Confidence intervals | 95% two-sided intervals accompany the reported HRs and ORR risk difference. |
| P-values | Reported for all three statistical analyses. |
| Binary endpoint analysis | ORR is analyzed using the Miettinen & Nurminen method. |
| Risk difference | ORR is summarized as a difference in percentages. |
| Different estimands | PFS, OS, and ORR quantify different aspects of treatment effect. |
| Safety denominators | Serious adverse events are reported across treatment-course categories with different denominators. |
21. Statistical Interpretation: Relative Effects Versus Absolute Effects
One of the most important lessons from this trial is that the reported HRs and the ORR risk difference live on different statistical scales.
The PFS HR of 0.50 and OS HR of 0.63 are relative time-to-event measures. They summarize treatment differences in event hazards under the Cox models.
The ORR risk difference of 16.6 is an absolute difference in response percentages. It is not a hazard ratio and should not be interpreted through the same language used for PFS or OS.
Relative measures and absolute measures answer different questions. A statistically complete interpretation therefore preserves the original effect measure rather than converting every result into a single common scale.
22. What the Primary Hazard Ratio Does — and Does Not — Mean
The primary PFS HR of 0.50 indicates that the estimated instantaneous rate of progression or death was approximately half as large in the pembrolizumab group as in the SOC chemotherapy group under the fitted Cox model.
It does not mean that exactly half of participants were progression-free, that half of participants benefited, or that each participant experienced a 50% reduction in risk.
The 95% CI of 0.37–0.68 describes uncertainty around the estimated hazard ratio. The interval does not describe the range of individual treatment responses.
The P-value of <0.001 describes the evidence against the relevant null hypothesis under the specified testing framework. It is not a substitute for the estimated HR or its confidence interval.
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Calculators
25. Sources
- ClinicalTrials.gov: NCT02142738, the official trial registry record.
- PubMed: PMID 27718847.
- PubMed: PMID 35073727.
- PubMed: PMID 34543477.
- PubMed: PMID 33872070.
- PubMed: PMID 32926507.
Continue through the Clinical Biostats statistical learning pathway
Explore tutorials and calculators covering the survival-analysis, regression, confidence-interval, and clinical-trial methods represented in KEYNOTE-024.
26. Record Summary
KEYNOTE-024 provides a compact example of how a randomized clinical trial can combine a time-to-event primary endpoint with secondary survival and binary-response analyses. The primary PFS analysis used an ITT population and a stratified Cox proportional-hazards model, producing an HR of 0.50 with a two-sided 95% CI of 0.37–0.68 and a P-value of <0.001. The secondary OS analysis reported an HR of 0.63 with a 95% CI of 0.47–0.86 and a P-value of 0.002, while ORR was summarized with a risk difference of 16.6 and a 95% CI of 6.0–27.0.
The statistical lesson is not simply that the reported P-values are small. The more important lesson is how the effect measures match the endpoint types: hazard ratios for time-to-event outcomes and a risk difference for the binary response outcome. Confidence intervals provide precision around those estimates, while the ITT population and prespecified stratification factors define the framework in which the randomized comparisons were made.