This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the trial data reported in that registry record.
1. Trial at a Glance
KEYNOTE-598 was a completed, randomized, parallel-group phase 3 trial evaluating pembrolizumab given with ipilimumab versus pembrolizumab given with placebo in participants with untreated metastatic non-small cell lung cancer. The registry reports 568 enrolled participants, two treatment arms, quadruple masking, and two primary time-to-event endpoints.
| Feature | KEYNOTE-598 |
|---|---|
| Phase | Phase 3 |
| Condition | Carcinoma, Non-Small-Cell Lung |
| Design | Randomized, parallel |
| Masking | Quadruple |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 568 |
| Primary endpoints | Overall Survival and Progression-free Survival per RECIST 1.1 based on BICR |
| Study start | 14 Dec 2017 |
| Primary completion | 01 Sep 2020 |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT03302234 |
2. Clinical Question
The central randomized comparison was between pembrolizumab plus ipilimumab and pembrolizumab plus placebo in participants with untreated metastatic non-small cell lung cancer. The two primary questions were framed through overall survival and progression-free survival, both measured as time-to-event outcomes.
Population
Participants with untreated metastatic non-small cell lung cancer.
Intervention
Pembrolizumab plus ipilimumab.
Comparator
Pembrolizumab plus placebo.
Primary question
How do the randomized treatment groups compare with respect to overall survival and progression-free survival?
3. Trial Design
Pembrolizumab + Ipilimumab
- Pembrolizumab
- Ipilimumab
- First Course serious adverse events reported as 146/282
Pembrolizumab + Placebo
- Pembrolizumab
- Placebo
- First Course serious adverse events reported as 114/281
- Second Course serious adverse events reported as 4/18
The registry identifies the allocation as randomized, the design model as parallel, and the masking as quadruple. The primary purpose is treatment. These design features matter statistically because randomization establishes the principal basis for comparing outcomes between treatment assignments, while masking is intended to reduce opportunities for treatment knowledge to influence trial conduct and outcome assessment.
4. Endpoints
The registry reports two primary time-to-event endpoints and three additional statistical analyses for secondary outcomes. The endpoint definitions below preserve the registry's terminology while explaining the statistical structure of each measure.
| Endpoint | Registry definition / time frame | Statistical structure |
|---|---|---|
| Overall Survival (OS) | OS was defined as the time from randomization to death due to any cause. Participants without documented death at the time of analysis were censored at the date of last known contact. Time frame: up to approximately 32 months, through data cut-off date 01 Sep 2020. | Time-to-event; Kaplan-Meier; Cox regression; stratified log-rank test. |
| Progression-free Survival (PFS) Per RECIST 1.1 Based on BICR | PFS was defined as the time from randomization to the first documented disease progression per RECIST 1.1 based on blinded independent central review or death due to any cause, whichever occurs first. Time frame: up to approximately 32 months, through data cut-off date 01 Sep 2020. | Time-to-event; stratified log-rank test; Cox regression. |
| Objective Response Rate (ORR) Per RECIST 1.1 Based on BICR | Percentage of participants with an objective response per RECIST 1.1 based on BICR. Time frame: up to approximately 32 months, data cut-off date 01 Sep 2020. | Binary outcome; score-based confidence interval for proportions; risk difference. |
| Time to True Deterioration (TTD) | Time to true deterioration in cough, pain in chest, and shortness of breath. Time frame: up to approximately 32 months, data cut-off date 01 Sep 2020. | Time-to-event; stratified log-rank test; Cox regression. |
| EORTC QLQ-C30 Global Health Status/Quality of Life | Change from baseline in EORTC QLQ-C30 Global Health Status/Quality of Life, Items 29 and 30, to Week 18. Time points: baseline and Week 18. | Continuous longitudinal outcome; constrained longitudinal data analysis. |
Analysis populations
| Outcome | Analysis population |
|---|---|
| Overall Survival | All randomized participants. |
| Progression-free Survival | All randomized participants. |
| Objective Response Rate | All randomized participants. |
| Time to True Deterioration | All participants randomized who received at least one dose of study treatment and had at least one EORTC QLQ-LC13 and EORTC QLQ-C30 available. |
| QLQ-C30 Global Health Status/Quality of Life | All randomized participants who received at least one dose of study treatment and had at least one EORTC QLQ-C30 assessment available. |
5. Primary Results
Overall Survival
The primary overall-survival analysis compared pembrolizumab plus ipilimumab with pembrolizumab plus placebo among all randomized participants. The reported method was a log-rank test, with the hazard ratio based on a Cox regression model using Efron's method of tie handling and treatment as a covariate, stratified by ECOG, geographic region of the enrolling site, and predominant tumor history.
Hazard ratio for death
95% CI: 0.85–1.37 · P = 0.74156
Pembrolizumab + ipilimumab vs pembrolizumab + placebo
| Primary endpoint | Estimate | 95% CI | P-value | Analysis population |
|---|---|---|---|---|
| Overall Survival | HR 1.08 | 0.85–1.37 | 0.74156 | All randomized participants |
The estimated OS hazard ratio of 1.08 means that the fitted model estimated the instantaneous rate of death in the pembrolizumab-plus-ipilimumab group at 1.08 times the corresponding rate in the pembrolizumab-plus-placebo group over the analyzed follow-up. Expressed descriptively, the point estimate is above 1 rather than below 1.
The hazard ratio is not a statement that 8% more participants died, nor does it mean that every participant experienced an 8% higher risk. A hazard ratio is a model-based relative measure of event rates over time, not an absolute risk difference or a probability that an individual patient will experience an event.
The 95% confidence interval of 0.85–1.37 describes the statistical uncertainty around the estimated hazard ratio. Because the interval spans 1, the reported interval includes both values below and above the no-difference value. The interval should be interpreted as uncertainty about the estimated treatment effect under the specified analysis framework, not as a range containing the effects experienced by individual participants.
The p-value of 0.74156 describes the compatibility of the observed data with the statistical testing framework; it does not measure the size or clinical importance of the effect. A p-value should therefore be read alongside the hazard ratio and confidence interval rather than substituted for them.
The analysis was stratified by ECOG, geographic region of the enrolling site, and predominant tumor history, and the Cox model used Efron's method for tied event times. As with other Cox-model hazard ratios, interpretation depends on the model and the time-to-event structure, including the relevance of the proportional-hazards assumption.
Progression-free Survival
The second primary endpoint was progression-free survival per RECIST 1.1 based on blinded independent central review. The analysis compared the randomized groups using a stratified log-rank test. The registry reports a one-sided p-value for the log-rank test and a two-sided 95% confidence interval for the hazard ratio. The Cox model used Efron's method for tied event times and stratified by ECOG, geographic region of the enrolling site, and predominant tumor history.
Hazard ratio for progression or death
95% CI: 0.86–1.30 · P = 0.71720
Pembrolizumab + ipilimumab vs pembrolizumab + placebo
| Primary endpoint | Estimate | 95% CI | P-value | Analysis population |
|---|---|---|---|---|
| Progression-free Survival | HR 1.06 | 0.86–1.30 | 0.71720 | All randomized participants |
The PFS hazard ratio of 1.06 is the estimated relative event rate for disease progression or death in the pembrolizumab-plus-ipilimumab group compared with the pembrolizumab-plus-placebo group. The point estimate is close to 1, so the estimated relative difference is correspondingly close to the model's no-difference value.
The 95% confidence interval of 0.86–1.30 spans 1. This indicates that the reported estimate is accompanied by uncertainty extending on both sides of the no-difference value. It does not establish that every possible treatment effect in that interval is equally plausible, nor does it describe individual patient outcomes.
The reported p-value of 0.71720 is a hypothesis-testing quantity, not an effect-size measure. It should not be interpreted as the probability that the null hypothesis is true, or as the probability that one treatment is effective or ineffective.
The registry specifically identifies a one-sided p-value for the stratified log-rank test while reporting a two-sided 95% confidence interval. These are different aspects of the statistical presentation and should not be silently treated as though they used the same tail convention.
Because PFS combines disease progression and death into a time-to-event endpoint, the result also depends on censoring rules and the handling of progression assessments. The BICR framework provides a centralized assessment basis, while the Cox-model interpretation remains subject to the assumptions of the model.
6. Secondary Endpoint Results
Objective Response Rate
Objective response rate was analyzed as a binary endpoint among all randomized participants. The reported method was the Miettinen & Nurminen method, categorized here as a score-based confidence-interval approach for proportions. The effect measure was the difference in percentage, represented as a risk difference.
Difference in objective response rate
95% CI: −8.2 to 8.1 · P = 0.50644
Percentage-point difference: pembrolizumab + ipilimumab minus pembrolizumab + placebo
| Secondary endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Objective Response Rate per RECIST 1.1 based on BICR | Risk difference | −0.1 | −8.2 to 8.1 | 0.50644 |
The estimated risk difference of −0.1 represents the difference in percentage between the two randomized groups as reported by the registry. The sign indicates the direction of the estimate under the specified ordering of the comparison; it is not a hazard ratio and should not be interpreted as a relative risk.
The 95% confidence interval of −8.2 to 8.1 spans zero, the no-difference value for a risk difference. The interval therefore conveys uncertainty in the estimated percentage-point difference and includes effects in either direction.
The p-value of 0.50644 is evidence from the specified testing framework, not a measure of how large or clinically meaningful the observed percentage difference is. For a binary endpoint, the risk difference itself and its confidence interval are more directly informative about the magnitude and precision of the group difference.
The registry notes that the p-value was one-sided and based on the Miettinen & Nurminen method stratified by ECOG, geographic region of the enrolling site, and predominant tumor history, while the reported confidence interval was two-sided. That distinction is important when interpreting the inferential framework.
Time to True Deterioration
The registry also reports a time-to-event analysis for time to true deterioration in cough, pain in chest, and shortness of breath. The analysis population was restricted to randomized participants who received at least one dose of study treatment and had at least one EORTC QLQ-LC13 and EORTC QLQ-C30 available.
| Secondary endpoint | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| Time to True Deterioration in cough, pain in chest, and shortness of breath | HR 0.9815 | 0.7386–1.3042 | 0.9112 | Log-rank test; Cox regression |
The registry specifies a two-sided p-value based on a log-rank test stratified by ECOG, geographic region of the enrolling site, and predominant tumor histology. The Cox regression model used Efron's method of tie handling with treatment as a covariate and the same stratification factors.
The hazard ratio of 0.9815 is close to 1. Under the fitted time-to-event model, the estimated instantaneous rate of true deterioration was therefore close to the corresponding rate in the comparison group.
The 95% confidence interval of 0.7386–1.3042 spans 1, so the interval includes values representing lower and higher estimated event rates. The width of the interval is also a reminder that a point estimate alone does not describe the precision of a time-to-event comparison.
The p-value of 0.9112 is not an effect-size measure and should not be interpreted as the probability that the treatments are identical. The HR and confidence interval provide the more direct description of the estimated magnitude and uncertainty.
Change From Baseline in EORTC QLQ-C30 Global Health Status/Quality of Life to Week 18
The registry analyzed change from baseline in the EORTC QLQ-C30 Global Health Status/Quality of Life scale score at Week 18 using a constrained longitudinal data analysis model. The response variable was the PRO score, with treatment-by-time interaction and the stratification factors ECOG, geographic region of the enrolling site, and predominant tumor histology included as covariates.
Difference in least-squares means
95% CI: −3.96 to 3.12 · P = 0.8151
Change from baseline to Week 18
| Secondary endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Change from baseline in EORTC QLQ-C30 Global Health Status/Quality of Life to Week 18 | Difference in LS Means | −0.42 | −3.96 to 3.12 | 0.8151 |
The estimated difference in least-squares means was −0.42. This is an adjusted model-based difference between treatment groups for the specified longitudinal outcome and time point; it is not a hazard ratio or a raw unadjusted difference in observed scores.
The 95% confidence interval of −3.96 to 3.12 spans zero, the no-difference value for a mean difference. The interval therefore includes both a negative and a positive adjusted difference.
The p-value of 0.8151 does not measure the magnitude of the quality-of-life difference. The appropriate effect-size information is the estimated difference together with its confidence interval.
The use of cLDA is important because repeated assessments are modeled jointly rather than treating the Week 18 value as an isolated endpoint. The model also incorporates treatment-by-time interaction and prespecified stratification factors as covariates.
7. Safety
The registry provides serious adverse-event counts by treatment course and arm. These data are presented exactly as reported rather than converted into percentages or reconstructed into a different denominator.
| Treatment group / course | Serious adverse events affected / at risk |
|---|---|
| Pembrolizumab + Ipilimumab — First Course | 146/282 |
| Pembrolizumab + Placebo — First Course | 114/281 |
| Pembrolizumab + Placebo — Second Course | 4/18 |
These are serious-adverse-event results rather than efficacy endpoints. They should therefore be interpreted separately from the OS, PFS, response, and patient-reported outcome analyses. The denominators also differ across the reported treatment-course categories, so the counts should not be combined into a single treatment-group comparison.
8. Statistical Methodology
Kaplan-Meier estimation
The registry identifies Kaplan-Meier methodology for the OS endpoint. For a time-to-event outcome, Kaplan-Meier estimation provides an estimate of the event-free survival function while allowing participants who have not experienced the event at the end of observed follow-up to contribute information until censoring.
where di is the number of events at time ti and ni is the number at risk immediately before that time.
The registry specifically states that the median survival and associated 95% confidence intervals were reported using the Kaplan-Meier method for OS. The numerical median values are not included in the ClinicalTrials.gov record, so they are not reproduced here.
Stratified log-rank test
The primary OS and PFS comparisons used log-rank testing. The analyses incorporated stratification factors, allowing the treatment comparison to account for the prespecified grouping variables identified in the registry analysis notes.
For OS, the Cox model was stratified by ECOG, geographic region of the enrolling site, and predominant tumor history. For PFS, the same stratification factors were used. The TTD analysis used ECOG, geographic region of the enrolling site, and predominant tumor histology.
Cox regression and hazard ratios
The OS, PFS, and TTD analyses use Cox regression to estimate hazard ratios. The registry specifies Efron's method of tie handling in these Cox models.
The hazard ratio is a relative time-to-event measure. It is not the same as a risk ratio, risk difference, median difference, or percentage of participants who benefit.
Score-based confidence intervals for proportions
The ORR analysis used the Miettinen & Nurminen method. This is a score-based approach for inference on differences between proportions. In this trial, the effect measure is the difference in percentage, represented as a risk difference.
This choice illustrates an important distinction between a binary endpoint and a time-to-event endpoint. ORR reduces the response outcome to whether a participant meets the response definition, whereas OS and PFS retain the timing of events and censoring information.
Constrained longitudinal data analysis
The Week 18 EORTC QLQ-C30 analysis used a constrained longitudinal data analysis model. The registry specifies PRO scores as the response variable, treatment-by-time interaction, and ECOG, geographic region of the enrolling site, and predominant tumor histology as covariates.
A longitudinal model is useful when measurements are collected at multiple time points because it can represent the trajectory of the outcome rather than treating each time point as an unrelated analysis. The constraint in cLDA provides a structured relationship between treatment groups at baseline while allowing treatment effects to be estimated over time.
Stratified analysis
Stratification appears repeatedly in the registry analyses. Its purpose is to incorporate important design factors into the comparison rather than ignoring the structure used in the randomized trial. The exact stratification variables differ by endpoint where the registry specifies predominant tumor history versus predominant tumor histology.
9. Statistical Methods Explained
Why was a log-rank test used for OS and PFS?
OS and PFS are time-to-event endpoints. A log-rank test is designed to compare survival experience between groups over follow-up while incorporating the timing of events and censoring. That is fundamentally different from comparing only the proportion of participants who experienced an event by a single fixed date.
What does an OS hazard ratio of 1.08 mean?
The estimated OS hazard ratio of 1.08 indicates that the fitted model estimated the instantaneous rate of death in the pembrolizumab-plus-ipilimumab group at 1.08 times the rate in the pembrolizumab-plus-placebo group. It does not mean that exactly 8% more participants died, nor does it describe an individual participant's probability of death.
Why is the confidence interval more informative than the point estimate alone?
A hazard ratio such as 1.08 is only one estimate from the observed data. The 95% CI of 0.85–1.37 shows how much statistical uncertainty surrounds that estimate under the specified analysis framework. The same principle applies to the PFS HR of 1.06 with its 95% CI of 0.86–1.30 and to the ORR risk difference of −0.1 with its 95% CI of −8.2 to 8.1.
Why is the p-value not an effect-size measure?
A p-value summarizes evidence under a specified hypothesis-testing framework. It does not tell us how large an effect is. Two studies can have the same p-value with very different effect estimates, and a small p-value does not automatically imply a clinically large effect. The estimated effect and its confidence interval must be considered separately.
Why does the PFS analysis mention one-sided testing?
The registry analysis notes identify the PFS p-value as one-sided, while the reported confidence interval is two-sided at the 95% level. The tail convention matters because it defines how the hypothesis test treats the direction of the treatment effect. It should therefore be reported explicitly rather than inferred from the confidence interval alone.
Why use cLDA for the quality-of-life endpoint?
The quality-of-life outcome was measured longitudinally, with baseline and Week 18 assessments. cLDA models these repeated observations jointly and includes treatment-by-time interaction and the specified stratification factors as covariates. This approach is distinct from simply comparing two unadjusted means at Week 18.
10. One-Sided Testing and Two-Sided Confidence Intervals
One of the more instructive features of the registry analysis is that the PFS and ORR entries explicitly identify one-sided p-values while reporting two-sided 95% confidence intervals.
| Endpoint | P-value convention reported | Confidence interval |
|---|---|---|
| Overall Survival | Hypothesis type listed as other / not stated | 95%, two-sided |
| Progression-free Survival | One-sided p-value based on stratified log-rank test | 95%, two-sided |
| Objective Response Rate | One-sided p-value based on Miettinen & Nurminen method | 95%, two-sided |
| Time to True Deterioration | Two-sided p-value | 95%, two-sided |
| QLQ-C30 Global Health Status/Quality of Life | Hypothesis type listed as other / not stated | 95%, two-sided |
The presence of a one-sided p-value does not turn the confidence interval into a one-sided interval. These are separate inferential quantities. A careful statistical report should preserve the convention actually stated in the registry rather than harmonizing the entries into a single assumed testing framework.
11. Analysis Populations and Censoring
The two primary endpoints were analyzed in all randomized participants. This is important because the randomized analysis population preserves treatment assignment as the organizing principle of the comparison.
For OS, participants without documented death at the time of analysis were censored at the date of last known contact. This is a standard time-to-event structure: participants can contribute observed follow-up even when the event of interest has not occurred by the analysis date.
PFS uses a composite event definition: the first documented disease progression per RECIST 1.1 based on BICR or death due to any cause, whichever occurs first. This means PFS is not simply a measure of radiographic progression; death is also an event for this endpoint.
12. Stratification and Covariate Adjustment
The registry analysis notes repeatedly identify stratification by baseline or site-level factors. For OS and PFS, the Cox model included treatment as a covariate and was stratified by ECOG, geographic region of the enrolling site, and predominant tumor history. For TTD and the quality-of-life analysis, the registry specifies predominant tumor histology along with ECOG and geographic region.
| Analysis | Adjustment / stratification information reported |
|---|---|
| Overall Survival | ECOG; geographic region of the enrolling site; predominant tumor history. |
| Progression-free Survival | ECOG; geographic region of the enrolling site; predominant tumor history. |
| Objective Response Rate | ECOG; geographic region of the enrolling site; predominant tumor history. |
| Time to True Deterioration | ECOG; geographic region of the enrolling site; predominant tumor histology. |
| QLQ-C30 Global Health Status/Quality of Life | Treatment-by-time interaction plus ECOG, geographic region of the enrolling site, and predominant tumor histology as covariates. |
Stratification is particularly useful in a randomized trial when important design factors are known in advance. Rather than pretending that every randomized sample is perfectly balanced within every factor, the analysis can incorporate those factors into the statistical comparison.
13. Interpreting the Primary Results Together
| Endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Overall Survival | Hazard ratio | 1.08 | 0.85–1.37 | 0.74156 |
| Progression-free Survival | Hazard ratio | 1.06 | 0.86–1.30 | 0.71720 |
The two primary endpoint estimates are both close to 1, and both reported 95% confidence intervals include 1. The p-values are also reported as 0.74156 for OS and 0.71720 for PFS. Taken together, these values describe the statistical results of the two primary comparisons without requiring the reader to collapse them into a single summary number.
An important statistical point is that OS and PFS are related but not interchangeable. OS records death from any cause. PFS records the earlier of documented progression or death. A treatment effect can therefore differ between the endpoints because they capture different events and different portions of the clinical course.
14. Why the Confidence Intervals Matter
OS
The HR of 1.08 is accompanied by a 95% CI of 0.85–1.37. The interval shows uncertainty extending below and above the no-difference value of 1.
PFS
The HR of 1.06 is accompanied by a 95% CI of 0.86–1.30. The same basic interpretation applies: the interval spans the no-difference value.
ORR
The risk difference of −0.1 has a 95% CI of −8.2 to 8.1, spanning zero, the no-difference value for a difference in percentages.
Quality of life
The LS-mean difference of −0.42 has a 95% CI of −3.96 to 3.12, also spanning zero.
Confidence intervals should be read in the scale of the effect measure. For OS and PFS, the neutral value is a hazard ratio of 1. For ORR and the quality-of-life analysis, the neutral value is a difference of 0. This distinction is fundamental when moving between time-to-event, binary, and continuous outcomes.
15. Limitations
- Registry-level numerical scope: This page is intentionally limited to the numerical results reported in the ClinicalTrials.gov-derived trial data. Additional results not present in that data are not inferred.
- Hazard-ratio interpretation: Cox-model hazard ratios are model-based relative measures. Their interpretation depends on the time-to-event structure and the assumptions underlying the model.
- Different analysis populations: The primary OS and PFS analyses include all randomized participants, whereas TTD and QLQ-C30 analyses use more restricted populations.
- Different outcome types: OS and PFS are time-to-event outcomes, ORR is binary, and the QLQ-C30 endpoint is continuous and longitudinal. Their estimates cannot be directly compared as though they were the same measure.
- One-sided versus two-sided inference: The registry explicitly reports one-sided p-values for PFS and ORR while reporting two-sided 95% confidence intervals. This should be preserved when interpreting the results.
- Safety denominators: Serious adverse-event counts are reported by treatment course with different denominators. They should not be recombined into an overall safety estimate.
- Missing-data assumptions: The ClinicalTrials.gov record identifies the cLDA method but do not provide a detailed missing-data or imputation strategy. No additional imputation procedure is inferred here.
- Multiplicity: The ClinicalTrials.gov record does not provide a detailed multiplicity or alpha-allocation scheme. No unreported multiplicity procedure is attributed to the trial.
- Interim analysis: The ClinicalTrials.gov record does not report an interim-analysis schedule or alpha-spending procedure. None is inferred.
- Crossover: The ClinicalTrials.gov record does not report a crossover scheme. No crossover effect on OS is therefore assumed.
- Non-inferiority: The ClinicalTrials.gov record does not identify a non-inferiority design or margin. The primary analyses are not described here using non-inferiority logic.
16. Why This Trial Matters Statistically
KEYNOTE-598 is a useful teaching case because the registry results bring several core biostatistical methods together within one randomized phase 3 trial. The primary outcomes are time-to-event endpoints, while the secondary analyses demonstrate binary and longitudinal methods.
| Concept | How it appears in KEYNOTE-598 |
|---|---|
| Randomization | The study uses randomized allocation in a parallel-group phase 3 design. |
| Blinding | The registry identifies quadruple masking. |
| Time-to-event endpoints | OS and PFS are the two registered primary endpoints. |
| Kaplan-Meier estimation | Reported for OS, with median survival and associated 95% CIs described in the registry endpoint definition. |
| Log-rank testing | Used for the primary OS and PFS comparisons and for TTD. |
| Hazard ratio | Used for OS, PFS, and TTD. |
| Cox regression | Used for time-to-event effect estimation, with Efron's method for tied event times. |
| Stratified analysis | Stratification by ECOG, geographic region, and tumor-history or histology factors is specified in the analyses. |
| Risk difference | Used for the ORR comparison. |
| Miettinen-Nurminen method | Used for the ORR analysis and its score-based confidence interval. |
| cLDA | Used for the longitudinal EORTC QLQ-C30 Global Health Status/Quality of Life analysis. |
| One-sided testing | Explicitly identified in the registry analysis notes for PFS and ORR. |
| Different analysis populations | Primary efficacy analyses use all randomized participants, while some secondary analyses require treatment exposure and available patient-reported outcome assessments. |
17. Statistical Methods Explained in More Depth
Hazard ratio versus risk difference
The OS and PFS analyses use hazard ratios because their endpoints incorporate event timing and censoring. ORR uses a risk difference because the outcome is binary. These effect measures live on different scales and answer different questions.
Neither measure should be substituted for the other. The appropriate effect measure follows from the endpoint definition.
Why censoring matters
For OS, participants without documented death at analysis were censored at the date of last known contact. Censoring allows information from participants who have not experienced the event to be used up to the point at which their outcome is known. The validity of survival analysis depends on assumptions concerning the relationship between censoring and the event process.
Why BICR matters for PFS
PFS was defined using RECIST 1.1 based on blinded independent central review. Central review provides a standardized assessment framework for radiographic progression. This is particularly relevant for an endpoint in which the timing of progression determines the time-to-event measurement.
Why the analysis population matters
The primary OS and PFS analyses were based on all randomized participants. This preserves the randomized comparison. In contrast, the TTD analysis requires at least one dose of study treatment and available EORTC QLQ-LC13 and QLQ-C30 information, while the QLQ-C30 analysis requires treatment exposure and at least one QLQ-C30 assessment. These restrictions can affect which participants contribute information to the secondary endpoints.
18. Primary Endpoint Comparison at a Glance
The graphic above is a visual representation of the two reported point estimates on a common display scale. It is not a Kaplan-Meier curve and does not represent survival probabilities, cumulative incidence, or event counts. The exact numerical estimates and confidence intervals remain the authoritative statistical quantities.
19. What the Primary Results Do — and Do Not — Establish
The registry reports an OS hazard ratio of 1.08 and a PFS hazard ratio of 1.06 for pembrolizumab plus ipilimumab versus pembrolizumab plus placebo. Both estimates are accompanied by two-sided 95% confidence intervals that span 1.
The OS interval is 0.85–1.37, while the PFS interval is 0.86–1.30. These intervals quantify uncertainty around the respective model-based effect estimates. They do not describe individual treatment responses or guarantee that the true effect lies within the interval.
The reported p-values are 0.74156 for OS and 0.71720 for PFS. These are outputs of the specified hypothesis-testing procedures. They are not probabilities that the null hypothesis is true and are not measures of treatment magnitude.
OS and PFS measure different clinical events. The appropriate interpretation therefore considers the endpoint definition, analysis population, effect measure, confidence interval, p-value convention, censoring structure, and model assumptions together.
20. Secondary Endpoint Perspective
| Endpoint | Outcome type | Effect measure | Estimate | 95% CI |
|---|---|---|---|---|
| Overall Survival | Time-to-event | Hazard ratio | 1.08 | 0.85–1.37 |
| Progression-free Survival | Time-to-event | Hazard ratio | 1.06 | 0.86–1.30 |
| Objective Response Rate | Binary | Risk difference | −0.1 | −8.2 to 8.1 |
| Time to True Deterioration | Time-to-event | Hazard ratio | 0.9815 | 0.7386–1.3042 |
| QLQ-C30 Global Health Status/Quality of Life | Continuous longitudinal | Mean difference | −0.42 | −3.96 to 3.12 |
This table illustrates a central principle in clinical-trial statistics: the statistical method follows the structure of the endpoint. Time-to-event outcomes use survival-analysis tools; binary response uses methods for proportions; and repeated quality-of-life measurements use a longitudinal model.
21. Related Tutorials
Learn more about the methods used in this trial:
22. Related Calculators
23. Sources
- ClinicalTrials.gov: NCT03302234 — KEYNOTE-598.
- PubMed: PMID 33513313.
Continue through Clinical Biostats
Explore statistical tutorials and calculators covering the methods used to design, analyze, and interpret randomized clinical trials.
24. Record Summary
KEYNOTE-598 provides a compact example of several important clinical-trial statistical methods. The randomized phase 3 design compares pembrolizumab plus ipilimumab with pembrolizumab plus placebo in untreated metastatic non-small cell lung cancer, with OS and PFS as primary time-to-event endpoints. The registry analyses use stratified log-rank testing and Cox regression for the time-to-event outcomes, a score-based Miettinen-Nurminen approach for the binary ORR endpoint, and constrained longitudinal data analysis for the Week 18 quality-of-life endpoint.
The primary OS estimate was HR 1.08 with a 95% CI of 0.85–1.37 and p-value 0.74156. The primary PFS estimate was HR 1.06 with a 95% CI of 0.86–1.30 and p-value 0.71720. The secondary analyses similarly illustrate how effect measures and inferential methods should be matched to endpoint structure: risk difference for ORR, hazard ratio for TTD, and adjusted mean difference for longitudinal quality of life.