← Clinical Trials
Triple-Negative Breast Cancer Phase 3 Completed NCT02819518

KEYNOTE-355: Complete Statistical Analysis of Pembrolizumab Plus Chemotherapy in Triple-Negative Breast Cancer

An independent statistical review of the randomized phase 3 KEYNOTE-355 trial comparing pembrolizumab plus chemotherapy with placebo plus chemotherapy in previously untreated locally recurrent inoperable or metastatic triple-negative breast cancer.

Enrollment: 882  ·  Start: July 27, 2016  ·  Primary completion: June 15, 2021
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

KEYNOTE-355 was a randomized, quadruple-masked, parallel-group phase 3 trial evaluating pembrolizumab plus chemotherapy versus placebo plus chemotherapy for previously untreated locally recurrent inoperable or metastatic triple-negative breast cancer.

882
Enrollment
ClinicalTrials.gov record
5
Arms
Registered trial arms
0.82
PFS HR
95% CI 0.70–0.98
0.73
OS HR, CPS ≥10
95% CI 0.55–0.95
FeatureKEYNOTE-355
TrialKEYNOTE-355
ClinicalTrials.gov identifierNCT02819518
PhasePhase 3
ConditionTriple Negative Breast Cancer (TNBC)
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment882
StatusCompleted
Lead sponsorMerck Sharp & Dohme LLC
Sponsor typeIndustry
StartJuly 27, 2016
Primary completionJune 15, 2021

2. Clinical Question

The central statistical question was whether adding pembrolizumab to chemotherapy improves time-to-event outcomes compared with placebo plus chemotherapy in previously untreated locally recurrent inoperable or metastatic triple-negative breast cancer.

Population

Participants with previously untreated locally recurrent inoperable or metastatic triple-negative breast cancer.

Intervention

Pembrolizumab combined with chemotherapy. Registered chemotherapy interventions included nab-paclitaxel, paclitaxel, gemcitabine, and carboplatin.

Comparator

Placebo plus chemotherapy.

Primary question

Does pembrolizumab plus chemotherapy improve progression-free survival and overall survival relative to placebo plus chemotherapy?

3. Trial Design

01
Randomize882 enrolled
02
Parallel groupsRandomized allocation
03
MaskedQuadruple masking
04
AssessPFS, OS, response, safety
05
AnalyzeStratified time-to-event and categorical methods
Allocation
Randomized. The primary efficacy analyses reported in the registry analyzed participants according to the treatment arm to which they were randomly assigned, regardless of whether they received treatment.
Masking
Quadruple. The registry identifies the study as quadruple-masked.
Model
Parallel. The trial used a parallel-group design rather than a crossover or factorial model.
Primary purpose
Treatment. The registered primary purpose was treatment.
PEMBROLIZUMAB-CHEMOTHERAPY

Pembrolizumab + chemotherapy

  • Pembrolizumab
  • Nab-paclitaxel
  • Paclitaxel
  • Gemcitabine
  • Carboplatin
CONTROL

Placebo + chemotherapy

  • Placebo intervention included Normale Saline Solution
  • Comparator chemotherapy regimen
  • Primary comparisons in the registry: pembrolizumab + chemotherapy versus placebo + chemotherapy

The registry data identify 5 arms overall while the posted formal efficacy analyses compare two aggregate treatment strategies: Part 2 pembrolizumab plus chemotherapy versus Part 2 placebo plus chemotherapy. The ClinicalTrials.gov record does not provide arm-specific enrollment counts for all 5 registered arms.

4. Endpoints

The registry lists eight primary endpoints. Two concern adverse events in Part 1; six concern PFS and OS in Part 2. The formal statistical analyses reported in the ClinicalTrials.gov record are for the six Part 2 time-to-event endpoints.

Primary endpointTime frameTypeFormal analysis reported
Part 1: Percentage of Participants Who Experienced an Adverse Event (AE) - All Participants Up to approximately 39 months Binary No formal statistical comparison reported
Part 1: Percentage of Participants Who Discontinued Study Drug Due to an AE - All Participants Up to approximately 39 months Binary No formal statistical comparison reported
Part 2: Progression-Free Survival (PFS) - All Participants Up to approximately 53 months Time-to-event Log-rank; Cox regression
Part 2: PFS - Participants With PD-L1 CPS ≥1 Tumors Up to approximately 53 months Time-to-event Log-rank; Cox regression
Part 2: PFS - Participants With PD-L1 CPS ≥10 Tumors Up to approximately 53 months Time-to-event Log-rank; Cox regression
Part 2: Overall Survival (OS) - All Participants Up to approximately 53 months Time-to-event Log-rank; Cox regression
Part 2: OS - Participants With PD-L1 CPS ≥1 Tumors Up to approximately 53 months Time-to-event Log-rank; Cox regression
Part 2: OS - Participants With PD-L1 CPS ≥10 Tumors Up to approximately 53 months Time-to-event Log-rank; Cox regression

Registry definitions for the time-to-event endpoints

Progression-free survival was defined as the time from randomization to the first documented progressive disease per RECIST 1.1 based on assessments by blinded independent central review, or death due to any cause, whichever occurred first. The registry definition states that progressive disease was defined as at least a 20% increase in the sum of diameters of target lesions; for the CPS ≥1 and CPS ≥10 definitions, the registry text additionally states that the sum must have demonstrated an absolute increase of at least 5 mm.

Overall survival was defined as the time from randomization to death due to any cause. Participants without documented death at the time of the analysis were censored at the date of last follow-up.

5. Statistical Methodology

Primary time-to-event framework

The six formal primary efficacy analyses use the log-rank test together with a hazard-ratio effect measure. The analysis notes specify Cox regression models with treatment as a covariate and stratification by prespecified trial factors.

Core survival-analysis structure
Randomization → event or censoring → Kaplan-Meier estimation → log-rank comparison → Cox hazard ratio

The registry's normalized methods are the log-rank test and stratified analysis through the Cox model. The ClinicalTrials.gov record does not provide separate Kaplan-Meier numerical estimates such as median PFS or median OS, so those quantities are not reported here.

Stratified Cox regression

For the all-participant PFS analysis, the Cox model included treatment as a covariate and was stratified by chemotherapy on study, tumor PD-L1 status, and prior treatment with the same class of chemotherapy in the (neo)adjuvant setting. For the CPS ≥1 and CPS ≥10 PFS analyses, the registry-reported analysis notes specify stratification by chemotherapy on study and prior treatment with the same class of chemotherapy in the (neo)adjuvant setting.

The corresponding OS analyses use the same structure: the all-participant OS model is stratified by chemotherapy on study, tumor PD-L1 status, and prior treatment with the same class of chemotherapy in the (neo)adjuvant setting; the CPS-defined OS analyses are stratified by chemotherapy on study and prior treatment with the same class of chemotherapy in the (neo)adjuvant setting.

Analysis population

For the primary all-participant efficacy analyses, participants were analyzed in the treatment arm to which they were randomly assigned, regardless of whether they received treatment. The CPS ≥1 and CPS ≥10 analyses similarly use randomized treatment assignment among participants meeting the corresponding PD-L1 criterion.

Secondary binary endpoints

The registry-reported secondary analyses of objective response rate and disease control rate use the Cochran-Mantel-Haenszel test. The reported effect measure is the difference in the percentage of participants between treatment groups, with two-sided 95% confidence intervals.

The analysis notes identify the Miettinen & Nurminen method for the reported confidence intervals, with stratification by trial factors. For ORR and DCR in the all-participant analyses, these include chemotherapy on study, tumor PD-L1 status where specified, and prior treatment with the same class of chemotherapy in the (neo)adjuvant setting.

Superiority framework

All registry-reported formal efficacy analyses are labeled as superiority hypotheses. The confidence intervals are two-sided 95% intervals, and the reported P-values are two-sided.

6. Primary Results: Progression-Free Survival

PFS — All Participants

Pembrolizumab + chemotherapy vs placebo + chemotherapy

HR 0.82

95% CI: 0.70–0.98   ·   P = 0.0120

Analysis time frame: up to approximately 53 months

The reported hazard ratio of 0.82 means that the estimated instantaneous rate of progression or death was 18% lower in the pembrolizumab-plus-chemotherapy group relative to the placebo-plus-chemotherapy group under the fitted Cox model.

Clinical Biostats interpretation

What the estimate means: HR 0.82 is a relative time-to-event measure. It summarizes the estimated treatment-group difference in the instantaneous event rate within the Cox model.

What it does not mean: it does not mean that 18% of participants avoided progression, that each participant had exactly an 18% reduction in risk, or that median PFS was reduced or increased by 18%.

Precision: the 95% CI of 0.70–0.98 describes uncertainty around the estimated hazard ratio under the model and sampling framework. Because the interval lies below 1, the reported estimate is compatible with a lower event hazard for the pembrolizumab combination.

The P-value: P = 0.0120 addresses evidence against the superiority null hypothesis under the specified statistical framework. It does not measure the magnitude or clinical importance of the treatment effect.

Important caution: the HR relies on a Cox-model framework. A single HR is most straightforward to interpret when the proportional-hazards structure is reasonably appropriate over the relevant follow-up. The ClinicalTrials.gov record does not provide a test or diagnostic for that assumption.

PFS — PD-L1 CPS ≥1

Participants with PD-L1 CPS ≥1 tumors

HR 0.75

95% CI: 0.62–0.91   ·   P = 0.0016

Analysis time frame: up to approximately 53 months

Clinical Biostats interpretation

HR 0.75 corresponds to a 25% lower estimated instantaneous rate of progression or death under the fitted model for the pembrolizumab-plus-chemotherapy group relative to placebo plus chemotherapy.

The 95% CI of 0.62–0.91 indicates the precision of this estimate; it does not describe the range of individual patient outcomes. P = 0.0016 is evidence against the specified superiority null hypothesis, but it is not an effect-size measure.

This analysis is restricted to participants with PD-L1 CPS ≥1 tumors. It should therefore not be read as an estimate for participants outside that analysis population.

PFS — PD-L1 CPS ≥10

Participants with PD-L1 CPS ≥10 tumors

HR 0.66

95% CI: 0.50–0.88   ·   P = 0.0018

Analysis time frame: up to approximately 53 months

Clinical Biostats interpretation

HR 0.66 corresponds to a 34% lower estimated instantaneous rate of progression or death under the fitted Cox model for pembrolizumab plus chemotherapy versus placebo plus chemotherapy.

The 95% CI of 0.50–0.88 quantifies uncertainty around the estimated relative hazard. It does not indicate that individual treatment effects range from 50% to 88%, and it does not provide an absolute probability of benefit.

P = 0.0018 is a hypothesis-test result. It should not be interpreted as the probability that the treatment works, nor as the probability that the observed effect is due to chance.

7. Primary Results: Overall Survival

OS — All Participants

Pembrolizumab + chemotherapy vs placebo + chemotherapy

HR 0.89

95% CI: 0.76–1.05   ·   P = 0.0797

Analysis time frame: up to approximately 53 months

Clinical Biostats interpretation

What the estimate means: HR 0.89 corresponds to an estimated 11% lower instantaneous rate of death in the pembrolizumab-plus-chemotherapy group relative to placebo plus chemotherapy under the fitted model.

What it does not mean: it does not mean that 11% of patients survived because of treatment, nor does it translate directly into an 11% increase in survival time.

Precision: the 95% CI of 0.76–1.05 includes 1.00. Thus, the interval is compatible with a range of relative effects that includes no difference in hazard under the model.

P-value: P = 0.0797 is not an effect-size measure. Under the reported two-sided superiority framework, it does not provide the conventional statistical evidence against the null hypothesis at a 0.05 threshold.

Interpretation caution: OS is susceptible to subsequent treatment and other post-randomization events. The ClinicalTrials.gov record does not provide information on subsequent therapy or crossover sufficient to quantify their influence on this estimate.

OS — PD-L1 CPS ≥1

Participants with PD-L1 CPS ≥1 tumors

HR 0.86

95% CI: 0.72–1.04   ·   P = 0.0563

Analysis time frame: up to approximately 53 months

Clinical Biostats interpretation

HR 0.86 corresponds to a 14% lower estimated instantaneous rate of death under the fitted model for pembrolizumab plus chemotherapy relative to placebo plus chemotherapy among participants with PD-L1 CPS ≥1 tumors.

The 95% CI of 0.72–1.04 crosses 1.00, so the interval includes a no-difference hazard ratio. P = 0.0563 is close to 0.05, but a P-value should not be converted into a graded measure of treatment effectiveness. The confidence interval gives substantially more information about the uncertainty of the estimated effect.

OS — PD-L1 CPS ≥10

Participants with PD-L1 CPS ≥10 tumors

HR 0.73

95% CI: 0.55–0.95   ·   P = 0.0093

Analysis time frame: up to approximately 53 months

Clinical Biostats interpretation

HR 0.73 corresponds to a 27% lower estimated instantaneous rate of death under the fitted Cox model for pembrolizumab plus chemotherapy relative to placebo plus chemotherapy among participants with PD-L1 CPS ≥10 tumors.

The 95% CI of 0.55–0.95 describes uncertainty around that relative hazard estimate and remains below 1.00. P = 0.0093 provides evidence against the reported superiority null hypothesis under the stated analysis framework, but does not quantify either the probability of treatment benefit or its clinical magnitude.

The CPS ≥10 population is a defined analysis subset. A smaller hazard ratio in one subgroup than another does not, by itself, establish that treatment effects differ between subgroups; that requires a formal interaction or heterogeneity analysis.

8. Primary Results at a Glance

Primary endpointHazard ratio95% CIP-value
PFS — All Participants0.820.70–0.980.0120
PFS — PD-L1 CPS ≥10.750.62–0.910.0016
PFS — PD-L1 CPS ≥100.660.50–0.880.0018
OS — All Participants0.890.76–1.050.0797
OS — PD-L1 CPS ≥10.860.72–1.040.0563
OS — PD-L1 CPS ≥100.730.55–0.950.0093

The six primary analyses show why it is important to examine the estimate, confidence interval, P-value, and analysis population together. The reported PFS hazard ratios are all below 1. For OS, the all-participant and CPS ≥1 estimates have confidence intervals that include 1.00, whereas the CPS ≥10 estimate is 0.73 with a 95% CI of 0.55–0.95.

9. Secondary Results: Objective Response Rate

Objective response rate was analyzed as a binary endpoint using the Cochran-Mantel-Haenszel method. The reported effect measure is the difference in ORR percentage versus control.

ORR — All Participants

Difference in ORR versus control

3.8 percentage points

95% CI: −3.2 to 10.6 percentage points   ·   P = 0.1413

Analysis time frame: up to approximately 53 months

ORR — PD-L1 CPS ≥1

Difference in ORR versus control

6.1 percentage points

95% CI: −2.1 to 14.0 percentage points   ·   P = 0.0725

Analysis time frame: up to approximately 53 months

ORR — PD-L1 CPS ≥10

Difference in ORR versus control

12.1 percentage points

95% CI: 0.4 to 23.4 percentage points   ·   P = 0.0213

Analysis time frame: up to approximately 53 months

How to interpret the ORR analyses

An ORR difference is an absolute percentage-point difference, not a hazard ratio. For example, an estimate of 12.1 means that the reported response percentage was estimated to be 12.1 percentage points higher in the pembrolizumab-plus-chemotherapy group than in the comparator group for the specified CPS ≥10 population.

The confidence interval expresses uncertainty around the percentage-point difference. In the all-participant and CPS ≥1 analyses, the intervals extend below 0; in the CPS ≥10 analysis, the registry-reported interval ranges from 0.4 to 23.4 percentage points.

The P-value tests the relevant superiority hypothesis under the specified categorical-data method. It does not measure how large or clinically important an ORR difference is.

10. Secondary Results: Disease Control Rate

EndpointDifference vs control95% CIP-value
DCR — All Participants4.7 percentage points−2.4 to 11.80.0966
DCR — PD-L1 CPS ≥15.0 percentage points−3.2 to 13.10.1164
DCR — PD-L1 CPS ≥1010.8 percentage points−0.7 to 22.30.0327

These estimates were obtained using the Cochran-Mantel-Haenszel method, with the reported confidence intervals based on the Miettinen & Nurminen method and the specified stratification factors.

Clinical Biostats interpretation

DCR is a binary response-related endpoint, so its effect measure is naturally expressed as a difference in percentages rather than a hazard ratio. The CPS ≥10 analysis reports a 10.8-percentage-point difference with a 95% CI of −0.7 to 22.3 and P = 0.0327.

The distinction between statistical significance and effect size remains important: P = 0.0327 does not say that the probability of treatment benefit is 96.73%, nor does it establish a 10.8% relative improvement. The estimate itself is 10.8 percentage points.

11. Secondary Results Summary

EndpointAnalysis methodEffect measureEstimate95% CIP-value
ORR — All Participants Cochran-Mantel-Haenszel Difference in ORR 3.8 −3.2 to 10.6 0.1413
ORR — CPS ≥1 Cochran-Mantel-Haenszel Difference in ORR 6.1 −2.1 to 14.0 0.0725
ORR — CPS ≥10 Cochran-Mantel-Haenszel Difference in ORR 12.1 0.4 to 23.4 0.0213
DCR — All Participants Cochran-Mantel-Haenszel Difference in DCR 4.7 −2.4 to 11.8 0.0966
DCR — CPS ≥1 Cochran-Mantel-Haenszel Difference in DCR 5.0 −3.2 to 13.1 0.1164
DCR — CPS ≥10 Cochran-Mantel-Haenszel Difference in DCR 10.8 −0.7 to 22.3 0.0327

12. Safety Results

The ClinicalTrials.gov record includes serious adverse-event counts by arm or analysis part. They do not provide a complete arm-level table for all adverse events or for discontinuations due to adverse events. Accordingly, the safety information below is limited to the serious adverse-event figures contained in the ClinicalTrials.gov record.

Analysis part / armSerious adverse events affectedParticipants at risk
Part 1: Pembrolizumab + Nab-Paclitaxel313
Part 1: Pembrolizumab + Paclitaxel410
Part 1: Pembrolizumab + Gemcitabine/Carb911
Part 2: Pembrolizumab + Chemotherapy (Fi)169562
Part 2: Pembrolizumab + Chemotherapy (Se)312
Part 2: Placebo + Chemotherapy68281

The registry's two Part 1 primary safety endpoints are the percentage of participants experiencing an adverse event and the percentage discontinuing study drug due to an adverse event, both with a time frame of up to approximately 39 months. The ClinicalTrials.gov record does not contain formal comparative estimates, confidence intervals, or P-values for either endpoint.

Safety interpretation: the serious-adverse-event counts above should not be converted into formal between-arm conclusions beyond what the ClinicalTrials.gov record supports. A conventional comparison of binary safety outcomes would generally use event proportions with an appropriate confidence interval and, where prespecified, a suitable categorical-data test or model. The ClinicalTrials.gov record does not report such a formal comparison for these primary safety endpoints.

13. Statistical Methods Explained

Why was a log-rank test used?

PFS and OS are time-to-event endpoints. Not every participant experiences the event during the observation period, so simply comparing proportions would discard information from censored participants. The log-rank test compares the observed and expected event experience between randomized groups over follow-up while accommodating right censoring.

What does a hazard ratio of 0.82 mean?

An HR of 0.82 means that the fitted Cox model estimates the instantaneous rate of progression or death to be 82% of the corresponding rate in the comparator group, or approximately 18% lower. It is not an 18-percentage-point difference in the probability of progression, and it is not an 18% increase in median survival.

Why was a stratified Cox model used?

Stratification allows the survival comparison to account for specified trial factors without requiring one common baseline hazard across all strata. In KEYNOTE-355, the registry-reported analysis notes identify chemotherapy on study and prior treatment with the same class of chemotherapy in the (neo)adjuvant setting for the CPS-defined analyses, with tumor PD-L1 status additionally included for the all-participant analyses.

Why does the confidence interval matter?

A point estimate such as HR 0.66 is only one estimate from the observed trial data. The 95% confidence interval describes the uncertainty around that estimate under the model and repeated-sampling framework. A narrow interval indicates greater precision than a wide interval, all else equal. It does not describe the range of treatment effects for individual patients.

Why does the P-value not measure treatment effect size?

A P-value is a measure of compatibility between the observed data and a specified null hypothesis under the statistical model. It is influenced by the amount of information as well as the observed effect. Effect size is described by the hazard ratio or percentage-point difference, while the confidence interval supplies information about precision.

Why are the PD-L1 CPS ≥1 and CPS ≥10 analyses separate?

They represent different prespecified analysis populations. CPS ≥10 is a subset defined by a higher PD-L1 threshold than CPS ≥1. The resulting estimates describe treatment effects in different populations and should not be compared as though they were independent randomized trials.

Why does a subgroup estimate not prove treatment-effect heterogeneity?

Two subgroup hazard ratios can differ numerically simply because of sampling variation. To conclude that the treatment effect genuinely differs between subgroups, a formal interaction or heterogeneity analysis is generally required. The ClinicalTrials.gov record does not report such an interaction test.

14. Stratification and Covariate Adjustment

The analysis notes repeatedly identify stratified analysis and covariate adjustment. The Cox models included treatment as a covariate and used stratification factors that reflected important trial-design variables.

AnalysisStratification factors specified in registry-reported analysis notes
PFS — All Participants Chemotherapy on study; tumor PD-L1 status; prior treatment with same class of chemotherapy in the (neo)adjuvant setting
PFS — CPS ≥1 Chemotherapy on study; prior treatment with same class of chemotherapy in the (neo)adjuvant setting
PFS — CPS ≥10 Chemotherapy on study; prior treatment with same class of chemotherapy in the (neo)adjuvant setting
OS — All Participants Chemotherapy on study; tumor PD-L1 status; prior treatment with same class of chemotherapy in the (neo)adjuvant setting
OS — CPS ≥1 Chemotherapy on study; prior treatment with same class of chemotherapy in the (neo)adjuvant setting
OS — CPS ≥10 Chemotherapy on study; prior treatment with same class of chemotherapy in the (neo)adjuvant setting

For the secondary ORR and DCR analyses, the registry-reported Miettinen & Nurminen descriptions use stratification by chemotherapy on study and prior treatment with the same class of chemotherapy in the (neo)adjuvant setting, with tumor PD-L1 status additionally specified for the all-participant analyses.

15. Multiplicity and Multiple Primary Analyses

KEYNOTE-355 has multiple registered primary efficacy endpoints: PFS and OS are each evaluated in the all-participant population and in PD-L1 CPS-defined populations. The ClinicalTrials.gov record identifies the hypothesis type as superiority and provide formal estimates and P-values for all six time-to-event primary analyses.

Endpoint familyAnalysis populationsFormal estimate reported
PFSAll Participants; CPS ≥1; CPS ≥10Yes
OSAll Participants; CPS ≥1; CPS ≥10Yes
Part 1 safetyAll ParticipantsNo formal comparison reported

The ClinicalTrials.gov record does not provide the complete multiplicity-adjustment algorithm, alpha allocation, gatekeeping sequence, or interim-analysis boundaries. Those details should therefore not be reconstructed from the observed P-values. In particular, the six reported P-values should not automatically be interpreted as six independent tests performed at an unadjusted 0.05 threshold.

Statistical discipline: the numerical results on this page are restricted to the registry analyses. Because the data do not specify an alpha-spending or multiplicity-control procedure, no additional claim about familywise error control is made here.

16. Interim Analysis, Non-Inferiority, Crossover, and Bayesian Methods

Non-inferiority margin

The ClinicalTrials.gov record identifies the hypothesis type as superiority. No non-inferiority margin is reported.

Crossover

No crossover procedure or crossover-adjusted analysis is reported in the ClinicalTrials.gov record.

Factorial design

The registered design model is parallel. No factorial analysis is reported.

Bayesian methods

No Bayesian analysis is identified in the registry-reported statistical methods. The normalized methods are the log-rank test and Cochran-Mantel-Haenszel test.

The absence of a registry-reported detail is important. A statistical analysis page should not infer an interim-monitoring scheme, crossover correction, Bayesian prior, or non-inferiority margin merely because such methods are common in other clinical trials.

17. Missing Data, Censoring, and Analysis Populations

For PFS, participants were followed until the first documented progression or death, whichever occurred first. Participants without an event contribute censored information according to the applicable time-to-event framework. For OS, participants without documented death at the time of analysis were censored at their last follow-up date.

The registry-reported analysis population definitions emphasize randomized treatment assignment regardless of whether treatment was received. This is an intention-to-treat-style principle that preserves the randomized comparison for efficacy analyses.

Why censoring matters
Observed follow-up = event time when event occurs; otherwise information ends at censoring time

Censoring allows participants with incomplete event observation to contribute partial follow-up information. The validity of standard survival methods depends on assumptions about the censoring process and its relationship to event risk.

The ClinicalTrials.gov record does not specify a separate missing-data imputation strategy for the primary time-to-event analyses. Therefore, no particular imputation method is attributed to KEYNOTE-355 here.

18. Understanding the Hazard Ratios Together

Primary hazard-ratio estimates
PFS — All
0.82
PFS — CPS ≥1
0.75
PFS — CPS ≥10
0.66
OS — All
0.89
OS — CPS ≥1
0.86
OS — CPS ≥10
0.73

The visual comparison is useful for orientation, but the confidence intervals are essential. The PFS estimates range from 0.66 to 0.82, while the OS estimates range from 0.73 to 0.89. The fact that these point estimates differ does not establish that the treatment effect is different across endpoint families or PD-L1 populations. Formal comparisons would require appropriate interaction or heterogeneity analyses.

19. What a Hazard Ratio Does — and Does Not — Mean

Example: HR 0.66

An HR of 0.66 means that the estimated instantaneous rate of the event was 66% of the comparator rate under the fitted Cox model, corresponding to a 34% lower estimated hazard. This is a relative model-based measure.

It does not mean that 34% of participants were protected from the event, that every participant experienced a 34% reduction in risk, or that survival time increased by 34%.

Why the confidence interval matters

For the CPS ≥10 PFS analysis, the HR is 0.66 with a 95% CI of 0.50–0.88. The interval conveys uncertainty around the estimate. It does not describe the range of treatment effects across individual patients.

Why absolute measures would add information

A hazard ratio is relative. Absolute event probabilities, median event times, or survival probabilities at clinically relevant time points can provide a different perspective. Those quantities are not reported in the ClinicalTrials.gov record used for this page and therefore are not added here.

20. P-Values and Confidence Intervals

The six primary analyses illustrate why P-values should be read alongside effect estimates and confidence intervals.

EndpointEstimate95% CI includes null?P-value
PFS — AllHR 0.82No0.0120
PFS — CPS ≥1HR 0.75No0.0016
PFS — CPS ≥10HR 0.66No0.0018
OS — AllHR 0.89Yes0.0797
OS — CPS ≥1HR 0.86Yes0.0563
OS — CPS ≥10HR 0.73No0.0093

For hazard ratios, the conventional null value is 1.00. A confidence interval entirely below 1 indicates that the reported interval does not include the no-difference hazard ratio. That is a different statement from saying that the treatment effect is large, clinically important, or certain to apply to every patient.

21. Limitations

22. Why This Trial Matters Statistically

KEYNOTE-355 is a useful teaching example because it combines randomized treatment assignment, masked trial conduct, stratified time-to-event modeling, PD-L1-defined analysis populations, categorical response endpoints, and multiple primary efficacy analyses.

Statistical conceptHow it appears in KEYNOTE-355
RandomizationThe registered allocation is randomized.
BlindingThe registered masking is quadruple.
Parallel designThe design model is parallel rather than crossover or factorial.
Intention-to-treat principlePrimary efficacy populations are analyzed according to randomized treatment assignment.
Time-to-event analysisPFS and OS are analyzed with log-rank testing and Cox regression.
Hazard ratioThe principal effect measure for the six formal primary time-to-event analyses.
Stratified analysisCox models incorporate specified stratification factors.
Categorical-data analysisORR and DCR use the Cochran-Mantel-Haenszel method.
Confidence intervalsPrimary hazard ratios and secondary percentage-point differences include 95% confidence intervals.
PD-L1 analysis populationsPrimary PFS and OS endpoints are evaluated in all participants and CPS-defined populations.
Superiority testingThe registry-reported formal analyses are labeled superiority hypotheses.
Safety analysisSerious adverse-event counts are reported separately from efficacy outcomes.

23. Planned Analysis for Primary Safety Endpoints

The registry lists two Part 1 primary safety endpoints but the ClinicalTrials.gov record does not provide formal comparative estimates for them.

EndpointRegistry time frameTypical statistical approachWhat is reported
Percentage of Participants Who Experienced an AE — All Participants Up to approximately 39 months Compare binary AE proportions between treatment groups, with an appropriate confidence interval and prespecified categorical-data test or model. Results are posted in the registry, but no formal comparative analysis is reported in the provided statistical-analysis records.
Percentage of Participants Who Discontinued Study Drug Due to an AE — All Participants Up to approximately 39 months Compare the proportions discontinuing because of an AE using an appropriate binary-outcome method. Results are posted in the registry, but no formal comparative analysis is reported in the provided statistical-analysis records.

Because the ClinicalTrials.gov record does not contain the corresponding event proportions, confidence intervals, or P-values, this page does not manufacture those quantities. The serious-adverse-event counts reported in the safety section are a separate piece of information and should not be substituted for the missing primary safety analyses.

24. Overall Statistical Reading of the Trial

The registry-reported primary efficacy results show a consistent pattern for PFS, with hazard ratios of 0.82 among all participants, 0.75 among participants with PD-L1 CPS ≥1 tumors, and 0.66 among participants with CPS ≥10 tumors. The corresponding OS estimates are 0.89, 0.86, and 0.73, respectively.

The important statistical distinction is that these are different estimands in different analysis populations. The all-participant result addresses the overall randomized population analyzed in Part 2. The CPS ≥1 and CPS ≥10 results address progressively more restricted populations. They should not be combined into one average treatment effect, and the numerical differences among their HRs should not be interpreted as treatment-effect modification without formal interaction evidence.

The secondary binary outcomes tell a related but different story. ORR and DCR are expressed as percentage-point differences and were analyzed using the Cochran-Mantel-Haenszel framework rather than a time-to-event model. This distinction illustrates a central principle of clinical-trial statistics: the analysis method should follow the endpoint structure.

Statistical takeaway: the strongest reading of a trial result comes from examining the endpoint definition, analysis population, effect measure, confidence interval, P-value, stratification, and censoring framework together. No single statistic — including a hazard ratio or P-value — captures the entire evidentiary structure of a randomized clinical trial.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Calculators

27. Sources

Continue with the underlying statistical methods

Explore the survival-analysis, categorical-data, confidence-interval, and clinical-trial methods that provide the statistical framework for interpreting KEYNOTE-355.

28. Record Summary

KEYNOTE-355 provides a detailed example of randomized clinical-trial analysis in which multiple primary time-to-event endpoints are evaluated across the overall population and PD-L1-defined analysis populations. The ClinicalTrials.gov record uses stratified Cox regression and log-rank testing for PFS and OS, while ORR and DCR use Cochran-Mantel-Haenszel methods with percentage-point differences as the effect measure.

The most informative statistical reading therefore combines the hazard ratio, 95% confidence interval, P-value, endpoint definition, analysis population, and stratification framework. The registry data also illustrate why safety outcomes, response outcomes, and time-to-event outcomes should not be reduced to a single statistic.

Clinical Biostats methodology: This analysis deliberately limits numerical trial claims to the ClinicalTrials.gov record. Where the registry summary does not provide an estimate, confidence interval, P-value, multiplicity procedure, or other statistical detail, that information is not reconstructed or imported from outside sources.