← Clinical Trials
Advanced Hepatocellular Carcinoma Phase 3 Completed NCT03062358

KEYNOTE-394: Complete Statistical Analysis of Pembrolizumab in Advanced Hepatocellular Carcinoma

An independent statistical analysis of the randomized phase 3 KEYNOTE-394 trial evaluating pembrolizumab or placebo given with best supportive care in Asian participants with previously treated advanced hepatocellular carcinoma.

Trial start: April 27, 2017  ·  Primary completion: June 30, 2021  ·  Enrollment: 453
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

KEYNOTE-394 was a randomized, double-masked, parallel-group phase 3 trial in Asian participants with previously treated advanced hepatocellular carcinoma. The registry reports an enrollment of 453 participants and compares pembrolizumab plus best supportive care with placebo plus best supportive care.

453
Enrollment
Participants
2
Arms
Parallel-group design
0.79
OS HR
95% CI 0.63–0.99
0.0180
OS p-value
One-sided
FeatureKEYNOTE-394
Trial nameKEYNOTE-394
NCT identifierNCT03062358
PhasePhase 3
StatusCompleted
ConditionCarcinoma, Hepatocellular
PopulationAsian participants with previously treated advanced hepatocellular carcinoma
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
Enrollment453
Lead sponsorMerck Sharp & Dohme LLC
Sponsor typeIndustry

2. Clinical Question

The clinical question is whether pembrolizumab plus best supportive care produces a superior overall-survival outcome compared with placebo plus best supportive care in Asian participants with previously treated advanced hepatocellular carcinoma.

Population

Asian participants with previously treated advanced hepatocellular carcinoma.

Intervention

Pembrolizumab given with best supportive care.

Comparator

Placebo given with best supportive care.

Primary question

Does pembrolizumab plus best supportive care improve overall survival relative to placebo plus best supportive care?

3. Trial Design

01
Randomize453 enrolled
02
Two groupsParallel design
03
MaskedDouble masking
04
AssessOS and secondary endpoints
05
AnalyzeUp to approximately 4 years
INTERVENTION

Pembrolizumab + BSC

  • Pembrolizumab
  • Best supportive care (BSC)
  • Compared with the randomized placebo + BSC group
CONTROL

Placebo + BSC

  • Placebo
  • Best supportive care (BSC)
  • Reference group for the randomized comparisons

The registry identifies the study as randomized, parallel, and double-masked. The primary purpose is treatment. The ClinicalTrials.gov record does not report a factorial structure, crossover procedure, non-inferiority margin, or Bayesian analysis.

Trial timeline

April 27, 2017

Study start

The registry lists April 27, 2017 as the study start date.

June 30, 2021

Primary completion

The registry lists June 30, 2021 as the primary completion date.

Completed

Registry status

The trial status is listed as completed.

4. Endpoints

The registry identifies Overall Survival (OS) as the single registered primary endpoint. The primary endpoint is a time-to-event outcome, and the registry also reports four secondary statistical analyses.

EndpointRoleTime frameDefinition / analysis
Overall Survival (OS) Primary Up to approximately 4 years OS is the time from randomization to death due to any cause, based on the Kaplan-Meier method for censored data.
Progression Free Survival (PFS) Per RECIST 1.1 Secondary Up to approximately 4 years Time-to-event endpoint analyzed with a Cox proportional-hazards model.
Objective Response Rate (ORR) Per RECIST 1.1 Secondary Up to approximately 4 years Binary endpoint; effect reported as percent difference.
Disease Control Rate (DCR) Per RECIST 1.1 Secondary Up to approximately 4 years Participants achieving CR, PR, or SD for ≥5 weeks prior to evidence of disease progression.
Time To Progression (TTP) Per RECIST 1.1 Secondary Up to approximately 4 years Time-to-event endpoint analyzed with a Cox proportional-hazards model.
Registry definition of OS: the primary endpoint is explicitly defined as the time from randomization to death due to any cause, with Kaplan-Meier methodology used for censored data. This is distinct from PFS and TTP because the OS event is death from any cause.

5. Statistical Methodology

Primary analysis: Cox proportional-hazards model

The primary OS comparison used a Cox proportional-hazards model. Participants were analyzed in the treatment group to which they were randomized. The reported effect measure was the hazard ratio.

Primary OS model
HR = estimated hazard in Pembrolizumab + BSC ÷ estimated hazard in Placebo + BSC

The registry analysis notes describe treatment as a covariate and specify Efron's method for handling tied event times. The model was stratified by macrovascular invasion, α-Fetoprotein, and region, with all cells corresponding to macrovascular invasion = Yes combined.

Stratified analysis

The OS analysis was stratified by macrovascular invasion (Yes vs. No), α-Fetoprotein (ng/mL) (< 200 vs. ≥ 200), and region (China vs. ex-China). This means the treatment comparison was constructed while accounting for these prespecified stratification factors in the reported Cox analysis.

Kaplan-Meier estimation

The registry defines OS using the Kaplan-Meier method for censored data. Kaplan-Meier estimation is appropriate for time-to-event outcomes because it allows participants who have not experienced the event by their last observation to contribute information up to the point of censoring.

Conceptual Kaplan-Meier estimator
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at time ti, while ni is the number at risk immediately before that event time.

Score-based confidence intervals for binary endpoints

ORR and DCR were binary endpoints. The registry analysis identifies a score-based CI for proportions, specifically the Miettinen-Nurminen approach, as the reported method. The analysis was stratified by macrovascular invasion, α-Fetoprotein, and region.

For ORR and DCR, the reported effect measure was a percent difference, corresponding to a risk difference on the proportion scale. A positive difference indicates a higher observed proportion in the pembrolizumab plus BSC group relative to placebo plus BSC.

Analysis populations

For the reported efficacy analyses, the analysis population was all randomized participants based on the treatment group to which they were randomized. This is an intention-to-treat-style approach and preserves the randomized comparison even if subsequent treatment exposure differs between groups.

6. Statistical Methods Explained

Why was a Cox proportional-hazards model used for OS?

OS records the time from randomization until death and therefore contains both an event indicator and a follow-up time. The Cox model summarizes the relative event rate through a hazard ratio while accommodating censoring. It is particularly useful when not every participant experiences the event during the observation period.

What does an OS hazard ratio of 0.79 mean?

An HR of 0.79 means that the estimated hazard of death in the pembrolizumab plus BSC group was 0.79 times the estimated hazard in the placebo plus BSC group under the fitted Cox model. Equivalently, 1 − 0.79 = 0.21, so the estimate corresponds to an approximately 21% lower estimated hazard of death. This is a relative model-based measure; it does not mean that 21% of participants avoided death or that each participant experienced exactly a 21% reduction in risk.

Why does stratification matter?

The reported Cox analysis was stratified by macrovascular invasion, α-Fetoprotein, and region. Stratification allows the baseline hazard structure to differ across the specified strata while estimating the treatment effect across those strata. It therefore avoids treating these factors as if they had a single common baseline hazard relationship in the model.

Why is the ORR analysis different from the OS analysis?

ORR is binary: a participant either meets the prespecified response definition or does not. OS is a time-to-event outcome, where both event occurrence and follow-up duration matter. The registry therefore reports a score-based method for the proportion comparison for ORR and a Cox proportional-hazards model for OS.

What does a risk difference of 11.4 mean?

The ORR estimate is a percent difference of 11.4 between the randomized groups. On the proportion scale, this represents an estimated 11.4-percentage-point difference in response rates between pembrolizumab plus BSC and placebo plus BSC. It is not a hazard ratio and should not be interpreted as a relative 11.4% change in the instantaneous risk of response.

Why can a confidence interval and p-value tell different parts of the story?

The confidence interval describes the statistical precision around the effect estimate, whereas the p-value addresses evidence against the specified null hypothesis under the analysis framework. A p-value does not tell us how large or clinically important an effect is. The effect estimate and its confidence interval are therefore essential for understanding the magnitude and uncertainty of the treatment comparison.

7. Primary Result: Overall Survival

The registry reports a formal statistical comparison of OS through approximately 4 years. All randomized participants were analyzed according to their randomized treatment group. The reported Cox analysis used stratification by macrovascular invasion, α-Fetoprotein, and region.

Hazard ratio for overall survival

0.79

95% CI: 0.63–0.99   ·   P = 0.0180   ·   One-sided p-value

Comparison: Pembrolizumab + BSC vs Placebo + BSC

Primary endpointPembrolizumab + BSCPlacebo + BSCEffect estimateP-value
Overall Survival Randomized analysis population Randomized analysis population HR 0.79 (95% CI 0.63–0.99) 0.0180
Clinical Biostats interpretation

The estimated HR of 0.79 indicates an approximately 21% lower estimated hazard of death for pembrolizumab plus BSC relative to placebo plus BSC under the reported Cox model. The 21% figure is obtained directly from the hazard-ratio scale as 1 − 0.79.

The HR does not mean that 21% of participants benefited, that survival increased by 21%, or that every participant had a 21% lower probability of death. A hazard ratio is a relative time-to-event measure based on a statistical model.

The 95% CI of 0.63–0.99 describes uncertainty around the estimated hazard ratio under the model and sampling framework. The interval is relatively close to 1 at its upper boundary, so the estimate should not be treated as if it were known with arbitrary precision.

The p-value of 0.0180 is evidence against the specified null hypothesis under the reported one-sided testing framework. It does not measure the size of the treatment effect, the probability that the treatment works, or the probability that the null hypothesis is true.

Because the analysis uses a Cox proportional-hazards model, interpretation of a single HR also depends on the proportional-hazards framework. The ClinicalTrials.gov record does not provide a separate assessment of that assumption.

8. Secondary Results

Progression-Free Survival

PFS was analyzed among all randomized participants according to randomized treatment group. The reported Cox model used the same general stratification structure described for the OS analysis.

Hazard ratio for progression or death

0.74

95% CI: 0.60–0.92   ·   P = 0.0032   ·   One-sided p-value

Endpoint: PFS per RECIST 1.1, up to approximately 4 years

An HR of 0.74 corresponds to an approximately 26% lower estimated hazard of progression or death for pembrolizumab plus BSC relative to placebo plus BSC under the reported Cox model. As with OS, this is not an absolute reduction in the probability of progression or death and does not imply that every participant experiences the same relative effect.

Objective Response Rate

ORR was analyzed as a binary endpoint using a score-based method for proportions. The reported percent difference was based on the Miettinen-Nurminen method and was stratified by macrovascular invasion, α-Fetoprotein, and region.

Difference in objective response rate

11.4

95% CI: 6.7–16.0   ·   P = 0.00004

Effect measure: Percent Difference / Risk Difference

The registry specifies the hypotheses as H0: difference in percentage = 0 versus H1: difference in percentage > 0. Thus, the positive estimate of 11.4 represents an 11.4-percentage-point difference in ORR between the randomized groups.

Disease Control Rate

DCR used an analysis population of all randomized participants according to randomized treatment group who achieved CR, PR, or SD for ≥5 weeks prior to evidence of disease progression.

Difference in disease control rate

5.4

95% CI: -4.1–14.8   ·   P = 0.13281

Effect measure: Percent Difference / Risk Difference

The point estimate is positive, but the 95% CI extends from a negative value to a positive value. Under the reported testing framework, the one-sided p-value is 0.13281. The ClinicalTrials.gov record therefore provide an estimate of the between-group difference but do not provide evidence against the stated null hypothesis at conventional significance levels.

Time To Progression

TTP was analyzed using a Cox proportional-hazards model among all randomized participants according to randomized treatment group.

Hazard ratio for time to progression

0.72

95% CI: 0.58–0.90   ·   P = 0.0019   ·   One-sided p-value

Endpoint: TTP per RECIST 1.1, up to approximately 4 years

An HR of 0.72 corresponds to an approximately 28% lower estimated hazard of progression under the reported Cox model. TTP differs conceptually from PFS because the endpoint is time to progression rather than the composite time-to-event outcome described as PFS.

9. Results Summary

EndpointRoleEffect95% CIP-valueMethod
Overall Survival Primary HR 0.79 0.63–0.99 0.0180 Cox proportional-hazards model
Progression Free Survival Secondary HR 0.74 0.60–0.92 0.0032 Cox proportional-hazards model
Objective Response Rate Secondary Difference 11.4 6.7–16.0 0.00004 Miettinen-Nurminen score-based method
Disease Control Rate Secondary Difference 5.4 -4.1–14.8 0.13281 Miettinen-Nurminen score-based method
Time To Progression Secondary HR 0.72 0.58–0.90 0.0019 Cox proportional-hazards model
Important statistical distinction: the confidence intervals for the hazard ratios are two-sided 95% intervals, while the registry reports one-sided p-values for the time-to-event analyses. For ORR and DCR, the registry likewise reports two-sided 95% confidence intervals alongside one-sided p-values for testing. These are different statistical quantities and should not be treated as interchangeable.

10. Safety

The ClinicalTrials.gov record reports serious adverse events by treatment exposure group. The counts are presented as affected participants divided by the corresponding number at risk exactly as reported.

Safety groupSerious adverse eventsAffected / at risk
Pembrolizumab First Course Serious adverse events 76 / 299
Placebo First Course Serious adverse events 31 / 153
Pembrolizumab Second Course Serious adverse events 1 / 12

These safety figures should be interpreted using the exposure labels reported by the registry rather than being converted into randomized-arm event rates. In particular, the registry separately identifies first-course and second-course pembrolizumab exposure.

Safety denominator matters: 76/299, 31/153, and 1/12 are affected-participant counts over stated numbers at risk for the reported safety groups. They should not be substituted for the randomized efficacy population or combined without additional information about exposure and treatment assignment.

11. Stratification and Covariate Adjustment

The statistical analysis notes provide unusually useful detail about how the time-to-event comparisons were constructed. The Cox analyses were stratified by three factors:

Stratification factor
Macrovascular invasion
Yes vs. No, with all cells corresponding to macrovascular invasion = Yes combined.
Stratification factor
α-Fetoprotein
ng/mL < 200 vs. ≥ 200.
Stratification factor
Region
China vs. ex-China.
Model covariate
Treatment
Pembrolizumab + BSC versus placebo + BSC.

Stratification is important because the treatment effect is estimated while respecting differences in the baseline hazard across the specified strata. It is different from simply adding every stratification variable as an ordinary covariate with one common coefficient.

12. Confidence Intervals and Statistical Precision

The reported confidence intervals provide more information than the point estimates alone. For the primary OS endpoint, the estimated HR is 0.79 and the 95% CI is 0.63–0.99. For the secondary time-to-event endpoints, the corresponding intervals are 0.60–0.92 for PFS and 0.58–0.90 for TTP.

OS

The interval 0.63–0.99 spans a range of plausible model-based hazard-ratio values around the estimate of 0.79 under the stated statistical framework.

PFS

The interval 0.60–0.92 accompanies the estimated HR of 0.74 and quantifies uncertainty around the relative treatment effect.

ORR

The 95% CI of 6.7–16.0 describes uncertainty around the estimated 11.4-percentage-point response-rate difference.

DCR

The 95% CI of -4.1–14.8 is wider around the estimate of 5.4 and includes both negative and positive values.

A confidence interval is not a prediction interval for individual patients. For the hazard-ratio endpoints, it describes uncertainty in a model-based relative treatment effect. For ORR and DCR, it describes uncertainty around the estimated between-group difference in proportions.

13. One-Sided Tests and Two-Sided Confidence Intervals

One of the more instructive features of the registry record is the combination of one-sided p-values with two-sided 95% confidence intervals. These quantities answer related but distinct questions.

ComponentReported frameworkWhat it addresses
OS p-valueOne-sided, 0.0180Evidence against the specified null in the superiority direction
OS CITwo-sided 95%, 0.63–0.99Precision and uncertainty around the HR estimate
PFS p-valueOne-sided, 0.0032Evidence against the specified null in the superiority direction
PFS CITwo-sided 95%, 0.60–0.92Precision and uncertainty around the HR estimate
ORR p-valueOne-sided, 0.00004Testing whether the response-rate difference exceeds zero in the specified direction
ORR CITwo-sided 95%, 6.7–16.0Precision around the estimated response-rate difference
DCR p-valueOne-sided, 0.13281Testing whether the DCR difference exceeds zero in the specified direction
DCR CITwo-sided 95%, -4.1–14.8Precision around the estimated DCR difference

The presence of a p-value should therefore never replace examination of the corresponding effect estimate and confidence interval.

14. Multiplicity, Interim Analysis, and Other Design Features

The ClinicalTrials.gov record identifies one registered primary endpoint and report five statistical analyses: one primary analysis and four secondary analyses. The primary hypothesis type is recorded as superiority.

Design topicWhat the ClinicalTrials.gov record supports
Primary endpoint count1 registered primary endpoint: Overall Survival
Hypothesis typeSuperiority for the primary OS analysis
Interim analysisThe ClinicalTrials.gov record does not report an interim-analysis procedure.
Alpha-spendingThe ClinicalTrials.gov record does not report an alpha-spending procedure.
Multiplicity adjustmentThe ClinicalTrials.gov record does not specify a multiplicity-adjustment strategy across the reported endpoints.
Non-inferiority marginNot reported in the ClinicalTrials.gov record.
CrossoverNot reported in the ClinicalTrials.gov record.
Factorial designNot reported; the design model is parallel.
Bayesian methodsNot reported.
Missing-data imputationNot reported in the ClinicalTrials.gov record.

This distinction is important. The registry provides formal endpoint-level analyses, but it does not provide enough information in the ClinicalTrials.gov record to reconstruct a full multiplicity hierarchy, interim-monitoring plan, missing-data strategy, or Bayesian component. Those features should not be inferred from the observed p-values.

15. What the Primary Hazard Ratio Does — and Does Not — Mean

Effect size

The primary OS estimate of HR 0.79 means the fitted Cox model estimates the instantaneous hazard of death in the pembrolizumab plus BSC group at approximately 79% of the corresponding hazard in the placebo plus BSC group.

What it does not mean

It does not mean that 21% of participants survived when they otherwise would have died, that every participant had a 21% lower probability of death, or that the median survival time was reduced or increased by 21%. None of those quantities is reported by the HR itself.

Precision

The 95% CI of 0.63–0.99 shows that the point estimate should be interpreted together with uncertainty. The upper endpoint is close to 1, so the estimated magnitude should not be treated as exact.

P-value

The one-sided p = 0.0180 evaluates evidence against the stated null hypothesis under the reported testing framework. It does not quantify clinical importance or the probability that the observed effect will be reproduced in every future population.

16. Comparing Relative and Absolute Measures

The KEYNOTE-394 registry results illustrate why a statistical analysis should present the appropriate effect measure for each endpoint rather than forcing all outcomes onto one scale.

Endpoint typeEffect measureInterpretation scale
Overall SurvivalHazard ratioRelative time-to-event effect
Progression Free SurvivalHazard ratioRelative time-to-event effect
Objective Response RateRisk difference / percent differenceAbsolute difference in response proportions
Disease Control RateRisk difference / percent differenceAbsolute difference in disease-control proportions
Time To ProgressionHazard ratioRelative time-to-event effect

A hazard ratio and a risk difference cannot be directly compared as though they were different versions of the same number. The HR describes a relative event-rate relationship over time, while the risk difference describes an absolute difference in proportions.

17. Analysis Population and Randomization

All reported efficacy analyses in the ClinicalTrials.gov record uses the population of all randomized participants based on the treatment group to which they were randomized. This is especially important for causal interpretation.

Randomized assignment

Randomization establishes the treatment groups before outcome analysis and provides the foundation for comparing outcomes by assigned treatment.

Analysis by randomized group

Analyzing participants according to randomized treatment preserves that original comparison rather than redefining groups according to later exposure.

Time-to-event censoring

Participants who have not experienced the event can still contribute follow-up information through their censoring time.

Safety is different

The ClinicalTrials.gov record uses explicitly reported exposure groups and denominators rather than the efficacy analysis population.

18. Limitations

19. Why This Trial Matters Statistically

KEYNOTE-394 is a useful teaching case because the registry results connect several core clinical-trial concepts in a single randomized phase 3 analysis: a time-to-event primary endpoint, Kaplan-Meier estimation, stratified Cox regression, hazard ratios, confidence intervals, one-sided hypothesis testing, and score-based methods for binary outcomes.

ConceptHow it appears in KEYNOTE-394
RandomizationThe trial is randomized with two parallel groups.
BlindingThe registry identifies the masking as double.
Kaplan-Meier estimationOS is defined using the Kaplan-Meier method for censored data.
Hazard ratioOS, PFS, and TTP are reported using hazard ratios from Cox models.
Cox proportional-hazards modelUsed for the reported time-to-event comparisons.
Stratified analysisModels are stratified by macrovascular invasion, α-Fetoprotein, and region.
Confidence intervalsTwo-sided 95% intervals accompany the reported effect estimates.
One-sided p-valuesReported for the formal testing of OS, PFS, ORR, DCR, and TTP analyses.
Risk differenceORR and DCR are expressed as percent differences between groups.
Score-based CIThe Miettinen-Nurminen method is used for the binary endpoint comparisons.
ITT-style efficacy analysisAll randomized participants are analyzed according to randomized treatment group.

20. Related Tutorials

Learn more about the methods used in this trial:

21. Related Statistical Calculators

22. Sources

Continue through the Clinical Biostats statistical library

Explore the statistical concepts behind randomized trials, survival analysis, confidence intervals, and clinical-trial effect measures.

23. Record Summary

KEYNOTE-394 provides a compact example of how a randomized phase 3 clinical trial can combine time-to-event analysis and binary-outcome analysis within the same statistical framework. The primary endpoint, overall survival, was analyzed using Kaplan-Meier methodology and a stratified Cox proportional-hazards model, producing an HR of 0.79 with a two-sided 95% CI of 0.63–0.99 and a reported one-sided p-value of 0.0180. Secondary analyses extended the same framework to PFS and TTP while using a score-based method for the ORR and DCR proportion comparisons.

The most important statistical lesson is that these estimates operate on different scales. Hazard ratios describe relative time-to-event effects, while percent differences describe absolute differences in binary outcome proportions. Confidence intervals describe precision, p-values address the stated hypothesis tests, and neither should be interpreted as a direct measure of individual patient benefit.

Clinical Biostats methodology: A trial-results page should distinguish the reported registry evidence from statistical interpretation. Where the ClinicalTrials.gov record does not provide a result, analysis detail, or design feature, the page does not infer it from the trial's reputation or from general clinical knowledge.