This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
KEYNOTE-394 was a randomized, double-masked, parallel-group phase 3 trial in Asian participants with previously treated advanced hepatocellular carcinoma. The registry reports an enrollment of 453 participants and compares pembrolizumab plus best supportive care with placebo plus best supportive care.
| Feature | KEYNOTE-394 |
|---|---|
| Trial name | KEYNOTE-394 |
| NCT identifier | NCT03062358 |
| Phase | Phase 3 |
| Status | Completed |
| Condition | Carcinoma, Hepatocellular |
| Population | Asian participants with previously treated advanced hepatocellular carcinoma |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Double |
| Primary purpose | Treatment |
| Enrollment | 453 |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | Industry |
2. Clinical Question
The clinical question is whether pembrolizumab plus best supportive care produces a superior overall-survival outcome compared with placebo plus best supportive care in Asian participants with previously treated advanced hepatocellular carcinoma.
Population
Asian participants with previously treated advanced hepatocellular carcinoma.
Intervention
Pembrolizumab given with best supportive care.
Comparator
Placebo given with best supportive care.
Primary question
Does pembrolizumab plus best supportive care improve overall survival relative to placebo plus best supportive care?
3. Trial Design
Pembrolizumab + BSC
- Pembrolizumab
- Best supportive care (BSC)
- Compared with the randomized placebo + BSC group
Placebo + BSC
- Placebo
- Best supportive care (BSC)
- Reference group for the randomized comparisons
The registry identifies the study as randomized, parallel, and double-masked. The primary purpose is treatment. The ClinicalTrials.gov record does not report a factorial structure, crossover procedure, non-inferiority margin, or Bayesian analysis.
Trial timeline
Study start
The registry lists April 27, 2017 as the study start date.
Primary completion
The registry lists June 30, 2021 as the primary completion date.
Registry status
The trial status is listed as completed.
4. Endpoints
The registry identifies Overall Survival (OS) as the single registered primary endpoint. The primary endpoint is a time-to-event outcome, and the registry also reports four secondary statistical analyses.
| Endpoint | Role | Time frame | Definition / analysis |
|---|---|---|---|
| Overall Survival (OS) | Primary | Up to approximately 4 years | OS is the time from randomization to death due to any cause, based on the Kaplan-Meier method for censored data. |
| Progression Free Survival (PFS) Per RECIST 1.1 | Secondary | Up to approximately 4 years | Time-to-event endpoint analyzed with a Cox proportional-hazards model. |
| Objective Response Rate (ORR) Per RECIST 1.1 | Secondary | Up to approximately 4 years | Binary endpoint; effect reported as percent difference. |
| Disease Control Rate (DCR) Per RECIST 1.1 | Secondary | Up to approximately 4 years | Participants achieving CR, PR, or SD for ≥5 weeks prior to evidence of disease progression. |
| Time To Progression (TTP) Per RECIST 1.1 | Secondary | Up to approximately 4 years | Time-to-event endpoint analyzed with a Cox proportional-hazards model. |
5. Statistical Methodology
Primary analysis: Cox proportional-hazards model
The primary OS comparison used a Cox proportional-hazards model. Participants were analyzed in the treatment group to which they were randomized. The reported effect measure was the hazard ratio.
The registry analysis notes describe treatment as a covariate and specify Efron's method for handling tied event times. The model was stratified by macrovascular invasion, α-Fetoprotein, and region, with all cells corresponding to macrovascular invasion = Yes combined.
Stratified analysis
The OS analysis was stratified by macrovascular invasion (Yes vs. No), α-Fetoprotein (ng/mL) (< 200 vs. ≥ 200), and region (China vs. ex-China). This means the treatment comparison was constructed while accounting for these prespecified stratification factors in the reported Cox analysis.
Kaplan-Meier estimation
The registry defines OS using the Kaplan-Meier method for censored data. Kaplan-Meier estimation is appropriate for time-to-event outcomes because it allows participants who have not experienced the event by their last observation to contribute information up to the point of censoring.
Here, di represents the number of events at time ti, while ni is the number at risk immediately before that event time.
Score-based confidence intervals for binary endpoints
ORR and DCR were binary endpoints. The registry analysis identifies a score-based CI for proportions, specifically the Miettinen-Nurminen approach, as the reported method. The analysis was stratified by macrovascular invasion, α-Fetoprotein, and region.
For ORR and DCR, the reported effect measure was a percent difference, corresponding to a risk difference on the proportion scale. A positive difference indicates a higher observed proportion in the pembrolizumab plus BSC group relative to placebo plus BSC.
Analysis populations
For the reported efficacy analyses, the analysis population was all randomized participants based on the treatment group to which they were randomized. This is an intention-to-treat-style approach and preserves the randomized comparison even if subsequent treatment exposure differs between groups.
6. Statistical Methods Explained
Why was a Cox proportional-hazards model used for OS?
OS records the time from randomization until death and therefore contains both an event indicator and a follow-up time. The Cox model summarizes the relative event rate through a hazard ratio while accommodating censoring. It is particularly useful when not every participant experiences the event during the observation period.
What does an OS hazard ratio of 0.79 mean?
An HR of 0.79 means that the estimated hazard of death in the pembrolizumab plus BSC group was 0.79 times the estimated hazard in the placebo plus BSC group under the fitted Cox model. Equivalently, 1 − 0.79 = 0.21, so the estimate corresponds to an approximately 21% lower estimated hazard of death. This is a relative model-based measure; it does not mean that 21% of participants avoided death or that each participant experienced exactly a 21% reduction in risk.
Why does stratification matter?
The reported Cox analysis was stratified by macrovascular invasion, α-Fetoprotein, and region. Stratification allows the baseline hazard structure to differ across the specified strata while estimating the treatment effect across those strata. It therefore avoids treating these factors as if they had a single common baseline hazard relationship in the model.
Why is the ORR analysis different from the OS analysis?
ORR is binary: a participant either meets the prespecified response definition or does not. OS is a time-to-event outcome, where both event occurrence and follow-up duration matter. The registry therefore reports a score-based method for the proportion comparison for ORR and a Cox proportional-hazards model for OS.
What does a risk difference of 11.4 mean?
The ORR estimate is a percent difference of 11.4 between the randomized groups. On the proportion scale, this represents an estimated 11.4-percentage-point difference in response rates between pembrolizumab plus BSC and placebo plus BSC. It is not a hazard ratio and should not be interpreted as a relative 11.4% change in the instantaneous risk of response.
Why can a confidence interval and p-value tell different parts of the story?
The confidence interval describes the statistical precision around the effect estimate, whereas the p-value addresses evidence against the specified null hypothesis under the analysis framework. A p-value does not tell us how large or clinically important an effect is. The effect estimate and its confidence interval are therefore essential for understanding the magnitude and uncertainty of the treatment comparison.
7. Primary Result: Overall Survival
The registry reports a formal statistical comparison of OS through approximately 4 years. All randomized participants were analyzed according to their randomized treatment group. The reported Cox analysis used stratification by macrovascular invasion, α-Fetoprotein, and region.
Hazard ratio for overall survival
95% CI: 0.63–0.99 · P = 0.0180 · One-sided p-value
Comparison: Pembrolizumab + BSC vs Placebo + BSC
| Primary endpoint | Pembrolizumab + BSC | Placebo + BSC | Effect estimate | P-value |
|---|---|---|---|---|
| Overall Survival | Randomized analysis population | Randomized analysis population | HR 0.79 (95% CI 0.63–0.99) | 0.0180 |
The estimated HR of 0.79 indicates an approximately 21% lower estimated hazard of death for pembrolizumab plus BSC relative to placebo plus BSC under the reported Cox model. The 21% figure is obtained directly from the hazard-ratio scale as 1 − 0.79.
The HR does not mean that 21% of participants benefited, that survival increased by 21%, or that every participant had a 21% lower probability of death. A hazard ratio is a relative time-to-event measure based on a statistical model.
The 95% CI of 0.63–0.99 describes uncertainty around the estimated hazard ratio under the model and sampling framework. The interval is relatively close to 1 at its upper boundary, so the estimate should not be treated as if it were known with arbitrary precision.
The p-value of 0.0180 is evidence against the specified null hypothesis under the reported one-sided testing framework. It does not measure the size of the treatment effect, the probability that the treatment works, or the probability that the null hypothesis is true.
Because the analysis uses a Cox proportional-hazards model, interpretation of a single HR also depends on the proportional-hazards framework. The ClinicalTrials.gov record does not provide a separate assessment of that assumption.
8. Secondary Results
Progression-Free Survival
PFS was analyzed among all randomized participants according to randomized treatment group. The reported Cox model used the same general stratification structure described for the OS analysis.
Hazard ratio for progression or death
95% CI: 0.60–0.92 · P = 0.0032 · One-sided p-value
Endpoint: PFS per RECIST 1.1, up to approximately 4 years
An HR of 0.74 corresponds to an approximately 26% lower estimated hazard of progression or death for pembrolizumab plus BSC relative to placebo plus BSC under the reported Cox model. As with OS, this is not an absolute reduction in the probability of progression or death and does not imply that every participant experiences the same relative effect.
Objective Response Rate
ORR was analyzed as a binary endpoint using a score-based method for proportions. The reported percent difference was based on the Miettinen-Nurminen method and was stratified by macrovascular invasion, α-Fetoprotein, and region.
Difference in objective response rate
95% CI: 6.7–16.0 · P = 0.00004
Effect measure: Percent Difference / Risk Difference
The registry specifies the hypotheses as H0: difference in percentage = 0 versus H1: difference in percentage > 0. Thus, the positive estimate of 11.4 represents an 11.4-percentage-point difference in ORR between the randomized groups.
Disease Control Rate
DCR used an analysis population of all randomized participants according to randomized treatment group who achieved CR, PR, or SD for ≥5 weeks prior to evidence of disease progression.
Difference in disease control rate
95% CI: -4.1–14.8 · P = 0.13281
Effect measure: Percent Difference / Risk Difference
The point estimate is positive, but the 95% CI extends from a negative value to a positive value. Under the reported testing framework, the one-sided p-value is 0.13281. The ClinicalTrials.gov record therefore provide an estimate of the between-group difference but do not provide evidence against the stated null hypothesis at conventional significance levels.
Time To Progression
TTP was analyzed using a Cox proportional-hazards model among all randomized participants according to randomized treatment group.
Hazard ratio for time to progression
95% CI: 0.58–0.90 · P = 0.0019 · One-sided p-value
Endpoint: TTP per RECIST 1.1, up to approximately 4 years
An HR of 0.72 corresponds to an approximately 28% lower estimated hazard of progression under the reported Cox model. TTP differs conceptually from PFS because the endpoint is time to progression rather than the composite time-to-event outcome described as PFS.
9. Results Summary
| Endpoint | Role | Effect | 95% CI | P-value | Method |
|---|---|---|---|---|---|
| Overall Survival | Primary | HR 0.79 | 0.63–0.99 | 0.0180 | Cox proportional-hazards model |
| Progression Free Survival | Secondary | HR 0.74 | 0.60–0.92 | 0.0032 | Cox proportional-hazards model |
| Objective Response Rate | Secondary | Difference 11.4 | 6.7–16.0 | 0.00004 | Miettinen-Nurminen score-based method |
| Disease Control Rate | Secondary | Difference 5.4 | -4.1–14.8 | 0.13281 | Miettinen-Nurminen score-based method |
| Time To Progression | Secondary | HR 0.72 | 0.58–0.90 | 0.0019 | Cox proportional-hazards model |
10. Safety
The ClinicalTrials.gov record reports serious adverse events by treatment exposure group. The counts are presented as affected participants divided by the corresponding number at risk exactly as reported.
| Safety group | Serious adverse events | Affected / at risk |
|---|---|---|
| Pembrolizumab First Course | Serious adverse events | 76 / 299 |
| Placebo First Course | Serious adverse events | 31 / 153 |
| Pembrolizumab Second Course | Serious adverse events | 1 / 12 |
These safety figures should be interpreted using the exposure labels reported by the registry rather than being converted into randomized-arm event rates. In particular, the registry separately identifies first-course and second-course pembrolizumab exposure.
11. Stratification and Covariate Adjustment
The statistical analysis notes provide unusually useful detail about how the time-to-event comparisons were constructed. The Cox analyses were stratified by three factors:
Yes vs. No, with all cells corresponding to macrovascular invasion = Yes combined.
ng/mL < 200 vs. ≥ 200.
China vs. ex-China.
Pembrolizumab + BSC versus placebo + BSC.
Stratification is important because the treatment effect is estimated while respecting differences in the baseline hazard across the specified strata. It is different from simply adding every stratification variable as an ordinary covariate with one common coefficient.
12. Confidence Intervals and Statistical Precision
The reported confidence intervals provide more information than the point estimates alone. For the primary OS endpoint, the estimated HR is 0.79 and the 95% CI is 0.63–0.99. For the secondary time-to-event endpoints, the corresponding intervals are 0.60–0.92 for PFS and 0.58–0.90 for TTP.
OS
The interval 0.63–0.99 spans a range of plausible model-based hazard-ratio values around the estimate of 0.79 under the stated statistical framework.
PFS
The interval 0.60–0.92 accompanies the estimated HR of 0.74 and quantifies uncertainty around the relative treatment effect.
ORR
The 95% CI of 6.7–16.0 describes uncertainty around the estimated 11.4-percentage-point response-rate difference.
DCR
The 95% CI of -4.1–14.8 is wider around the estimate of 5.4 and includes both negative and positive values.
A confidence interval is not a prediction interval for individual patients. For the hazard-ratio endpoints, it describes uncertainty in a model-based relative treatment effect. For ORR and DCR, it describes uncertainty around the estimated between-group difference in proportions.
13. One-Sided Tests and Two-Sided Confidence Intervals
One of the more instructive features of the registry record is the combination of one-sided p-values with two-sided 95% confidence intervals. These quantities answer related but distinct questions.
| Component | Reported framework | What it addresses |
|---|---|---|
| OS p-value | One-sided, 0.0180 | Evidence against the specified null in the superiority direction |
| OS CI | Two-sided 95%, 0.63–0.99 | Precision and uncertainty around the HR estimate |
| PFS p-value | One-sided, 0.0032 | Evidence against the specified null in the superiority direction |
| PFS CI | Two-sided 95%, 0.60–0.92 | Precision and uncertainty around the HR estimate |
| ORR p-value | One-sided, 0.00004 | Testing whether the response-rate difference exceeds zero in the specified direction |
| ORR CI | Two-sided 95%, 6.7–16.0 | Precision around the estimated response-rate difference |
| DCR p-value | One-sided, 0.13281 | Testing whether the DCR difference exceeds zero in the specified direction |
| DCR CI | Two-sided 95%, -4.1–14.8 | Precision around the estimated DCR difference |
The presence of a p-value should therefore never replace examination of the corresponding effect estimate and confidence interval.
14. Multiplicity, Interim Analysis, and Other Design Features
The ClinicalTrials.gov record identifies one registered primary endpoint and report five statistical analyses: one primary analysis and four secondary analyses. The primary hypothesis type is recorded as superiority.
| Design topic | What the ClinicalTrials.gov record supports |
|---|---|
| Primary endpoint count | 1 registered primary endpoint: Overall Survival |
| Hypothesis type | Superiority for the primary OS analysis |
| Interim analysis | The ClinicalTrials.gov record does not report an interim-analysis procedure. |
| Alpha-spending | The ClinicalTrials.gov record does not report an alpha-spending procedure. |
| Multiplicity adjustment | The ClinicalTrials.gov record does not specify a multiplicity-adjustment strategy across the reported endpoints. |
| Non-inferiority margin | Not reported in the ClinicalTrials.gov record. |
| Crossover | Not reported in the ClinicalTrials.gov record. |
| Factorial design | Not reported; the design model is parallel. |
| Bayesian methods | Not reported. |
| Missing-data imputation | Not reported in the ClinicalTrials.gov record. |
This distinction is important. The registry provides formal endpoint-level analyses, but it does not provide enough information in the ClinicalTrials.gov record to reconstruct a full multiplicity hierarchy, interim-monitoring plan, missing-data strategy, or Bayesian component. Those features should not be inferred from the observed p-values.
15. What the Primary Hazard Ratio Does — and Does Not — Mean
The primary OS estimate of HR 0.79 means the fitted Cox model estimates the instantaneous hazard of death in the pembrolizumab plus BSC group at approximately 79% of the corresponding hazard in the placebo plus BSC group.
It does not mean that 21% of participants survived when they otherwise would have died, that every participant had a 21% lower probability of death, or that the median survival time was reduced or increased by 21%. None of those quantities is reported by the HR itself.
The 95% CI of 0.63–0.99 shows that the point estimate should be interpreted together with uncertainty. The upper endpoint is close to 1, so the estimated magnitude should not be treated as exact.
The one-sided p = 0.0180 evaluates evidence against the stated null hypothesis under the reported testing framework. It does not quantify clinical importance or the probability that the observed effect will be reproduced in every future population.
16. Comparing Relative and Absolute Measures
The KEYNOTE-394 registry results illustrate why a statistical analysis should present the appropriate effect measure for each endpoint rather than forcing all outcomes onto one scale.
| Endpoint type | Effect measure | Interpretation scale |
|---|---|---|
| Overall Survival | Hazard ratio | Relative time-to-event effect |
| Progression Free Survival | Hazard ratio | Relative time-to-event effect |
| Objective Response Rate | Risk difference / percent difference | Absolute difference in response proportions |
| Disease Control Rate | Risk difference / percent difference | Absolute difference in disease-control proportions |
| Time To Progression | Hazard ratio | Relative time-to-event effect |
A hazard ratio and a risk difference cannot be directly compared as though they were different versions of the same number. The HR describes a relative event-rate relationship over time, while the risk difference describes an absolute difference in proportions.
17. Analysis Population and Randomization
All reported efficacy analyses in the ClinicalTrials.gov record uses the population of all randomized participants based on the treatment group to which they were randomized. This is especially important for causal interpretation.
Randomized assignment
Randomization establishes the treatment groups before outcome analysis and provides the foundation for comparing outcomes by assigned treatment.
Analysis by randomized group
Analyzing participants according to randomized treatment preserves that original comparison rather than redefining groups according to later exposure.
Time-to-event censoring
Participants who have not experienced the event can still contribute follow-up information through their censoring time.
Safety is different
The ClinicalTrials.gov record uses explicitly reported exposure groups and denominators rather than the efficacy analysis population.
18. Limitations
- Limited registry detail: the ClinicalTrials.gov record does not report median OS, median PFS, median TTP, event counts for the efficacy endpoints, or time-specific survival estimates.
- No subgroup estimates reported: although the Cox models are stratified by macrovascular invasion, α-Fetoprotein, and region, the ClinicalTrials.gov record does not provide treatment-effect estimates for individual subgroups.
- No baseline table reported: the ClinicalTrials.gov record does not include a detailed baseline demographic or disease-characteristic comparison.
- Proportional-hazards assumption: Cox-model interpretation depends on the proportional-hazards framework. The ClinicalTrials.gov record does not report a formal assessment of that assumption.
- Multiplicity information: five statistical analyses are reported, but the ClinicalTrials.gov record does not specify a complete multiplicity-control strategy across the primary and secondary endpoints.
- One-sided testing: the reported p-values are one-sided, whereas the confidence intervals are two-sided 95% intervals. These quantities should not be treated as interchangeable.
- Safety denominators: the serious-adverse-event groups are reported using exposure-specific denominators, including a separate second-course pembrolizumab group, and therefore should not be interpreted as simple randomized-arm efficacy denominators.
- Unreported design features: the ClinicalTrials.gov record does not provide an interim-analysis plan, alpha-spending procedure, non-inferiority margin, crossover strategy, Bayesian method, or detailed missing-data/imputation method.
- Population scope: the trial population is specifically described as Asian participants with previously treated advanced hepatocellular carcinoma; the ClinicalTrials.gov record does not establish how the results generalize beyond that population.
19. Why This Trial Matters Statistically
KEYNOTE-394 is a useful teaching case because the registry results connect several core clinical-trial concepts in a single randomized phase 3 analysis: a time-to-event primary endpoint, Kaplan-Meier estimation, stratified Cox regression, hazard ratios, confidence intervals, one-sided hypothesis testing, and score-based methods for binary outcomes.
| Concept | How it appears in KEYNOTE-394 |
|---|---|
| Randomization | The trial is randomized with two parallel groups. |
| Blinding | The registry identifies the masking as double. |
| Kaplan-Meier estimation | OS is defined using the Kaplan-Meier method for censored data. |
| Hazard ratio | OS, PFS, and TTP are reported using hazard ratios from Cox models. |
| Cox proportional-hazards model | Used for the reported time-to-event comparisons. |
| Stratified analysis | Models are stratified by macrovascular invasion, α-Fetoprotein, and region. |
| Confidence intervals | Two-sided 95% intervals accompany the reported effect estimates. |
| One-sided p-values | Reported for the formal testing of OS, PFS, ORR, DCR, and TTP analyses. |
| Risk difference | ORR and DCR are expressed as percent differences between groups. |
| Score-based CI | The Miettinen-Nurminen method is used for the binary endpoint comparisons. |
| ITT-style efficacy analysis | All randomized participants are analyzed according to randomized treatment group. |
20. Related Tutorials
Learn more about the methods used in this trial:
21. Related Statistical Calculators
22. Sources
- ClinicalTrials.gov: NCT03062358 — KEYNOTE-394.
- PubMed: PMID 40486134.
Continue through the Clinical Biostats statistical library
Explore the statistical concepts behind randomized trials, survival analysis, confidence intervals, and clinical-trial effect measures.
23. Record Summary
KEYNOTE-394 provides a compact example of how a randomized phase 3 clinical trial can combine time-to-event analysis and binary-outcome analysis within the same statistical framework. The primary endpoint, overall survival, was analyzed using Kaplan-Meier methodology and a stratified Cox proportional-hazards model, producing an HR of 0.79 with a two-sided 95% CI of 0.63–0.99 and a reported one-sided p-value of 0.0180. Secondary analyses extended the same framework to PFS and TTP while using a score-based method for the ORR and DCR proportion comparisons.
The most important statistical lesson is that these estimates operate on different scales. Hazard ratios describe relative time-to-event effects, while percent differences describe absolute differences in binary outcome proportions. Confidence intervals describe precision, p-values address the stated hypothesis tests, and neither should be interpreted as a direct measure of individual patient benefit.