This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
KEYNOTE-811 was a randomized, quadruple-masked, parallel phase 3 treatment trial with 738 enrolled participants. The registered primary endpoints were progression-free survival assessed by blinded independent central review and overall survival, both analyzed as time-to-event outcomes.
| Feature | KEYNOTE-811 |
|---|---|
| Trial | KEYNOTE-811 |
| Phase | Phase 3 |
| Status | Completed |
| Population | HER2-positive advanced gastric or gastroesophageal junction adenocarcinoma |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 738 |
| Primary endpoints | Progression-free survival per RECIST 1.1 assessed by BICR; overall survival |
| Results posted | Yes |
| Statistical analyses posted | 3 |
| ClinicalTrials.gov | NCT03615326 |
| Lead sponsor | Merck Sharp & Dohme LLC |
2. Clinical Question
The primary statistical question was whether adding pembrolizumab to trastuzumab plus chemotherapy was superior to trastuzumab plus chemotherapy alone for the registered time-to-event endpoints in the global cohort.
Population
Participants with HER2-positive advanced gastric or gastroesophageal junction adenocarcinoma enrolled in the phase 3 KEYNOTE-811 trial.
Intervention
Pembrolizumab in combination with trastuzumab plus chemotherapy, represented in the registry as the pembrolizumab first-course treatment strategy.
Comparator
Standard of care consisting of trastuzumab plus chemotherapy, represented in the registry as the SOC arm.
Primary question
Does pembrolizumab in combination with trastuzumab plus chemotherapy provide superior PFS and OS compared with trastuzumab plus chemotherapy alone?
3. Trial Design
Pembrolizumab + Standard of Care
- Pembrolizumab
- Trastuzumab
- Chemotherapy
Standard of Care
- Placebo
- Trastuzumab
- Chemotherapy
The registered intervention list includes pembrolizumab, placebo, cisplatin, 5-FU, oxaliplatin, capecitabine, S-1, and trastuzumab. The statistical analyses reported on this page compare the global pembrolizumab first-course group with the global standard-of-care group, as specified in the posted analyses.
4. Analysis Populations and Cohorts
The posted primary analyses were defined around the Global Cohort. Participants in the Japan-specific SOX Cohort were excluded from the efficacy analysis according to the statistical analysis plan described in the registry.
| Population / cohort | Role in the posted analysis |
|---|---|
| Global Pembrolizumab First Course | Included in the primary PFS and OS efficacy analyses. |
| Global Standard of Care | Comparator included in the primary PFS and OS efficacy analyses. |
| Japan-specific SOX Cohort were not included in the efficacy analysis. | Not included in the efficacy analysis according to the statistical analysis plan. |
| Randomized Global Cohort participants | The analysis population described for the posted primary endpoint comparisons. |
5. Primary Endpoints
| Endpoint | Registry definition | Time frame | Endpoint type |
|---|---|---|---|
| Progression Free Survival (PFS) Per RECIST 1.1 Assessed by BICR | PFS is defined as the time from randomization to the first documented disease progression per RECIST 1.1 as assessed by BICR or death due to any cause, whichever occurs first. Per RECIST 1.1, progressive disease is defined as at least a 20% increase in the sum of diameters of target lesions, taking as reference the smallest sum on study. | Up to 46 months | Time-to-event |
| Overall Survival (OS) | OS is defined as the time from randomization to death due to any cause. | Up to 63 months | Time-to-event |
The registry specifies two primary endpoints and identifies both as time-to-event outcomes. This is important statistically: neither endpoint is adequately represented by a simple comparison of percentages at one arbitrary time point. The primary framework instead uses the ordering and timing of events over follow-up.
6. Statistical Methodology
Stratified log-rank testing
The posted PFS and OS analyses used a log-rank test. The analysis notes specify a one-sided p-value based on a log-rank test stratified by geographic region, PD-L1 status at baseline, and chemotherapy regimen.
Randomization → death for OS
The log-rank framework compares the observed timing of events between randomized groups while accounting for the fact that participants can have different lengths of follow-up.
Stratified Cox regression
The hazard ratio and its 95% confidence interval were estimated using a stratified Cox regression model with Efron's method of tie handling and treatment as a covariate. The model was stratified by geographic region, PD-L1 status (positive versus negative) at baseline, and chemotherapy regimen (FP or CAPOX).
An HR below 1 indicates a lower estimated instantaneous event rate in the pembrolizumab first-course group under the fitted time-to-event model.
Score-based confidence interval for ORR
The secondary ORR analysis used the Miettinen and Nurminen method, reported in the registry as a stratified method for estimating the difference in percentage and its 95% confidence interval. Stratification used geographic region, PD-L1 status at baseline, and chemotherapy regimen.
One-sided hypothesis testing
The primary PFS and OS analysis notes specify a one-sided p-value from the stratified log-rank test. This is distinct from the confidence intervals reported for the hazard ratios, which are explicitly identified as 95% two-sided confidence intervals.
7. Primary Result: Progression-Free Survival
The first primary endpoint was PFS per RECIST 1.1 assessed by BICR, with a registered time frame of up to 46 months.
Hazard ratio for progression or death
95% CI: 0.61–0.87 · One-sided P = 0.0002
Global Pembrolizumab + Standard of Care First Course vs Global Standard of Care
| Feature | Posted PFS analysis |
|---|---|
| Analysis population | All randomized Global Cohort participants in the Pembrolizumab First Course arm and SOC arm |
| Endpoint | PFS per RECIST 1.1 assessed by BICR |
| Time frame | Up to 46 months |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.73 |
| 95% CI | 0.61–0.87 |
| P-value | 0.0002 |
| Hypothesis | Superiority |
The estimated HR of 0.73 means that the fitted analysis estimates approximately a 27% lower instantaneous hazard of progression or death for the pembrolizumab first-course group relative to the standard-of-care group, because 1 − 0.73 = 0.27.
The HR does not mean that 27% of participants avoided progression, that every participant experienced exactly a 27% reduction in risk, or that PFS was increased by 27% in months. It is a relative time-to-event measure derived from the hazard model.
The 95% CI of 0.61–0.87 describes statistical uncertainty around the estimated hazard ratio. Because the interval lies below 1, the posted confidence interval is consistent with a lower estimated hazard in the pembrolizumab first-course group under the model.
The P = 0.0002 is evidence against the prespecified null hypothesis under the reported one-sided testing framework. It is not a measure of how large the treatment effect is. The magnitude of effect is described by the HR and its confidence interval.
Interpretation also depends on the proportional-hazards framework underlying the Cox HR. A single HR summarizes relative instantaneous hazards over the analyzed follow-up; it should not automatically be translated into a constant percentage difference in cumulative risk at every time point.
8. Primary Result: Overall Survival
The second primary endpoint was overall survival, defined as the time from randomization to death due to any cause, with a registered time frame of up to 63 months.
Hazard ratio for death
95% CI: 0.67–0.94 · One-sided P = 0.0040
Global Pembrolizumab + Standard of Care First Course vs Global Standard of Care First Course
| Feature | Posted OS analysis |
|---|---|
| Analysis population | All randomized Global Cohort participants in the Pembrolizumab First Course arm and SOC arm |
| Endpoint | Overall survival |
| Definition | Time from randomization to death due to any cause |
| Time frame | Up to 63 months |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.80 |
| 95% CI | 0.67–0.94 |
| P-value | 0.0040 |
| Hypothesis | Superiority |
The estimated HR of 0.80 corresponds to an approximately 20% lower estimated instantaneous hazard of death in the pembrolizumab first-course group relative to the standard-of-care group, because 1 − 0.80 = 0.20.
This does not mean that 20% of participants were saved, that survival increased by 20%, or that an individual participant's probability of death was reduced by exactly 20%. The HR is a relative time-to-event parameter from the fitted analysis.
The 95% CI of 0.67–0.94 quantifies uncertainty around the estimated HR. Its upper bound remains below 1, so the interval is consistent with a lower estimated hazard of death for the pembrolizumab first-course group under the reported model.
The P = 0.0040 reflects the reported one-sided log-rank testing framework. It addresses statistical evidence against the null hypothesis; it does not quantify the size, clinical importance, or durability of the treatment effect.
As with PFS, interpretation of a Cox HR requires care about the underlying hazard structure and censoring. The HR should not be treated as though it were an absolute survival probability or a universal percentage reduction in risk for every participant.
9. Secondary Endpoint: Objective Response Rate
Objective response rate was a posted secondary endpoint assessed per RECIST 1.1 by BICR, with a time frame of up to 63 months. The analysis compared the percentage of participants with an objective response between the global pembrolizumab first-course group and the global standard-of-care group.
Difference in objective response rate
95% CI: 5.6–19.4 · P = 0.00020
Stratified Miettinen and Nurminen method
| Feature | Posted ORR analysis |
|---|---|
| Analysis population | All randomized Global Cohort participants in the Pembrolizumab First Course arm and SOC arm |
| Endpoint | Objective Response Rate per RECIST 1.1 assessed by BICR |
| Time frame | Up to 63 months |
| Method | Stratified Miettinen and Nurminen method |
| Effect measure | Difference in percentage / risk difference |
| Estimate | 12.6 |
| 95% CI | 5.6–19.4 |
| P-value | 0.00020 |
| Hypothesis | Superiority |
The estimated risk difference of 12.6 percentage points means that the percentage of participants meeting the registry's objective-response definition was estimated to be 12.6 percentage points higher in the pembrolizumab first-course group than in the standard-of-care group, using the posted stratified analysis.
A risk difference is an absolute contrast rather than a relative one. It should not be interpreted as a 12.6% relative increase in response, and it does not describe duration of response or survival.
The 95% CI of 5.6–19.4 percentage points expresses uncertainty around the estimated difference. The interval remains above zero, which is consistent with a positive difference under the reported estimation framework.
The P = 0.00020 addresses the statistical comparison under the reported hypothesis-testing framework. A small p-value does not mean the treatment effect is necessarily large; the estimated difference and its confidence interval provide the information about magnitude and precision.
The Miettinen and Nurminen approach is particularly relevant because response is binary. Unlike a time-to-event HR, the estimand here is a difference between response proportions, with stratification incorporated into the confidence interval and comparison.
10. Comparing the Three Posted Efficacy Analyses
| Endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| PFS | Hazard ratio | 0.73 | 0.61–0.87 | 0.0002 |
| OS | Hazard ratio | 0.80 | 0.67–0.94 | 0.0040 |
| ORR | Risk difference | 12.6 percentage points | 5.6–19.4 | 0.00020 |
These estimates should not be treated as interchangeable. PFS and OS are time-to-event endpoints, so their principal effect measure is a hazard ratio. ORR is binary, so the posted analysis reports an absolute difference in percentages. The three estimates therefore describe different aspects of the randomized comparison.
PFS
Addresses the timing of disease progression or death and uses a hazard ratio as the principal relative effect measure.
OS
Addresses time from randomization to death from any cause and uses a hazard ratio for the primary comparison.
ORR
Addresses whether a participant achieved an objective response and uses a risk difference rather than a hazard ratio.
Why all three matter
Time-to-event and binary endpoints provide complementary information and should not be collapsed into a single effect statistic.
11. Statistical Methods Explained
Why was a log-rank test used for PFS and OS?
PFS and OS are time-to-event outcomes. Participants can have different follow-up times, and some may not experience the event during observation. A log-rank test is designed to compare the event-time distributions between groups while accounting for censoring. In KEYNOTE-811, the posted primary analyses used a stratified version of this framework.
What does an HR of 0.73 mean?
An HR of 0.73 means that the fitted model estimates the instantaneous hazard in the pembrolizumab first-course group at approximately 73% of the corresponding hazard in the standard-of-care group. Equivalently, 1 − 0.73 = 0.27, so the estimated relative hazard is 27% lower. It is not a 27% absolute reduction in the probability of progression or death.
Why is the OS HR of 0.80 not the same type of number as the ORR difference of 12.6?
The OS HR compares event hazards over time. The ORR estimate is a difference between response percentages. An HR of 0.80 and a risk difference of 12.6 percentage points have different mathematical meanings and cannot be directly compared by magnitude.
Why was the Miettinen and Nurminen method used for ORR?
ORR is a binary endpoint. The Miettinen and Nurminen method provides a score-based confidence interval for the difference between proportions and can incorporate stratification. The posted analysis therefore uses a method appropriate to the binary nature of the response endpoint rather than applying a survival-analysis method to it.
What does the 95% confidence interval tell us?
The confidence interval describes statistical uncertainty around an estimated treatment effect under the specified analysis framework. For the PFS HR, the interval is 0.61–0.87; for the OS HR it is 0.67–0.94; and for the ORR difference it is 5.6–19.4 percentage points. A confidence interval is not a range containing the true effect with a stated probability after the data have been observed, nor is it a range of effects expected for individual patients.
Why is the one-sided p-value different from the two-sided confidence interval?
The registry explicitly reports one-sided p-values for the primary log-rank tests while reporting 95% two-sided confidence intervals for the hazard ratios. The p-value is tied to the prespecified hypothesis-testing direction; the confidence interval provides a two-sided description of estimation uncertainty. They should therefore be interpreted according to their respective statistical roles.
Why does stratification matter?
The posted analyses stratified the primary time-to-event testing and Cox modeling by geographic region, PD-L1 status at baseline, and chemotherapy regimen. Stratification allows the comparison to account for these prespecified factors without treating their effects as though they were identical across all strata. It also aligns the efficacy analysis with important aspects of the trial's randomized design.
12. Covariate Adjustment and Stratification
The registry's analysis notes identify covariate adjustment and stratified analysis as concepts in both primary endpoint analyses. The posted Cox model used treatment as a covariate and was stratified by geographic region, baseline PD-L1 status, and chemotherapy regimen.
| Stratification factor | Registry specification | Statistical role |
|---|---|---|
| Geographic region | Included in stratification | Allows baseline hazard structures to differ across geographic strata. |
| PD-L1 status | Positive vs negative at baseline | Allows the time-to-event comparison to be stratified by baseline PD-L1 status. |
| Chemotherapy regimen | FP or CAPOX | Allows the analysis to account for the chemotherapy regimen stratum. |
Stratification does not mean that the treatment effect is separately estimated as a distinct primary result within every stratum. Rather, the stratified analysis uses the prespecified strata when estimating and comparing the treatment groups.
13. One-Sided Testing and Superiority
The registered primary analyses are identified as superiority hypotheses. The analysis notes specify one-sided p-values from stratified log-rank tests.
For PFS and OS, the favorable direction corresponds to a hazard ratio below 1. The posted results are therefore naturally interpreted relative to the null value HR = 1.
A one-sided test is directional: it asks whether the data provide evidence in the prespecified favorable direction rather than merely asking whether the groups differ in either direction. That testing convention should be kept separate from the two-sided confidence intervals used to quantify uncertainty.
14. Safety Results
The trial data provide serious adverse event counts by arm. These are reported as affected participants divided by the number at risk in the corresponding arm or cohort.
| Arm / cohort | Serious adverse events | Interpretation of reported quantity |
|---|---|---|
| Global Pembrolizumab + Standard of Care | 163 / 350 | 163 affected participants among 350 at risk |
| Global Standard of Care | 159 / 346 | 159 affected participants among 346 at risk |
| Japan Pembrolizumab + Trastuzumab + S-1 | 10 / 20 | 10 affected participants among 20 at risk |
| Japan Trastuzumab + S-1 Plus Oxaliplatin | 9 / 20 | 9 affected participants among 20 at risk |
| Global Pembrolizumab + Standard of Care | 0 / 11 | 0 affected participants among 11 at risk |
The registry data contain multiple serious-adverse-event entries, including global and Japan-specific cohorts. They should therefore not be combined into a single overall safety rate without knowing the exact population represented by each registry entry.
15. Why Censoring Matters for PFS and OS
Both primary endpoints are time-to-event outcomes. In a typical time-to-event analysis, a participant who has not experienced the relevant event by the end of available follow-up does not simply disappear from the analysis. Instead, the participant contributes information up to the point at which follow-up ends or censoring occurs.
For PFS, the event is the first qualifying progression or death. For OS, the event is death from any cause.
This is why a time-to-event analysis can use different follow-up durations across participants. It also explains why a hazard ratio cannot be reconstructed simply by dividing two percentages at one time point.
16. Proportional-Hazards Interpretation
The primary effect measure for PFS and OS was the hazard ratio from a stratified Cox regression model. This makes the proportional-hazards framework an important interpretive consideration.
An HR below 1 represents a lower estimated instantaneous event rate in the pembrolizumab first-course group relative to the standard-of-care group. The model summarizes relative event rates over the analyzed follow-up.
The HR is not an absolute risk reduction, a relative reduction in the probability of an event at a particular time, a ratio of median survival times, or a statement that every individual participant experienced the same proportional benefit.
If the relative hazards change substantially over time, a single HR can compress a more complicated time-varying treatment effect into one summary number. The posted trial data provide the HRs and their confidence intervals but do not provide enough underlying event-time data to independently evaluate the proportional-hazards assumption here.
17. Missing Data and Imputation
The ClinicalTrials.gov record identifies the analysis populations, endpoints, stratification factors, and statistical methods, but they do not provide a missing-data or imputation strategy for the primary analyses.
For PFS and OS, censoring is intrinsic to the time-to-event framework and should not automatically be described as conventional missing-data imputation. For ORR, missing or unevaluable response assessments can affect the estimand and analysis population, but the ClinicalTrials.gov record does not specify an imputation rule. Accordingly, no additional imputation method is attributed to KEYNOTE-811 on this page.
18. Multiplicity and the Two Primary Endpoints
KEYNOTE-811 registered two primary endpoints: PFS and OS. Both have posted formal analyses and both are designated superiority hypotheses.
| Primary endpoint | Formal analysis | Effect measure | Hypothesis |
|---|---|---|---|
| PFS | Yes | Hazard ratio | Superiority |
| OS | Yes | Hazard ratio | Superiority |
Two primary endpoints create a multiplicity question because more than one confirmatory outcome is being evaluated. The ClinicalTrials.gov record identifies the endpoints and their p-values but do not provide a complete alpha-allocation or hierarchical testing procedure. Therefore, this page does not infer an unreported multiplicity adjustment.
19. Interim Analysis
The ClinicalTrials.gov record identifies the posted analyses and their methods but do not report an interim-analysis schedule, alpha-spending function, stopping boundary, or information fraction.
Because those design details are not included in the ClinicalTrials.gov record, no interim-analysis procedure is attributed to KEYNOTE-811 here. In general, an interim efficacy analysis requires prespecified control of the type I error if repeated looks at the accumulating data are used for confirmatory decision-making.
20. Non-Inferiority, Equivalence, and Crossover
The primary hypotheses in the statistical analyses posted on ClinicalTrials.gov are explicitly superiority. No non-inferiority margin or equivalence margin is reported in the trial data.
Non-inferiority
No non-inferiority margin is provided, and the primary analyses are not identified as non-inferiority analyses.
Equivalence
No equivalence margin or equivalence hypothesis is provided in the ClinicalTrials.gov record.
Crossover
The ClinicalTrials.gov record does not report a crossover design or crossover analysis for the primary endpoints.
Factorial design
The trial is identified as a parallel design; no factorial analysis is reported in the ClinicalTrials.gov record.
21. Results by Endpoint: What Can and Cannot Be Concluded
| Question | Supported by the ClinicalTrials.gov record? | Reason |
|---|---|---|
| Is there a posted PFS treatment-effect estimate? | Yes | HR 0.73 with 95% CI 0.61–0.87 and P = 0.0002. |
| Is there a posted OS treatment-effect estimate? | Yes | HR 0.80 with 95% CI 0.67–0.94 and P = 0.0040. |
| Is there a posted ORR comparison? | Yes | Risk difference 12.6 percentage points with 95% CI 5.6–19.4 and P = 0.00020. |
| Are median PFS and OS values available? | No | They are not included in the ClinicalTrials.gov record. |
| Are subgroup-specific efficacy estimates available? | No | The statistical analyses posted on ClinicalTrials.gov do not provide subgroup estimates. |
| Are complete baseline characteristics available? | No | No baseline table is included in the ClinicalTrials.gov record. |
| Is a complete adverse-event profile available? | No | Only specified serious-adverse-event affected/at-risk counts are provided. |
22. Clinical Biostats Interpretation of the Effect Measures
The PFS HR of 0.73 is a relative time-to-event estimate. Its interpretation is tied to progression or death and to the stratified Cox model. It should not be converted into a median PFS difference because no median PFS values are provided in the ClinicalTrials.gov record.
The OS HR of 0.80 summarizes the relative hazard of death under the reported model. It does not provide an absolute survival probability at any particular time. Such an absolute interpretation would require survival estimates that are not included in the ClinicalTrials.gov record.
The ORR risk difference of 12.6 percentage points is directly interpretable as an absolute difference in the proportion of participants meeting the response definition. Its 95% CI of 5.6–19.4 percentage points provides the corresponding uncertainty interval.
The PFS, OS, and ORR p-values provide evidence against the corresponding null hypotheses under the reported statistical frameworks. They do not rank the three endpoints, quantify clinical importance, or replace the effect estimates and confidence intervals.
23. Limitations
- Global versus Japan-specific cohorts: the primary efficacy analyses are explicitly described for the Global Cohort, while participants in the Japan-specific SOX Cohort were not included in the efficacy analysis according to the statistical analysis plan.
- Incomplete numerical outcome detail: the ClinicalTrials.gov record provides HRs, confidence intervals, and p-values but do not provide median PFS or OS, Kaplan-Meier time-point estimates, event counts, or survival probabilities.
- Subgroup information: no subgroup-specific efficacy estimates are reported, so consistency of the treatment effect across baseline subgroups cannot be evaluated from these data.
- Safety scope: the ClinicalTrials.gov record is limited to selected serious-adverse-event affected/at-risk counts and does not provide a complete adverse-event profile.
- Multiplicity: two primary endpoints are registered, but the ClinicalTrials.gov record does not provide the complete alpha-allocation or hierarchical testing procedure.
- Interim monitoring: the ClinicalTrials.gov record does not report an interim-analysis schedule, stopping boundary, or alpha-spending procedure.
- Missing data: the ClinicalTrials.gov record does not identify an imputation strategy for the posted endpoints.
- Hazard-ratio interpretation: a Cox HR is model-based and should not be interpreted as a constant percentage reduction in individual risk at every point in time.
- Limited reconstruction: summary statistics alone are insufficient to independently reconstruct the Kaplan-Meier curves or verify model assumptions from patient-level event-time data.
24. Why This Trial Matters Statistically
KEYNOTE-811 is a useful statistical teaching case because its posted results combine randomized treatment allocation, masking, two primary time-to-event endpoints, blinded central assessment for PFS, stratified log-rank testing, stratified Cox regression, one-sided superiority testing, and a score-based confidence interval method for a binary response endpoint.
| Concept | How it appears in KEYNOTE-811 |
|---|---|
| Randomization | The trial is registered as randomized. |
| Masking | The registry identifies the study as quadruple masked. |
| Parallel design | The design model is registered as parallel. |
| Time-to-event endpoints | PFS and OS are the two primary endpoints. |
| Blinded central review | PFS is assessed by BICR. |
| Log-rank testing | Used for the posted primary PFS and OS analyses. |
| Hazard ratio | Used for the PFS and OS treatment-effect estimates. |
| Cox regression | Stratified Cox regression was used to estimate HRs and 95% CIs. |
| Stratification | Geographic region, baseline PD-L1 status, and chemotherapy regimen were used as strata. |
| One-sided testing | The posted primary log-rank p-values are one-sided. |
| Binary endpoint analysis | ORR was analyzed with a stratified Miettinen and Nurminen method. |
| Risk difference | The ORR treatment effect was reported as a difference in percentage. |
| Confidence intervals | 95% two-sided CIs were reported for the primary HR estimates and ORR difference. |
| Superiority testing | The primary PFS and OS analyses are identified as superiority hypotheses. |
25. Statistical Methods: A Deeper View
Why the PFS analysis is more than a response-rate comparison
PFS records when progression or death occurs, not merely whether it eventually occurs. A participant who remains event-free for a longer period contributes a different pattern of information from a participant who experiences an event early. The Kaplan-Meier and Cox frameworks are designed around that time dimension.
Why BICR is statistically relevant
The PFS endpoint was assessed by blinded independent central review. Independent blinded assessment can reduce the influence of knowledge of treatment assignment on radiologic endpoint determination. It does not eliminate all sources of uncertainty, but it establishes a more controlled framework for evaluating progression.
Why the stratified Cox model is different from simply fitting an unadjusted HR
The posted analysis did not merely compare crude event rates. Treatment was included as a covariate while baseline geographic region, PD-L1 status, and chemotherapy regimen were used as strata. This permits the baseline hazard structure to vary by stratum while estimating the treatment comparison across those strata.
Why ORR requires a different statistical method
ORR is binary rather than time-to-event. A participant either meets the response criterion or does not within the specified assessment framework. Consequently, a proportion-based method is appropriate. The posted Miettinen and Nurminen approach targets the difference in percentages and its confidence interval rather than a hazard ratio.
Why statistical significance and clinical importance are separate concepts
The p-values indicate evidence against null hypotheses within the reported statistical framework. They do not by themselves establish whether an effect is clinically meaningful. Clinical interpretation requires attention to the effect measure, its uncertainty, the endpoint itself, and the context in which the treatment comparison was conducted.
26. Sources
- ClinicalTrials.gov: NCT03615326 — KEYNOTE-811.
- PubMed: PMID 37871604.
- PubMed: PMID 39282917.
- PubMed: PMID 34912120.
- PubMed: PMID 33167735.
27. Related Tutorials
Learn more about the methods used in this trial:
28. Related Calculators
Continue through the Clinical Biostats statistical library
Connect the trial's endpoints and methods to deeper tutorials and statistical calculation tools.
29. Record Summary
KEYNOTE-811 provides a useful example of a modern randomized phase 3 statistical framework. The trial enrolled 738 participants and registered two primary time-to-event endpoints: PFS assessed by BICR and OS. The posted primary analyses used stratified log-rank tests, with hazard ratios and 95% two-sided confidence intervals estimated from stratified Cox regression. The PFS analysis reported an HR of 0.73 (95% CI 0.61–0.87; one-sided P = 0.0002), while the OS analysis reported an HR of 0.80 (95% CI 0.67–0.94; one-sided P = 0.0040).
The posted secondary ORR analysis illustrates a different statistical problem. Rather than a hazard ratio, it reported a 12.6 percentage-point risk difference with a 95% CI of 5.6–19.4 and P = 0.00020 using the stratified Miettinen and Nurminen method. The distinction is important: time-to-event endpoints and binary response endpoints require different estimands and statistical methods.
The most informative statistical reading therefore combines the effect estimates, confidence intervals, p-values, analysis populations, stratification factors, and endpoint-specific methods. Just as importantly, the interpretation should remain within the limits of the available registry data: median survival values, subgroup estimates, complete baseline characteristics, detailed missing-data procedures, and interim-analysis specifications are not reported here and are not inferred.