← Clinical Trials
Gastric / GEJ Adenocarcinoma Phase 3 Completed NCT03615326

KEYNOTE-811: Complete Statistical Analysis of Pembrolizumab Plus Trastuzumab and Chemotherapy in HER2+ Gastric or GEJ Adenocarcinoma

An independent statistical review of the randomized phase 3 KEYNOTE-811 trial evaluating pembrolizumab plus trastuzumab and chemotherapy versus standard of care in participants with HER2-positive advanced gastric or gastroesophageal junction adenocarcinoma.

Trial start: 2018-10-05  ·  Primary completion: 2024-03-20  ·  Enrollment: 738
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

KEYNOTE-811 was a randomized, quadruple-masked, parallel phase 3 treatment trial with 738 enrolled participants. The registered primary endpoints were progression-free survival assessed by blinded independent central review and overall survival, both analyzed as time-to-event outcomes.

738
Enrolled
4-arm trial design
2
Primary endpoints
PFS and OS
0.73
PFS HR
95% CI 0.61–0.87
0.80
OS HR
95% CI 0.67–0.94
FeatureKEYNOTE-811
TrialKEYNOTE-811
PhasePhase 3
StatusCompleted
PopulationHER2-positive advanced gastric or gastroesophageal junction adenocarcinoma
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment738
Primary endpointsProgression-free survival per RECIST 1.1 assessed by BICR; overall survival
Results postedYes
Statistical analyses posted3
ClinicalTrials.govNCT03615326
Lead sponsorMerck Sharp & Dohme LLC

2. Clinical Question

The primary statistical question was whether adding pembrolizumab to trastuzumab plus chemotherapy was superior to trastuzumab plus chemotherapy alone for the registered time-to-event endpoints in the global cohort.

Population

Participants with HER2-positive advanced gastric or gastroesophageal junction adenocarcinoma enrolled in the phase 3 KEYNOTE-811 trial.

Intervention

Pembrolizumab in combination with trastuzumab plus chemotherapy, represented in the registry as the pembrolizumab first-course treatment strategy.

Comparator

Standard of care consisting of trastuzumab plus chemotherapy, represented in the registry as the SOC arm.

Primary question

Does pembrolizumab in combination with trastuzumab plus chemotherapy provide superior PFS and OS compared with trastuzumab plus chemotherapy alone?

3. Trial Design

01
Randomize 738 enrolled
02
Parallel groups Randomized treatment comparison
03
Quadruple masked Study masking
04
Assess PFS, OS, response, safety
05
Analyze Stratified time-to-event methods
Allocation
Randomized. Randomization provides the design foundation for comparing outcomes between treatment strategies.
Masking
Quadruple masking. The registry identifies the study as quadruple masked.
Model
Parallel. Participants were evaluated within parallel randomized treatment groups rather than a crossover or factorial framework.
Primary purpose
Treatment. The registered primary purpose was treatment.
GLOBAL PEMBROLIZUMAB FIRST COURSE

Pembrolizumab + Standard of Care

  • Pembrolizumab
  • Trastuzumab
  • Chemotherapy
GLOBAL CONTROL

Standard of Care

  • Placebo
  • Trastuzumab
  • Chemotherapy

The registered intervention list includes pembrolizumab, placebo, cisplatin, 5-FU, oxaliplatin, capecitabine, S-1, and trastuzumab. The statistical analyses reported on this page compare the global pembrolizumab first-course group with the global standard-of-care group, as specified in the posted analyses.

4. Analysis Populations and Cohorts

The posted primary analyses were defined around the Global Cohort. Participants in the Japan-specific SOX Cohort were excluded from the efficacy analysis according to the statistical analysis plan described in the registry.

Population / cohortRole in the posted analysis
Global Pembrolizumab First CourseIncluded in the primary PFS and OS efficacy analyses.
Global Standard of CareComparator included in the primary PFS and OS efficacy analyses.
Japan-specific SOX Cohort were not included in the efficacy analysis.Not included in the efficacy analysis according to the statistical analysis plan.
Randomized Global Cohort participantsThe analysis population described for the posted primary endpoint comparisons.
Why this distinction matters: the enrollment number of 738 describes the trial as a whole, whereas the posted efficacy analyses are specifically described for randomized participants in the Global Cohort. The registry therefore should not be read as though every enrolled participant necessarily contributed to every reported efficacy comparison.

5. Primary Endpoints

EndpointRegistry definitionTime frameEndpoint type
Progression Free Survival (PFS) Per RECIST 1.1 Assessed by BICR PFS is defined as the time from randomization to the first documented disease progression per RECIST 1.1 as assessed by BICR or death due to any cause, whichever occurs first. Per RECIST 1.1, progressive disease is defined as at least a 20% increase in the sum of diameters of target lesions, taking as reference the smallest sum on study. Up to 46 months Time-to-event
Overall Survival (OS) OS is defined as the time from randomization to death due to any cause. Up to 63 months Time-to-event

The registry specifies two primary endpoints and identifies both as time-to-event outcomes. This is important statistically: neither endpoint is adequately represented by a simple comparison of percentages at one arbitrary time point. The primary framework instead uses the ordering and timing of events over follow-up.

6. Statistical Methodology

Stratified log-rank testing

The posted PFS and OS analyses used a log-rank test. The analysis notes specify a one-sided p-value based on a log-rank test stratified by geographic region, PD-L1 status at baseline, and chemotherapy regimen.

Primary time-to-event comparison
Randomization → progression or death for PFS
Randomization → death for OS

The log-rank framework compares the observed timing of events between randomized groups while accounting for the fact that participants can have different lengths of follow-up.

Stratified Cox regression

The hazard ratio and its 95% confidence interval were estimated using a stratified Cox regression model with Efron's method of tie handling and treatment as a covariate. The model was stratified by geographic region, PD-L1 status (positive versus negative) at baseline, and chemotherapy regimen (FP or CAPOX).

Hazard-ratio framework
HR = estimated hazard in the pembrolizumab first-course group ÷ estimated hazard in the standard-of-care group

An HR below 1 indicates a lower estimated instantaneous event rate in the pembrolizumab first-course group under the fitted time-to-event model.

Score-based confidence interval for ORR

The secondary ORR analysis used the Miettinen and Nurminen method, reported in the registry as a stratified method for estimating the difference in percentage and its 95% confidence interval. Stratification used geographic region, PD-L1 status at baseline, and chemotherapy regimen.

One-sided hypothesis testing

The primary PFS and OS analysis notes specify a one-sided p-value from the stratified log-rank test. This is distinct from the confidence intervals reported for the hazard ratios, which are explicitly identified as 95% two-sided confidence intervals.

Do not mix these quantities: a one-sided p-value and a two-sided 95% confidence interval are not contradictory. They answer different statistical questions and reflect different conventions in the prespecified testing and estimation framework.

7. Primary Result: Progression-Free Survival

The first primary endpoint was PFS per RECIST 1.1 assessed by BICR, with a registered time frame of up to 46 months.

Hazard ratio for progression or death

0.73

95% CI: 0.61–0.87   ·   One-sided P = 0.0002

Global Pembrolizumab + Standard of Care First Course vs Global Standard of Care

FeaturePosted PFS analysis
Analysis populationAll randomized Global Cohort participants in the Pembrolizumab First Course arm and SOC arm
EndpointPFS per RECIST 1.1 assessed by BICR
Time frameUp to 46 months
MethodLog-rank test
Effect measureHazard ratio
Estimate0.73
95% CI0.61–0.87
P-value0.0002
HypothesisSuperiority
Clinical Biostats interpretation

The estimated HR of 0.73 means that the fitted analysis estimates approximately a 27% lower instantaneous hazard of progression or death for the pembrolizumab first-course group relative to the standard-of-care group, because 1 − 0.73 = 0.27.

The HR does not mean that 27% of participants avoided progression, that every participant experienced exactly a 27% reduction in risk, or that PFS was increased by 27% in months. It is a relative time-to-event measure derived from the hazard model.

The 95% CI of 0.61–0.87 describes statistical uncertainty around the estimated hazard ratio. Because the interval lies below 1, the posted confidence interval is consistent with a lower estimated hazard in the pembrolizumab first-course group under the model.

The P = 0.0002 is evidence against the prespecified null hypothesis under the reported one-sided testing framework. It is not a measure of how large the treatment effect is. The magnitude of effect is described by the HR and its confidence interval.

Interpretation also depends on the proportional-hazards framework underlying the Cox HR. A single HR summarizes relative instantaneous hazards over the analyzed follow-up; it should not automatically be translated into a constant percentage difference in cumulative risk at every time point.

8. Primary Result: Overall Survival

The second primary endpoint was overall survival, defined as the time from randomization to death due to any cause, with a registered time frame of up to 63 months.

Hazard ratio for death

0.80

95% CI: 0.67–0.94   ·   One-sided P = 0.0040

Global Pembrolizumab + Standard of Care First Course vs Global Standard of Care First Course

FeaturePosted OS analysis
Analysis populationAll randomized Global Cohort participants in the Pembrolizumab First Course arm and SOC arm
EndpointOverall survival
DefinitionTime from randomization to death due to any cause
Time frameUp to 63 months
MethodLog-rank test
Effect measureHazard ratio
Estimate0.80
95% CI0.67–0.94
P-value0.0040
HypothesisSuperiority
Clinical Biostats interpretation

The estimated HR of 0.80 corresponds to an approximately 20% lower estimated instantaneous hazard of death in the pembrolizumab first-course group relative to the standard-of-care group, because 1 − 0.80 = 0.20.

This does not mean that 20% of participants were saved, that survival increased by 20%, or that an individual participant's probability of death was reduced by exactly 20%. The HR is a relative time-to-event parameter from the fitted analysis.

The 95% CI of 0.67–0.94 quantifies uncertainty around the estimated HR. Its upper bound remains below 1, so the interval is consistent with a lower estimated hazard of death for the pembrolizumab first-course group under the reported model.

The P = 0.0040 reflects the reported one-sided log-rank testing framework. It addresses statistical evidence against the null hypothesis; it does not quantify the size, clinical importance, or durability of the treatment effect.

As with PFS, interpretation of a Cox HR requires care about the underlying hazard structure and censoring. The HR should not be treated as though it were an absolute survival probability or a universal percentage reduction in risk for every participant.

9. Secondary Endpoint: Objective Response Rate

Objective response rate was a posted secondary endpoint assessed per RECIST 1.1 by BICR, with a time frame of up to 63 months. The analysis compared the percentage of participants with an objective response between the global pembrolizumab first-course group and the global standard-of-care group.

Difference in objective response rate

12.6 percentage points

95% CI: 5.6–19.4   ·   P = 0.00020

Stratified Miettinen and Nurminen method

FeaturePosted ORR analysis
Analysis populationAll randomized Global Cohort participants in the Pembrolizumab First Course arm and SOC arm
EndpointObjective Response Rate per RECIST 1.1 assessed by BICR
Time frameUp to 63 months
MethodStratified Miettinen and Nurminen method
Effect measureDifference in percentage / risk difference
Estimate12.6
95% CI5.6–19.4
P-value0.00020
HypothesisSuperiority
Clinical Biostats interpretation

The estimated risk difference of 12.6 percentage points means that the percentage of participants meeting the registry's objective-response definition was estimated to be 12.6 percentage points higher in the pembrolizumab first-course group than in the standard-of-care group, using the posted stratified analysis.

A risk difference is an absolute contrast rather than a relative one. It should not be interpreted as a 12.6% relative increase in response, and it does not describe duration of response or survival.

The 95% CI of 5.6–19.4 percentage points expresses uncertainty around the estimated difference. The interval remains above zero, which is consistent with a positive difference under the reported estimation framework.

The P = 0.00020 addresses the statistical comparison under the reported hypothesis-testing framework. A small p-value does not mean the treatment effect is necessarily large; the estimated difference and its confidence interval provide the information about magnitude and precision.

The Miettinen and Nurminen approach is particularly relevant because response is binary. Unlike a time-to-event HR, the estimand here is a difference between response proportions, with stratification incorporated into the confidence interval and comparison.

10. Comparing the Three Posted Efficacy Analyses

EndpointEffect measureEstimate95% CIP-value
PFSHazard ratio0.730.61–0.870.0002
OSHazard ratio0.800.67–0.940.0040
ORRRisk difference12.6 percentage points5.6–19.40.00020

These estimates should not be treated as interchangeable. PFS and OS are time-to-event endpoints, so their principal effect measure is a hazard ratio. ORR is binary, so the posted analysis reports an absolute difference in percentages. The three estimates therefore describe different aspects of the randomized comparison.

PFS

Addresses the timing of disease progression or death and uses a hazard ratio as the principal relative effect measure.

OS

Addresses time from randomization to death from any cause and uses a hazard ratio for the primary comparison.

ORR

Addresses whether a participant achieved an objective response and uses a risk difference rather than a hazard ratio.

Why all three matter

Time-to-event and binary endpoints provide complementary information and should not be collapsed into a single effect statistic.

11. Statistical Methods Explained

Why was a log-rank test used for PFS and OS?

PFS and OS are time-to-event outcomes. Participants can have different follow-up times, and some may not experience the event during observation. A log-rank test is designed to compare the event-time distributions between groups while accounting for censoring. In KEYNOTE-811, the posted primary analyses used a stratified version of this framework.

What does an HR of 0.73 mean?

An HR of 0.73 means that the fitted model estimates the instantaneous hazard in the pembrolizumab first-course group at approximately 73% of the corresponding hazard in the standard-of-care group. Equivalently, 1 − 0.73 = 0.27, so the estimated relative hazard is 27% lower. It is not a 27% absolute reduction in the probability of progression or death.

Why is the OS HR of 0.80 not the same type of number as the ORR difference of 12.6?

The OS HR compares event hazards over time. The ORR estimate is a difference between response percentages. An HR of 0.80 and a risk difference of 12.6 percentage points have different mathematical meanings and cannot be directly compared by magnitude.

Why was the Miettinen and Nurminen method used for ORR?

ORR is a binary endpoint. The Miettinen and Nurminen method provides a score-based confidence interval for the difference between proportions and can incorporate stratification. The posted analysis therefore uses a method appropriate to the binary nature of the response endpoint rather than applying a survival-analysis method to it.

What does the 95% confidence interval tell us?

The confidence interval describes statistical uncertainty around an estimated treatment effect under the specified analysis framework. For the PFS HR, the interval is 0.61–0.87; for the OS HR it is 0.67–0.94; and for the ORR difference it is 5.6–19.4 percentage points. A confidence interval is not a range containing the true effect with a stated probability after the data have been observed, nor is it a range of effects expected for individual patients.

Why is the one-sided p-value different from the two-sided confidence interval?

The registry explicitly reports one-sided p-values for the primary log-rank tests while reporting 95% two-sided confidence intervals for the hazard ratios. The p-value is tied to the prespecified hypothesis-testing direction; the confidence interval provides a two-sided description of estimation uncertainty. They should therefore be interpreted according to their respective statistical roles.

Why does stratification matter?

The posted analyses stratified the primary time-to-event testing and Cox modeling by geographic region, PD-L1 status at baseline, and chemotherapy regimen. Stratification allows the comparison to account for these prespecified factors without treating their effects as though they were identical across all strata. It also aligns the efficacy analysis with important aspects of the trial's randomized design.

12. Covariate Adjustment and Stratification

The registry's analysis notes identify covariate adjustment and stratified analysis as concepts in both primary endpoint analyses. The posted Cox model used treatment as a covariate and was stratified by geographic region, baseline PD-L1 status, and chemotherapy regimen.

Stratification factorRegistry specificationStatistical role
Geographic regionIncluded in stratificationAllows baseline hazard structures to differ across geographic strata.
PD-L1 statusPositive vs negative at baselineAllows the time-to-event comparison to be stratified by baseline PD-L1 status.
Chemotherapy regimenFP or CAPOXAllows the analysis to account for the chemotherapy regimen stratum.

Stratification does not mean that the treatment effect is separately estimated as a distinct primary result within every stratum. Rather, the stratified analysis uses the prespecified strata when estimating and comparing the treatment groups.

13. One-Sided Testing and Superiority

The registered primary analyses are identified as superiority hypotheses. The analysis notes specify one-sided p-values from stratified log-rank tests.

Direction of the primary hypothesis
Superiority hypothesis → pembrolizumab + SOC has a favorable time-to-event comparison versus SOC alone

For PFS and OS, the favorable direction corresponds to a hazard ratio below 1. The posted results are therefore naturally interpreted relative to the null value HR = 1.

A one-sided test is directional: it asks whether the data provide evidence in the prespecified favorable direction rather than merely asking whether the groups differ in either direction. That testing convention should be kept separate from the two-sided confidence intervals used to quantify uncertainty.

14. Safety Results

The trial data provide serious adverse event counts by arm. These are reported as affected participants divided by the number at risk in the corresponding arm or cohort.

Arm / cohortSerious adverse eventsInterpretation of reported quantity
Global Pembrolizumab + Standard of Care163 / 350163 affected participants among 350 at risk
Global Standard of Care159 / 346159 affected participants among 346 at risk
Japan Pembrolizumab + Trastuzumab + S-110 / 2010 affected participants among 20 at risk
Japan Trastuzumab + S-1 Plus Oxaliplatin9 / 209 affected participants among 20 at risk
Global Pembrolizumab + Standard of Care0 / 110 affected participants among 11 at risk

The registry data contain multiple serious-adverse-event entries, including global and Japan-specific cohorts. They should therefore not be combined into a single overall safety rate without knowing the exact population represented by each registry entry.

Safety interpretation: the serious-adverse-event counts are descriptive arm-level safety information. They are not a substitute for a complete adverse-event table, and the ClinicalTrials.gov record do not supply enough information to construct additional safety comparisons beyond the reported affected/at-risk counts.

15. Why Censoring Matters for PFS and OS

Both primary endpoints are time-to-event outcomes. In a typical time-to-event analysis, a participant who has not experienced the relevant event by the end of available follow-up does not simply disappear from the analysis. Instead, the participant contributes information up to the point at which follow-up ends or censoring occurs.

Conceptual survival function
S(t) = probability of remaining event-free beyond time t

For PFS, the event is the first qualifying progression or death. For OS, the event is death from any cause.

This is why a time-to-event analysis can use different follow-up durations across participants. It also explains why a hazard ratio cannot be reconstructed simply by dividing two percentages at one time point.

16. Proportional-Hazards Interpretation

The primary effect measure for PFS and OS was the hazard ratio from a stratified Cox regression model. This makes the proportional-hazards framework an important interpretive consideration.

What the model-based HR means

An HR below 1 represents a lower estimated instantaneous event rate in the pembrolizumab first-course group relative to the standard-of-care group. The model summarizes relative event rates over the analyzed follow-up.

What the HR does not mean

The HR is not an absolute risk reduction, a relative reduction in the probability of an event at a particular time, a ratio of median survival times, or a statement that every individual participant experienced the same proportional benefit.

Why proportional hazards matter

If the relative hazards change substantially over time, a single HR can compress a more complicated time-varying treatment effect into one summary number. The posted trial data provide the HRs and their confidence intervals but do not provide enough underlying event-time data to independently evaluate the proportional-hazards assumption here.

17. Missing Data and Imputation

The ClinicalTrials.gov record identifies the analysis populations, endpoints, stratification factors, and statistical methods, but they do not provide a missing-data or imputation strategy for the primary analyses.

For PFS and OS, censoring is intrinsic to the time-to-event framework and should not automatically be described as conventional missing-data imputation. For ORR, missing or unevaluable response assessments can affect the estimand and analysis population, but the ClinicalTrials.gov record does not specify an imputation rule. Accordingly, no additional imputation method is attributed to KEYNOTE-811 on this page.

Methodological boundary: absence of a stated imputation method in the ClinicalTrials.gov record is not evidence that no missing-data procedures existed in the complete statistical analysis plan. It means only that such a procedure is not reported in the material used for this page.

18. Multiplicity and the Two Primary Endpoints

KEYNOTE-811 registered two primary endpoints: PFS and OS. Both have posted formal analyses and both are designated superiority hypotheses.

Primary endpointFormal analysisEffect measureHypothesis
PFSYesHazard ratioSuperiority
OSYesHazard ratioSuperiority

Two primary endpoints create a multiplicity question because more than one confirmatory outcome is being evaluated. The ClinicalTrials.gov record identifies the endpoints and their p-values but do not provide a complete alpha-allocation or hierarchical testing procedure. Therefore, this page does not infer an unreported multiplicity adjustment.

Statistical caution: the reported p-values should be interpreted in the context of the prespecified statistical analysis plan. A reader should not assume that simply observing two p-values establishes a particular familywise-error procedure unless that procedure is documented.

19. Interim Analysis

The ClinicalTrials.gov record identifies the posted analyses and their methods but do not report an interim-analysis schedule, alpha-spending function, stopping boundary, or information fraction.

Because those design details are not included in the ClinicalTrials.gov record, no interim-analysis procedure is attributed to KEYNOTE-811 here. In general, an interim efficacy analysis requires prespecified control of the type I error if repeated looks at the accumulating data are used for confirmatory decision-making.

20. Non-Inferiority, Equivalence, and Crossover

The primary hypotheses in the statistical analyses posted on ClinicalTrials.gov are explicitly superiority. No non-inferiority margin or equivalence margin is reported in the trial data.

Non-inferiority

No non-inferiority margin is provided, and the primary analyses are not identified as non-inferiority analyses.

Equivalence

No equivalence margin or equivalence hypothesis is provided in the ClinicalTrials.gov record.

Crossover

The ClinicalTrials.gov record does not report a crossover design or crossover analysis for the primary endpoints.

Factorial design

The trial is identified as a parallel design; no factorial analysis is reported in the ClinicalTrials.gov record.

21. Results by Endpoint: What Can and Cannot Be Concluded

QuestionSupported by the ClinicalTrials.gov record?Reason
Is there a posted PFS treatment-effect estimate?YesHR 0.73 with 95% CI 0.61–0.87 and P = 0.0002.
Is there a posted OS treatment-effect estimate?YesHR 0.80 with 95% CI 0.67–0.94 and P = 0.0040.
Is there a posted ORR comparison?YesRisk difference 12.6 percentage points with 95% CI 5.6–19.4 and P = 0.00020.
Are median PFS and OS values available?NoThey are not included in the ClinicalTrials.gov record.
Are subgroup-specific efficacy estimates available?NoThe statistical analyses posted on ClinicalTrials.gov do not provide subgroup estimates.
Are complete baseline characteristics available?NoNo baseline table is included in the ClinicalTrials.gov record.
Is a complete adverse-event profile available?NoOnly specified serious-adverse-event affected/at-risk counts are provided.

22. Clinical Biostats Interpretation of the Effect Measures

PFS: relative effect

The PFS HR of 0.73 is a relative time-to-event estimate. Its interpretation is tied to progression or death and to the stratified Cox model. It should not be converted into a median PFS difference because no median PFS values are provided in the ClinicalTrials.gov record.

OS: relative effect

The OS HR of 0.80 summarizes the relative hazard of death under the reported model. It does not provide an absolute survival probability at any particular time. Such an absolute interpretation would require survival estimates that are not included in the ClinicalTrials.gov record.

ORR: absolute effect

The ORR risk difference of 12.6 percentage points is directly interpretable as an absolute difference in the proportion of participants meeting the response definition. Its 95% CI of 5.6–19.4 percentage points provides the corresponding uncertainty interval.

P-values: evidence, not magnitude

The PFS, OS, and ORR p-values provide evidence against the corresponding null hypotheses under the reported statistical frameworks. They do not rank the three endpoints, quantify clinical importance, or replace the effect estimates and confidence intervals.

23. Limitations

24. Why This Trial Matters Statistically

KEYNOTE-811 is a useful statistical teaching case because its posted results combine randomized treatment allocation, masking, two primary time-to-event endpoints, blinded central assessment for PFS, stratified log-rank testing, stratified Cox regression, one-sided superiority testing, and a score-based confidence interval method for a binary response endpoint.

ConceptHow it appears in KEYNOTE-811
RandomizationThe trial is registered as randomized.
MaskingThe registry identifies the study as quadruple masked.
Parallel designThe design model is registered as parallel.
Time-to-event endpointsPFS and OS are the two primary endpoints.
Blinded central reviewPFS is assessed by BICR.
Log-rank testingUsed for the posted primary PFS and OS analyses.
Hazard ratioUsed for the PFS and OS treatment-effect estimates.
Cox regressionStratified Cox regression was used to estimate HRs and 95% CIs.
StratificationGeographic region, baseline PD-L1 status, and chemotherapy regimen were used as strata.
One-sided testingThe posted primary log-rank p-values are one-sided.
Binary endpoint analysisORR was analyzed with a stratified Miettinen and Nurminen method.
Risk differenceThe ORR treatment effect was reported as a difference in percentage.
Confidence intervals95% two-sided CIs were reported for the primary HR estimates and ORR difference.
Superiority testingThe primary PFS and OS analyses are identified as superiority hypotheses.

25. Statistical Methods: A Deeper View

Why the PFS analysis is more than a response-rate comparison

PFS records when progression or death occurs, not merely whether it eventually occurs. A participant who remains event-free for a longer period contributes a different pattern of information from a participant who experiences an event early. The Kaplan-Meier and Cox frameworks are designed around that time dimension.

Why BICR is statistically relevant

The PFS endpoint was assessed by blinded independent central review. Independent blinded assessment can reduce the influence of knowledge of treatment assignment on radiologic endpoint determination. It does not eliminate all sources of uncertainty, but it establishes a more controlled framework for evaluating progression.

Why the stratified Cox model is different from simply fitting an unadjusted HR

The posted analysis did not merely compare crude event rates. Treatment was included as a covariate while baseline geographic region, PD-L1 status, and chemotherapy regimen were used as strata. This permits the baseline hazard structure to vary by stratum while estimating the treatment comparison across those strata.

Why ORR requires a different statistical method

ORR is binary rather than time-to-event. A participant either meets the response criterion or does not within the specified assessment framework. Consequently, a proportion-based method is appropriate. The posted Miettinen and Nurminen approach targets the difference in percentages and its confidence interval rather than a hazard ratio.

Why statistical significance and clinical importance are separate concepts

The p-values indicate evidence against null hypotheses within the reported statistical framework. They do not by themselves establish whether an effect is clinically meaningful. Clinical interpretation requires attention to the effect measure, its uncertainty, the endpoint itself, and the context in which the treatment comparison was conducted.

26. Sources

Source boundary: numerical trial results and trial-specific methodological statements on this page are restricted to the registry-reported KEYNOTE-811 trial data. Where a requested trial characteristic was not contained in those data, it has not been reconstructed from outside sources.

27. Related Tutorials

Learn more about the methods used in this trial:

28. Related Calculators

Continue through the Clinical Biostats statistical library

Connect the trial's endpoints and methods to deeper tutorials and statistical calculation tools.

29. Record Summary

KEYNOTE-811 provides a useful example of a modern randomized phase 3 statistical framework. The trial enrolled 738 participants and registered two primary time-to-event endpoints: PFS assessed by BICR and OS. The posted primary analyses used stratified log-rank tests, with hazard ratios and 95% two-sided confidence intervals estimated from stratified Cox regression. The PFS analysis reported an HR of 0.73 (95% CI 0.61–0.87; one-sided P = 0.0002), while the OS analysis reported an HR of 0.80 (95% CI 0.67–0.94; one-sided P = 0.0040).

The posted secondary ORR analysis illustrates a different statistical problem. Rather than a hazard ratio, it reported a 12.6 percentage-point risk difference with a 95% CI of 5.6–19.4 and P = 0.00020 using the stratified Miettinen and Nurminen method. The distinction is important: time-to-event endpoints and binary response endpoints require different estimands and statistical methods.

The most informative statistical reading therefore combines the effect estimates, confidence intervals, p-values, analysis populations, stratification factors, and endpoint-specific methods. Just as importantly, the interpretation should remain within the limits of the available registry data: median survival values, subgroup estimates, complete baseline characteristics, detailed missing-data procedures, and interim-analysis specifications are not reported here and are not inferred.

Clinical Biostats methodology: A trial-results page should distinguish the numerical result from the statistical meaning of that result. For KEYNOTE-811, the central teaching points are the interpretation of hazard ratios, stratified survival analysis, one-sided superiority testing, confidence intervals, and the difference between time-to-event and binary endpoint estimands.