← Clinical Trials
Endometrial Cancer Phase 3 Completed NCT03517449

KEYNOTE-775: Complete Statistical Analysis of Lenvatinib Plus Pembrolizumab in Advanced Endometrial Cancer

An independent statistical analysis of the randomized phase 3 KEYNOTE-775 trial evaluating lenvatinib plus pembrolizumab versus treatment of physician's choice in participants with advanced endometrial cancer.

Randomized  ·  Parallel design  ·  827 enrolled  ·  Results posted
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

KEYNOTE-775 was a completed randomized phase 3 parallel-group trial in advanced endometrial cancer. The registry reports 827 enrolled participants, two treatment arms, no masking, and a primary purpose of treatment. The trial compared lenvatinib plus pembrolizumab with treatment of physician's choice consisting of doxorubicin or paclitaxel.

827
Enrolled
Phase 3
2
Arms
Parallel design
0.56
All-comer PFS HR
95% CI 0.47–0.66
0.65
All-comer OS HR
95% CI 0.55–0.77
FeatureKEYNOTE-775
Trial nameKEYNOTE-775
NCT IDNCT03517449
PhasePhase 3
StatusCompleted
ConditionEndometrial Neoplasms
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment827
Lead sponsorEisai Inc.
Sponsor typeIndustry
Start2018-06-11
Primary completion2022-03-01
Results postedYes

2. Clinical Question

The primary statistical question was whether lenvatinib 20 mg plus pembrolizumab 200 mg produced superior progression-free survival and overall survival compared with treatment of physician's choice consisting of doxorubicin or paclitaxel. The registry reports these primary comparisons separately in mismatch repair proficient (pMMR) participants and in all-comer participants.

Population

Participants with advanced endometrial cancer enrolled in the randomized phase 3 trial. Primary analyses were reported both for pMMR participants and for all-comer participants.

Intervention

Lenvatinib 20 mg plus pembrolizumab 200 mg.

Comparator

Treatment of physician's choice: doxorubicin or paclitaxel.

Primary question

Does the lenvatinib plus pembrolizumab regimen improve PFS and OS relative to treatment of physician's choice?

3. Trial Design

01
Enroll827 participants
02
RandomizeTwo parallel arms
03
TreatTwo treatment strategies
04
AssessPFS, OS, response, QLQ-C30
05
AnalyzeITT-based analyses
Allocation
Randomized allocation in a parallel-group phase 3 design.
Masking
The registry reports no masking.
Primary purpose
Treatment.
Hypothesis type
Superiority for the reported primary and secondary statistical analyses.
INTERVENTION

Lenvatinib + Pembrolizumab

  • Lenvatinib 20 mg
  • Pembrolizumab 200 mg
COMPARATOR

Treatment of Physician's Choice

  • Doxorubicin or paclitaxel
What the design tells us statistically: randomization establishes the treatment comparison at the level of assignment, while the parallel design allows the two randomized groups to be followed under their assigned strategies. Because the registry reports no masking, treatment assignment was not masked in the trial design described here.

4. Endpoints

The registry lists four primary endpoints. Two concern progression-free survival and two concern overall survival, each reported separately in pMMR participants and in all-comer participants.

Primary endpointTime frameEndpoint typeRegistry definition / description
PFS in pMMR participants Up to approximately 27 months Time-to-event PFS was defined as the time from the date of randomization to the date of the first documentation of disease progression, as determined by Blinded Independent Central Review (BICR) per RECIST version 1.1 or death due to any cause, whichever occurred first.
PFS in all-comer participants Up to approximately 27 months Binary in the registry endpoint classification; analyzed statistically as a time-to-event outcome PFS was defined as the time from the date of randomization to the date of the first documentation of disease progression, as determined by BICR per RECIST version 1.1 or death due to any cause, whichever occurred first.
OS in pMMR participants Up to approximately 43 months Time-to-event OS was defined as the time from the date of randomization to the date of death due to any cause. Participants who were lost to follow-up and those who were alive at the date of data cut-off were censored at the date the participant was last known alive, or date of data cut-off, whichever occurred first.
OS in all-comer participants Up to approximately 43 months Binary in the registry endpoint classification; analyzed statistically as a time-to-event outcome OS was defined as the time from the date of randomization to the date of death due to any cause. Participants who were lost to follow-up and those who were alive at the date of data cut-off were censored at the date the participant was last known alive, or date of data cut-off, whichever occurred first.

Secondary endpoints represented in the posted analyses

EndpointTime frameMethodEffect measure
Objective Response Rate (ORR) in pMMR participants Up to approximately 80 months Miettinen & Nurminen method Difference in percentage / risk difference
ORR in all-comer participants Up to approximately 80 months Miettinen & Nurminen method Difference in percentage / risk difference
Change from baseline in EORTC QLQ-C30 in pMMR participants Baseline, Week 12 cLDA model Difference in LS Means / mean difference
Change from baseline in EORTC QLQ-C30 in all-comer participants Baseline, Week 12 cLDA model Difference in LS Means / mean difference

5. Statistical Methodology

Intention-to-treat analysis

The registry states that each posted efficacy analysis used an ITT population consisting of all randomized participants. For pMMR-specific analyses, the reported data were restricted to pMMR participants within that ITT framework.

The ITT principle is important because treatment comparisons remain tied to randomized assignment rather than being redefined according to treatment received after randomization. This preserves the principal design advantage created by randomization.

Log-rank test and time-to-event analysis

All four primary analyses used the log-rank method as reported by the registry. PFS and OS are naturally time-to-event outcomes because the analysis records not only whether an event occurred, but also when progression, death, or censoring occurred.

Conceptual comparison
Log-rank test  →  comparison of event-time distributions
Hazard ratio  →  relative treatment effect from a time-to-event model

The registry analysis notes for all four primary analyses identify regression using the Cox method. Thus, the reported hazard ratios are model-based measures accompanying the time-to-event comparison.

Hazard ratio

The four primary endpoints use the hazard ratio as the effect measure. A hazard ratio below 1 indicates a lower estimated instantaneous event rate in the lenvatinib plus pembrolizumab group relative to treatment of physician's choice, within the statistical model used for the analysis.

Conceptual interpretation
HR < 1  →  lower estimated instantaneous event rate in the intervention group

For PFS, the event is progression or death. For OS, the event is death from any cause. The hazard ratio is not an absolute risk difference and does not mean that every participant experiences the same proportional change in risk.

Score-based confidence intervals for response

The posted ORR analyses used the Miettinen & Nurminen method. This is a score-based confidence-interval approach for differences in proportions, related to the Newcombe and Wilson methods.

For this trial, the reported effect measure is the difference in percentage, or risk difference, between the randomized treatment groups. This puts response in an absolute rather than relative scale.

Constrained longitudinal data analysis

The change-from-baseline QLQ-C30 analyses used a constrained longitudinal data analysis (cLDA) model. The reported effect measure was the difference in least-squares means between the treatment groups, with a 95% two-sided confidence interval.

The cLDA framework is designed for longitudinal outcomes because measurements from the same participant at different times are related. The model therefore addresses the repeated-measures structure rather than treating each observed score as if it came from an independent participant.

6. Primary Results: Progression-Free Survival

PFS in pMMR Participants

Hazard ratio for progression or death

0.60

95% CI: 0.50–0.72   ·   P < 0.0001

Time frame: Up to approximately 27 months

Analysis featureReported result
PopulationITT population; data reported for pMMR participants
ComparisonLenvatinib 20 mg + pembrolizumab 200 mg vs TPC: doxorubicin or paclitaxel
MethodLog-rank test; regression, Cox method
Effect measureHazard ratio
Estimate0.60
95% CI0.50–0.72
P-value< 0.0001
HypothesisSuperiority
Clinical Biostats interpretation

An HR of 0.60 means that the estimated instantaneous rate of the PFS event—progression or death—was 0.60 times that of the treatment-of-physician's-choice group under the reported time-to-event model. Expressed as a simple relative interpretation, this corresponds to an estimated 40% lower hazard.

The HR does not mean that 40% of participants avoided progression, nor does it mean that each individual participant had exactly a 40% reduction in risk. It is a relative model-based measure over the analyzed follow-up.

The 95% CI of 0.50–0.72 describes statistical uncertainty around the estimated hazard ratio. Because the entire interval is below 1, the interval is consistent with a lower estimated event hazard for the intervention group under the stated model.

The P-value of < 0.0001 addresses evidence against the relevant null hypothesis under the statistical testing framework. It is not a measure of the magnitude or clinical importance of the treatment effect. The result should also be interpreted in light of censoring and the assumptions underlying the Cox hazard-ratio analysis, including the proportional-hazards interpretation.

PFS in All-comer Participants

Hazard ratio for progression or death

0.56

95% CI: 0.47–0.66   ·   P < 0.0001

Time frame: Up to approximately 27 months

Analysis featureReported result
PopulationITT population consisting of all randomized participants
ComparisonLenvatinib 20 mg + pembrolizumab 200 mg vs TPC: doxorubicin or paclitaxel
MethodLog-rank test; regression, Cox method
Effect measureHazard ratio
Estimate0.56
95% CI0.47–0.66
P-value< 0.0001
HypothesisSuperiority
Clinical Biostats interpretation

The all-comer PFS HR of 0.56 corresponds to an estimated 44% lower instantaneous hazard of progression or death in the intervention group relative to treatment of physician's choice, under the reported Cox-model framework.

The 95% CI of 0.47–0.66 expresses the precision of the estimated relative effect. It does not describe the range of outcomes that individual participants might experience.

The P-value of < 0.0001 indicates strong statistical evidence against the null hypothesis used for this superiority analysis, but it does not quantify effect size. The HR and its confidence interval are the quantities that describe the estimated treatment effect and its statistical uncertainty.

Because PFS is a censored time-to-event endpoint, interpretation depends on the event definitions, censoring rules, and the assumptions associated with the time-to-event model. A hazard ratio should not automatically be translated into a fixed difference in median PFS or a fixed percentage of patients who benefit.

7. Primary Results: Overall Survival

OS in pMMR Participants

Hazard ratio for death

0.70

95% CI: 0.58–0.83   ·   P < 0.0001

Time frame: Up to approximately 43 months

Analysis featureReported result
PopulationITT population; data reported for pMMR participants
ComparisonLenvatinib 20 mg + pembrolizumab 200 mg vs TPC: doxorubicin or paclitaxel
MethodLog-rank test; regression, Cox method
Effect measureHazard ratio
Estimate0.70
95% CI0.58–0.83
P-value< 0.0001
HypothesisSuperiority
Clinical Biostats interpretation

An OS HR of 0.70 means that the estimated instantaneous rate of death was 0.70 times that in the treatment-of-physician's-choice group under the reported Cox model. As a simple relative interpretation, this corresponds to an estimated 30% lower hazard.

This does not mean that 30% of participants survived or that every participant experienced a 30% reduction in the probability of death. The hazard ratio describes a relative event-rate comparison, not an individual treatment guarantee.

The 95% CI of 0.58–0.83 quantifies statistical uncertainty around the estimated HR. The interval remains below 1, indicating that the reported interval is compatible with a lower estimated death hazard for the intervention group.

The P-value of < 0.0001 measures evidence against the null hypothesis under the specified testing framework; it does not measure the size of the survival benefit. OS also involves censoring for participants who are alive at the data cut-off or lost to follow-up according to the registry definition.

OS in All-comer Participants

Hazard ratio for death

0.65

95% CI: 0.55–0.77   ·   P < 0.0001

Time frame: Up to approximately 43 months

Analysis featureReported result
PopulationITT population consisting of all randomized participants
ComparisonLenvatinib 20 mg + pembrolizumab 200 mg vs TPC: doxorubicin or paclitaxel
MethodLog-rank test; regression, Cox method
Effect measureHazard ratio
Estimate0.65
95% CI0.55–0.77
P-value< 0.0001
HypothesisSuperiority
Clinical Biostats interpretation

The all-comer OS HR of 0.65 corresponds to an estimated 35% lower instantaneous hazard of death in the intervention group relative to treatment of physician's choice, under the reported Cox-model framework.

The 95% CI of 0.55–0.77 describes uncertainty around that estimate. It does not describe a range of survival probabilities for individual participants.

The P-value of < 0.0001 indicates strong statistical evidence against the null hypothesis for the reported superiority comparison. It should not be confused with the magnitude of the effect, which is represented by the HR and its confidence interval.

Because OS is defined from randomization to death from any cause, it is a time-to-event endpoint. Participants who remain alive are censored according to the registry's stated rule. Interpretation therefore depends on both observed deaths and the censoring process.

8. Primary Results at a Glance

Primary endpointPopulationHR95% CIP-value
PFSpMMR0.600.50–0.72< 0.0001
PFSAll-comer0.560.47–0.66< 0.0001
OSpMMR0.700.58–0.83< 0.0001
OSAll-comer0.650.55–0.77< 0.0001

The four reported primary analyses all have hazard-ratio estimates below 1, with two-sided 95% confidence intervals entirely below 1 and P-values reported as less than 0.0001. The pMMR and all-comer analyses address different analysis populations, so their estimates should be read as separate prespecified population-level summaries rather than as repeated measurements of exactly the same population.

9. Secondary Results: Objective Response Rate

ORR in pMMR Participants

Difference in response percentage

15.2

95% CI: 9.1–21.4   ·   P < 0.0001

Time frame: Up to approximately 80 months

Analysis featureReported result
PopulationITT population; data reported for pMMR participants
ComparisonLenvatinib 20 mg + pembrolizumab 200 mg vs TPC: doxorubicin or paclitaxel
MethodMiettinen & Nurminen method
Effect measureDifference in percentage / risk difference
Estimate15.2
95% CI9.1–21.4
P-value< 0.0001
HypothesisSuperiority

A risk difference of 15.2 percentage points is an absolute comparison of response rates rather than a relative measure such as a risk ratio. The positive estimate indicates a higher response percentage in the intervention group according to the reported comparison.

ORR in All-comer Participants

Difference in response percentage

17.2

95% CI: 11.5–22.9   ·   P < 0.0001

Time frame: Up to approximately 80 months

Analysis featureReported result
PopulationITT population consisting of all randomized participants
ComparisonLenvatinib 20 mg + pembrolizumab 200 mg vs TPC: doxorubicin or paclitaxel
MethodMiettinen & Nurminen method
Effect measureDifference in percentage / risk difference
Estimate17.2
95% CI11.5–22.9
P-value< 0.0001
HypothesisSuperiority

The all-comer ORR comparison is an absolute difference in response percentage. Its 95% CI of 11.5–22.9 provides the statistical uncertainty reported for this difference, while the P-value of < 0.0001 addresses evidence against the null hypothesis rather than the magnitude of the difference itself.

Clinical Biostats interpretation

Response-rate differences and hazard ratios answer different questions. The ORR analyses describe a binary outcome—whether a participant met the response definition—whereas PFS and OS account for the timing of events and censoring. A risk difference of 17.2 percentage points therefore cannot be converted directly into an HR, and an HR cannot be interpreted as a response-rate difference.

The Miettinen-Nurminen method provides a score-based interval for the comparison of proportions. The confidence interval describes uncertainty around the estimated between-group difference; it does not describe the variability of response within an individual participant.

10. Secondary Results: Quality of Life

EORTC QLQ-C30 in pMMR Participants

Analysis featureReported result
EndpointChange from baseline in EORTC QLQ-C30
Time frameBaseline, Week 12
PopulationITT population; data reported for pMMR participants
MethodConstrained longitudinal data analysis (cLDA)
Effect measureDifference in LS Means / mean difference
Estimate1.16
95% CI-2.49–4.81
P-value= 0.5316
HypothesisSuperiority
Clinical Biostats interpretation

The estimated difference in least-squares means was 1.16 points, with a 95% CI from -2.49 to 4.81. The interval includes zero, so the reported estimate is compatible with both a negative and a positive between-group difference within the uncertainty of this analysis.

The P-value of 0.5316 does not establish a statistically significant difference under the reported superiority test. Importantly, this does not prove that the treatment groups are identical. A non-small P-value is not a test of equivalence or non-inferiority.

EORTC QLQ-C30 in All-comer Participants

Analysis featureReported result
EndpointChange from baseline in EORTC QLQ-C30
Time frameBaseline, Week 12
PopulationITT population; participants evaluable for this outcome measure
MethodConstrained longitudinal data analysis (cLDA)
Effect measureDifference in LS Means / mean difference
Estimate1.01
95% CI-2.28–4.31
P-value= 0.5460
HypothesisSuperiority
Clinical Biostats interpretation

The estimated mean difference was 1.01 points, with a 95% CI from -2.28 to 4.31. As with the pMMR analysis, the interval crosses zero, so the reported data do not isolate a single direction of the between-group difference with high statistical precision.

The P-value of 0.5460 is not evidence of a statistically significant superiority difference under the reported analysis. It also should not be interpreted as evidence that the treatment groups are equivalent, because equivalence requires a prespecified equivalence margin and an appropriate equivalence analysis.

11. Safety Results

The registry provides serious adverse-event counts by treatment course and treatment group. These data are reported as affected participants divided by participants at risk. The ClinicalTrials.gov record does not provide a broader adverse-event table, so the analysis below is restricted to the serious adverse-event figures provided.

Treatment / courseSerious adverse eventsInterpretation of denominator
First Course: Lenvatinib + Pembrolizumab237/406237 affected out of 406 at risk
First Course: TPC Doxorubicin or Paclitaxel121/388121 affected out of 388 at risk
Second Course: Lenvatinib + Pembrolizumab4/214 affected out of 21 at risk
TPC Crossover1/61 affected out of 6 at risk
Safety denominator matters: these serious-adverse-event figures use the specific denominators reported by the registry for each course/group. They should not be combined into a single treatment-group rate or compared as though all four rows represented the same analysis population.

The first-course rows provide the principal randomized treatment-group safety counts reported in the ClinicalTrials.gov record. The second-course and crossover rows have much smaller denominators and represent different treatment contexts. Consequently, the latter rows should be interpreted descriptively rather than treated as directly comparable randomized-arm estimates.

12. How the Primary Endpoints Were Analyzed

EndpointPopulationMethodEffect measureHypothesis
PFS, pMMR ITT; pMMR participants Log-rank; Cox regression Hazard ratio Superiority
PFS, all-comer ITT; all randomized participants Log-rank; Cox regression Hazard ratio Superiority
OS, pMMR ITT; pMMR participants Log-rank; Cox regression Hazard ratio Superiority
OS, all-comer ITT; all randomized participants Log-rank; Cox regression Hazard ratio Superiority

The registry therefore combines a nonparametric time-to-event comparison—the log-rank test—with a model-based effect estimate from Cox regression. This is a common statistical pairing: the log-rank test addresses the treatment-group comparison across the observed event-time distributions, while the Cox model provides the hazard-ratio estimate and its confidence interval.

13. Statistical Methods Explained

Why was a log-rank test used for PFS and OS?

PFS and OS are time-to-event endpoints. Participants can experience the event at different times, while some participants remain event-free or alive when follow-up ends and are therefore censored. The log-rank test is designed to compare the survival experience of two groups across follow-up rather than reducing the outcome to a single binary event indicator.

What does a hazard ratio of 0.56 mean for all-comer PFS?

The reported HR of 0.56 means that the estimated instantaneous hazard of progression or death was 0.56 times the corresponding hazard in the treatment-of-physician's-choice group under the reported Cox-model analysis. A simple descriptive transformation is a 44% lower estimated hazard. It does not mean that 44% of participants avoided progression or death.

Why is the confidence interval important?

A point estimate such as 0.56 is only one estimate of the treatment effect. The 95% CI of 0.47–0.66 shows the statistical uncertainty around that estimate under the analysis framework. The width of the interval conveys precision; it is not a range containing the individual effects experienced by participants.

Why doesn't a P-value measure treatment-effect size?

The P-value addresses how compatible the observed data are with a specified null hypothesis under the statistical model and testing framework. It depends on both the estimated effect and the amount of information in the analysis. Effect size is better described using the hazard ratio, risk difference, mean difference, and their confidence intervals.

Why was a score-based CI used for ORR?

ORR is a binary outcome, so each participant is classified according to whether the response criterion was met. The registry reports the Miettinen & Nurminen method, a score-based approach for comparing proportions. This is appropriate to the categorical structure of the response endpoint and produces an interval for the between-group difference in response percentages.

Why was cLDA used for QLQ-C30?

The QLQ-C30 analysis involves measurements at baseline and Week 12. A cLDA model accounts for the longitudinal structure of repeated measurements and reports the between-group difference in least-squares means. The model therefore uses the information in the repeated observations rather than treating the measurements as unrelated observations.

Why is ITT important in this trial?

The registry states that the efficacy analyses used an ITT population consisting of all randomized participants. An ITT analysis maintains the original randomized comparison rather than redefining the groups according to later treatment exposure. This is particularly important when interpreting randomized treatment effects because the validity of the comparison begins with the randomization process.

14. Understanding the Four Primary Hazard Ratios

Reported primary hazard-ratio estimates
PFS · pMMR
0.60
PFS · All-comer
0.56
OS · pMMR
0.70
OS · All-comer
0.65

The graphic presents the four reported point estimates on the same visual scale. It is an educational display of the registry-reported hazard ratios, not a reconstruction of Kaplan-Meier curves or a replacement for the reported confidence intervals.

EndpointHR95% CISimple relative interpretation
PFS, pMMR0.600.50–0.72Approximately 40% lower estimated hazard
PFS, all-comer0.560.47–0.66Approximately 44% lower estimated hazard
OS, pMMR0.700.58–0.83Approximately 30% lower estimated hazard
OS, all-comer0.650.55–0.77Approximately 35% lower estimated hazard

These percentage statements are simple arithmetic interpretations of the reported HRs and describe relative hazard, not absolute event reduction. They should not be read as probabilities that an individual participant will benefit.

15. Interpreting pMMR and All-comer Analyses

The registry reports each of the four primary endpoints separately for pMMR participants and for all-comer participants. This distinction is statistically important because the population being summarized changes between analyses.

pMMR analysis

The analysis population is the ITT population, with data reported for pMMR participants. This produces a treatment-effect estimate specific to that reported subgroup.

All-comer analysis

The analysis population consists of all randomized participants. The estimate therefore summarizes the randomized population without the pMMR restriction.

Why estimates differ

The pMMR and all-comer estimates need not be numerically identical because they summarize different analysis populations.

What cannot be concluded

A numerical difference between two subgroup estimates does not by itself demonstrate that the treatment effect differs between populations. A formal interaction analysis would be needed for that question.

16. Confidence Intervals, Effect Size, and Statistical Evidence

KEYNOTE-775 provides several useful examples of how effect estimates and statistical evidence should be separated conceptually.

MeasureExample from this trialWhat it tells us
Hazard ratioAll-comer OS HR 0.65Relative event-hazard estimate from the time-to-event analysis
Confidence interval0.55–0.77Statistical uncertainty around the HR estimate
P-value< 0.0001Evidence against the relevant null hypothesis under the testing framework
Risk differenceAll-comer ORR difference 17.2Absolute difference in response percentage
Mean differenceAll-comer QLQ-C30 difference 1.01Difference in least-squares means for the longitudinal outcome

A statistically small P-value and a large treatment effect are related but distinct concepts. A P-value does not tell the reader how large the effect is. Conversely, a treatment-effect estimate without a confidence interval does not show how precisely that effect was estimated.

A practical reading sequence

For each endpoint, first identify the analysis population, then the effect measure, then the point estimate, then the 95% confidence interval, and finally the P-value. This prevents a P-value from becoming the only piece of statistical information considered.

17. Limitations

18. Why This Trial Matters Statistically

KEYNOTE-775 is a useful teaching case because it combines several common clinical-trial analysis problems in a single randomized phase 3 design: time-to-event endpoints, categorical response outcomes, longitudinal quality-of-life measurements, ITT analysis, Cox regression, log-rank testing, score-based confidence intervals, and cLDA.

Statistical conceptHow it appears in KEYNOTE-775
RandomizationRandomized allocation in a two-arm parallel phase 3 trial
Intention-to-treat analysisPrimary efficacy analyses use the ITT population
Time-to-event endpointsPFS and OS
Log-rank testReported method for all four primary analyses
Cox regressionReported analysis notes for the primary hazard-ratio analyses
Hazard ratioEffect measure for PFS and OS
Confidence intervals95% two-sided intervals for primary and secondary effect estimates
Risk differenceDifference in percentage for ORR
Miettinen-Nurminen methodReported method for ORR analyses
cLDAReported method for QLQ-C30 change from baseline
Longitudinal analysisBaseline and Week 12 QLQ-C30 measurements
Analysis populationspMMR-specific and all-comer ITT analyses

19. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The four reported primary analyses produced HR estimates below 1, with two-sided 95% confidence intervals entirely below 1 and P-values reported as < 0.0001. The ORR analyses reported positive risk differences, while the QLQ-C30 mean-difference intervals included zero.

What the data directly support

The reported analyses provide estimated relative effects for PFS and OS, absolute response-rate differences, and between-group longitudinal mean differences for QLQ-C30. Each measure describes a different endpoint and should be interpreted on its own scale.

It is useful to resist collapsing these results into one overall numerical summary. PFS, OS, ORR, and QLQ-C30 measure different dimensions of a randomized clinical-trial outcome. The statistical evidence is strongest when each endpoint is interpreted according to its design, analysis population, effect measure, and uncertainty interval.

20. Record-Level Statistical Summary

DomainKEYNOTE-775
Trial phasePhase 3
Enrollment827
AllocationRandomized
Design modelParallel
MaskingNone
Primary endpointsFour: PFS in pMMR; PFS in all-comer; OS in pMMR; OS in all-comer
Primary endpoint typeTime-to-event / registry classification includes binary
Primary hypothesisSuperiority
Primary methodsLog-rank test; Cox regression
Primary effect measureHazard ratio
Secondary response methodMiettinen & Nurminen method
Secondary longitudinal methodConstrained longitudinal data analysis
Primary analyses posted4
Statistical analyses posted8
Outcome measures posted17

21. Sources

Continue with the statistical methods

Explore the statistical concepts behind randomized clinical trials, time-to-event endpoints, confidence intervals, longitudinal models, and categorical-data comparisons.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Calculators

24. Related Clinical Trial Results

25. Record Summary

KEYNOTE-775 provides a compact example of several important clinical-trial statistical methods. The trial was a randomized phase 3 parallel-group study with 827 enrolled participants and two treatment strategies. Its four primary analyses evaluated PFS and OS in pMMR and all-comer populations using log-rank testing and Cox regression, with hazard ratio as the effect measure. All four reported HR estimates were below 1, with 95% confidence intervals entirely below 1 and P-values reported as < 0.0001.

The secondary analyses demonstrate a different statistical structure. ORR was analyzed as a binary endpoint using the Miettinen & Nurminen method and reported as a risk difference, while change from baseline in EORTC QLQ-C30 was analyzed using cLDA and reported as a difference in least-squares means. The QLQ-C30 confidence intervals included zero and the corresponding P-values were 0.5316 and 0.5460.

The most important statistical lesson is that these estimates should not be collapsed into a single measure of treatment effect. Hazard ratios, risk differences, and longitudinal mean differences answer different questions. Interpreting each result requires attention to its endpoint definition, analysis population, statistical method, effect measure, confidence interval, and P-value.

Clinical Biostats methodology: A trial-results page should distinguish reported numerical evidence from statistical interpretation. The objective is not simply to repeat endpoint results, but to explain what each estimate measures, what its uncertainty means, and which conclusions the reported analysis can and cannot support.