This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
KEYNOTE-775 was a completed randomized phase 3 parallel-group trial in advanced endometrial cancer. The registry reports 827 enrolled participants, two treatment arms, no masking, and a primary purpose of treatment. The trial compared lenvatinib plus pembrolizumab with treatment of physician's choice consisting of doxorubicin or paclitaxel.
| Feature | KEYNOTE-775 |
|---|---|
| Trial name | KEYNOTE-775 |
| NCT ID | NCT03517449 |
| Phase | Phase 3 |
| Status | Completed |
| Condition | Endometrial Neoplasms |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 827 |
| Lead sponsor | Eisai Inc. |
| Sponsor type | Industry |
| Start | 2018-06-11 |
| Primary completion | 2022-03-01 |
| Results posted | Yes |
2. Clinical Question
The primary statistical question was whether lenvatinib 20 mg plus pembrolizumab 200 mg produced superior progression-free survival and overall survival compared with treatment of physician's choice consisting of doxorubicin or paclitaxel. The registry reports these primary comparisons separately in mismatch repair proficient (pMMR) participants and in all-comer participants.
Population
Participants with advanced endometrial cancer enrolled in the randomized phase 3 trial. Primary analyses were reported both for pMMR participants and for all-comer participants.
Intervention
Lenvatinib 20 mg plus pembrolizumab 200 mg.
Comparator
Treatment of physician's choice: doxorubicin or paclitaxel.
Primary question
Does the lenvatinib plus pembrolizumab regimen improve PFS and OS relative to treatment of physician's choice?
3. Trial Design
Lenvatinib + Pembrolizumab
- Lenvatinib 20 mg
- Pembrolizumab 200 mg
Treatment of Physician's Choice
- Doxorubicin or paclitaxel
4. Endpoints
The registry lists four primary endpoints. Two concern progression-free survival and two concern overall survival, each reported separately in pMMR participants and in all-comer participants.
| Primary endpoint | Time frame | Endpoint type | Registry definition / description |
|---|---|---|---|
| PFS in pMMR participants | Up to approximately 27 months | Time-to-event | PFS was defined as the time from the date of randomization to the date of the first documentation of disease progression, as determined by Blinded Independent Central Review (BICR) per RECIST version 1.1 or death due to any cause, whichever occurred first. |
| PFS in all-comer participants | Up to approximately 27 months | Binary in the registry endpoint classification; analyzed statistically as a time-to-event outcome | PFS was defined as the time from the date of randomization to the date of the first documentation of disease progression, as determined by BICR per RECIST version 1.1 or death due to any cause, whichever occurred first. |
| OS in pMMR participants | Up to approximately 43 months | Time-to-event | OS was defined as the time from the date of randomization to the date of death due to any cause. Participants who were lost to follow-up and those who were alive at the date of data cut-off were censored at the date the participant was last known alive, or date of data cut-off, whichever occurred first. |
| OS in all-comer participants | Up to approximately 43 months | Binary in the registry endpoint classification; analyzed statistically as a time-to-event outcome | OS was defined as the time from the date of randomization to the date of death due to any cause. Participants who were lost to follow-up and those who were alive at the date of data cut-off were censored at the date the participant was last known alive, or date of data cut-off, whichever occurred first. |
Secondary endpoints represented in the posted analyses
| Endpoint | Time frame | Method | Effect measure |
|---|---|---|---|
| Objective Response Rate (ORR) in pMMR participants | Up to approximately 80 months | Miettinen & Nurminen method | Difference in percentage / risk difference |
| ORR in all-comer participants | Up to approximately 80 months | Miettinen & Nurminen method | Difference in percentage / risk difference |
| Change from baseline in EORTC QLQ-C30 in pMMR participants | Baseline, Week 12 | cLDA model | Difference in LS Means / mean difference |
| Change from baseline in EORTC QLQ-C30 in all-comer participants | Baseline, Week 12 | cLDA model | Difference in LS Means / mean difference |
5. Statistical Methodology
Intention-to-treat analysis
The registry states that each posted efficacy analysis used an ITT population consisting of all randomized participants. For pMMR-specific analyses, the reported data were restricted to pMMR participants within that ITT framework.
The ITT principle is important because treatment comparisons remain tied to randomized assignment rather than being redefined according to treatment received after randomization. This preserves the principal design advantage created by randomization.
Log-rank test and time-to-event analysis
All four primary analyses used the log-rank method as reported by the registry. PFS and OS are naturally time-to-event outcomes because the analysis records not only whether an event occurred, but also when progression, death, or censoring occurred.
Hazard ratio → relative treatment effect from a time-to-event model
The registry analysis notes for all four primary analyses identify regression using the Cox method. Thus, the reported hazard ratios are model-based measures accompanying the time-to-event comparison.
Hazard ratio
The four primary endpoints use the hazard ratio as the effect measure. A hazard ratio below 1 indicates a lower estimated instantaneous event rate in the lenvatinib plus pembrolizumab group relative to treatment of physician's choice, within the statistical model used for the analysis.
For PFS, the event is progression or death. For OS, the event is death from any cause. The hazard ratio is not an absolute risk difference and does not mean that every participant experiences the same proportional change in risk.
Score-based confidence intervals for response
The posted ORR analyses used the Miettinen & Nurminen method. This is a score-based confidence-interval approach for differences in proportions, related to the Newcombe and Wilson methods.
For this trial, the reported effect measure is the difference in percentage, or risk difference, between the randomized treatment groups. This puts response in an absolute rather than relative scale.
Constrained longitudinal data analysis
The change-from-baseline QLQ-C30 analyses used a constrained longitudinal data analysis (cLDA) model. The reported effect measure was the difference in least-squares means between the treatment groups, with a 95% two-sided confidence interval.
The cLDA framework is designed for longitudinal outcomes because measurements from the same participant at different times are related. The model therefore addresses the repeated-measures structure rather than treating each observed score as if it came from an independent participant.
6. Primary Results: Progression-Free Survival
PFS in pMMR Participants
Hazard ratio for progression or death
95% CI: 0.50–0.72 · P < 0.0001
Time frame: Up to approximately 27 months
| Analysis feature | Reported result |
|---|---|
| Population | ITT population; data reported for pMMR participants |
| Comparison | Lenvatinib 20 mg + pembrolizumab 200 mg vs TPC: doxorubicin or paclitaxel |
| Method | Log-rank test; regression, Cox method |
| Effect measure | Hazard ratio |
| Estimate | 0.60 |
| 95% CI | 0.50–0.72 |
| P-value | < 0.0001 |
| Hypothesis | Superiority |
An HR of 0.60 means that the estimated instantaneous rate of the PFS event—progression or death—was 0.60 times that of the treatment-of-physician's-choice group under the reported time-to-event model. Expressed as a simple relative interpretation, this corresponds to an estimated 40% lower hazard.
The HR does not mean that 40% of participants avoided progression, nor does it mean that each individual participant had exactly a 40% reduction in risk. It is a relative model-based measure over the analyzed follow-up.
The 95% CI of 0.50–0.72 describes statistical uncertainty around the estimated hazard ratio. Because the entire interval is below 1, the interval is consistent with a lower estimated event hazard for the intervention group under the stated model.
The P-value of < 0.0001 addresses evidence against the relevant null hypothesis under the statistical testing framework. It is not a measure of the magnitude or clinical importance of the treatment effect. The result should also be interpreted in light of censoring and the assumptions underlying the Cox hazard-ratio analysis, including the proportional-hazards interpretation.
PFS in All-comer Participants
Hazard ratio for progression or death
95% CI: 0.47–0.66 · P < 0.0001
Time frame: Up to approximately 27 months
| Analysis feature | Reported result |
|---|---|
| Population | ITT population consisting of all randomized participants |
| Comparison | Lenvatinib 20 mg + pembrolizumab 200 mg vs TPC: doxorubicin or paclitaxel |
| Method | Log-rank test; regression, Cox method |
| Effect measure | Hazard ratio |
| Estimate | 0.56 |
| 95% CI | 0.47–0.66 |
| P-value | < 0.0001 |
| Hypothesis | Superiority |
The all-comer PFS HR of 0.56 corresponds to an estimated 44% lower instantaneous hazard of progression or death in the intervention group relative to treatment of physician's choice, under the reported Cox-model framework.
The 95% CI of 0.47–0.66 expresses the precision of the estimated relative effect. It does not describe the range of outcomes that individual participants might experience.
The P-value of < 0.0001 indicates strong statistical evidence against the null hypothesis used for this superiority analysis, but it does not quantify effect size. The HR and its confidence interval are the quantities that describe the estimated treatment effect and its statistical uncertainty.
Because PFS is a censored time-to-event endpoint, interpretation depends on the event definitions, censoring rules, and the assumptions associated with the time-to-event model. A hazard ratio should not automatically be translated into a fixed difference in median PFS or a fixed percentage of patients who benefit.
7. Primary Results: Overall Survival
OS in pMMR Participants
Hazard ratio for death
95% CI: 0.58–0.83 · P < 0.0001
Time frame: Up to approximately 43 months
| Analysis feature | Reported result |
|---|---|
| Population | ITT population; data reported for pMMR participants |
| Comparison | Lenvatinib 20 mg + pembrolizumab 200 mg vs TPC: doxorubicin or paclitaxel |
| Method | Log-rank test; regression, Cox method |
| Effect measure | Hazard ratio |
| Estimate | 0.70 |
| 95% CI | 0.58–0.83 |
| P-value | < 0.0001 |
| Hypothesis | Superiority |
An OS HR of 0.70 means that the estimated instantaneous rate of death was 0.70 times that in the treatment-of-physician's-choice group under the reported Cox model. As a simple relative interpretation, this corresponds to an estimated 30% lower hazard.
This does not mean that 30% of participants survived or that every participant experienced a 30% reduction in the probability of death. The hazard ratio describes a relative event-rate comparison, not an individual treatment guarantee.
The 95% CI of 0.58–0.83 quantifies statistical uncertainty around the estimated HR. The interval remains below 1, indicating that the reported interval is compatible with a lower estimated death hazard for the intervention group.
The P-value of < 0.0001 measures evidence against the null hypothesis under the specified testing framework; it does not measure the size of the survival benefit. OS also involves censoring for participants who are alive at the data cut-off or lost to follow-up according to the registry definition.
OS in All-comer Participants
Hazard ratio for death
95% CI: 0.55–0.77 · P < 0.0001
Time frame: Up to approximately 43 months
| Analysis feature | Reported result |
|---|---|
| Population | ITT population consisting of all randomized participants |
| Comparison | Lenvatinib 20 mg + pembrolizumab 200 mg vs TPC: doxorubicin or paclitaxel |
| Method | Log-rank test; regression, Cox method |
| Effect measure | Hazard ratio |
| Estimate | 0.65 |
| 95% CI | 0.55–0.77 |
| P-value | < 0.0001 |
| Hypothesis | Superiority |
The all-comer OS HR of 0.65 corresponds to an estimated 35% lower instantaneous hazard of death in the intervention group relative to treatment of physician's choice, under the reported Cox-model framework.
The 95% CI of 0.55–0.77 describes uncertainty around that estimate. It does not describe a range of survival probabilities for individual participants.
The P-value of < 0.0001 indicates strong statistical evidence against the null hypothesis for the reported superiority comparison. It should not be confused with the magnitude of the effect, which is represented by the HR and its confidence interval.
Because OS is defined from randomization to death from any cause, it is a time-to-event endpoint. Participants who remain alive are censored according to the registry's stated rule. Interpretation therefore depends on both observed deaths and the censoring process.
8. Primary Results at a Glance
| Primary endpoint | Population | HR | 95% CI | P-value |
|---|---|---|---|---|
| PFS | pMMR | 0.60 | 0.50–0.72 | < 0.0001 |
| PFS | All-comer | 0.56 | 0.47–0.66 | < 0.0001 |
| OS | pMMR | 0.70 | 0.58–0.83 | < 0.0001 |
| OS | All-comer | 0.65 | 0.55–0.77 | < 0.0001 |
The four reported primary analyses all have hazard-ratio estimates below 1, with two-sided 95% confidence intervals entirely below 1 and P-values reported as less than 0.0001. The pMMR and all-comer analyses address different analysis populations, so their estimates should be read as separate prespecified population-level summaries rather than as repeated measurements of exactly the same population.
9. Secondary Results: Objective Response Rate
ORR in pMMR Participants
Difference in response percentage
95% CI: 9.1–21.4 · P < 0.0001
Time frame: Up to approximately 80 months
| Analysis feature | Reported result |
|---|---|
| Population | ITT population; data reported for pMMR participants |
| Comparison | Lenvatinib 20 mg + pembrolizumab 200 mg vs TPC: doxorubicin or paclitaxel |
| Method | Miettinen & Nurminen method |
| Effect measure | Difference in percentage / risk difference |
| Estimate | 15.2 |
| 95% CI | 9.1–21.4 |
| P-value | < 0.0001 |
| Hypothesis | Superiority |
A risk difference of 15.2 percentage points is an absolute comparison of response rates rather than a relative measure such as a risk ratio. The positive estimate indicates a higher response percentage in the intervention group according to the reported comparison.
ORR in All-comer Participants
Difference in response percentage
95% CI: 11.5–22.9 · P < 0.0001
Time frame: Up to approximately 80 months
| Analysis feature | Reported result |
|---|---|
| Population | ITT population consisting of all randomized participants |
| Comparison | Lenvatinib 20 mg + pembrolizumab 200 mg vs TPC: doxorubicin or paclitaxel |
| Method | Miettinen & Nurminen method |
| Effect measure | Difference in percentage / risk difference |
| Estimate | 17.2 |
| 95% CI | 11.5–22.9 |
| P-value | < 0.0001 |
| Hypothesis | Superiority |
The all-comer ORR comparison is an absolute difference in response percentage. Its 95% CI of 11.5–22.9 provides the statistical uncertainty reported for this difference, while the P-value of < 0.0001 addresses evidence against the null hypothesis rather than the magnitude of the difference itself.
Response-rate differences and hazard ratios answer different questions. The ORR analyses describe a binary outcome—whether a participant met the response definition—whereas PFS and OS account for the timing of events and censoring. A risk difference of 17.2 percentage points therefore cannot be converted directly into an HR, and an HR cannot be interpreted as a response-rate difference.
The Miettinen-Nurminen method provides a score-based interval for the comparison of proportions. The confidence interval describes uncertainty around the estimated between-group difference; it does not describe the variability of response within an individual participant.
10. Secondary Results: Quality of Life
EORTC QLQ-C30 in pMMR Participants
| Analysis feature | Reported result |
|---|---|
| Endpoint | Change from baseline in EORTC QLQ-C30 |
| Time frame | Baseline, Week 12 |
| Population | ITT population; data reported for pMMR participants |
| Method | Constrained longitudinal data analysis (cLDA) |
| Effect measure | Difference in LS Means / mean difference |
| Estimate | 1.16 |
| 95% CI | -2.49–4.81 |
| P-value | = 0.5316 |
| Hypothesis | Superiority |
The estimated difference in least-squares means was 1.16 points, with a 95% CI from -2.49 to 4.81. The interval includes zero, so the reported estimate is compatible with both a negative and a positive between-group difference within the uncertainty of this analysis.
The P-value of 0.5316 does not establish a statistically significant difference under the reported superiority test. Importantly, this does not prove that the treatment groups are identical. A non-small P-value is not a test of equivalence or non-inferiority.
EORTC QLQ-C30 in All-comer Participants
| Analysis feature | Reported result |
|---|---|
| Endpoint | Change from baseline in EORTC QLQ-C30 |
| Time frame | Baseline, Week 12 |
| Population | ITT population; participants evaluable for this outcome measure |
| Method | Constrained longitudinal data analysis (cLDA) |
| Effect measure | Difference in LS Means / mean difference |
| Estimate | 1.01 |
| 95% CI | -2.28–4.31 |
| P-value | = 0.5460 |
| Hypothesis | Superiority |
The estimated mean difference was 1.01 points, with a 95% CI from -2.28 to 4.31. As with the pMMR analysis, the interval crosses zero, so the reported data do not isolate a single direction of the between-group difference with high statistical precision.
The P-value of 0.5460 is not evidence of a statistically significant superiority difference under the reported analysis. It also should not be interpreted as evidence that the treatment groups are equivalent, because equivalence requires a prespecified equivalence margin and an appropriate equivalence analysis.
11. Safety Results
The registry provides serious adverse-event counts by treatment course and treatment group. These data are reported as affected participants divided by participants at risk. The ClinicalTrials.gov record does not provide a broader adverse-event table, so the analysis below is restricted to the serious adverse-event figures provided.
| Treatment / course | Serious adverse events | Interpretation of denominator |
|---|---|---|
| First Course: Lenvatinib + Pembrolizumab | 237/406 | 237 affected out of 406 at risk |
| First Course: TPC Doxorubicin or Paclitaxel | 121/388 | 121 affected out of 388 at risk |
| Second Course: Lenvatinib + Pembrolizumab | 4/21 | 4 affected out of 21 at risk |
| TPC Crossover | 1/6 | 1 affected out of 6 at risk |
The first-course rows provide the principal randomized treatment-group safety counts reported in the ClinicalTrials.gov record. The second-course and crossover rows have much smaller denominators and represent different treatment contexts. Consequently, the latter rows should be interpreted descriptively rather than treated as directly comparable randomized-arm estimates.
12. How the Primary Endpoints Were Analyzed
| Endpoint | Population | Method | Effect measure | Hypothesis |
|---|---|---|---|---|
| PFS, pMMR | ITT; pMMR participants | Log-rank; Cox regression | Hazard ratio | Superiority |
| PFS, all-comer | ITT; all randomized participants | Log-rank; Cox regression | Hazard ratio | Superiority |
| OS, pMMR | ITT; pMMR participants | Log-rank; Cox regression | Hazard ratio | Superiority |
| OS, all-comer | ITT; all randomized participants | Log-rank; Cox regression | Hazard ratio | Superiority |
The registry therefore combines a nonparametric time-to-event comparison—the log-rank test—with a model-based effect estimate from Cox regression. This is a common statistical pairing: the log-rank test addresses the treatment-group comparison across the observed event-time distributions, while the Cox model provides the hazard-ratio estimate and its confidence interval.
13. Statistical Methods Explained
Why was a log-rank test used for PFS and OS?
PFS and OS are time-to-event endpoints. Participants can experience the event at different times, while some participants remain event-free or alive when follow-up ends and are therefore censored. The log-rank test is designed to compare the survival experience of two groups across follow-up rather than reducing the outcome to a single binary event indicator.
What does a hazard ratio of 0.56 mean for all-comer PFS?
The reported HR of 0.56 means that the estimated instantaneous hazard of progression or death was 0.56 times the corresponding hazard in the treatment-of-physician's-choice group under the reported Cox-model analysis. A simple descriptive transformation is a 44% lower estimated hazard. It does not mean that 44% of participants avoided progression or death.
Why is the confidence interval important?
A point estimate such as 0.56 is only one estimate of the treatment effect. The 95% CI of 0.47–0.66 shows the statistical uncertainty around that estimate under the analysis framework. The width of the interval conveys precision; it is not a range containing the individual effects experienced by participants.
Why doesn't a P-value measure treatment-effect size?
The P-value addresses how compatible the observed data are with a specified null hypothesis under the statistical model and testing framework. It depends on both the estimated effect and the amount of information in the analysis. Effect size is better described using the hazard ratio, risk difference, mean difference, and their confidence intervals.
Why was a score-based CI used for ORR?
ORR is a binary outcome, so each participant is classified according to whether the response criterion was met. The registry reports the Miettinen & Nurminen method, a score-based approach for comparing proportions. This is appropriate to the categorical structure of the response endpoint and produces an interval for the between-group difference in response percentages.
Why was cLDA used for QLQ-C30?
The QLQ-C30 analysis involves measurements at baseline and Week 12. A cLDA model accounts for the longitudinal structure of repeated measurements and reports the between-group difference in least-squares means. The model therefore uses the information in the repeated observations rather than treating the measurements as unrelated observations.
Why is ITT important in this trial?
The registry states that the efficacy analyses used an ITT population consisting of all randomized participants. An ITT analysis maintains the original randomized comparison rather than redefining the groups according to later treatment exposure. This is particularly important when interpreting randomized treatment effects because the validity of the comparison begins with the randomization process.
14. Understanding the Four Primary Hazard Ratios
The graphic presents the four reported point estimates on the same visual scale. It is an educational display of the registry-reported hazard ratios, not a reconstruction of Kaplan-Meier curves or a replacement for the reported confidence intervals.
| Endpoint | HR | 95% CI | Simple relative interpretation |
|---|---|---|---|
| PFS, pMMR | 0.60 | 0.50–0.72 | Approximately 40% lower estimated hazard |
| PFS, all-comer | 0.56 | 0.47–0.66 | Approximately 44% lower estimated hazard |
| OS, pMMR | 0.70 | 0.58–0.83 | Approximately 30% lower estimated hazard |
| OS, all-comer | 0.65 | 0.55–0.77 | Approximately 35% lower estimated hazard |
These percentage statements are simple arithmetic interpretations of the reported HRs and describe relative hazard, not absolute event reduction. They should not be read as probabilities that an individual participant will benefit.
15. Interpreting pMMR and All-comer Analyses
The registry reports each of the four primary endpoints separately for pMMR participants and for all-comer participants. This distinction is statistically important because the population being summarized changes between analyses.
pMMR analysis
The analysis population is the ITT population, with data reported for pMMR participants. This produces a treatment-effect estimate specific to that reported subgroup.
All-comer analysis
The analysis population consists of all randomized participants. The estimate therefore summarizes the randomized population without the pMMR restriction.
Why estimates differ
The pMMR and all-comer estimates need not be numerically identical because they summarize different analysis populations.
What cannot be concluded
A numerical difference between two subgroup estimates does not by itself demonstrate that the treatment effect differs between populations. A formal interaction analysis would be needed for that question.
16. Confidence Intervals, Effect Size, and Statistical Evidence
KEYNOTE-775 provides several useful examples of how effect estimates and statistical evidence should be separated conceptually.
| Measure | Example from this trial | What it tells us |
|---|---|---|
| Hazard ratio | All-comer OS HR 0.65 | Relative event-hazard estimate from the time-to-event analysis |
| Confidence interval | 0.55–0.77 | Statistical uncertainty around the HR estimate |
| P-value | < 0.0001 | Evidence against the relevant null hypothesis under the testing framework |
| Risk difference | All-comer ORR difference 17.2 | Absolute difference in response percentage |
| Mean difference | All-comer QLQ-C30 difference 1.01 | Difference in least-squares means for the longitudinal outcome |
A statistically small P-value and a large treatment effect are related but distinct concepts. A P-value does not tell the reader how large the effect is. Conversely, a treatment-effect estimate without a confidence interval does not show how precisely that effect was estimated.
For each endpoint, first identify the analysis population, then the effect measure, then the point estimate, then the 95% confidence interval, and finally the P-value. This prevents a P-value from becoming the only piece of statistical information considered.
17. Limitations
- Registry-defined scope: this analysis is restricted to the numerical results and methodological information reported in the ClinicalTrials.gov-derived trial data. Results not present in that data are not reconstructed from external sources.
- No median PFS or OS: the ClinicalTrials.gov record does not report median survival times, so no median PFS or OS comparison is presented here.
- No subgroup forest plot: the ClinicalTrials.gov record does not provide additional subgroup hazard ratios or interaction tests beyond the pMMR and all-comer primary analyses.
- No baseline table: baseline demographic and disease characteristics are not included in the ClinicalTrials.gov record, so balance between randomized groups cannot be assessed from this record.
- No multiplicity details: the ClinicalTrials.gov record identifies four primary endpoints and superiority testing but do not provide a multiplicity-adjustment strategy or alpha-allocation scheme. No such procedure is inferred here.
- No interim-analysis details: the ClinicalTrials.gov record does not report an interim monitoring plan or alpha-spending procedure. None is assumed.
- No imputation details: the ClinicalTrials.gov record does not specify a missing-data or imputation strategy. No method is attributed to the trial beyond the reported cLDA analysis.
- Proportional-hazards assumption: the Cox-model hazard ratio is a model-based summary. Its usual interpretation requires attention to the relationship of hazards over time; the ClinicalTrials.gov record does not provide a formal proportional-hazards diagnostic.
- Different endpoint populations: pMMR and all-comer analyses summarize different populations, so their numerical estimates should not be treated as repeated estimates from an identical sample.
- Safety denominators: serious adverse-event counts are reported for distinct treatment courses and denominators. They should not be pooled into a single rate.
- Quality-of-life analysis population: the registry text notes that the overall number analyzed can refer to participants evaluable for the outcome measure. This differs conceptually from simply assuming that every randomized participant contributed a Week 12 measurement.
18. Why This Trial Matters Statistically
KEYNOTE-775 is a useful teaching case because it combines several common clinical-trial analysis problems in a single randomized phase 3 design: time-to-event endpoints, categorical response outcomes, longitudinal quality-of-life measurements, ITT analysis, Cox regression, log-rank testing, score-based confidence intervals, and cLDA.
| Statistical concept | How it appears in KEYNOTE-775 |
|---|---|
| Randomization | Randomized allocation in a two-arm parallel phase 3 trial |
| Intention-to-treat analysis | Primary efficacy analyses use the ITT population |
| Time-to-event endpoints | PFS and OS |
| Log-rank test | Reported method for all four primary analyses |
| Cox regression | Reported analysis notes for the primary hazard-ratio analyses |
| Hazard ratio | Effect measure for PFS and OS |
| Confidence intervals | 95% two-sided intervals for primary and secondary effect estimates |
| Risk difference | Difference in percentage for ORR |
| Miettinen-Nurminen method | Reported method for ORR analyses |
| cLDA | Reported method for QLQ-C30 change from baseline |
| Longitudinal analysis | Baseline and Week 12 QLQ-C30 measurements |
| Analysis populations | pMMR-specific and all-comer ITT analyses |
19. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The four reported primary analyses produced HR estimates below 1, with two-sided 95% confidence intervals entirely below 1 and P-values reported as < 0.0001. The ORR analyses reported positive risk differences, while the QLQ-C30 mean-difference intervals included zero.
What the data directly support
The reported analyses provide estimated relative effects for PFS and OS, absolute response-rate differences, and between-group longitudinal mean differences for QLQ-C30. Each measure describes a different endpoint and should be interpreted on its own scale.
It is useful to resist collapsing these results into one overall numerical summary. PFS, OS, ORR, and QLQ-C30 measure different dimensions of a randomized clinical-trial outcome. The statistical evidence is strongest when each endpoint is interpreted according to its design, analysis population, effect measure, and uncertainty interval.
20. Record-Level Statistical Summary
| Domain | KEYNOTE-775 |
|---|---|
| Trial phase | Phase 3 |
| Enrollment | 827 |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary endpoints | Four: PFS in pMMR; PFS in all-comer; OS in pMMR; OS in all-comer |
| Primary endpoint type | Time-to-event / registry classification includes binary |
| Primary hypothesis | Superiority |
| Primary methods | Log-rank test; Cox regression |
| Primary effect measure | Hazard ratio |
| Secondary response method | Miettinen & Nurminen method |
| Secondary longitudinal method | Constrained longitudinal data analysis |
| Primary analyses posted | 4 |
| Statistical analyses posted | 8 |
| Outcome measures posted | 17 |
21. Sources
- ClinicalTrials.gov: NCT03517449 — KEYNOTE-775.
- PubMed record: PMID 42167105.
- PubMed record: PMID 38302725.
- PubMed record: PMID 37523661.
- PubMed record: PMID 37086595.
- PubMed record: PMID 37058687.
Continue with the statistical methods
Explore the statistical concepts behind randomized clinical trials, time-to-event endpoints, confidence intervals, longitudinal models, and categorical-data comparisons.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Calculators
24. Related Clinical Trial Results
25. Record Summary
KEYNOTE-775 provides a compact example of several important clinical-trial statistical methods. The trial was a randomized phase 3 parallel-group study with 827 enrolled participants and two treatment strategies. Its four primary analyses evaluated PFS and OS in pMMR and all-comer populations using log-rank testing and Cox regression, with hazard ratio as the effect measure. All four reported HR estimates were below 1, with 95% confidence intervals entirely below 1 and P-values reported as < 0.0001.
The secondary analyses demonstrate a different statistical structure. ORR was analyzed as a binary endpoint using the Miettinen & Nurminen method and reported as a risk difference, while change from baseline in EORTC QLQ-C30 was analyzed using cLDA and reported as a difference in least-squares means. The QLQ-C30 confidence intervals included zero and the corresponding P-values were 0.5316 and 0.5460.
The most important statistical lesson is that these estimates should not be collapsed into a single measure of treatment effect. Hazard ratios, risk differences, and longitudinal mean differences answer different questions. Interpreting each result requires attention to its endpoint definition, analysis population, statistical method, effect measure, confidence interval, and P-value.