This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
KEYNOTE-A18 was a randomized, parallel, triple-masked phase 3 trial evaluating chemoradiotherapy with pembrolizumab versus chemoradiotherapy with placebo in participants with uterine cervical neoplasms. The registry reports 1060 participants enrolled and two primary time-to-event endpoints.
| Feature | KEYNOTE-A18 |
|---|---|
| Phase | Phase 3 |
| Condition | Uterine Cervical Neoplasms |
| Brief trial title | Study of Chemoradiotherapy With or Without Pembrolizumab (MK-3475) For The Treatment of Locally Advanced Cervical Cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Triple |
| Primary purpose | Treatment |
| Enrollment | 1060 |
| Primary endpoints | Progression-Free Survival (PFS) and Overall Survival (OS) |
| Primary endpoint type | Time-to-event |
| Trial status | Completed |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | Industry |
| Results posted | Yes |
2. Clinical Question
The central statistical question is whether adding pembrolizumab to chemoradiotherapy changes progression-free survival and overall survival compared with chemoradiotherapy containing placebo for pembrolizumab in the registered population.
Population
Participants in a phase 3 trial for the treatment of locally advanced cervical cancer, with the registry condition listed as uterine cervical neoplasms.
Intervention
Pembrolizumab with cisplatin, external beam radiotherapy (EBRT), and brachytherapy.
Comparator
Placebo for pembrolizumab with cisplatin, external beam radiotherapy (EBRT), and brachytherapy.
Primary question
Does chemoradiotherapy with pembrolizumab produce a different time-to-event outcome from chemoradiotherapy with placebo for pembrolizumab?
3. Trial Design
Pembrolizumab + CCRT
- Pembrolizumab
- Cisplatin
- External Beam Radiotherapy (EBRT)
- Brachytherapy
Placebo + CCRT
- Placebo for pembrolizumab
- Cisplatin
- External Beam Radiotherapy (EBRT)
- Brachytherapy
4. Randomization, Stratification, and Analysis Populations
The primary efficacy analyses were performed according to the randomized treatment assignment. For both PFS and OS, the analysis population is explicitly defined as all randomized participants analyzed in the treatment group to which they were randomized.
| Analysis population | Definition / role |
|---|---|
| Primary efficacy population | All randomized participants analyzed in the treatment group to which they were randomized. |
| PFS PD-L1-positive population | All randomized participants in the randomized treatment group; PD-L1-positive participants were analyzed. |
| Response population | All randomized participants in the randomized treatment group who had measurable disease at baseline. |
| PRO population | Randomized participants with at least 1 patient-reported outcome assessment available and who received at least 1 dose of study medication. |
The registry reports that the primary PFS and OS analyses were stratified by planned type of external beam radiation therapy (EBRT), stage at screening, and planned total radiotherapy dose. This is important because stratification preserves information from factors used in the randomized comparison and can improve precision when the analysis reflects those strata.
5. Primary Endpoints
| Endpoint | Registry definition | Time frame | Posted primary analysis |
|---|---|---|---|
| Progression-Free Survival (PFS) Per RECIST 1.1 as Assessed by the Investigator | PFS is the time from randomization to the first documented progressive disease (PD) or death due to any cause, whichever occurs first. Per RECIST 1.1, or by histopathologic confirmation of suspected disease progression, PD is defined as ≥20% increase in the sum of diameters of target lesions. In addition to the relative increase of 20%, the sum must also demonstrate an absolute increase. | Up to approximately 55 months | Yes |
| Overall Survival (OS) | OS is the time from randomization to death due to any cause. | Up to approximately 55 months | Yes |
Both primary endpoints are time-to-event outcomes. That structure matters because not every participant necessarily experiences the event before the analysis cutoff. Participants without an observed event contribute information through their available follow-up and are then handled as censored observations under the survival-analysis framework.
6. Statistical Methodology
Kaplan-Meier estimation
The registry specifically states that the Kaplan-Meier nonparametric product-limit method for censored data was used to estimate OS. Kaplan-Meier estimation is also the standard descriptive framework for presenting a time-to-event endpoint such as PFS.
Here, di represents the number of events at event time ti, while ni is the number at risk immediately before that time.
Stratified log-rank testing
The formal primary PFS and OS comparisons used a stratified log-rank test. The analysis was stratified by planned EBRT type, stage at screening, and planned total radiotherapy dose.
The log-rank test compares the observed and expected numbers of events across treatment groups over follow-up. Stratification allows the comparison to account for the prespecified strata rather than treating the trial as though those factors had not been incorporated into the design.
Cox regression and hazard ratios
For both primary endpoints, the registry reports a Cox regression model with Efron's method of tie handling and treatment as a covariate. The primary analyses were stratified by planned EBRT type, stage at screening, and planned total radiotherapy dose.
An HR below 1 indicates a lower estimated instantaneous event rate in the pembrolizumab group relative to the placebo group under the fitted time-to-event model. It is not an absolute risk difference or a probability that an individual patient will benefit.
Score-based confidence intervals for proportions
For the reported complete-response and objective-response comparisons, the registry analysis text identifies the Miettinen & Nurminen method and describes a score-based approach for the difference in percentages. These methods are designed for confidence intervals around differences in binomial proportions and avoid relying on a simple normal approximation alone.
Constrained longitudinal data analysis
The patient-reported outcome analyses used a constrained longitudinal data analysis (cLDA) model. The response was the PRO score, with covariates for treatment-by-study-visit interaction and stratification by planned type of EBRT.
A cLDA model evaluates repeated measurements jointly rather than analyzing each visit as an isolated endpoint. In this trial, the reported effect measure was the difference in least-squares means between pembrolizumab and placebo.
7. Primary Results: Progression-Free Survival
PFS hazard ratio
95% CI: 0.59–0.87 · P = 0.0004
Stratified log-rank analysis; Cox regression with Efron's method of tie handling.
The reported PFS HR of 0.72 means that, under the fitted Cox model, the estimated instantaneous rate of progression or death was approximately 72% of that in the placebo-combination group. Equivalently, 1 − 0.72 = 0.28, so the estimate corresponds to an approximately 28% lower estimated hazard in relative terms.
The HR does not mean that 28% of participants avoided progression, nor does it mean that every participant experienced exactly a 28% reduction in risk. It is a model-based relative comparison of event hazards over the analyzed follow-up.
The 95% CI of 0.59–0.87 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. Because the entire interval is below 1, the interval is consistent with a lower estimated hazard in the pembrolizumab group. It does not describe the range of outcomes for individual patients.
The p-value of 0.0004 addresses evidence against the null hypothesis under the reported testing framework; it does not measure the size or clinical importance of the treatment effect. The magnitude of the effect is conveyed by the HR and its confidence interval.
The interpretation also depends on the Cox-model framework and its proportional-hazards interpretation. A single HR is most straightforward when the relative hazards are reasonably represented by a common ratio over time; the ClinicalTrials.gov record does not provide a formal diagnostic for that assumption.
| Primary PFS feature | Reported result |
|---|---|
| Analysis population | All randomized participants analyzed in the treatment group to which they were randomized |
| Groups compared | Chemoradiotherapy + Pembrolizumab vs Chemoradiotherapy + Placebo for Pembrolizumab |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.72 |
| 95% CI | 0.59–0.87 |
| P-value | 0.0004 |
| Hypothesis type | Superiority |
| Time frame | Up to approximately 55 months |
8. Primary Results: Overall Survival
OS hazard ratio
95% CI: 0.57–0.94 · P = 0.0076
Stratified log-rank analysis; Cox regression with Efron's method of tie handling.
The reported OS HR of 0.73 means that, under the fitted Cox model, the estimated instantaneous rate of death was approximately 73% of that in the placebo-combination group. In relative terms, 1 − 0.73 = 0.27, corresponding to an approximately 27% lower estimated hazard of death.
This does not mean that 27% of participants survived, that 27% of participants were cured, or that each individual experienced a 27% reduction in mortality risk. The HR summarizes a relative time-to-event comparison across the analysis period.
The 95% CI of 0.57–0.94 quantifies uncertainty around the estimated HR. Its upper bound remains below 1, so the interval is compatible with a lower estimated hazard of death for pembrolizumab relative to placebo under the reported model.
The p-value of 0.0076 measures the evidence against the null hypothesis under the stated statistical test. It does not measure effect size, probability that the treatment works, or the magnitude of clinical benefit.
As with PFS, the HR is model-based and should not be treated as an absolute survival difference. The ClinicalTrials.gov record does not provide median OS or time-specific OS rates, so those quantities are not added here.
| Primary OS feature | Reported result |
|---|---|
| Analysis population | All randomized participants analyzed in the treatment group to which they were randomized |
| Groups compared | Chemoradiotherapy + Pembrolizumab vs Chemoradiotherapy + Placebo for Pembrolizumab |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.73 |
| 95% CI | 0.57–0.94 |
| P-value | 0.0076 |
| Hypothesis type | Superiority |
| Time frame | Up to approximately 55 months |
9. Secondary Time-to-Event Results
The registry contains several secondary survival analyses. These are useful for understanding the consistency of the reported treatment effect across alternative assessments and analysis populations, but they should be distinguished from the two primary endpoints.
| Secondary endpoint | Effect | 95% CI | P-value | Method |
|---|---|---|---|---|
| PFS per RECIST 1.1, BICR | HR 0.72 | 0.58–0.89 | 0.0011 | Log-rank; stratified |
| PFS in PD-L1-positive participants, investigator assessment | HR 0.73 | 0.59–0.89 | 0.0010 | Log-rank |
| PFS in PD-L1-positive participants, BICR | HR 0.72 | 0.58–0.90 | 0.0016 | Log-rank |
| OS in PD-L1-positive participants | HR 0.73 | 0.56–0.95 | 0.0083 | Log-rank |
| PFS after next-line treatment (PFS 2) | HR 0.69 | 0.54–0.88 | 0.0015 | Log-rank; stratified |
The secondary estimates are all below 1, but they answer different questions. The BICR PFS analysis changes the outcome-assessment source; the PD-L1-positive analyses restrict the analysis population; and PFS 2 extends the time-to-event concept to the period following discontinuation of study treatment and subsequent treatment.
Consistency across these analyses can be descriptively informative, but each estimate has its own uncertainty and analysis population. A subgroup-specific HR should not automatically be interpreted as a treatment-effect difference between subgroup categories unless a formal interaction or heterogeneity analysis supports that conclusion.
10. Fixed-Time PFS and OS Results
Several secondary outcomes report differences in survival rates at prespecified time points. These are absolute percentage-point differences, not hazard ratios.
| Endpoint | Difference | 95% CI | Direction |
|---|---|---|---|
| PFS per RECIST 1.1 at Month 24, investigator assessment | 10.9 percentage points | 5.1–16.7 | Pembrolizumab minus placebo |
| PFS per RECIST 1.1 at Month 24, BICR | 7.9 percentage points | 2.3–13.5 | Pembrolizumab minus placebo |
| OS at Month 36 | 7.4 percentage points | 2.3–12.5 | Pembrolizumab minus placebo |
11. Response Results
Complete Response at Week 12
| Assessment | Risk difference | 95% CI | P-value | Method |
|---|---|---|---|---|
| Investigator | 3.3 percentage points | -2.5 to 9.1 | 0.1307 | Score-based CI for proportions; Miettinen & Nurminen method, stratified |
| BICR | 0.7 percentage points | -5.2 to 6.7 | 0.4082 | Score-based CI for proportions; Miettinen & Nurminen method, stratified |
The investigator-assessed CR rate difference was 3.3 percentage points, with a 95% CI from -2.5 to 9.1. The corresponding BICR estimate was 0.7 percentage points, with a 95% CI from -5.2 to 6.7.
Both confidence intervals include zero. Thus, the reported intervals encompass both a small negative difference and a positive difference. The p-values of 0.1307 and 0.4082 quantify the evidence under their respective testing frameworks; they do not establish the magnitude of any treatment effect.
Importantly, these CR analyses use participants with measurable disease at baseline, rather than simply treating the entire randomized population as the denominator for the response endpoint.
Objective Response Rate
| Assessment | Risk difference | 95% CI | P-value | Method |
|---|---|---|---|---|
| Investigator | 3.6 percentage points | -0.7 to 7.9 | 0.0496 | Score-based CI for proportions; Miettinen & Nurminen method, stratified |
| BICR | 2.2 percentage points | -1.5 to 5.9 | 0.1244 | Score-based CI for proportions; Miettinen & Nurminen method, stratified |
The investigator-assessed ORR difference is estimated at 3.6 percentage points, while the BICR estimate is 2.2 percentage points. The investigator-assessed 95% CI extends from -0.7 to 7.9, while the BICR CI extends from -1.5 to 5.9.
The p-value of 0.0496 for the investigator assessment is close to the conventional 0.05 reference level, but a p-value should not be treated as an effect-size measure. The corresponding confidence interval is essential because it communicates the precision and range of values compatible with the analysis.
The BICR analysis has a p-value of 0.1244 and a confidence interval that includes zero. Differences between investigator and BICR assessments also illustrate why outcome-assessment procedures matter in clinical-trial statistics.
12. Patient-Reported Outcomes and cLDA
Three secondary endpoints evaluated change from baseline in patient-reported outcome measures from baseline to week 36. These analyses used constrained longitudinal data analysis and reported differences in least-squares means.
| Patient-reported outcome | Time frame | Difference in LS means | 95% CI | P-value |
|---|---|---|---|---|
| EORTC QLQ-C30 Global Health Status Score | Baseline and week 36 | 0.07 | -2.66 to 2.80 | 0.9593 |
| EORTC QLQ-C30 Physical Function Score | Baseline and week 36 | 0.68 | -1.36 to 2.71 | 0.5145 |
| EORTC QLQ-CX24 Score | Baseline and Week 36 | 0.75 | -0.65 to 2.15 | 0.2953 |
All three estimated differences in least-squares means are relatively close to zero compared with their corresponding confidence intervals. The global health status estimate is 0.07 with a 95% CI of -2.66 to 2.80; physical function is 0.68 with a 95% CI of -1.36 to 2.71; and the QLQ-CX24 estimate is 0.75 with a 95% CI of -0.65 to 2.15.
The p-values are 0.9593, 0.5145, and 0.2953, respectively. These p-values do not measure the size of the differences. The confidence intervals provide the more useful description of the uncertainty surrounding each estimated between-group difference.
Because cLDA models repeated measurements jointly, its interpretation differs from simply subtracting a treatment-group mean at week 36 from a control-group mean. The model incorporates the longitudinal structure and treatment-by-visit interaction specified in the registry analysis.
13. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm using affected participants over participants at risk. The reported counts are:
| Safety measure | Pembrolizumab + CCRT | Placebo + CCRT |
|---|---|---|
| Serious adverse events | 175 / 528 | 153 / 530 |
The registry reports 175 of 528 participants affected in the pembrolizumab-plus-CCRT group and 153 of 530 in the placebo-plus-CCRT group. These are the reported affected/at-risk counts and are therefore not silently converted here into separately rounded percentages.
Safety events are different from time-to-event efficacy endpoints. A serious adverse event count does not require the same survival-model interpretation as PFS or OS, and the ClinicalTrials.gov record does not provide a formal between-arm statistical analysis for this safety measure.
14. Statistical Methods Explained
Why was a stratified log-rank test used?
PFS and OS are time-to-event outcomes, so simply comparing proportions at a single time would discard information about when events occurred and about censoring. The log-rank test uses the event history across follow-up. In KEYNOTE-A18, the primary analyses were additionally stratified by planned EBRT type, stage at screening, and planned total radiotherapy dose.
What does a hazard ratio of 0.72 mean?
An HR of 0.72 means that the estimated instantaneous rate of the relevant event in the pembrolizumab group was 72% of that in the placebo group under the fitted model. For PFS, the event is progression or death; for OS, it is death. It does not mean that 72% of participants experienced an event or that the treatment reduces every individual's probability of an event by exactly 28%.
Why is the confidence interval important?
The point estimate is only one summary of the randomized comparison. A 95% confidence interval provides information about statistical precision. For example, the PFS HR of 0.72 has a 95% CI of 0.59–0.87, while the OS HR of 0.73 has a 95% CI of 0.57–0.94. The width of these intervals reflects uncertainty that is not captured by the point estimate alone.
Why does the p-value not measure effect size?
A p-value quantifies how compatible the observed data are with a specified null hypothesis under the testing framework. It is affected by both the observed effect and the amount of information available. The effect size is described by the HR, risk difference, or mean difference, while the confidence interval describes uncertainty around that estimate.
Why use BICR as well as investigator assessment?
The trial reports PFS and response outcomes using both investigator assessment and blinded independent central review. Independent review provides a separate assessment of radiologic outcomes and can reduce concerns that knowledge of treatment assignment could influence disease-assessment decisions. In a triple-masked randomized trial, assessment procedures remain an important part of the statistical design.
What does cLDA add to the patient-reported outcome analysis?
Patient-reported outcomes were measured longitudinally, so observations from the same participant are related rather than independent. A cLDA model evaluates the repeated scores together and estimates the treatment-by-visit structure. Here, the reported treatment effect is the difference in least-squares means, with a 95% confidence interval.
Why is a risk difference different from a hazard ratio?
A risk difference compares probabilities or percentages at a defined endpoint, whereas a hazard ratio compares event rates over time through a survival model. KEYNOTE-A18 illustrates both approaches: PFS and OS are summarized with HRs, while fixed-time PFS outcomes and response outcomes are reported as differences in percentages.
15. Understanding the Analysis Populations
One of the most important statistical distinctions in this trial is that different endpoints use different analysis populations. The primary PFS and OS analyses include all randomized participants and retain randomized treatment assignment. Response analyses are restricted to participants with measurable disease at baseline. Patient-reported outcomes require at least one available PRO assessment and at least one dose of study medication.
| Endpoint family | Population specified in the ClinicalTrials.gov record | Why it matters |
|---|---|---|
| Primary PFS / OS | All randomized participants, analyzed according to randomized treatment | Maintains the treatment comparison created by randomization. |
| CR / ORR | All randomized participants in their randomized treatment group, with measurable disease at baseline | Response can only be assessed within the specified measurable-disease population. |
| PD-L1 analyses | Randomized participants who were PD-L1 positive | These are subgroup analyses rather than the overall randomized population. |
| PRO analyses | Randomized participants with at least 1 PRO assessment and at least 1 dose | Longitudinal patient-reported outcomes require an available PRO measurement. |
This distinction prevents an important statistical error: treating every reported estimate as though it came from the same denominator and analysis population. The interpretation of a result depends partly on who was included in that particular analysis.
16. Stratified Analysis and Covariate Adjustment
The primary PFS and OS analyses were stratified by three prespecified trial factors: planned type of EBRT, stage at screening, and planned total radiotherapy dose. The registry also describes the Cox analyses as using treatment as a covariate, with Efron's method for handling tied event times.
Stratification
Stratification allows the time-to-event comparison to account for the designated trial strata when comparing observed and expected events.
Covariate adjustment
The Cox model includes treatment as a covariate while the analysis incorporates the reported stratification structure.
Tied event times
Efron's method is used for handling tied event times in the Cox regression model.
Interpretive consequence
The reported HR is a model-based adjusted time-to-event comparison, not a simple ratio of two crude event proportions.
17. Confidence Intervals Across Endpoint Types
KEYNOTE-A18 provides a useful comparison of confidence intervals for several different effect measures.
| Endpoint type | Effect measure | Example estimate | 95% CI |
|---|---|---|---|
| Time-to-event | Hazard ratio | PFS HR 0.72 | 0.59–0.87 |
| Fixed-time survival | Difference in survival rate | PFS difference at Month 24: 10.9 | 5.1–16.7 |
| Binary response | Risk difference | Investigator CR difference: 3.3 | -2.5–9.1 |
| Continuous PRO | Mean difference / difference in LS means | Global health status: 0.07 | -2.66–2.80 |
The numerical scales cannot be compared directly. A confidence interval centered around zero is natural for a difference measure, while a hazard-ratio interval is centered around the null value of 1. Correct interpretation therefore requires knowing both the endpoint and the effect measure.
18. Primary Endpoint Interpretation: Relative Effects vs Absolute Effects
The primary PFS and OS results are expressed as hazard ratios, while several secondary outcomes provide fixed-time differences in survival percentages. These measures complement one another.
Relative effect
PFS HR 0.72 and OS HR 0.73 summarize the relative time-to-event comparison under the Cox model.
Absolute effect
The Month 24 PFS differences of 10.9 and 7.9 percentage points provide a fixed-time absolute comparison.
Response effect
Risk differences of 3.3, 0.7, 3.6, and 2.2 percentage points summarize differences in response rates.
PRO effect
Differences in LS means describe between-group differences in longitudinal patient-reported outcome scores.
These measures answer different questions. A hazard ratio summarizes the relative event rate over time; a fixed-time risk difference asks how much the estimated survival probability differs at one specified time; a response risk difference compares binary response probabilities; and a mean difference compares continuous outcomes.
19. Limitations and Interpretation Issues
- Hazard-ratio interpretation: Cox HRs are model-based relative measures. A single HR is most directly interpretable when the proportional-hazards framework is a reasonable representation of the data over the analyzed period.
- Different analysis populations: primary survival analyses, response analyses, PD-L1-positive analyses, and PRO analyses do not use identical populations.
- Investigator versus BICR assessment: the trial reports both assessment approaches for selected efficacy endpoints. Differences between them should be interpreted as differences in outcome-assessment procedures rather than as independent randomized treatment comparisons.
- Subgroup interpretation: PD-L1-positive analyses describe a restricted population. A subgroup estimate should not be interpreted as proof of treatment-effect heterogeneity without a formal comparison between subgroup effects.
- Fixed-time endpoints: survival-rate differences at Month 24 or Month 36 describe a specific time point and should not be substituted for the overall time-to-event analysis.
- Response denominators: CR and ORR analyses are based on participants with measurable disease at baseline, so their estimates should not automatically be interpreted as rates in the entire randomized population.
- PRO analysis population: the cLDA analyses require an available PRO assessment and at least one dose of study medication, which differs from the all-randomized population used for the primary efficacy analyses.
- Safety interpretation: the ClinicalTrials.gov record contains serious adverse-event counts by arm but does not provide a formal statistical comparison for that endpoint.
- Registry method detail: the ClinicalTrials.gov record identifies the formal method for several analyses but not for the reported fixed-time survival-rate differences. A specific unreported statistical procedure is therefore not attributed to those results.
20. Why This Trial Matters Statistically
KEYNOTE-A18 provides a compact example of several important clinical-trial methods operating within the same randomized study.
| Concept | How it appears in KEYNOTE-A18 |
|---|---|
| Randomization | Randomized parallel-group phase 3 design with 1060 enrolled participants |
| Blinding | Triple-masked design |
| Time-to-event analysis | PFS and OS are primary time-to-event endpoints |
| Kaplan-Meier estimation | OS is estimated using the Kaplan-Meier nonparametric product-limit method for censored data |
| Log-rank testing | Primary PFS and OS comparisons use log-rank testing |
| Hazard ratio | Primary PFS and OS effects are reported as HRs |
| Cox regression | Primary analyses use Cox regression with Efron's method of tie handling |
| Stratified analysis | Primary analyses are stratified by EBRT type, stage at screening, and planned total radiotherapy dose |
| Risk difference | CR and ORR are reported as differences in percentages |
| Score-based CI | Miettinen & Nurminen methodology is reported for response comparisons |
| Longitudinal analysis | PRO endpoints use constrained longitudinal data analysis |
| Multiple assessment sources | Selected PFS and response endpoints are assessed by investigator and BICR |
21. A Statistical Reading of the Complete Results
The primary efficacy evidence consists of two time-to-event comparisons. For PFS, the reported HR is 0.72 with a 95% CI of 0.59–0.87 and P = 0.0004. For OS, the reported HR is 0.73 with a 95% CI of 0.57–0.94 and P = 0.0076.
The secondary time-to-event analyses show similar HR estimates: BICR-assessed PFS has an HR of 0.72, investigator-assessed PFS among PD-L1-positive participants has an HR of 0.73, BICR-assessed PFS among PD-L1-positive participants has an HR of 0.72, OS among PD-L1-positive participants has an HR of 0.73, and PFS 2 has an HR of 0.69. Each estimate has a confidence interval and p-value that should be interpreted in the context of its particular endpoint and analysis population.
The fixed-time results provide an additional absolute perspective. The reported difference in PFS survival rate at Month 24 is 10.9 percentage points by investigator assessment and 7.9 percentage points by BICR. The reported OS difference at Month 36 is 7.4 percentage points. These estimates are complementary to, rather than interchangeable with, the hazard ratios.
The response analyses illustrate a different statistical scale. Investigator-assessed CR has a risk difference of 3.3 percentage points, while BICR-assessed CR has a difference of 0.7 percentage points. Investigator-assessed ORR has a difference of 3.6 percentage points, and BICR-assessed ORR has a difference of 2.2 percentage points. Their confidence intervals and p-values show why a point estimate should not be interpreted without its uncertainty interval.
Finally, the PRO analyses show how a trial can combine survival, binary, and continuous longitudinal outcomes within one statistical program. The cLDA estimates are close to zero for all three reported PRO measures, with confidence intervals spanning both negative and positive differences. This does not change the interpretation of the primary survival analyses; it provides a separate assessment of patient-reported outcomes.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Statistical Calculators
24. Sources
- ClinicalTrials.gov: NCT04221945 — KEYNOTE-A18.
- PubMed: PMID 38521086.
- PubMed: PMID 39288779.
- PubMed: PMID 40592026.
Continue with the statistical methods
Explore the underlying survival, categorical-data, longitudinal, and clinical-trial methods used to interpret randomized evidence.
25. Record Summary
KEYNOTE-A18 illustrates a modern randomized phase 3 statistical program built around two primary time-to-event endpoints and supported by multiple secondary endpoint types. The primary PFS analysis reported an HR of 0.72 with a 95% CI of 0.59–0.87 and P = 0.0004; the primary OS analysis reported an HR of 0.73 with a 95% CI of 0.57–0.94 and P = 0.0076.
The trial also demonstrates why statistical interpretation cannot be reduced to a single p-value. PFS and OS use hazard ratios; fixed-time survival outcomes use differences in survival rates; response outcomes use risk differences with score-based confidence intervals; and patient-reported outcomes use differences in least-squares means from a cLDA model. The analysis populations also differ by endpoint, making the definition of each denominator part of the statistical result.