← Clinical Trials
Locally Advanced Cervical Cancer Phase 3 Completed NCT04221945

KEYNOTE-A18: Complete Statistical Analysis of Chemoradiotherapy With Pembrolizumab in Locally Advanced Cervical Cancer

An independent statistical review of the randomized phase 3 KEYNOTE-A18 trial evaluating chemoradiotherapy with pembrolizumab versus chemoradiotherapy with placebo for the treatment of uterine cervical neoplasms.

Trial start: May 12, 2020  ·  Primary completion: January 7, 2025  ·  Enrollment: 1060
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

KEYNOTE-A18 was a randomized, parallel, triple-masked phase 3 trial evaluating chemoradiotherapy with pembrolizumab versus chemoradiotherapy with placebo in participants with uterine cervical neoplasms. The registry reports 1060 participants enrolled and two primary time-to-event endpoints.

1060
Enrollment
ClinicalTrials.gov record
2
Arms
Parallel design
0.72
PFS HR
95% CI 0.59–0.87
0.73
OS HR
95% CI 0.57–0.94
FeatureKEYNOTE-A18
PhasePhase 3
ConditionUterine Cervical Neoplasms
Brief trial titleStudy of Chemoradiotherapy With or Without Pembrolizumab (MK-3475) For The Treatment of Locally Advanced Cervical Cancer
AllocationRandomized
Design modelParallel
MaskingTriple
Primary purposeTreatment
Enrollment1060
Primary endpointsProgression-Free Survival (PFS) and Overall Survival (OS)
Primary endpoint typeTime-to-event
Trial statusCompleted
Lead sponsorMerck Sharp & Dohme LLC
Sponsor typeIndustry
Results postedYes

2. Clinical Question

The central statistical question is whether adding pembrolizumab to chemoradiotherapy changes progression-free survival and overall survival compared with chemoradiotherapy containing placebo for pembrolizumab in the registered population.

Population

Participants in a phase 3 trial for the treatment of locally advanced cervical cancer, with the registry condition listed as uterine cervical neoplasms.

Intervention

Pembrolizumab with cisplatin, external beam radiotherapy (EBRT), and brachytherapy.

Comparator

Placebo for pembrolizumab with cisplatin, external beam radiotherapy (EBRT), and brachytherapy.

Primary question

Does chemoradiotherapy with pembrolizumab produce a different time-to-event outcome from chemoradiotherapy with placebo for pembrolizumab?

3. Trial Design

01
Randomize 1060 enrolled
02
Two arms Pembrolizumab or placebo
03
CCRT Cisplatin + EBRT + brachytherapy
04
Assessment PFS, OS, response and PROs
05
Follow-up Primary endpoints up to approximately 55 months
ARM A · 528 at risk for reported serious AEs

Pembrolizumab + CCRT

  • Pembrolizumab
  • Cisplatin
  • External Beam Radiotherapy (EBRT)
  • Brachytherapy
ARM B · 530 at risk for reported serious AEs

Placebo + CCRT

  • Placebo for pembrolizumab
  • Cisplatin
  • External Beam Radiotherapy (EBRT)
  • Brachytherapy
Masking is part of the design. The registry classifies the trial as triple-masked. That is relevant statistically because knowledge of treatment assignment can influence aspects of treatment administration, assessment, reporting, or outcome ascertainment. The ClinicalTrials.gov record does not specify the three masked roles, so they are not inferred here.

4. Randomization, Stratification, and Analysis Populations

The primary efficacy analyses were performed according to the randomized treatment assignment. For both PFS and OS, the analysis population is explicitly defined as all randomized participants analyzed in the treatment group to which they were randomized.

Analysis populationDefinition / role
Primary efficacy populationAll randomized participants analyzed in the treatment group to which they were randomized.
PFS PD-L1-positive populationAll randomized participants in the randomized treatment group; PD-L1-positive participants were analyzed.
Response populationAll randomized participants in the randomized treatment group who had measurable disease at baseline.
PRO populationRandomized participants with at least 1 patient-reported outcome assessment available and who received at least 1 dose of study medication.

The registry reports that the primary PFS and OS analyses were stratified by planned type of external beam radiation therapy (EBRT), stage at screening, and planned total radiotherapy dose. This is important because stratification preserves information from factors used in the randomized comparison and can improve precision when the analysis reflects those strata.

5. Primary Endpoints

EndpointRegistry definitionTime framePosted primary analysis
Progression-Free Survival (PFS) Per RECIST 1.1 as Assessed by the Investigator PFS is the time from randomization to the first documented progressive disease (PD) or death due to any cause, whichever occurs first. Per RECIST 1.1, or by histopathologic confirmation of suspected disease progression, PD is defined as ≥20% increase in the sum of diameters of target lesions. In addition to the relative increase of 20%, the sum must also demonstrate an absolute increase. Up to approximately 55 months Yes
Overall Survival (OS) OS is the time from randomization to death due to any cause. Up to approximately 55 months Yes

Both primary endpoints are time-to-event outcomes. That structure matters because not every participant necessarily experiences the event before the analysis cutoff. Participants without an observed event contribute information through their available follow-up and are then handled as censored observations under the survival-analysis framework.

6. Statistical Methodology

Kaplan-Meier estimation

The registry specifically states that the Kaplan-Meier nonparametric product-limit method for censored data was used to estimate OS. Kaplan-Meier estimation is also the standard descriptive framework for presenting a time-to-event endpoint such as PFS.

Kaplan-Meier survival function
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at event time ti, while ni is the number at risk immediately before that time.

Stratified log-rank testing

The formal primary PFS and OS comparisons used a stratified log-rank test. The analysis was stratified by planned EBRT type, stage at screening, and planned total radiotherapy dose.

The log-rank test compares the observed and expected numbers of events across treatment groups over follow-up. Stratification allows the comparison to account for the prespecified strata rather than treating the trial as though those factors had not been incorporated into the design.

Cox regression and hazard ratios

For both primary endpoints, the registry reports a Cox regression model with Efron's method of tie handling and treatment as a covariate. The primary analyses were stratified by planned EBRT type, stage at screening, and planned total radiotherapy dose.

Hazard-ratio interpretation
HR = estimated hazard in pembrolizumab + CCRT ÷ estimated hazard in placebo + CCRT

An HR below 1 indicates a lower estimated instantaneous event rate in the pembrolizumab group relative to the placebo group under the fitted time-to-event model. It is not an absolute risk difference or a probability that an individual patient will benefit.

Score-based confidence intervals for proportions

For the reported complete-response and objective-response comparisons, the registry analysis text identifies the Miettinen & Nurminen method and describes a score-based approach for the difference in percentages. These methods are designed for confidence intervals around differences in binomial proportions and avoid relying on a simple normal approximation alone.

Constrained longitudinal data analysis

The patient-reported outcome analyses used a constrained longitudinal data analysis (cLDA) model. The response was the PRO score, with covariates for treatment-by-study-visit interaction and stratification by planned type of EBRT.

What cLDA contributes
Outcome = treatment × visit structure + longitudinal covariance + stratification

A cLDA model evaluates repeated measurements jointly rather than analyzing each visit as an isolated endpoint. In this trial, the reported effect measure was the difference in least-squares means between pembrolizumab and placebo.

7. Primary Results: Progression-Free Survival

PFS hazard ratio

0.72

95% CI: 0.59–0.87   ·   P = 0.0004

Stratified log-rank analysis; Cox regression with Efron's method of tie handling.

Clinical Biostats interpretation

The reported PFS HR of 0.72 means that, under the fitted Cox model, the estimated instantaneous rate of progression or death was approximately 72% of that in the placebo-combination group. Equivalently, 1 − 0.72 = 0.28, so the estimate corresponds to an approximately 28% lower estimated hazard in relative terms.

The HR does not mean that 28% of participants avoided progression, nor does it mean that every participant experienced exactly a 28% reduction in risk. It is a model-based relative comparison of event hazards over the analyzed follow-up.

The 95% CI of 0.59–0.87 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. Because the entire interval is below 1, the interval is consistent with a lower estimated hazard in the pembrolizumab group. It does not describe the range of outcomes for individual patients.

The p-value of 0.0004 addresses evidence against the null hypothesis under the reported testing framework; it does not measure the size or clinical importance of the treatment effect. The magnitude of the effect is conveyed by the HR and its confidence interval.

The interpretation also depends on the Cox-model framework and its proportional-hazards interpretation. A single HR is most straightforward when the relative hazards are reasonably represented by a common ratio over time; the ClinicalTrials.gov record does not provide a formal diagnostic for that assumption.

Primary PFS featureReported result
Analysis populationAll randomized participants analyzed in the treatment group to which they were randomized
Groups comparedChemoradiotherapy + Pembrolizumab vs Chemoradiotherapy + Placebo for Pembrolizumab
MethodLog-rank test
Effect measureHazard ratio
Estimate0.72
95% CI0.59–0.87
P-value0.0004
Hypothesis typeSuperiority
Time frameUp to approximately 55 months

8. Primary Results: Overall Survival

OS hazard ratio

0.73

95% CI: 0.57–0.94   ·   P = 0.0076

Stratified log-rank analysis; Cox regression with Efron's method of tie handling.

Clinical Biostats interpretation

The reported OS HR of 0.73 means that, under the fitted Cox model, the estimated instantaneous rate of death was approximately 73% of that in the placebo-combination group. In relative terms, 1 − 0.73 = 0.27, corresponding to an approximately 27% lower estimated hazard of death.

This does not mean that 27% of participants survived, that 27% of participants were cured, or that each individual experienced a 27% reduction in mortality risk. The HR summarizes a relative time-to-event comparison across the analysis period.

The 95% CI of 0.57–0.94 quantifies uncertainty around the estimated HR. Its upper bound remains below 1, so the interval is compatible with a lower estimated hazard of death for pembrolizumab relative to placebo under the reported model.

The p-value of 0.0076 measures the evidence against the null hypothesis under the stated statistical test. It does not measure effect size, probability that the treatment works, or the magnitude of clinical benefit.

As with PFS, the HR is model-based and should not be treated as an absolute survival difference. The ClinicalTrials.gov record does not provide median OS or time-specific OS rates, so those quantities are not added here.

Primary OS featureReported result
Analysis populationAll randomized participants analyzed in the treatment group to which they were randomized
Groups comparedChemoradiotherapy + Pembrolizumab vs Chemoradiotherapy + Placebo for Pembrolizumab
MethodLog-rank test
Effect measureHazard ratio
Estimate0.73
95% CI0.57–0.94
P-value0.0076
Hypothesis typeSuperiority
Time frameUp to approximately 55 months

9. Secondary Time-to-Event Results

The registry contains several secondary survival analyses. These are useful for understanding the consistency of the reported treatment effect across alternative assessments and analysis populations, but they should be distinguished from the two primary endpoints.

Secondary endpoint Effect 95% CI P-value Method
PFS per RECIST 1.1, BICR HR 0.72 0.58–0.89 0.0011 Log-rank; stratified
PFS in PD-L1-positive participants, investigator assessment HR 0.73 0.59–0.89 0.0010 Log-rank
PFS in PD-L1-positive participants, BICR HR 0.72 0.58–0.90 0.0016 Log-rank
OS in PD-L1-positive participants HR 0.73 0.56–0.95 0.0083 Log-rank
PFS after next-line treatment (PFS 2) HR 0.69 0.54–0.88 0.0015 Log-rank; stratified
How to read these secondary HRs

The secondary estimates are all below 1, but they answer different questions. The BICR PFS analysis changes the outcome-assessment source; the PD-L1-positive analyses restrict the analysis population; and PFS 2 extends the time-to-event concept to the period following discontinuation of study treatment and subsequent treatment.

Consistency across these analyses can be descriptively informative, but each estimate has its own uncertainty and analysis population. A subgroup-specific HR should not automatically be interpreted as a treatment-effect difference between subgroup categories unless a formal interaction or heterogeneity analysis supports that conclusion.

10. Fixed-Time PFS and OS Results

Several secondary outcomes report differences in survival rates at prespecified time points. These are absolute percentage-point differences, not hazard ratios.

Reported survival-rate differences
PFS at Month 24, Investigator
10.9
PFS at Month 24, BICR
7.9
OS at Month 36
7.4
EndpointDifference95% CIDirection
PFS per RECIST 1.1 at Month 24, investigator assessment10.9 percentage points5.1–16.7Pembrolizumab minus placebo
PFS per RECIST 1.1 at Month 24, BICR7.9 percentage points2.3–13.5Pembrolizumab minus placebo
OS at Month 367.4 percentage points2.3–12.5Pembrolizumab minus placebo
Method-reporting distinction: the ClinicalTrials.gov record identifies the effect measure and confidence interval for these fixed-time outcomes but list the formal method as not reported. A fixed-time survival percentage is ordinarily obtained from a time-to-event estimator such as Kaplan-Meier, with an appropriate confidence interval for the estimated survival probability. The ClinicalTrials.gov record does not identify the formal comparison method for these three analyses, so no specific test is attributed to them here.

11. Response Results

Complete Response at Week 12

AssessmentRisk difference95% CIP-valueMethod
Investigator 3.3 percentage points -2.5 to 9.1 0.1307 Score-based CI for proportions; Miettinen & Nurminen method, stratified
BICR 0.7 percentage points -5.2 to 6.7 0.4082 Score-based CI for proportions; Miettinen & Nurminen method, stratified
Clinical Biostats interpretation

The investigator-assessed CR rate difference was 3.3 percentage points, with a 95% CI from -2.5 to 9.1. The corresponding BICR estimate was 0.7 percentage points, with a 95% CI from -5.2 to 6.7.

Both confidence intervals include zero. Thus, the reported intervals encompass both a small negative difference and a positive difference. The p-values of 0.1307 and 0.4082 quantify the evidence under their respective testing frameworks; they do not establish the magnitude of any treatment effect.

Importantly, these CR analyses use participants with measurable disease at baseline, rather than simply treating the entire randomized population as the denominator for the response endpoint.

Objective Response Rate

AssessmentRisk difference95% CIP-valueMethod
Investigator 3.6 percentage points -0.7 to 7.9 0.0496 Score-based CI for proportions; Miettinen & Nurminen method, stratified
BICR 2.2 percentage points -1.5 to 5.9 0.1244 Score-based CI for proportions; Miettinen & Nurminen method, stratified
Why the confidence interval matters

The investigator-assessed ORR difference is estimated at 3.6 percentage points, while the BICR estimate is 2.2 percentage points. The investigator-assessed 95% CI extends from -0.7 to 7.9, while the BICR CI extends from -1.5 to 5.9.

The p-value of 0.0496 for the investigator assessment is close to the conventional 0.05 reference level, but a p-value should not be treated as an effect-size measure. The corresponding confidence interval is essential because it communicates the precision and range of values compatible with the analysis.

The BICR analysis has a p-value of 0.1244 and a confidence interval that includes zero. Differences between investigator and BICR assessments also illustrate why outcome-assessment procedures matter in clinical-trial statistics.

12. Patient-Reported Outcomes and cLDA

Three secondary endpoints evaluated change from baseline in patient-reported outcome measures from baseline to week 36. These analyses used constrained longitudinal data analysis and reported differences in least-squares means.

Patient-reported outcomeTime frameDifference in LS means95% CIP-value
EORTC QLQ-C30 Global Health Status Score Baseline and week 36 0.07 -2.66 to 2.80 0.9593
EORTC QLQ-C30 Physical Function Score Baseline and week 36 0.68 -1.36 to 2.71 0.5145
EORTC QLQ-CX24 Score Baseline and Week 36 0.75 -0.65 to 2.15 0.2953
Clinical Biostats interpretation

All three estimated differences in least-squares means are relatively close to zero compared with their corresponding confidence intervals. The global health status estimate is 0.07 with a 95% CI of -2.66 to 2.80; physical function is 0.68 with a 95% CI of -1.36 to 2.71; and the QLQ-CX24 estimate is 0.75 with a 95% CI of -0.65 to 2.15.

The p-values are 0.9593, 0.5145, and 0.2953, respectively. These p-values do not measure the size of the differences. The confidence intervals provide the more useful description of the uncertainty surrounding each estimated between-group difference.

Because cLDA models repeated measurements jointly, its interpretation differs from simply subtracting a treatment-group mean at week 36 from a control-group mean. The model incorporates the longitudinal structure and treatment-by-visit interaction specified in the registry analysis.

13. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm using affected participants over participants at risk. The reported counts are:

Safety measurePembrolizumab + CCRTPlacebo + CCRT
Serious adverse events175 / 528153 / 530
Serious adverse events — reported affected / at risk
Pembrolizumab + CCRT
175 / 528
Placebo + CCRT
153 / 530
How to interpret the safety counts

The registry reports 175 of 528 participants affected in the pembrolizumab-plus-CCRT group and 153 of 530 in the placebo-plus-CCRT group. These are the reported affected/at-risk counts and are therefore not silently converted here into separately rounded percentages.

Safety events are different from time-to-event efficacy endpoints. A serious adverse event count does not require the same survival-model interpretation as PFS or OS, and the ClinicalTrials.gov record does not provide a formal between-arm statistical analysis for this safety measure.

14. Statistical Methods Explained

Why was a stratified log-rank test used?

PFS and OS are time-to-event outcomes, so simply comparing proportions at a single time would discard information about when events occurred and about censoring. The log-rank test uses the event history across follow-up. In KEYNOTE-A18, the primary analyses were additionally stratified by planned EBRT type, stage at screening, and planned total radiotherapy dose.

What does a hazard ratio of 0.72 mean?

An HR of 0.72 means that the estimated instantaneous rate of the relevant event in the pembrolizumab group was 72% of that in the placebo group under the fitted model. For PFS, the event is progression or death; for OS, it is death. It does not mean that 72% of participants experienced an event or that the treatment reduces every individual's probability of an event by exactly 28%.

Why is the confidence interval important?

The point estimate is only one summary of the randomized comparison. A 95% confidence interval provides information about statistical precision. For example, the PFS HR of 0.72 has a 95% CI of 0.59–0.87, while the OS HR of 0.73 has a 95% CI of 0.57–0.94. The width of these intervals reflects uncertainty that is not captured by the point estimate alone.

Why does the p-value not measure effect size?

A p-value quantifies how compatible the observed data are with a specified null hypothesis under the testing framework. It is affected by both the observed effect and the amount of information available. The effect size is described by the HR, risk difference, or mean difference, while the confidence interval describes uncertainty around that estimate.

Why use BICR as well as investigator assessment?

The trial reports PFS and response outcomes using both investigator assessment and blinded independent central review. Independent review provides a separate assessment of radiologic outcomes and can reduce concerns that knowledge of treatment assignment could influence disease-assessment decisions. In a triple-masked randomized trial, assessment procedures remain an important part of the statistical design.

What does cLDA add to the patient-reported outcome analysis?

Patient-reported outcomes were measured longitudinally, so observations from the same participant are related rather than independent. A cLDA model evaluates the repeated scores together and estimates the treatment-by-visit structure. Here, the reported treatment effect is the difference in least-squares means, with a 95% confidence interval.

Why is a risk difference different from a hazard ratio?

A risk difference compares probabilities or percentages at a defined endpoint, whereas a hazard ratio compares event rates over time through a survival model. KEYNOTE-A18 illustrates both approaches: PFS and OS are summarized with HRs, while fixed-time PFS outcomes and response outcomes are reported as differences in percentages.

15. Understanding the Analysis Populations

One of the most important statistical distinctions in this trial is that different endpoints use different analysis populations. The primary PFS and OS analyses include all randomized participants and retain randomized treatment assignment. Response analyses are restricted to participants with measurable disease at baseline. Patient-reported outcomes require at least one available PRO assessment and at least one dose of study medication.

Endpoint familyPopulation specified in the ClinicalTrials.gov recordWhy it matters
Primary PFS / OS All randomized participants, analyzed according to randomized treatment Maintains the treatment comparison created by randomization.
CR / ORR All randomized participants in their randomized treatment group, with measurable disease at baseline Response can only be assessed within the specified measurable-disease population.
PD-L1 analyses Randomized participants who were PD-L1 positive These are subgroup analyses rather than the overall randomized population.
PRO analyses Randomized participants with at least 1 PRO assessment and at least 1 dose Longitudinal patient-reported outcomes require an available PRO measurement.

This distinction prevents an important statistical error: treating every reported estimate as though it came from the same denominator and analysis population. The interpretation of a result depends partly on who was included in that particular analysis.

16. Stratified Analysis and Covariate Adjustment

The primary PFS and OS analyses were stratified by three prespecified trial factors: planned type of EBRT, stage at screening, and planned total radiotherapy dose. The registry also describes the Cox analyses as using treatment as a covariate, with Efron's method for handling tied event times.

Stratification

Stratification allows the time-to-event comparison to account for the designated trial strata when comparing observed and expected events.

Covariate adjustment

The Cox model includes treatment as a covariate while the analysis incorporates the reported stratification structure.

Tied event times

Efron's method is used for handling tied event times in the Cox regression model.

Interpretive consequence

The reported HR is a model-based adjusted time-to-event comparison, not a simple ratio of two crude event proportions.

17. Confidence Intervals Across Endpoint Types

KEYNOTE-A18 provides a useful comparison of confidence intervals for several different effect measures.

Endpoint typeEffect measureExample estimate95% CI
Time-to-eventHazard ratioPFS HR 0.720.59–0.87
Fixed-time survivalDifference in survival ratePFS difference at Month 24: 10.95.1–16.7
Binary responseRisk differenceInvestigator CR difference: 3.3-2.5–9.1
Continuous PROMean difference / difference in LS meansGlobal health status: 0.07-2.66–2.80

The numerical scales cannot be compared directly. A confidence interval centered around zero is natural for a difference measure, while a hazard-ratio interval is centered around the null value of 1. Correct interpretation therefore requires knowing both the endpoint and the effect measure.

18. Primary Endpoint Interpretation: Relative Effects vs Absolute Effects

The primary PFS and OS results are expressed as hazard ratios, while several secondary outcomes provide fixed-time differences in survival percentages. These measures complement one another.

Relative effect

PFS HR 0.72 and OS HR 0.73 summarize the relative time-to-event comparison under the Cox model.

Absolute effect

The Month 24 PFS differences of 10.9 and 7.9 percentage points provide a fixed-time absolute comparison.

Response effect

Risk differences of 3.3, 0.7, 3.6, and 2.2 percentage points summarize differences in response rates.

PRO effect

Differences in LS means describe between-group differences in longitudinal patient-reported outcome scores.

These measures answer different questions. A hazard ratio summarizes the relative event rate over time; a fixed-time risk difference asks how much the estimated survival probability differs at one specified time; a response risk difference compares binary response probabilities; and a mean difference compares continuous outcomes.

19. Limitations and Interpretation Issues

20. Why This Trial Matters Statistically

KEYNOTE-A18 provides a compact example of several important clinical-trial methods operating within the same randomized study.

ConceptHow it appears in KEYNOTE-A18
RandomizationRandomized parallel-group phase 3 design with 1060 enrolled participants
BlindingTriple-masked design
Time-to-event analysisPFS and OS are primary time-to-event endpoints
Kaplan-Meier estimationOS is estimated using the Kaplan-Meier nonparametric product-limit method for censored data
Log-rank testingPrimary PFS and OS comparisons use log-rank testing
Hazard ratioPrimary PFS and OS effects are reported as HRs
Cox regressionPrimary analyses use Cox regression with Efron's method of tie handling
Stratified analysisPrimary analyses are stratified by EBRT type, stage at screening, and planned total radiotherapy dose
Risk differenceCR and ORR are reported as differences in percentages
Score-based CIMiettinen & Nurminen methodology is reported for response comparisons
Longitudinal analysisPRO endpoints use constrained longitudinal data analysis
Multiple assessment sourcesSelected PFS and response endpoints are assessed by investigator and BICR

21. A Statistical Reading of the Complete Results

The primary efficacy evidence consists of two time-to-event comparisons. For PFS, the reported HR is 0.72 with a 95% CI of 0.59–0.87 and P = 0.0004. For OS, the reported HR is 0.73 with a 95% CI of 0.57–0.94 and P = 0.0076.

The secondary time-to-event analyses show similar HR estimates: BICR-assessed PFS has an HR of 0.72, investigator-assessed PFS among PD-L1-positive participants has an HR of 0.73, BICR-assessed PFS among PD-L1-positive participants has an HR of 0.72, OS among PD-L1-positive participants has an HR of 0.73, and PFS 2 has an HR of 0.69. Each estimate has a confidence interval and p-value that should be interpreted in the context of its particular endpoint and analysis population.

The fixed-time results provide an additional absolute perspective. The reported difference in PFS survival rate at Month 24 is 10.9 percentage points by investigator assessment and 7.9 percentage points by BICR. The reported OS difference at Month 36 is 7.4 percentage points. These estimates are complementary to, rather than interchangeable with, the hazard ratios.

The response analyses illustrate a different statistical scale. Investigator-assessed CR has a risk difference of 3.3 percentage points, while BICR-assessed CR has a difference of 0.7 percentage points. Investigator-assessed ORR has a difference of 3.6 percentage points, and BICR-assessed ORR has a difference of 2.2 percentage points. Their confidence intervals and p-values show why a point estimate should not be interpreted without its uncertainty interval.

Finally, the PRO analyses show how a trial can combine survival, binary, and continuous longitudinal outcomes within one statistical program. The cLDA estimates are close to zero for all three reported PRO measures, with confidence intervals spanning both negative and positive differences. This does not change the interpretation of the primary survival analyses; it provides a separate assessment of patient-reported outcomes.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Statistical Calculators

24. Sources

Continue with the statistical methods

Explore the underlying survival, categorical-data, longitudinal, and clinical-trial methods used to interpret randomized evidence.

25. Record Summary

KEYNOTE-A18 illustrates a modern randomized phase 3 statistical program built around two primary time-to-event endpoints and supported by multiple secondary endpoint types. The primary PFS analysis reported an HR of 0.72 with a 95% CI of 0.59–0.87 and P = 0.0004; the primary OS analysis reported an HR of 0.73 with a 95% CI of 0.57–0.94 and P = 0.0076.

The trial also demonstrates why statistical interpretation cannot be reduced to a single p-value. PFS and OS use hazard ratios; fixed-time survival outcomes use differences in survival rates; response outcomes use risk differences with score-based confidence intervals; and patient-reported outcomes use differences in least-squares means from a cLDA model. The analysis populations also differ by endpoint, making the definition of each denominator part of the statistical result.

Clinical Biostats methodology: A trial-results page should distinguish the reported numerical evidence from educational interpretation. For KEYNOTE-A18, the ClinicalTrials.gov record supports detailed interpretation of survival, response, patient-reported outcome, stratification, and safety analyses without adding unreported medians, subgroup estimates, treatment schedules, or other numerical results.