← Clinical Trials
Advanced Melanoma Phase 3 Randomized NCT01866319

KEYNOTE-006: Complete Statistical Analysis of Pembrolizumab in Advanced Melanoma

An independent statistical analysis of the randomized phase 3 KEYNOTE-006 trial comparing two pembrolizumab dosing schedules with ipilimumab in participants with advanced melanoma, using the statistical analyses and results posted on ClinicalTrials.gov.

Trial start: 28-Aug-2013  ·  Primary completion: 03-Mar-2015  ·  Status: Completed
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov record. The registry record is the official source for the trial design and posted analyses.

Registry record: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. View NCT01866319 on ClinicalTrials.gov.

1. Trial at a Glance

KEYNOTE-006 was a completed randomized phase 3 trial evaluating the safety and efficacy of two different pembrolizumab dosing schedules compared with ipilimumab in participants with advanced melanoma.

834
Enrolled
3 randomized arms
3
Arms
2 pembrolizumab schedules + ipilimumab
0.58
PFS HR
Q2W vs ipilimumab
0.63
12-Month OS HR
Q2W vs ipilimumab
FeatureKEYNOTE-006
PhasePhase 3
ConditionMelanoma
PopulationParticipants with advanced melanoma
DesignRandomized, parallel-group
AllocationRandomized
MaskingNone
Primary purposeTreatment
Enrollment834.0
InterventionsPembrolizumab and ipilimumab
Primary endpoints2 time-to-event endpoints
Results postedYes
Statistical analyses posted9
Lead sponsorMerck Sharp & Dohme LLC
Sponsor typeIndustry
ClinicalTrials.govNCT01866319

2. Clinical Question

The central statistical question was whether either pembrolizumab dosing schedule produced different progression-free survival and overall survival outcomes from ipilimumab in participants with advanced melanoma, while also comparing the two pembrolizumab schedules with each other.

Population

Participants with advanced melanoma enrolled in the randomized phase 3 trial.

Intervention

Pembrolizumab, evaluated using two dosing schedules: Q2W and Q3W.

Comparator

Ipilimumab.

Primary question

How do the two pembrolizumab schedules compare with ipilimumab for PFS and 12-month OS, and how do the two pembrolizumab schedules compare with each other?

3. Trial Design

01
Randomize834 participants
02
3 armsTwo pembrolizumab schedules + ipilimumab
03
Follow-upPFS and OS assessment
04
ResponseRECIST 1.1 / IRO
05
SafetySerious adverse events by arm
ARM · PEMBROLIZUMAB Q2W

Pembrolizumab every 2 weeks

  • Pembrolizumab biological intervention
  • Compared directly with ipilimumab for the primary analyses
  • Also compared directly with the Q3W pembrolizumab schedule
ARM · PEMBROLIZUMAB Q3W

Pembrolizumab every 3 weeks

  • Pembrolizumab biological intervention
  • Compared directly with ipilimumab for the primary analyses
  • Also compared directly with the Q2W pembrolizumab schedule
ARM · IPILIMUMAB

Ipilimumab

  • Ipilimumab biological intervention
  • Reference comparator for the two primary pembrolizumab comparisons
Design interpretation: The ClinicalTrials.gov record identifies the allocation as randomized and the design model as parallel, with no masking. The ClinicalTrials.gov record does not provide an allocation ratio or arm-level enrollment totals for all three groups, so those quantities are not inferred here.

4. Primary Endpoints

EndpointRegistry definition / time frameEndpoint type
Progression-free Survival (PFS) According to RECIST 1.1 as Assessed by Independent Radiology Plus Oncology Review (IRO) Time from randomization to the first documented disease progression, based on blinded Independent Radiology plus Oncology review (IRO) using RECIST 1.1, or death due to any cause, whichever occurred first. Time frame: up to approximately 12 months, through the first pre-specified statistical analysis cut-off date of 03-Sep-2014. Time-to-event
Percentage of Participants With Overall Survival (OS) at 12 Months OS was defined as the time from randomization to death due to any cause. The percentage of participants with OS at 12 months was estimated using a product-limit (Kaplan-Meier) method for censored data; data were censored at the date of cut-off. Time frame: Month 12. Time-to-event
Endpoint distinction: PFS is a time-to-event endpoint whose event is progression or death. The OS endpoint reports the percentage of participants alive at Month 12, with the registry describing a Kaplan-Meier product-limit estimate. These are related survival concepts but are not interchangeable outcomes.

5. Statistical Methodology

Intention-to-treat analysis

The primary analyses used the ITT population, defined in the ClinicalTrials.gov record as all participants as randomized to a study arm. This means the primary treatment comparisons are anchored to randomization rather than being restricted to participants who remained on treatment.

Stratified time-to-event testing

The PFS comparisons were performed using the log-rank test, while the OS comparisons used a Cox proportional-hazards model. For the reported primary comparisons, statistical testing was stratified by line of therapy (1st vs. 2nd), PD-L1 status (positive vs. negative), and ECOG performance status (0 vs. 1).

Hazard ratio framework
HR = estimated hazard in comparison group / estimated hazard in reference group

For these analyses, an HR below 1 indicates a lower estimated instantaneous event hazard in the first-named comparison group relative to the second-named group. The HR is a relative time-to-event measure; it is not a percentage of participants who benefit.

Score-based confidence intervals for response

Objective response rate was analyzed as a binary endpoint. The registry method is reported as Miettinen & Nurminen and normalized in the ClinicalTrials.gov record as a score-based confidence interval approach for proportions, including Miettinen-Nurminen, Newcombe, and Wilson methods. The effect measure was a risk difference, reported as a percent difference.

Superiority framework

All nine registry-reported statistical analyses are identified as superiority analyses. The registry therefore frames the reported comparisons as tests for evidence of a difference favoring one treatment group rather than as non-inferiority comparisons against a prespecified margin.

Analysis componentRegistry-supported approach
Primary efficacy populationITT; all participants as randomized
PFS comparisonLog-rank test
PFS effect measureHazard ratio
OS comparisonCox proportional-hazards model
OS effect measureHazard ratio
ORR comparisonScore-based CI for proportions
ORR effect measureRisk difference / percent difference
Hypothesis typeSuperiority
Stratification factorsLine of therapy, PD-L1 status, ECOG performance status

6. Primary Results: Progression-Free Survival

The registry reports three pairwise primary PFS comparisons, all based on the ITT population and the first pre-specified statistical analysis cut-off date of 03-Sep-2014.

Ipilimumab vs Pembrolizumab Q2W

PFS hazard ratio

0.58

95% CI: 0.46–0.72   ·   P < 0.00001

Two-sided superiority analysis using the log-rank test.

Clinical Biostats interpretation

An HR of 0.58 means that, under the time-to-event comparison represented by the reported analysis, the estimated instantaneous hazard of progression or death for pembrolizumab Q2W relative to ipilimumab was 0.58. Expressed as a simple relative-hazard interpretation, this corresponds to an estimated 42% lower hazard for the pembrolizumab Q2W group relative to ipilimumab.

The HR does not mean that 42% of participants avoided progression, that 42% more participants were cured, or that every participant experienced the same proportional reduction in risk. It is a relative time-to-event measure.

The 95% CI of 0.46–0.72 describes statistical uncertainty around the estimated HR. It is not a range containing the individual treatment effects experienced by patients.

The P-value of <0.00001 addresses evidence against the null comparison under the stated statistical framework. It does not measure the size or clinical importance of the treatment effect.

Because this is a time-to-event analysis, interpretation also depends on censoring and the assumptions underlying the survival-analysis framework. The ClinicalTrials.gov record identifies the analysis as stratified by line of therapy, PD-L1 status, and ECOG performance status.

Ipilimumab vs Pembrolizumab Q3W

PFS hazard ratio

0.58

95% CI: 0.47–0.72   ·   P < 0.00001

Two-sided superiority analysis using the log-rank test.

Clinical Biostats interpretation

An HR of 0.58 indicates an estimated instantaneous hazard of progression or death equal to 58% of that in the ipilimumab comparison group for pembrolizumab Q3W. As a relative-hazard statement, that corresponds to an estimated 42% lower hazard.

Again, this is not equivalent to a 42% absolute reduction in progression or death, nor does it specify how long an individual participant will remain progression-free. The confidence interval of 0.47–0.72 communicates the precision of the estimated relative effect, while the P-value of <0.00001 addresses statistical evidence rather than effect magnitude.

The registry specifies the ITT population and stratification by line of therapy, PD-L1 status, and ECOG performance status. The result therefore should be interpreted as the randomized comparison represented by that analysis rather than as an estimate restricted to treatment-adherent participants.

Pembrolizumab Q2W vs Pembrolizumab Q3W

PFS hazard ratio

0.97

95% CI: 0.77–1.21   ·   P = 0.75869

Two-sided superiority analysis using the log-rank test.

Clinical Biostats interpretation

An HR of 0.97 is close to 1.00, meaning the estimated instantaneous hazard of progression or death was similar between the two pembrolizumab dosing schedules in this comparison.

The 95% CI of 0.77–1.21 spans 1.00. Thus, the ClinicalTrials.gov record does not establish a statistically detectable superiority difference between the two schedules under this superiority analysis.

The P-value of 0.75869 is not a probability that the two dosing schedules are identical, and it does not prove equivalence. It simply indicates that the observed data do not provide strong evidence against the null comparison under the specified superiority framework.

This distinction is important: failure to demonstrate superiority is not the same statistical claim as demonstrating equivalence or non-inferiority. No non-inferiority margin is reported in the ClinicalTrials.gov record, so none is inferred here.

7. Primary Results: Overall Survival at 12 Months

The second primary endpoint was the percentage of participants with overall survival at Month 12. The registry defines OS as time from randomization to death due to any cause and states that the reported percentage was estimated using a product-limit Kaplan-Meier method with censoring at the date of cut-off.

Ipilimumab vs Pembrolizumab Q2W

OS hazard ratio at the Month 12 analysis

0.63

95% CI: 0.47–0.83   ·   P = 0.00052

Two-sided superiority analysis using a Cox proportional-hazards model.

Clinical Biostats interpretation

An HR of 0.63 means that the estimated instantaneous hazard of death for pembrolizumab Q2W relative to ipilimumab was 0.63 under the reported Cox model. As a relative-hazard interpretation, this corresponds to an estimated 37% lower hazard of death.

The result does not mean that 37% of participants were saved, that survival increased by exactly 37%, or that an individual patient's probability of survival changed by 37 percentage points. Those are different statistical quantities.

The 95% CI of 0.47–0.83 expresses uncertainty around the HR estimate. The P-value of 0.00052 measures evidence against the null hypothesis in the specified analysis; it is not an effect-size metric.

The Cox proportional-hazards model also carries a proportional-hazards interpretation. The ClinicalTrials.gov record does not provide a diagnostic assessment of that assumption, so the HR should be understood as the model-based summary reported by the trial rather than as a direct description of every point in time.

Ipilimumab vs Pembrolizumab Q3W

OS hazard ratio at the Month 12 analysis

0.69

95% CI: 0.52–0.90   ·   P = 0.00358

Two-sided superiority analysis using a Cox proportional-hazards model.

Clinical Biostats interpretation

An HR of 0.69 corresponds to an estimated instantaneous hazard of death equal to 69% of the ipilimumab comparison group, or an estimated 31% lower hazard under the reported model.

The 95% CI of 0.52–0.90 remains below 1.00, indicating that the reported interval is compatible with a lower hazard for pembrolizumab Q3W under this analysis. The P-value of 0.00358 quantifies evidence against the null comparison; it does not quantify how large or clinically important the effect is.

As with the other OS comparison, the estimate is based on a Cox model and the ITT population. The registry-reported analysis notes specify stratification by line of therapy, PD-L1 status, and ECOG performance status.

Pembrolizumab Q2W vs Pembrolizumab Q3W

OS hazard ratio at the Month 12 analysis

0.91

95% CI: 0.67–1.22   ·   P = 0.51319

Two-sided superiority analysis using a Cox proportional-hazards model.

Clinical Biostats interpretation

An HR of 0.91 is close to 1.00 and corresponds to an estimated hazard of death 9% lower for Q2W than Q3W under the reported Cox model. The estimate alone should not be interpreted as establishing a meaningful difference between schedules.

The 95% CI of 0.67–1.22 spans 1.00, and the P-value of 0.51319 does not provide strong evidence of superiority of one schedule over the other under the reported analysis.

Importantly, this does not establish equivalence. An equivalence conclusion would require an equivalence framework and prespecified margins, neither of which is reported in the ClinicalTrials.gov record.

8. Primary Results Summary

Primary endpointComparisonMethodEffect95% CIP-value
PFSIpilimumab vs Pembrolizumab Q2WLog-rankHR 0.580.46–0.72<0.00001
PFSIpilimumab vs Pembrolizumab Q3WLog-rankHR 0.580.47–0.72<0.00001
PFSPembrolizumab Q2W vs Q3WLog-rankHR 0.970.77–1.210.75869
OS at Month 12Ipilimumab vs Pembrolizumab Q2WCox modelHR 0.630.47–0.830.00052
OS at Month 12Ipilimumab vs Pembrolizumab Q3WCox modelHR 0.690.52–0.900.00358
OS at Month 12Pembrolizumab Q2W vs Q3WCox modelHR 0.910.67–1.220.51319
How to read the table: The two pembrolizumab-versus-ipilimumab comparisons show HR estimates below 1 for both primary endpoints. The direct Q2W-versus-Q3W comparisons are different questions: their HR estimates are closer to 1 and their confidence intervals include 1.00. These comparisons should not be collapsed into a single overall treatment effect.

9. Secondary Results: Objective Response Rate

Objective Response Rate (ORR) according to RECIST 1.1 as assessed by IRO was a secondary endpoint. The analysis population was the ITT population. The registry reports a risk difference, expressed as a percent difference, with score-based confidence intervals and superiority testing.

Ipilimumab vs Pembrolizumab Q2W

Risk difference in objective response rate

16.1 percentage points

95% CI: 7.8–24.5   ·   P = 0.00013

Clinical Biostats interpretation

The reported risk difference of 16.1 means the estimated percentage-point difference in objective response rate between the two groups was 16.1 percentage points in the direction of the first-named pembrolizumab Q2W comparison group relative to ipilimumab.

This is an absolute percentage-point difference, not a relative risk and not a hazard ratio. The 95% CI of 7.8–24.5 describes uncertainty around the estimated difference. The P-value of 0.00013 addresses statistical evidence against the null comparison, not the size of the response difference itself.

Ipilimumab vs Pembrolizumab Q3W

Risk difference in objective response rate

17.2 percentage points

95% CI: 9.5–25.6   ·   P = 0.00002

Clinical Biostats interpretation

The estimated response-rate difference was 17.2 percentage points for pembrolizumab Q3W relative to ipilimumab in the reported comparison. The confidence interval of 9.5–25.6 describes uncertainty around this absolute difference.

The P-value of 0.00002 is evidence against the null comparison under the reported superiority analysis. It does not mean that the probability of the treatment effect being exactly zero is 0.00002, nor does it quantify the clinical value of a response.

Pembrolizumab Q2W vs Pembrolizumab Q3W

Risk difference in objective response rate

−1.1 percentage points

95% CI: −10.6–8.6   ·   P = 0.82636

Clinical Biostats interpretation

The estimated risk difference of −1.1 percentage points is close to zero. The 95% CI of −10.6–8.6 includes zero, and the P-value of 0.82636 does not provide evidence of superiority of one pembrolizumab schedule over the other for ORR.

As with the PFS and OS schedule comparisons, absence of evidence for superiority is not proof of equivalence. A formal equivalence or non-inferiority conclusion would require a prespecified margin and the corresponding hypothesis-testing framework.

Secondary endpointComparisonEffect95% CIP-value
ORRIpilimumab vs Pembrolizumab Q2WRisk difference 16.17.8–24.50.00013
ORRIpilimumab vs Pembrolizumab Q3WRisk difference 17.29.5–25.60.00002
ORRPembrolizumab Q2W vs Q3WRisk difference −1.1−10.6–8.60.82636

10. Safety: Serious Adverse Events by Arm

The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected participants divided by participants at risk. These counts provide an arm-level safety summary without requiring an inference about treatment exposure or causality beyond the registry field.

ArmSerious adverse eventsAt risk
Ipilimumab77256
Pembrolizumab Q2W89278
Pembrolizumab Q3W90277
Serious adverse events: affected participants / at risk
Ipilimumab
77 / 256
Pembrolizumab Q2W
89 / 278
Pembrolizumab Q3W
90 / 277

The serious-adverse-event field should be kept separate from the efficacy results. An efficacy hazard ratio does not summarize safety, and a safety count does not establish a comparative causal effect without considering the full safety analysis, exposure, follow-up, event definitions, and appropriate statistical framework.

11. Statistical Methods Explained

Why was the log-rank test used for PFS?

PFS is a time-to-event endpoint because participants can experience progression or death at different times, while some participants may be censored before an event is observed. The log-rank test is designed to compare survival distributions between randomized groups while incorporating the timing of events and censoring rather than reducing follow-up to a simple binary outcome.

What does a hazard ratio of 0.58 mean?

An HR of 0.58 indicates that the estimated instantaneous event hazard in the first comparison group is 58% of the reference group's hazard under the reported time-to-event analysis. A convenient descriptive transformation is \(1-0.58=0.42\), or a 42% lower estimated hazard. This does not mean a 42% absolute reduction in the probability of an event.

Why was a Cox proportional-hazards model used for OS?

The registry reports Cox regression for the OS endpoint. A Cox model provides a way to estimate a relative hazard while accounting for the time-to-event structure and censoring. The resulting HR is model-based and should not be interpreted as a simple risk ratio.

Why were the analyses stratified?

The registry specifies stratification by line of therapy, PD-L1 status, and ECOG performance status. Stratification allows the time-to-event comparison to account for these prespecified factors rather than treating the entire randomized population as statistically homogeneous with respect to those variables.

What does the 95% confidence interval tell us?

A 95% confidence interval describes the statistical uncertainty around an estimated effect under the corresponding model and sampling framework. For example, the PFS HR of 0.58 for Q2W versus ipilimumab has a 95% CI of 0.46–0.72. The interval does not describe the range of effects that individual patients experienced.

Why is a risk difference used for ORR rather than a hazard ratio?

ORR is a binary endpoint: a participant either meets the prespecified response definition or does not. A risk difference therefore describes an absolute difference in the proportion responding. The registry used a score-based approach for confidence intervals rather than applying a time-to-event method to this endpoint.

Why does a nonsignificant Q2W-versus-Q3W comparison not prove equivalence?

The direct schedule comparisons were analyzed under a superiority hypothesis. A P-value such as 0.75869 indicates that the data do not provide evidence of superiority under that test. It does not demonstrate that the schedules are equivalent or that their effects differ by less than a clinically acceptable amount. Such conclusions require prespecified equivalence or non-inferiority margins.

12. Confidence Intervals, P-values, and Effect Size

Effect estimate

The HR or risk difference describes the magnitude and direction of the observed statistical comparison. It is the starting point for interpretation, not the entire conclusion.

Confidence interval

The confidence interval communicates precision. Narrower intervals generally indicate more precise estimation than wider intervals, all else equal.

P-value

The P-value addresses compatibility with a null hypothesis under the specified statistical model. It does not measure effect size or clinical importance.

Clinical meaning

A statistically detectable difference still needs clinical context. Relative effects, absolute effects, response outcomes, time-to-event outcomes, and safety describe different dimensions of a trial.

KEYNOTE-006 is particularly useful for showing why these quantities should be read together. The two pembrolizumab-versus-ipilimumab PFS comparisons both have HR estimates of 0.58, but their confidence intervals differ slightly. The direct schedule comparison has an HR of 0.97 with a wider interval spanning 1.00. The corresponding ORR analyses similarly distinguish a positive risk difference against ipilimumab from a near-zero difference between pembrolizumab schedules.

13. Intention-to-Treat Analysis and Randomization

The registry defines the primary analysis population as the ITT population, comprising all participants as randomized to a study arm. This is a central principle of randomized trial analysis because it preserves the treatment assignment created by randomization.

Why ITT matters
Analyze participants according to randomized assignment → preserve the randomized comparison

The ITT principle helps prevent post-randomization treatment decisions from redefining the primary efficacy comparison. It does not mean that every participant contributes identical amounts of follow-up information; survival methods explicitly account for differing event and censoring times.

The ClinicalTrials.gov record also identify stratification factors used in the statistical testing: line of therapy, PD-L1 status, and ECOG performance status. These factors are relevant to how the randomized comparison was statistically constructed.

14. Time-to-Event Endpoints: PFS and OS

Both primary endpoints are classified as time-to-event endpoints in the ClinicalTrials.gov record, but they represent different events.

FeaturePFSOS at 12 months
Starting pointRandomizationRandomization
EventFirst documented disease progression or deathDeath due to any cause
AssessmentIRO using RECIST 1.1Death-based endpoint
Registry time frameUp to approximately 12 months; cut-off 03-Sep-2014Month 12
Primary statistical comparisonLog-rank test; HRCox proportional-hazards model; HR

The distinction matters because progression can occur before death, while OS counts only death. Consequently, a treatment can affect PFS and OS differently, and the two endpoints should not be treated as interchangeable measures of treatment effect.

15. What the Hazard Ratios Do — and Do Not — Mean

PFS example

An HR of 0.58 for pembrolizumab Q2W versus ipilimumab means that the estimated instantaneous hazard of progression or death was 58% of the ipilimumab comparison hazard under the reported analysis. It does not mean that 42% of participants avoided progression or that the median PFS was reduced or increased by 42%.

OS example

An HR of 0.63 for pembrolizumab Q2W versus ipilimumab means an estimated instantaneous hazard of death equal to 63% of the ipilimumab comparison hazard under the reported Cox model. It does not mean that 37% of participants survived because of treatment or that individual survival probabilities changed by 37 percentage points.

Schedule comparison

An HR of 0.91 for Q2W versus Q3W is close to 1.00. Its 95% CI of 0.67–1.22 demonstrates why the point estimate alone is insufficient: the plausible statistical uncertainty includes both values below and above 1.00.

16. Multiplicity and Multiple Pairwise Comparisons

The statistical analyses posted on ClinicalTrials.gov contain three pairwise comparisons for each of the two primary endpoints: each pembrolizumab schedule versus ipilimumab, and Q2W versus Q3W. The same three pairwise structures are also reported for ORR as a secondary endpoint.

EndpointPairwise comparisons reportedNumber of analyses
PFSIpilimumab vs Q2W; Ipilimumab vs Q3W; Q2W vs Q3W3
OS at 12 monthsIpilimumab vs Q2W; Ipilimumab vs Q3W; Q2W vs Q3W3
ORRIpilimumab vs Q2W; Ipilimumab vs Q3W; Q2W vs Q3W3
Multiplicity caution: The ClinicalTrials.gov record identifies the analyses and their P-values but do not provide a complete multiplicity-adjustment procedure, alpha allocation, or hierarchical testing sequence. Therefore, the reported P-values should not be reinterpreted as a newly reconstructed familywise-error analysis. No unreported adjustment is inferred.

This is an important distinction in a multi-arm trial. Multiple statistical questions can be clinically useful, but the evidentiary interpretation of each P-value depends on the prespecified testing strategy. The ClinicalTrials.gov record does not provide enough information to reconstruct an additional multiplicity procedure beyond the reported analyses.

17. Stratified Analysis

The statistical testing for the reported primary comparisons was stratified by three factors:

Stratum 1
Line of therapy: 1st vs. 2nd
Stratum 2
PD-L1 status: positive vs. negative
Stratum 3
ECOG performance status: 0 vs. 1
Analysis implication
The treatment comparison was evaluated within a framework accounting for these prespecified strata.

Stratification is different from simply reporting subgroup-specific treatment effects. A stratified analysis incorporates the specified strata into the primary comparison, whereas a subgroup analysis asks whether effects appear different within particular subsets. The ClinicalTrials.gov record supports the former but do not provide a set of subgroup-specific estimates for these factors.

18. Missing Data, Censoring, and What the Registry Does Not Report

Time-to-event analyses necessarily involve censoring when a participant has not experienced the event by the last available observation. The registry definition for 12-month OS explicitly states that the percentage was estimated using a product-limit Kaplan-Meier method and that data were censored at the date of cut-off.

The ClinicalTrials.gov record does not specify a separate imputation strategy for missing efficacy observations, nor do they provide detailed rules for sensitivity analyses addressing missing data. Accordingly, no imputation method is inferred.

Data boundary: The ClinicalTrials.gov record does not report median PFS, median OS, detailed baseline characteristics, subgroup forest plots, crossover information, a non-inferiority margin, a Bayesian analysis, or an interim-analysis/alpha-spending procedure. Those topics are therefore not presented as features of this trial page.

19. Limitations

20. Why This Trial Matters Statistically

KEYNOTE-006 is a useful teaching case because its registry-posted analyses illustrate how a randomized multi-arm clinical trial can generate several related but distinct statistical questions.

ConceptHow it appears in KEYNOTE-006
RandomizationParticipants were randomized in a parallel-group phase 3 design.
Three-arm comparisonTwo pembrolizumab schedules were each compared with ipilimumab and with each other.
ITT analysisPrimary analyses used all participants as randomized.
Time-to-event analysisPFS and OS were classified as time-to-event endpoints.
Log-rank testUsed for the posted PFS comparisons.
Cox modelUsed for the posted OS comparisons.
Hazard ratioUsed to express relative effects for PFS and OS.
Risk differenceUsed for the ORR comparisons.
Confidence intervals95% two-sided intervals were reported for the posted primary and secondary effects.
Stratified analysisTesting was stratified by line of therapy, PD-L1 status, and ECOG performance status.
Superiority testingThe analyses posted on ClinicalTrials.gov were designated as superiority hypotheses.

The most instructive feature is the distinction between treatment-versus-comparator questions and schedule-versus-schedule questions. The pembrolizumab-versus-ipilimumab estimates are below 1 for both primary endpoints, while the direct Q2W-versus-Q3W comparisons have estimates close to 1 and confidence intervals that include the null value. These are separate statistical statements and should not be conflated.

21. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The posted primary analyses show HR estimates below 1 for PFS and 12-month OS when each pembrolizumab schedule is compared with ipilimumab. The direct schedule comparisons do not show statistical evidence of superiority under the reported tests.

Effect-size interpretation

The magnitude of an effect depends on the measure. HRs describe relative time-to-event hazards, while the ORR analyses report absolute percentage-point differences.

Uncertainty interpretation

The confidence intervals quantify uncertainty around each estimate. They are essential for understanding precision and should be read alongside the point estimate and P-value.

Safety interpretation

Serious adverse events are reported separately by arm. Safety counts answer a different question from the efficacy hazard ratios and should be interpreted independently.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Calculators

24. Sources

Continue through the Clinical Biostats statistical pathway

Explore the survival-analysis, clinical-trial, confidence-interval, and binary-outcome methods that appear in KEYNOTE-006.

25. Record Summary

KEYNOTE-006 provides a compact example of several core clinical-trial statistical methods: randomized parallel-group comparison, ITT analysis, stratified time-to-event testing, log-rank analysis, Cox proportional-hazards modeling, hazard ratios, Kaplan-Meier estimation, risk differences for binary response, and two-sided confidence intervals. The registry reports three primary PFS comparisons and three primary OS comparisons, together with three secondary ORR comparisons.

The statistical story is most clearly understood by keeping the comparison questions separate. Against ipilimumab, the two pembrolizumab schedules produced reported PFS HR estimates of 0.58, while the reported OS HRs at Month 12 were 0.63 for Q2W and 0.69 for Q3W. The direct Q2W-versus-Q3W comparisons produced HRs of 0.97 for PFS and 0.91 for OS, with confidence intervals spanning the respective null value. The ORR analyses show the same distinction between treatment-versus-comparator and schedule-versus-schedule questions.

Clinical Biostats methodology: A trial-results page should distinguish the reported statistical evidence from the interpretation of that evidence. For KEYNOTE-006, that means preserving the registry's endpoint definitions, analysis populations, methods, effect measures, confidence intervals, P-values, and stratification factors while avoiding unsupported claims about unreported trial features.