This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov record. The registry record is the official source for the trial design and posted analyses.
1. Trial at a Glance
KEYNOTE-006 was a completed randomized phase 3 trial evaluating the safety and efficacy of two different pembrolizumab dosing schedules compared with ipilimumab in participants with advanced melanoma.
| Feature | KEYNOTE-006 |
|---|---|
| Phase | Phase 3 |
| Condition | Melanoma |
| Population | Participants with advanced melanoma |
| Design | Randomized, parallel-group |
| Allocation | Randomized |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 834.0 |
| Interventions | Pembrolizumab and ipilimumab |
| Primary endpoints | 2 time-to-event endpoints |
| Results posted | Yes |
| Statistical analyses posted | 9 |
| Lead sponsor | Merck Sharp & Dohme LLC |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT01866319 |
2. Clinical Question
The central statistical question was whether either pembrolizumab dosing schedule produced different progression-free survival and overall survival outcomes from ipilimumab in participants with advanced melanoma, while also comparing the two pembrolizumab schedules with each other.
Population
Participants with advanced melanoma enrolled in the randomized phase 3 trial.
Intervention
Pembrolizumab, evaluated using two dosing schedules: Q2W and Q3W.
Comparator
Ipilimumab.
Primary question
How do the two pembrolizumab schedules compare with ipilimumab for PFS and 12-month OS, and how do the two pembrolizumab schedules compare with each other?
3. Trial Design
Pembrolizumab every 2 weeks
- Pembrolizumab biological intervention
- Compared directly with ipilimumab for the primary analyses
- Also compared directly with the Q3W pembrolizumab schedule
Pembrolizumab every 3 weeks
- Pembrolizumab biological intervention
- Compared directly with ipilimumab for the primary analyses
- Also compared directly with the Q2W pembrolizumab schedule
Ipilimumab
- Ipilimumab biological intervention
- Reference comparator for the two primary pembrolizumab comparisons
4. Primary Endpoints
| Endpoint | Registry definition / time frame | Endpoint type |
|---|---|---|
| Progression-free Survival (PFS) According to RECIST 1.1 as Assessed by Independent Radiology Plus Oncology Review (IRO) | Time from randomization to the first documented disease progression, based on blinded Independent Radiology plus Oncology review (IRO) using RECIST 1.1, or death due to any cause, whichever occurred first. Time frame: up to approximately 12 months, through the first pre-specified statistical analysis cut-off date of 03-Sep-2014. | Time-to-event |
| Percentage of Participants With Overall Survival (OS) at 12 Months | OS was defined as the time from randomization to death due to any cause. The percentage of participants with OS at 12 months was estimated using a product-limit (Kaplan-Meier) method for censored data; data were censored at the date of cut-off. Time frame: Month 12. | Time-to-event |
5. Statistical Methodology
Intention-to-treat analysis
The primary analyses used the ITT population, defined in the ClinicalTrials.gov record as all participants as randomized to a study arm. This means the primary treatment comparisons are anchored to randomization rather than being restricted to participants who remained on treatment.
Stratified time-to-event testing
The PFS comparisons were performed using the log-rank test, while the OS comparisons used a Cox proportional-hazards model. For the reported primary comparisons, statistical testing was stratified by line of therapy (1st vs. 2nd), PD-L1 status (positive vs. negative), and ECOG performance status (0 vs. 1).
For these analyses, an HR below 1 indicates a lower estimated instantaneous event hazard in the first-named comparison group relative to the second-named group. The HR is a relative time-to-event measure; it is not a percentage of participants who benefit.
Score-based confidence intervals for response
Objective response rate was analyzed as a binary endpoint. The registry method is reported as Miettinen & Nurminen and normalized in the ClinicalTrials.gov record as a score-based confidence interval approach for proportions, including Miettinen-Nurminen, Newcombe, and Wilson methods. The effect measure was a risk difference, reported as a percent difference.
Superiority framework
All nine registry-reported statistical analyses are identified as superiority analyses. The registry therefore frames the reported comparisons as tests for evidence of a difference favoring one treatment group rather than as non-inferiority comparisons against a prespecified margin.
| Analysis component | Registry-supported approach |
|---|---|
| Primary efficacy population | ITT; all participants as randomized |
| PFS comparison | Log-rank test |
| PFS effect measure | Hazard ratio |
| OS comparison | Cox proportional-hazards model |
| OS effect measure | Hazard ratio |
| ORR comparison | Score-based CI for proportions |
| ORR effect measure | Risk difference / percent difference |
| Hypothesis type | Superiority |
| Stratification factors | Line of therapy, PD-L1 status, ECOG performance status |
6. Primary Results: Progression-Free Survival
The registry reports three pairwise primary PFS comparisons, all based on the ITT population and the first pre-specified statistical analysis cut-off date of 03-Sep-2014.
Ipilimumab vs Pembrolizumab Q2W
PFS hazard ratio
95% CI: 0.46–0.72 · P < 0.00001
Two-sided superiority analysis using the log-rank test.
An HR of 0.58 means that, under the time-to-event comparison represented by the reported analysis, the estimated instantaneous hazard of progression or death for pembrolizumab Q2W relative to ipilimumab was 0.58. Expressed as a simple relative-hazard interpretation, this corresponds to an estimated 42% lower hazard for the pembrolizumab Q2W group relative to ipilimumab.
The HR does not mean that 42% of participants avoided progression, that 42% more participants were cured, or that every participant experienced the same proportional reduction in risk. It is a relative time-to-event measure.
The 95% CI of 0.46–0.72 describes statistical uncertainty around the estimated HR. It is not a range containing the individual treatment effects experienced by patients.
The P-value of <0.00001 addresses evidence against the null comparison under the stated statistical framework. It does not measure the size or clinical importance of the treatment effect.
Because this is a time-to-event analysis, interpretation also depends on censoring and the assumptions underlying the survival-analysis framework. The ClinicalTrials.gov record identifies the analysis as stratified by line of therapy, PD-L1 status, and ECOG performance status.
Ipilimumab vs Pembrolizumab Q3W
PFS hazard ratio
95% CI: 0.47–0.72 · P < 0.00001
Two-sided superiority analysis using the log-rank test.
An HR of 0.58 indicates an estimated instantaneous hazard of progression or death equal to 58% of that in the ipilimumab comparison group for pembrolizumab Q3W. As a relative-hazard statement, that corresponds to an estimated 42% lower hazard.
Again, this is not equivalent to a 42% absolute reduction in progression or death, nor does it specify how long an individual participant will remain progression-free. The confidence interval of 0.47–0.72 communicates the precision of the estimated relative effect, while the P-value of <0.00001 addresses statistical evidence rather than effect magnitude.
The registry specifies the ITT population and stratification by line of therapy, PD-L1 status, and ECOG performance status. The result therefore should be interpreted as the randomized comparison represented by that analysis rather than as an estimate restricted to treatment-adherent participants.
Pembrolizumab Q2W vs Pembrolizumab Q3W
PFS hazard ratio
95% CI: 0.77–1.21 · P = 0.75869
Two-sided superiority analysis using the log-rank test.
An HR of 0.97 is close to 1.00, meaning the estimated instantaneous hazard of progression or death was similar between the two pembrolizumab dosing schedules in this comparison.
The 95% CI of 0.77–1.21 spans 1.00. Thus, the ClinicalTrials.gov record does not establish a statistically detectable superiority difference between the two schedules under this superiority analysis.
The P-value of 0.75869 is not a probability that the two dosing schedules are identical, and it does not prove equivalence. It simply indicates that the observed data do not provide strong evidence against the null comparison under the specified superiority framework.
This distinction is important: failure to demonstrate superiority is not the same statistical claim as demonstrating equivalence or non-inferiority. No non-inferiority margin is reported in the ClinicalTrials.gov record, so none is inferred here.
7. Primary Results: Overall Survival at 12 Months
The second primary endpoint was the percentage of participants with overall survival at Month 12. The registry defines OS as time from randomization to death due to any cause and states that the reported percentage was estimated using a product-limit Kaplan-Meier method with censoring at the date of cut-off.
Ipilimumab vs Pembrolizumab Q2W
OS hazard ratio at the Month 12 analysis
95% CI: 0.47–0.83 · P = 0.00052
Two-sided superiority analysis using a Cox proportional-hazards model.
An HR of 0.63 means that the estimated instantaneous hazard of death for pembrolizumab Q2W relative to ipilimumab was 0.63 under the reported Cox model. As a relative-hazard interpretation, this corresponds to an estimated 37% lower hazard of death.
The result does not mean that 37% of participants were saved, that survival increased by exactly 37%, or that an individual patient's probability of survival changed by 37 percentage points. Those are different statistical quantities.
The 95% CI of 0.47–0.83 expresses uncertainty around the HR estimate. The P-value of 0.00052 measures evidence against the null hypothesis in the specified analysis; it is not an effect-size metric.
The Cox proportional-hazards model also carries a proportional-hazards interpretation. The ClinicalTrials.gov record does not provide a diagnostic assessment of that assumption, so the HR should be understood as the model-based summary reported by the trial rather than as a direct description of every point in time.
Ipilimumab vs Pembrolizumab Q3W
OS hazard ratio at the Month 12 analysis
95% CI: 0.52–0.90 · P = 0.00358
Two-sided superiority analysis using a Cox proportional-hazards model.
An HR of 0.69 corresponds to an estimated instantaneous hazard of death equal to 69% of the ipilimumab comparison group, or an estimated 31% lower hazard under the reported model.
The 95% CI of 0.52–0.90 remains below 1.00, indicating that the reported interval is compatible with a lower hazard for pembrolizumab Q3W under this analysis. The P-value of 0.00358 quantifies evidence against the null comparison; it does not quantify how large or clinically important the effect is.
As with the other OS comparison, the estimate is based on a Cox model and the ITT population. The registry-reported analysis notes specify stratification by line of therapy, PD-L1 status, and ECOG performance status.
Pembrolizumab Q2W vs Pembrolizumab Q3W
OS hazard ratio at the Month 12 analysis
95% CI: 0.67–1.22 · P = 0.51319
Two-sided superiority analysis using a Cox proportional-hazards model.
An HR of 0.91 is close to 1.00 and corresponds to an estimated hazard of death 9% lower for Q2W than Q3W under the reported Cox model. The estimate alone should not be interpreted as establishing a meaningful difference between schedules.
The 95% CI of 0.67–1.22 spans 1.00, and the P-value of 0.51319 does not provide strong evidence of superiority of one schedule over the other under the reported analysis.
Importantly, this does not establish equivalence. An equivalence conclusion would require an equivalence framework and prespecified margins, neither of which is reported in the ClinicalTrials.gov record.
8. Primary Results Summary
| Primary endpoint | Comparison | Method | Effect | 95% CI | P-value |
|---|---|---|---|---|---|
| PFS | Ipilimumab vs Pembrolizumab Q2W | Log-rank | HR 0.58 | 0.46–0.72 | <0.00001 |
| PFS | Ipilimumab vs Pembrolizumab Q3W | Log-rank | HR 0.58 | 0.47–0.72 | <0.00001 |
| PFS | Pembrolizumab Q2W vs Q3W | Log-rank | HR 0.97 | 0.77–1.21 | 0.75869 |
| OS at Month 12 | Ipilimumab vs Pembrolizumab Q2W | Cox model | HR 0.63 | 0.47–0.83 | 0.00052 |
| OS at Month 12 | Ipilimumab vs Pembrolizumab Q3W | Cox model | HR 0.69 | 0.52–0.90 | 0.00358 |
| OS at Month 12 | Pembrolizumab Q2W vs Q3W | Cox model | HR 0.91 | 0.67–1.22 | 0.51319 |
9. Secondary Results: Objective Response Rate
Objective Response Rate (ORR) according to RECIST 1.1 as assessed by IRO was a secondary endpoint. The analysis population was the ITT population. The registry reports a risk difference, expressed as a percent difference, with score-based confidence intervals and superiority testing.
Ipilimumab vs Pembrolizumab Q2W
Risk difference in objective response rate
95% CI: 7.8–24.5 · P = 0.00013
The reported risk difference of 16.1 means the estimated percentage-point difference in objective response rate between the two groups was 16.1 percentage points in the direction of the first-named pembrolizumab Q2W comparison group relative to ipilimumab.
This is an absolute percentage-point difference, not a relative risk and not a hazard ratio. The 95% CI of 7.8–24.5 describes uncertainty around the estimated difference. The P-value of 0.00013 addresses statistical evidence against the null comparison, not the size of the response difference itself.
Ipilimumab vs Pembrolizumab Q3W
Risk difference in objective response rate
95% CI: 9.5–25.6 · P = 0.00002
The estimated response-rate difference was 17.2 percentage points for pembrolizumab Q3W relative to ipilimumab in the reported comparison. The confidence interval of 9.5–25.6 describes uncertainty around this absolute difference.
The P-value of 0.00002 is evidence against the null comparison under the reported superiority analysis. It does not mean that the probability of the treatment effect being exactly zero is 0.00002, nor does it quantify the clinical value of a response.
Pembrolizumab Q2W vs Pembrolizumab Q3W
Risk difference in objective response rate
95% CI: −10.6–8.6 · P = 0.82636
The estimated risk difference of −1.1 percentage points is close to zero. The 95% CI of −10.6–8.6 includes zero, and the P-value of 0.82636 does not provide evidence of superiority of one pembrolizumab schedule over the other for ORR.
As with the PFS and OS schedule comparisons, absence of evidence for superiority is not proof of equivalence. A formal equivalence or non-inferiority conclusion would require a prespecified margin and the corresponding hypothesis-testing framework.
| Secondary endpoint | Comparison | Effect | 95% CI | P-value |
|---|---|---|---|---|
| ORR | Ipilimumab vs Pembrolizumab Q2W | Risk difference 16.1 | 7.8–24.5 | 0.00013 |
| ORR | Ipilimumab vs Pembrolizumab Q3W | Risk difference 17.2 | 9.5–25.6 | 0.00002 |
| ORR | Pembrolizumab Q2W vs Q3W | Risk difference −1.1 | −10.6–8.6 | 0.82636 |
10. Safety: Serious Adverse Events by Arm
The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected participants divided by participants at risk. These counts provide an arm-level safety summary without requiring an inference about treatment exposure or causality beyond the registry field.
| Arm | Serious adverse events | At risk |
|---|---|---|
| Ipilimumab | 77 | 256 |
| Pembrolizumab Q2W | 89 | 278 |
| Pembrolizumab Q3W | 90 | 277 |
The serious-adverse-event field should be kept separate from the efficacy results. An efficacy hazard ratio does not summarize safety, and a safety count does not establish a comparative causal effect without considering the full safety analysis, exposure, follow-up, event definitions, and appropriate statistical framework.
11. Statistical Methods Explained
Why was the log-rank test used for PFS?
PFS is a time-to-event endpoint because participants can experience progression or death at different times, while some participants may be censored before an event is observed. The log-rank test is designed to compare survival distributions between randomized groups while incorporating the timing of events and censoring rather than reducing follow-up to a simple binary outcome.
What does a hazard ratio of 0.58 mean?
An HR of 0.58 indicates that the estimated instantaneous event hazard in the first comparison group is 58% of the reference group's hazard under the reported time-to-event analysis. A convenient descriptive transformation is \(1-0.58=0.42\), or a 42% lower estimated hazard. This does not mean a 42% absolute reduction in the probability of an event.
Why was a Cox proportional-hazards model used for OS?
The registry reports Cox regression for the OS endpoint. A Cox model provides a way to estimate a relative hazard while accounting for the time-to-event structure and censoring. The resulting HR is model-based and should not be interpreted as a simple risk ratio.
Why were the analyses stratified?
The registry specifies stratification by line of therapy, PD-L1 status, and ECOG performance status. Stratification allows the time-to-event comparison to account for these prespecified factors rather than treating the entire randomized population as statistically homogeneous with respect to those variables.
What does the 95% confidence interval tell us?
A 95% confidence interval describes the statistical uncertainty around an estimated effect under the corresponding model and sampling framework. For example, the PFS HR of 0.58 for Q2W versus ipilimumab has a 95% CI of 0.46–0.72. The interval does not describe the range of effects that individual patients experienced.
Why is a risk difference used for ORR rather than a hazard ratio?
ORR is a binary endpoint: a participant either meets the prespecified response definition or does not. A risk difference therefore describes an absolute difference in the proportion responding. The registry used a score-based approach for confidence intervals rather than applying a time-to-event method to this endpoint.
Why does a nonsignificant Q2W-versus-Q3W comparison not prove equivalence?
The direct schedule comparisons were analyzed under a superiority hypothesis. A P-value such as 0.75869 indicates that the data do not provide evidence of superiority under that test. It does not demonstrate that the schedules are equivalent or that their effects differ by less than a clinically acceptable amount. Such conclusions require prespecified equivalence or non-inferiority margins.
12. Confidence Intervals, P-values, and Effect Size
Effect estimate
The HR or risk difference describes the magnitude and direction of the observed statistical comparison. It is the starting point for interpretation, not the entire conclusion.
Confidence interval
The confidence interval communicates precision. Narrower intervals generally indicate more precise estimation than wider intervals, all else equal.
P-value
The P-value addresses compatibility with a null hypothesis under the specified statistical model. It does not measure effect size or clinical importance.
Clinical meaning
A statistically detectable difference still needs clinical context. Relative effects, absolute effects, response outcomes, time-to-event outcomes, and safety describe different dimensions of a trial.
KEYNOTE-006 is particularly useful for showing why these quantities should be read together. The two pembrolizumab-versus-ipilimumab PFS comparisons both have HR estimates of 0.58, but their confidence intervals differ slightly. The direct schedule comparison has an HR of 0.97 with a wider interval spanning 1.00. The corresponding ORR analyses similarly distinguish a positive risk difference against ipilimumab from a near-zero difference between pembrolizumab schedules.
13. Intention-to-Treat Analysis and Randomization
The registry defines the primary analysis population as the ITT population, comprising all participants as randomized to a study arm. This is a central principle of randomized trial analysis because it preserves the treatment assignment created by randomization.
The ITT principle helps prevent post-randomization treatment decisions from redefining the primary efficacy comparison. It does not mean that every participant contributes identical amounts of follow-up information; survival methods explicitly account for differing event and censoring times.
The ClinicalTrials.gov record also identify stratification factors used in the statistical testing: line of therapy, PD-L1 status, and ECOG performance status. These factors are relevant to how the randomized comparison was statistically constructed.
14. Time-to-Event Endpoints: PFS and OS
Both primary endpoints are classified as time-to-event endpoints in the ClinicalTrials.gov record, but they represent different events.
| Feature | PFS | OS at 12 months |
|---|---|---|
| Starting point | Randomization | Randomization |
| Event | First documented disease progression or death | Death due to any cause |
| Assessment | IRO using RECIST 1.1 | Death-based endpoint |
| Registry time frame | Up to approximately 12 months; cut-off 03-Sep-2014 | Month 12 |
| Primary statistical comparison | Log-rank test; HR | Cox proportional-hazards model; HR |
The distinction matters because progression can occur before death, while OS counts only death. Consequently, a treatment can affect PFS and OS differently, and the two endpoints should not be treated as interchangeable measures of treatment effect.
15. What the Hazard Ratios Do — and Do Not — Mean
An HR of 0.58 for pembrolizumab Q2W versus ipilimumab means that the estimated instantaneous hazard of progression or death was 58% of the ipilimumab comparison hazard under the reported analysis. It does not mean that 42% of participants avoided progression or that the median PFS was reduced or increased by 42%.
An HR of 0.63 for pembrolizumab Q2W versus ipilimumab means an estimated instantaneous hazard of death equal to 63% of the ipilimumab comparison hazard under the reported Cox model. It does not mean that 37% of participants survived because of treatment or that individual survival probabilities changed by 37 percentage points.
An HR of 0.91 for Q2W versus Q3W is close to 1.00. Its 95% CI of 0.67–1.22 demonstrates why the point estimate alone is insufficient: the plausible statistical uncertainty includes both values below and above 1.00.
16. Multiplicity and Multiple Pairwise Comparisons
The statistical analyses posted on ClinicalTrials.gov contain three pairwise comparisons for each of the two primary endpoints: each pembrolizumab schedule versus ipilimumab, and Q2W versus Q3W. The same three pairwise structures are also reported for ORR as a secondary endpoint.
| Endpoint | Pairwise comparisons reported | Number of analyses |
|---|---|---|
| PFS | Ipilimumab vs Q2W; Ipilimumab vs Q3W; Q2W vs Q3W | 3 |
| OS at 12 months | Ipilimumab vs Q2W; Ipilimumab vs Q3W; Q2W vs Q3W | 3 |
| ORR | Ipilimumab vs Q2W; Ipilimumab vs Q3W; Q2W vs Q3W | 3 |
This is an important distinction in a multi-arm trial. Multiple statistical questions can be clinically useful, but the evidentiary interpretation of each P-value depends on the prespecified testing strategy. The ClinicalTrials.gov record does not provide enough information to reconstruct an additional multiplicity procedure beyond the reported analyses.
17. Stratified Analysis
The statistical testing for the reported primary comparisons was stratified by three factors:
Stratification is different from simply reporting subgroup-specific treatment effects. A stratified analysis incorporates the specified strata into the primary comparison, whereas a subgroup analysis asks whether effects appear different within particular subsets. The ClinicalTrials.gov record supports the former but do not provide a set of subgroup-specific estimates for these factors.
18. Missing Data, Censoring, and What the Registry Does Not Report
Time-to-event analyses necessarily involve censoring when a participant has not experienced the event by the last available observation. The registry definition for 12-month OS explicitly states that the percentage was estimated using a product-limit Kaplan-Meier method and that data were censored at the date of cut-off.
The ClinicalTrials.gov record does not specify a separate imputation strategy for missing efficacy observations, nor do they provide detailed rules for sensitivity analyses addressing missing data. Accordingly, no imputation method is inferred.
19. Limitations
- Registry-level detail: The statistical ClinicalTrials.gov record contains the posted methods and estimates but not the full statistical analysis plan, so additional design details should not be inferred.
- Multiple comparisons: Three pairwise comparisons are reported for each of the two primary endpoints and for ORR. The ClinicalTrials.gov record does not provide a complete multiplicity-adjustment strategy.
- Hazard-ratio interpretation: HRs are model-based relative measures and are not equivalent to risk ratios, risk differences, or survival probabilities.
- Proportional-hazards assumption: The ClinicalTrials.gov record reports Cox proportional-hazards analysis but do not report a diagnostic assessment of the proportional-hazards assumption.
- Censoring: Survival analyses rely on censoring information. The ClinicalTrials.gov record identifies censoring for the 12-month OS estimate but do not provide a detailed censoring analysis.
- Superiority versus equivalence: A nonsignificant Q2W-versus-Q3W superiority test does not establish equivalence or non-inferiority.
- Safety interpretation: Serious adverse-event counts by arm are descriptive registry data and should not be treated as a complete safety analysis.
- Endpoint scope: PFS and OS measure different clinical events and should not be combined into one generic measure of benefit.
20. Why This Trial Matters Statistically
KEYNOTE-006 is a useful teaching case because its registry-posted analyses illustrate how a randomized multi-arm clinical trial can generate several related but distinct statistical questions.
| Concept | How it appears in KEYNOTE-006 |
|---|---|
| Randomization | Participants were randomized in a parallel-group phase 3 design. |
| Three-arm comparison | Two pembrolizumab schedules were each compared with ipilimumab and with each other. |
| ITT analysis | Primary analyses used all participants as randomized. |
| Time-to-event analysis | PFS and OS were classified as time-to-event endpoints. |
| Log-rank test | Used for the posted PFS comparisons. |
| Cox model | Used for the posted OS comparisons. |
| Hazard ratio | Used to express relative effects for PFS and OS. |
| Risk difference | Used for the ORR comparisons. |
| Confidence intervals | 95% two-sided intervals were reported for the posted primary and secondary effects. |
| Stratified analysis | Testing was stratified by line of therapy, PD-L1 status, and ECOG performance status. |
| Superiority testing | The analyses posted on ClinicalTrials.gov were designated as superiority hypotheses. |
The most instructive feature is the distinction between treatment-versus-comparator questions and schedule-versus-schedule questions. The pembrolizumab-versus-ipilimumab estimates are below 1 for both primary endpoints, while the direct Q2W-versus-Q3W comparisons have estimates close to 1 and confidence intervals that include the null value. These are separate statistical statements and should not be conflated.
21. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The posted primary analyses show HR estimates below 1 for PFS and 12-month OS when each pembrolizumab schedule is compared with ipilimumab. The direct schedule comparisons do not show statistical evidence of superiority under the reported tests.
Effect-size interpretation
The magnitude of an effect depends on the measure. HRs describe relative time-to-event hazards, while the ORR analyses report absolute percentage-point differences.
Uncertainty interpretation
The confidence intervals quantify uncertainty around each estimate. They are essential for understanding precision and should be read alongside the point estimate and P-value.
Safety interpretation
Serious adverse events are reported separately by arm. Safety counts answer a different question from the efficacy hazard ratios and should be interpreted independently.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Calculators
24. Sources
- ClinicalTrials.gov: NCT01866319 — KEYNOTE-006.
- PubMed: PMID 28822576.
- PubMed: PMID 41121745.
- PubMed: PMID 39306585.
- PubMed: PMID 37348035.
- PubMed: PMID 34571336.
Continue through the Clinical Biostats statistical pathway
Explore the survival-analysis, clinical-trial, confidence-interval, and binary-outcome methods that appear in KEYNOTE-006.
25. Record Summary
KEYNOTE-006 provides a compact example of several core clinical-trial statistical methods: randomized parallel-group comparison, ITT analysis, stratified time-to-event testing, log-rank analysis, Cox proportional-hazards modeling, hazard ratios, Kaplan-Meier estimation, risk differences for binary response, and two-sided confidence intervals. The registry reports three primary PFS comparisons and three primary OS comparisons, together with three secondary ORR comparisons.
The statistical story is most clearly understood by keeping the comparison questions separate. Against ipilimumab, the two pembrolizumab schedules produced reported PFS HR estimates of 0.58, while the reported OS HRs at Month 12 were 0.63 for Q2W and 0.69 for Q3W. The direct Q2W-versus-Q3W comparisons produced HRs of 0.97 for PFS and 0.91 for OS, with confidence intervals spanning the respective null value. The ORR analyses show the same distinction between treatment-versus-comparator and schedule-versus-schedule questions.