← Clinical Trials
Metastatic Breast Cancer Phase 3 Completed NCT00567190

CLEOPATRA: Complete Statistical Analysis of Pertuzumab in Metastatic Breast Cancer

An independent statistical analysis of the randomized phase 3 CLEOPATRA trial comparing pertuzumab plus trastuzumab and docetaxel with placebo plus trastuzumab and docetaxel in previously untreated HER2-positive metastatic breast cancer.

Randomized phase 3  ·  Enrollment 808  ·  Primary completion 2011-05-13
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

CLEOPATRA was a randomized, parallel, triple-masked phase 3 treatment trial evaluating pertuzumab added to trastuzumab and docetaxel against placebo added to trastuzumab and docetaxel in previously untreated HER2-positive metastatic breast cancer.

808
Enrolled
Randomized trial
2
Arms
Parallel design
0.62
Primary PFS HR
95% CI 0.51–0.75
<0.0001
Primary PFS P-value
Stratified log-rank
FeatureCLEOPATRA
Trial nameCLEOPATRA
PhasePhase 3
ConditionMetastatic Breast Cancer
PopulationPreviously untreated HER2-positive metastatic breast cancer
DesignRandomized, parallel, triple-masked
AllocationRandomized
Primary purposeTreatment
Enrollment808
Primary endpointProgression-Free Survival (PFS) Determined by an Independent Review Facility
Primary endpoint typeTime-to-event
Trial statusCOMPLETED
Start2008-02-12
Primary completion2011-05-13
Lead sponsorGenentech, Inc.
Sponsor typeINDUSTRY
ClinicalTrials.govNCT00567190

2. Clinical Question

The central statistical question was whether adding pertuzumab to trastuzumab and docetaxel changed the time from randomization to progression or death compared with placebo plus trastuzumab and docetaxel in previously untreated HER2-positive metastatic breast cancer.

Population

Previously untreated HER2-positive metastatic breast cancer, as described by the trial's brief title.

Intervention

Pertuzumab + Trastuzumab + Docetaxel.

Comparator

Placebo + Trastuzumab + Docetaxel.

Primary question

Does adding pertuzumab change the distribution of independent-review-facility PFS compared with the placebo regimen?

3. Trial Design

01
Randomize808 participants
02
Two armsPertuzumab vs placebo
03
Combination therapyTrastuzumab + docetaxel
04
AssessTumor assessments every 9 weeks
05
AnalyzeTime-to-event and response outcomes
Allocation
RANDOMIZED
Design model
PARALLEL
Masking
TRIPLE
Primary purpose
TREATMENT
PERTUZUMAB ARM

Pertuzumab combination

  • Pertuzumab
  • Trastuzumab
  • Docetaxel
CONTROL ARM

Placebo combination

  • Placebo
  • Trastuzumab
  • Docetaxel

The registry describes the study as randomized, parallel, and triple-masked. These design features are important statistically because randomization establishes the treatment groups before outcomes are observed, the parallel structure maintains separate randomized treatment assignments, and masking is intended to reduce the influence of treatment knowledge on trial conduct and assessment.

4. Endpoints

The registered primary endpoint was a time-to-event measure based on independent review of tumor assessments. The registry also posted secondary efficacy, symptom, cardiac-function, and other statistical analyses.

EndpointRegistry definition / time frameType
Progression-Free Survival (PFS) Determined by an Independent Review Facility Tumor assessments every 9 weeks from randomization to IRF-determined PD or death from any cause, whichever occurred first. Time-to-event
Overall Survival From randomization to death from any cause, up to each respective analysis data cut-off date. Time-to-event
Progression-Free Survival (PFS) Determined by the Investigator Tumor assessments every 9 weeks from randomization to investigator-determined PD or death from any cause, whichever occurred first. Time-to-event
Objective Response Determined by an Independent Review Facility Tumor assessments every 9 weeks from Baseline until IRF-determined progressive disease (PD), death, or first administration. Binary
Duration of Objective Response Determined by an Independent Review Facility From initial IRF-confirmed objective response until IRF-determined progressive disease (PD), death, or first administration. Time-to-event
Time to Symptom Progression Every 9 weeks from Baseline until investigator-determined progressive disease, up to the primary completion date. Time-to-event
Baseline LVEF Value and Change in LVEF From Baseline at Maximum Absolute Decrease Value During the Treatment Period Every 9 weeks from the date of randomization until Treatment Discontinuation Visit. Binary
Primary endpoint definition: PFS was defined as the time from randomization to first documented radiographical progressive disease, as determined by an independent review facility using RECIST version 1.0, or death from any cause within 18 weeks of last tumor assessment, whichever occurred first. For target lesions, PD was defined as at least a 20% increase in the sum of the longest diameter.

5. Analysis Populations and Stratification

The primary PFS analyses were conducted in the Intent-to-Treat (ITT) Population: All randomized participants. The investigator-determined PFS analysis likewise used the ITT population. Objective response analyses used all randomized participants with IRF-determined measurable disease at baseline, while duration of response was restricted further to participants with an objective response. The LVEF analysis used the safety population.

Analysis populationDefinition / role
ITT populationAll randomized participants; used for the primary PFS analyses and several secondary efficacy analyses.
ITT with measurable diseaseAll randomized participants with IRF-determined measurable disease at baseline, defined as at least 1 target lesion; used for objective response analysis.
ITT with objective responseRandomized participants with measurable disease at baseline who had an objective response; used for duration of objective response.
Safety populationAll participants who received at least one dose of any study medication; used for the LVEF analysis and reflected in the registry-reported serious-adverse-event counts.

The primary stratified PFS analysis was stratified by prior treatment status and region. The same stratification factors were used in the reported stratified overall-survival, investigator-PFS, and symptom-progression analyses.

6. Primary Endpoint Results

Independent-Review-Facility Progression-Free Survival — Stratified Analysis

Hazard ratio for progression or death

0.62

95% CI: 0.51–0.75   ·   P < 0.0001

Stratified log-rank test; ITT population.

The registry's primary analysis compared the PFS survival distributions between the pertuzumab and placebo arms using a stratified log-rank test. The hazard ratio was estimated with a Cox proportional-hazards approach and compared the pertuzumab arm with the placebo arm.

Clinical Biostats interpretation

A hazard ratio of 0.62 means that, under the fitted proportional-hazards framework, the estimated instantaneous rate of progression or death was approximately 38% lower in the pertuzumab arm than in the placebo arm. This is a relative time-to-event measure, not a statement that 38% of patients avoided progression or death.

The 95% CI of 0.51–0.75 describes statistical uncertainty around the estimated hazard ratio. It does not describe the range of individual patient outcomes and should not be read as a prediction interval for individual benefit.

The P-value of <0.0001 addresses evidence against the null hypothesis that the PFS survival distributions were the same. It does not measure the magnitude of the treatment effect; the hazard ratio and confidence interval provide that information.

Because the effect measure is a Cox proportional-hazards estimate, interpretation also depends on the proportional-hazards model being a reasonable summary of the treatment effect over the analyzed follow-up. The registry result itself does not provide a diagnostic assessment of that assumption.

Independent-Review-Facility Progression-Free Survival — Unstratified Analysis

Unstratified hazard ratio

0.63

95% CI: 0.52–0.76   ·   P < 0.0001

Unstratified log-rank test; ITT population.

The unstratified analysis produced an HR of 0.63, compared with 0.62 for the stratified analysis. The registry identifies the unstratified hazard ratio as comparing the pertuzumab arm with the placebo arm.

Clinical Biostats interpretation

The unstratified estimate is close to the stratified estimate, with a 95% CI of 0.52–0.76. The two analyses therefore provide closely aligned descriptions of the randomized PFS comparison, while the stratified analysis explicitly accounts for prior treatment status and region.

The P-value of <0.0001 again concerns evidence against equality of the survival distributions; it should not be interpreted as a measure of how clinically important the treatment effect is.

Primary PFS analysisMethodHR95% CIP-value
Stratified Stratified log-rank; Cox proportional hazard 0.62 0.51–0.75 <0.0001
Unstratified Unstratified log-rank; Cox proportional hazard 0.63 0.52–0.76 <0.0001

7. Secondary Efficacy Results

Overall Survival

The registry contains four distinct interim/final or exploratory overall-survival analyses in addition to an end-of-study exploratory analysis. All were based on the ITT population and used stratified log-rank testing with a Cox proportional-hazards effect measure. The analyses were stratified by prior treatment status and region.

OS analysisHR95% CIP-valueRegistry interpretation
First Interim OS Analysis0.640.47–0.880.0050Pre-defined O'Brien-Fleming stopping boundary for the Lan-DeMets α-spending function: HR≤0.603, p≤0.0012.
Second Interim OS Analysis0.660.52–0.840.0008Pre-defined O'Brien-Fleming stopping boundary for the Lan-DeMets α-spending function: HR≤0.739, p≤0.0138.
Event-Driven Final OS Analysis0.680.56–0.840.0002Planned after a total of 385 deaths; considered exploratory because the confirmatory OS analysis had previously occurred at the second interim OS analysis.
End-of-Study OS Analysis0.690.58–0.82<0.0001Considered exploratory because the confirmatory OS analysis had previously occurred at the second interim OS analysis.
Clinical Biostats interpretation

The reported OS hazard ratios range from 0.64 to 0.69 across the analyses posted on ClinicalTrials.gov, all below 1. The registry therefore reports a consistently lower estimated hazard of death in the pertuzumab arm across these analyses.

The key statistical distinction is between the confirmatory second interim OS analysis and the later analyses. The second interim analysis had a prespecified O'Brien-Fleming stopping boundary implemented through the Lan-DeMets alpha-spending function. The registry explicitly describes the later event-driven final and end-of-study OS analyses as exploratory because the confirmatory OS analysis had already occurred.

Thus, the later P-values should not be interpreted as though they represented a newly established confirmatory hypothesis test independent of the earlier interim monitoring. Alpha spending exists precisely because repeated looks at accumulating time-to-event data require control of the overall type I error.

Investigator-Determined Progression-Free Survival

Hazard ratio for investigator-determined PFS

0.69

95% CI: 0.59–0.81   ·   P < 0.0001

Stratified log-rank test; ITT population.

Clinical Biostats interpretation

The investigator-determined PFS analysis estimated a 0.69 hazard ratio for pertuzumab versus placebo. In relative terms, this corresponds to approximately a 31% lower estimated instantaneous hazard of progression or death under the Cox model.

The 95% CI of 0.59–0.81 provides the uncertainty interval around that estimate. The P-value of <0.0001 addresses the treatment-comparison hypothesis and is not itself a measure of effect magnitude.

Objective Response — Independent Review Facility

Difference in objective response rates

10.83

95% CI: 4.2–17.5   ·   P = 0.0011

Difference calculated as pertuzumab arm minus placebo arm; Cochran-Mantel-Haenszel analysis.

The registry reports objective response as a binary endpoint based on complete response plus partial response. The analysis included randomized participants with IRF-determined measurable disease at baseline, defined as at least one target lesion. The difference in objective response rates was calculated as the pertuzumab rate minus the placebo rate and the 95% CI used the Hauck-Anderson method.

Clinical Biostats interpretation

An estimated response-rate difference of 10.83 means that the reported objective response rate was higher in the pertuzumab arm by 10.83 percentage points under the registry's definition of the effect measure.

The 95% CI of 4.2–17.5 expresses uncertainty around that between-arm difference. Unlike a hazard ratio, this measure is directly expressed on the percentage-point scale.

The P-value of 0.0011 evaluates evidence against the corresponding null comparison; it does not tell us that the effect is "0.0011 large." Effect magnitude is described by the difference and its confidence interval.

Objective Response — Odds Ratio

Odds ratio for objective response

1.79

95% CI: 1.26–2.54

Method not reported in the registry analysis entry.

Clinical Biostats interpretation

An odds ratio of 1.79 means that the odds of objective response were estimated to be 1.79 times as high in the pertuzumab arm as in the placebo arm. Odds are not probabilities, so the odds ratio should not be read as a 79% increase in the response rate itself.

The 95% CI of 1.26–2.54 describes uncertainty around the odds-ratio estimate. Because the interval is entirely above 1, it is consistent with higher odds of response in the pertuzumab arm under the statistical framework used for this comparison.

Duration of Objective Response

Hazard ratio for duration of response

0.66

95% CI: 0.51–0.85

Cox proportional-hazards effect measure; formal analysis method not reported.

This endpoint begins at initial IRF-confirmed objective response and therefore applies only to participants who achieved an objective response. It is consequently a selected population rather than the full randomized population.

Clinical Biostats interpretation

A hazard ratio of 0.66 corresponds to approximately a 34% lower estimated instantaneous hazard of progression or death after an objective response, under the Cox model.

The important methodological distinction is that this is a conditional-on-response analysis. It does not answer the same question as the ITT PFS analysis, because patients who never achieved an objective response cannot enter the duration-of-response risk set.

Time to Symptom Progression

Hazard ratio for symptom progression

0.97

95% CI: 0.81–1.16   ·   P = 0.7161

Stratified log-rank test; ITT population; only female participants included.

Clinical Biostats interpretation

The estimated hazard ratio of 0.97 is close to 1, indicating little estimated relative separation between the randomized arms for this endpoint in the reported analysis.

The 95% CI of 0.81–1.16 spans 1. The P-value of 0.7161 does not provide evidence against the null comparison. Importantly, a nonsignificant P-value does not prove that the treatment effects are identical; it indicates that this analysis did not provide strong statistical evidence of a difference under its specified test.

8. Cardiac Function Analysis

The registry also reports a safety-population analysis of baseline LVEF and change in LVEF from baseline at the maximum absolute decrease during treatment. The endpoint was expressed in percentage points of LVEF.

Wilcoxon test of maximum decrease in LVEF

P = 0.7174

Wilcoxon Rank Sum Test; safety population.

The analysis compared the placebo plus trastuzumab plus docetaxel group with the pertuzumab plus trastuzumab plus docetaxel group among participants with evaluable LVEF assessments.

Clinical Biostats interpretation

The Wilcoxon rank-sum test is a nonparametric method for comparing the distributions of an outcome between two independent groups. Here it was used for the maximum decrease in LVEF from baseline.

The reported P-value of 0.7174 does not provide evidence of a difference between the groups under this analysis. The registry does not provide an effect estimate or confidence interval for this analysis, so the magnitude and precision of any between-group difference cannot be quantified from the registry-reported statistical result.

9. Safety Results

The ClinicalTrials.gov record includes serious adverse events by arm as affected participants divided by participants at risk. These counts are presented directly rather than converted into additional rates.

GroupSerious adverse events affected / at risk
Placebo + Trastuzumab + Docetaxel116/396
Pertuzumab + Trastuzumab + Docetaxel160/408
Crossover From Placebo to Pertuzumab10/50

The denominators in these registry-reported safety counts differ from the overall randomized enrollment of 808. That distinction matters: the serious-adverse-event counts should not be treated as though 396 and 408 were the original randomized group sizes without further qualification.

Safety interpretation: the ClinicalTrials.gov record identifies affected and at-risk participants for serious adverse events, but they do not provide a complete adverse-event table, grade distribution, treatment-relatedness classification, or a formal between-arm hypothesis test for these serious-adverse-event counts. The counts should therefore be described rather than converted into an unsupported comparative safety conclusion.

10. Statistical Methodology

Kaplan-Meier estimation

The primary PFS endpoint and the other reported time-to-event endpoints are naturally analyzed with survival-analysis methods because follow-up continues until an event occurs or observation is censored. The primary endpoint definition explicitly identifies the Kaplan-Meier method.

Conceptual Kaplan-Meier form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at time ti, while ni is the number at risk immediately before that time.

The advantage of Kaplan-Meier estimation is that participants who have not experienced progression or death by the end of their available observation can still contribute information up to their censoring time.

Stratified log-rank test

The principal PFS analysis used a stratified log-rank test, with stratification by prior treatment status and region. The same general approach was used for the reported stratified overall-survival, investigator-PFS, and time-to-symptom-progression analyses.

A log-rank test compares the observed pattern of events between treatment groups over follow-up. Stratification allows the comparison to account for prespecified strata rather than treating all participants as though they came from one homogeneous risk set.

Cox proportional-hazards model

The reported hazard ratios were based on a Cox proportional-hazards effect measure. The hazard ratio compares the estimated instantaneous event rates between the randomized treatment groups under the model.

Interpretation of a hazard ratio
HR < 1  →  lower estimated instantaneous event rate in the pertuzumab group

For example, an HR of 0.62 corresponds to an estimated 38% lower instantaneous event rate under the model. It does not mean that 38% of patients avoid the event or that each individual experiences the same reduction.

Intention-to-treat analysis

The primary PFS analyses used the ITT population consisting of all randomized participants. An ITT analysis preserves the treatment comparison established by randomization by analyzing participants according to their randomized group rather than redefining groups based on treatment exposure or subsequent events.

This is particularly important for interpreting a randomized treatment effect. Once randomization has occurred, excluding participants after randomization can compromise the balance created by the randomization process.

Cochran-Mantel-Haenszel analysis

The objective-response difference was analyzed using a Mantel-Haenszel method, normalized in the ClinicalTrials.gov record as a Cochran-Mantel-Haenszel test. The analysis was stratified by prior treatment status and region.

The method combines information across strata while accounting for the stratification variables. The reported response-rate difference was defined as the pertuzumab rate minus the placebo rate, and the confidence interval used the Hauck-Anderson method.

Wilcoxon rank-sum test

The maximum decrease in LVEF from baseline was analyzed using the Wilcoxon Rank Sum Test. This nonparametric method compares the relative distributions of a continuous or ordinal outcome between two independent groups without requiring the same normal-distribution assumptions as a conventional two-sample t-test.

11. Interim Analysis and Alpha Spending

Interim monitoring is explicitly supported by the registry's overall-survival analyses. The second interim OS analysis used a pre-defined O'Brien-Fleming stopping boundary for a Lan-DeMets α-spending function.

First interim OS analysis

The pre-defined boundary was HR≤0.603 and p≤0.0012. The reported estimate was HR 0.64 with P = 0.0050.

Second interim OS analysis

The pre-defined boundary was HR≤0.739 and p≤0.0138. The reported estimate was HR 0.66 with P = 0.0008.

The statistical reason for alpha spending is straightforward: if investigators repeatedly inspect accumulating outcome data and apply an ordinary final-analysis threshold at every look, the chance of a false-positive conclusion can exceed the intended type I error rate. A Lan-DeMets spending function provides a flexible way to allocate the available type I error across information times.

O'Brien-Fleming principle
Early evidence threshold  →  more stringent  ·  Later evidence threshold  →  less stringent

The registry-reported second-interim boundary illustrates this principle: the analysis had a prespecified HR and P-value threshold rather than simply applying an unadjusted significance rule at the interim look.

The registry identifies the event-driven final OS analysis as occurring after a total of 385 deaths. It also states that this final analysis was exploratory because the confirmatory OS analysis had already occurred at the second interim analysis.

Multiplicity lesson: the confirmatory interpretation of an interim analysis depends on the prespecified monitoring plan. A later P-value cannot automatically be treated as though it came from a single, previously unseen final analysis when earlier looks at the accumulating data have already occurred.

12. Stratification and Why It Matters

Prior treatment status and region were used as stratification factors in the primary PFS analysis and several secondary time-to-event analyses. Stratification is especially relevant when the randomization scheme or analysis plan anticipates that baseline strata may be associated with the event process.

AnalysisStratificationPurpose in the reported analysis
Primary PFSPrior treatment status and regionStratified log-rank comparison and Cox hazard-ratio analysis.
Overall survivalPrior treatment status and regionStratified log-rank comparison and Cox hazard-ratio analysis.
Investigator PFSPrior treatment status and regionStratified log-rank comparison and Cox hazard-ratio analysis.
Objective responsePrior treatment status and regionCochran-Mantel-Haenszel analysis and stratified response-rate difference.

A stratified estimate is not simply an average of the separate stratum-specific hazard ratios. It is produced within the statistical framework of the stratified analysis. The key point is that the comparison respects the prespecified strata rather than discarding them from the analysis.

13. Statistical Methods Explained

Why was a log-rank test used for PFS?

PFS is a time-to-event endpoint. Some participants will progress or die during follow-up, while others may remain event-free when their observation ends. The log-rank test is designed to compare the survival distributions of two groups while incorporating the timing of events and censoring.

Why was the primary PFS analysis stratified?

The registry specifies prior treatment status and region as stratification factors. A stratified log-rank test compares treatment groups while accounting for these predefined strata. This is different from simply ignoring the strata and performing one pooled unstratified comparison.

What does a hazard ratio of 0.62 mean?

Under the Cox proportional-hazards framework, 0.62 means that the estimated instantaneous hazard of progression or death was 0.62 times that of the placebo group. Equivalently, it represents an approximately 38% lower estimated instantaneous hazard. It does not mean that 38% of patients benefited or that each patient had exactly a 38% reduction in risk.

Why is the confidence interval important?

The point estimate alone does not communicate statistical precision. The 95% CI of 0.51–0.75 around the primary PFS HR shows the uncertainty associated with estimating the treatment effect. A narrower interval would indicate greater precision; a wider interval would indicate less precision.

Why doesn't the P-value measure effect size?

A P-value measures how compatible the observed data are with a specified null hypothesis under the statistical model. It does not directly quantify the size of the treatment effect. In CLEOPATRA, the HR and its confidence interval describe the relative time-to-event effect, while the P-value addresses evidence against the null comparison.

Why does the interim OS analysis need an alpha-spending boundary?

Because the accumulating OS data were examined more than once, the statistical design needed to account for repeated opportunities to declare efficacy. The ClinicalTrials.gov record identifies an O'Brien-Fleming stopping boundary implemented through a Lan-DeMets alpha-spending function. The boundary protects the overall inferential framework against inflation caused by repeated looks.

Why is the duration-of-response analysis different from ITT PFS?

Duration of response begins only after a participant has achieved an objective response. The analysis therefore conditions on having responded. By contrast, ITT PFS starts at randomization and includes all randomized participants. These endpoints answer different questions and should not be treated as interchangeable measures of treatment effect.

14. Interpreting the Main Statistical Results Together

EndpointEffect measureEstimate95% CIP-value
Primary IRF PFS, stratifiedHazard ratio0.620.51–0.75<0.0001
Primary IRF PFS, unstratifiedHazard ratio0.630.52–0.76<0.0001
Overall survival, second interimHazard ratio0.660.52–0.840.0008
Investigator PFSHazard ratio0.690.59–0.81<0.0001
Objective responseDifference in response rates10.834.2–17.50.0011
Objective responseOdds ratio1.791.26–2.54Not reported
Duration of objective responseHazard ratio0.660.51–0.85Not reported
Time to symptom progressionHazard ratio0.970.81–1.160.7161

Several distinct statistical questions are represented in this table. The primary PFS analysis asks whether the time-to-progression-or-death distribution differs between randomized groups. Objective response asks a binary question about tumor response. Duration of response asks how long responses persist among responders. Time to symptom progression asks a separate patient-relevant time-to-event question.

The estimates should therefore not be collapsed into a single overall treatment statistic. A hazard ratio for PFS, an odds ratio for response, and a hazard ratio for duration of response have different estimands and different analysis populations.

15. What the Primary Hazard Ratio Does — and Does Not — Mean

The estimate

The primary stratified PFS hazard ratio of 0.62 indicates that the estimated instantaneous hazard of progression or death was approximately 38% lower in the pertuzumab arm than in the placebo arm under the reported Cox model.

What it does not mean

The HR does not mean that 38% of patients were protected from progression, that every patient experienced the same relative benefit, or that the median PFS was reduced or increased by 38%. It is a model-based relative measure of the event hazard.

The confidence interval

The 95% CI of 0.51–0.75 describes uncertainty around the estimated hazard ratio. It does not describe the range of treatment effects across individual patients.

The P-value

The P-value of <0.0001 addresses the null hypothesis that the PFS survival distributions in the two treatment groups were the same. It is not a measure of clinical magnitude.

The model assumption

Because the effect measure is a Cox proportional-hazards estimate, the interpretation of a single HR is most direct when the proportional-hazards framework is a reasonable representation of the event processes over time. The ClinicalTrials.gov record does not provide a separate proportional-hazards diagnostic.

16. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The randomized ITT comparison of independent-review-facility PFS produced a stratified HR of 0.62 with a 95% CI of 0.51–0.75 and P < 0.0001. The unstratified analysis produced a closely aligned HR of 0.63.

Broader evidence

The registry also reports HRs below 1 for overall survival and investigator-determined PFS, a positive objective-response difference, an objective-response odds ratio above 1, and a duration-of-response HR below 1.

These results address different endpoints and analysis populations. A statistical review should therefore preserve those distinctions rather than treating all estimates as interchangeable evidence of one common effect.

17. Interim Analysis, Confirmatory Testing, and Later Results

The overall-survival results illustrate why the timing and role of an analysis are part of the statistical result itself.

First Interim OS Analysis

Early efficacy assessment

HR 0.64 with 95% CI 0.47–0.88 and P = 0.0050. The pre-defined O'Brien-Fleming boundary was HR≤0.603 and p≤0.0012.

Second Interim OS Analysis

Confirmatory OS analysis

HR 0.66 with 95% CI 0.52–0.84 and P = 0.0008. The pre-defined boundary was HR≤0.739 and p≤0.0138.

Event-Driven Final OS Analysis

385-death analysis

The registry states that the analysis was planned after a total of 385 deaths and considered exploratory because the confirmatory OS analysis had previously occurred at the second interim analysis.

End-of-Study OS Analysis

Later exploratory analysis

HR 0.69 with 95% CI 0.58–0.82 and P < 0.0001. The registry again characterizes this analysis as exploratory.

The important lesson is that a later analysis can be statistically informative without having the same confirmatory status as the prespecified interim analysis. The label "exploratory" describes the role of the analysis in the trial's inferential sequence; it does not mean that the numerical estimate is unusable.

18. Limitations

19. Why This Trial Matters Statistically

CLEOPATRA is a useful statistical teaching case because its registry results combine randomized treatment allocation, a primary time-to-event endpoint, stratified survival analysis, independent review, multiple overall-survival looks, binary response analysis, duration-of-response analysis, symptom progression, and a nonparametric cardiac-function comparison.

ConceptHow it appears in CLEOPATRA
Randomization808 participants were enrolled in a randomized, parallel phase 3 trial.
Triple maskingThe registry identifies the trial as triple-masked.
ITT analysisThe primary PFS analyses used all randomized participants.
Kaplan-Meier estimationThe registered PFS definition identifies the Kaplan-Meier method.
Hazard ratioPrimary PFS and several secondary time-to-event analyses use hazard ratios.
Confidence intervalPrimary and several secondary hazard ratios have two-sided 95% CIs.
Log-rank testingStratified and unstratified log-rank analyses were reported for PFS.
Stratified analysisPrior treatment status and region were used for the reported stratified analyses.
Cochran-Mantel-Haenszel testingUsed for the objective-response comparison.
Odds ratioAn OR of 1.79 was reported for objective response.
Interim analysisMultiple OS analyses were reported, including two interim analyses.
Alpha spendingLan-DeMets alpha spending with O'Brien-Fleming stopping boundaries was reported for OS.
Nonparametric analysisWilcoxon Rank Sum Test was used for maximum LVEF decrease.

20. Related Tutorials

Learn more about the methods used in this trial:

21. Related Statistical Calculators

22. Sources

Continue through Clinical Biostats

Build deeper understanding of the statistical methods that appear across randomized clinical trials, from survival analysis and hazard ratios to response analysis and interim monitoring.

23. Record Summary

CLEOPATRA provides a compact example of several central principles in clinical-trial statistics. The primary endpoint was a time-to-event outcome assessed by an independent review facility, analyzed in the ITT population using a stratified log-rank test and a Cox proportional-hazards effect measure. The reported primary PFS estimate was HR 0.62 (95% CI 0.51–0.75; P < 0.0001), with a closely aligned unstratified analysis of HR 0.63 (95% CI 0.52–0.76; P < 0.0001).

The secondary results illustrate why clinical-trial statistics cannot be reduced to a single P-value. Overall survival was examined at multiple analysis times under an interim-monitoring framework using Lan-DeMets alpha spending and O'Brien-Fleming stopping boundaries. Objective response was analyzed with both a response-rate difference and an odds ratio. Duration of response used a hazard ratio among responders, while time to symptom progression produced a different hazard-ratio estimate and P-value. LVEF was analyzed separately using a Wilcoxon rank-sum approach in the safety population.

The most informative statistical reading therefore combines the effect estimate, confidence interval, hypothesis test, analysis population, endpoint definition, and timing of the analysis. These elements together explain what the CLEOPATRA registry results establish statistically and what they do not establish.

Clinical Biostats methodology: A trial-results page should not merely repeat a study abstract. The goal is to reconstruct the statistical story of the trial while clearly separating reported numerical evidence from statistical interpretation and preserving the distinctions among endpoints, analysis populations, and inferential frameworks.