This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
CLEOPATRA was a randomized, parallel, triple-masked phase 3 treatment trial evaluating pertuzumab added to trastuzumab and docetaxel against placebo added to trastuzumab and docetaxel in previously untreated HER2-positive metastatic breast cancer.
| Feature | CLEOPATRA |
|---|---|
| Trial name | CLEOPATRA |
| Phase | Phase 3 |
| Condition | Metastatic Breast Cancer |
| Population | Previously untreated HER2-positive metastatic breast cancer |
| Design | Randomized, parallel, triple-masked |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 808 |
| Primary endpoint | Progression-Free Survival (PFS) Determined by an Independent Review Facility |
| Primary endpoint type | Time-to-event |
| Trial status | COMPLETED |
| Start | 2008-02-12 |
| Primary completion | 2011-05-13 |
| Lead sponsor | Genentech, Inc. |
| Sponsor type | INDUSTRY |
| ClinicalTrials.gov | NCT00567190 |
2. Clinical Question
The central statistical question was whether adding pertuzumab to trastuzumab and docetaxel changed the time from randomization to progression or death compared with placebo plus trastuzumab and docetaxel in previously untreated HER2-positive metastatic breast cancer.
Population
Previously untreated HER2-positive metastatic breast cancer, as described by the trial's brief title.
Intervention
Pertuzumab + Trastuzumab + Docetaxel.
Comparator
Placebo + Trastuzumab + Docetaxel.
Primary question
Does adding pertuzumab change the distribution of independent-review-facility PFS compared with the placebo regimen?
3. Trial Design
Pertuzumab combination
- Pertuzumab
- Trastuzumab
- Docetaxel
Placebo combination
- Placebo
- Trastuzumab
- Docetaxel
The registry describes the study as randomized, parallel, and triple-masked. These design features are important statistically because randomization establishes the treatment groups before outcomes are observed, the parallel structure maintains separate randomized treatment assignments, and masking is intended to reduce the influence of treatment knowledge on trial conduct and assessment.
4. Endpoints
The registered primary endpoint was a time-to-event measure based on independent review of tumor assessments. The registry also posted secondary efficacy, symptom, cardiac-function, and other statistical analyses.
| Endpoint | Registry definition / time frame | Type |
|---|---|---|
| Progression-Free Survival (PFS) Determined by an Independent Review Facility | Tumor assessments every 9 weeks from randomization to IRF-determined PD or death from any cause, whichever occurred first. | Time-to-event |
| Overall Survival | From randomization to death from any cause, up to each respective analysis data cut-off date. | Time-to-event |
| Progression-Free Survival (PFS) Determined by the Investigator | Tumor assessments every 9 weeks from randomization to investigator-determined PD or death from any cause, whichever occurred first. | Time-to-event |
| Objective Response Determined by an Independent Review Facility | Tumor assessments every 9 weeks from Baseline until IRF-determined progressive disease (PD), death, or first administration. | Binary |
| Duration of Objective Response Determined by an Independent Review Facility | From initial IRF-confirmed objective response until IRF-determined progressive disease (PD), death, or first administration. | Time-to-event |
| Time to Symptom Progression | Every 9 weeks from Baseline until investigator-determined progressive disease, up to the primary completion date. | Time-to-event |
| Baseline LVEF Value and Change in LVEF From Baseline at Maximum Absolute Decrease Value During the Treatment Period | Every 9 weeks from the date of randomization until Treatment Discontinuation Visit. | Binary |
5. Analysis Populations and Stratification
The primary PFS analyses were conducted in the Intent-to-Treat (ITT) Population: All randomized participants. The investigator-determined PFS analysis likewise used the ITT population. Objective response analyses used all randomized participants with IRF-determined measurable disease at baseline, while duration of response was restricted further to participants with an objective response. The LVEF analysis used the safety population.
| Analysis population | Definition / role |
|---|---|
| ITT population | All randomized participants; used for the primary PFS analyses and several secondary efficacy analyses. |
| ITT with measurable disease | All randomized participants with IRF-determined measurable disease at baseline, defined as at least 1 target lesion; used for objective response analysis. |
| ITT with objective response | Randomized participants with measurable disease at baseline who had an objective response; used for duration of objective response. |
| Safety population | All participants who received at least one dose of any study medication; used for the LVEF analysis and reflected in the registry-reported serious-adverse-event counts. |
The primary stratified PFS analysis was stratified by prior treatment status and region. The same stratification factors were used in the reported stratified overall-survival, investigator-PFS, and symptom-progression analyses.
6. Primary Endpoint Results
Independent-Review-Facility Progression-Free Survival — Stratified Analysis
Hazard ratio for progression or death
95% CI: 0.51–0.75 · P < 0.0001
Stratified log-rank test; ITT population.
The registry's primary analysis compared the PFS survival distributions between the pertuzumab and placebo arms using a stratified log-rank test. The hazard ratio was estimated with a Cox proportional-hazards approach and compared the pertuzumab arm with the placebo arm.
A hazard ratio of 0.62 means that, under the fitted proportional-hazards framework, the estimated instantaneous rate of progression or death was approximately 38% lower in the pertuzumab arm than in the placebo arm. This is a relative time-to-event measure, not a statement that 38% of patients avoided progression or death.
The 95% CI of 0.51–0.75 describes statistical uncertainty around the estimated hazard ratio. It does not describe the range of individual patient outcomes and should not be read as a prediction interval for individual benefit.
The P-value of <0.0001 addresses evidence against the null hypothesis that the PFS survival distributions were the same. It does not measure the magnitude of the treatment effect; the hazard ratio and confidence interval provide that information.
Because the effect measure is a Cox proportional-hazards estimate, interpretation also depends on the proportional-hazards model being a reasonable summary of the treatment effect over the analyzed follow-up. The registry result itself does not provide a diagnostic assessment of that assumption.
Independent-Review-Facility Progression-Free Survival — Unstratified Analysis
Unstratified hazard ratio
95% CI: 0.52–0.76 · P < 0.0001
Unstratified log-rank test; ITT population.
The unstratified analysis produced an HR of 0.63, compared with 0.62 for the stratified analysis. The registry identifies the unstratified hazard ratio as comparing the pertuzumab arm with the placebo arm.
The unstratified estimate is close to the stratified estimate, with a 95% CI of 0.52–0.76. The two analyses therefore provide closely aligned descriptions of the randomized PFS comparison, while the stratified analysis explicitly accounts for prior treatment status and region.
The P-value of <0.0001 again concerns evidence against equality of the survival distributions; it should not be interpreted as a measure of how clinically important the treatment effect is.
| Primary PFS analysis | Method | HR | 95% CI | P-value |
|---|---|---|---|---|
| Stratified | Stratified log-rank; Cox proportional hazard | 0.62 | 0.51–0.75 | <0.0001 |
| Unstratified | Unstratified log-rank; Cox proportional hazard | 0.63 | 0.52–0.76 | <0.0001 |
7. Secondary Efficacy Results
Overall Survival
The registry contains four distinct interim/final or exploratory overall-survival analyses in addition to an end-of-study exploratory analysis. All were based on the ITT population and used stratified log-rank testing with a Cox proportional-hazards effect measure. The analyses were stratified by prior treatment status and region.
| OS analysis | HR | 95% CI | P-value | Registry interpretation |
|---|---|---|---|---|
| First Interim OS Analysis | 0.64 | 0.47–0.88 | 0.0050 | Pre-defined O'Brien-Fleming stopping boundary for the Lan-DeMets α-spending function: HR≤0.603, p≤0.0012. |
| Second Interim OS Analysis | 0.66 | 0.52–0.84 | 0.0008 | Pre-defined O'Brien-Fleming stopping boundary for the Lan-DeMets α-spending function: HR≤0.739, p≤0.0138. |
| Event-Driven Final OS Analysis | 0.68 | 0.56–0.84 | 0.0002 | Planned after a total of 385 deaths; considered exploratory because the confirmatory OS analysis had previously occurred at the second interim OS analysis. |
| End-of-Study OS Analysis | 0.69 | 0.58–0.82 | <0.0001 | Considered exploratory because the confirmatory OS analysis had previously occurred at the second interim OS analysis. |
The reported OS hazard ratios range from 0.64 to 0.69 across the analyses posted on ClinicalTrials.gov, all below 1. The registry therefore reports a consistently lower estimated hazard of death in the pertuzumab arm across these analyses.
The key statistical distinction is between the confirmatory second interim OS analysis and the later analyses. The second interim analysis had a prespecified O'Brien-Fleming stopping boundary implemented through the Lan-DeMets alpha-spending function. The registry explicitly describes the later event-driven final and end-of-study OS analyses as exploratory because the confirmatory OS analysis had already occurred.
Thus, the later P-values should not be interpreted as though they represented a newly established confirmatory hypothesis test independent of the earlier interim monitoring. Alpha spending exists precisely because repeated looks at accumulating time-to-event data require control of the overall type I error.
Investigator-Determined Progression-Free Survival
Hazard ratio for investigator-determined PFS
95% CI: 0.59–0.81 · P < 0.0001
Stratified log-rank test; ITT population.
The investigator-determined PFS analysis estimated a 0.69 hazard ratio for pertuzumab versus placebo. In relative terms, this corresponds to approximately a 31% lower estimated instantaneous hazard of progression or death under the Cox model.
The 95% CI of 0.59–0.81 provides the uncertainty interval around that estimate. The P-value of <0.0001 addresses the treatment-comparison hypothesis and is not itself a measure of effect magnitude.
Objective Response — Independent Review Facility
Difference in objective response rates
95% CI: 4.2–17.5 · P = 0.0011
Difference calculated as pertuzumab arm minus placebo arm; Cochran-Mantel-Haenszel analysis.
The registry reports objective response as a binary endpoint based on complete response plus partial response. The analysis included randomized participants with IRF-determined measurable disease at baseline, defined as at least one target lesion. The difference in objective response rates was calculated as the pertuzumab rate minus the placebo rate and the 95% CI used the Hauck-Anderson method.
An estimated response-rate difference of 10.83 means that the reported objective response rate was higher in the pertuzumab arm by 10.83 percentage points under the registry's definition of the effect measure.
The 95% CI of 4.2–17.5 expresses uncertainty around that between-arm difference. Unlike a hazard ratio, this measure is directly expressed on the percentage-point scale.
The P-value of 0.0011 evaluates evidence against the corresponding null comparison; it does not tell us that the effect is "0.0011 large." Effect magnitude is described by the difference and its confidence interval.
Objective Response — Odds Ratio
Odds ratio for objective response
95% CI: 1.26–2.54
Method not reported in the registry analysis entry.
An odds ratio of 1.79 means that the odds of objective response were estimated to be 1.79 times as high in the pertuzumab arm as in the placebo arm. Odds are not probabilities, so the odds ratio should not be read as a 79% increase in the response rate itself.
The 95% CI of 1.26–2.54 describes uncertainty around the odds-ratio estimate. Because the interval is entirely above 1, it is consistent with higher odds of response in the pertuzumab arm under the statistical framework used for this comparison.
Duration of Objective Response
Hazard ratio for duration of response
95% CI: 0.51–0.85
Cox proportional-hazards effect measure; formal analysis method not reported.
This endpoint begins at initial IRF-confirmed objective response and therefore applies only to participants who achieved an objective response. It is consequently a selected population rather than the full randomized population.
A hazard ratio of 0.66 corresponds to approximately a 34% lower estimated instantaneous hazard of progression or death after an objective response, under the Cox model.
The important methodological distinction is that this is a conditional-on-response analysis. It does not answer the same question as the ITT PFS analysis, because patients who never achieved an objective response cannot enter the duration-of-response risk set.
Time to Symptom Progression
Hazard ratio for symptom progression
95% CI: 0.81–1.16 · P = 0.7161
Stratified log-rank test; ITT population; only female participants included.
The estimated hazard ratio of 0.97 is close to 1, indicating little estimated relative separation between the randomized arms for this endpoint in the reported analysis.
The 95% CI of 0.81–1.16 spans 1. The P-value of 0.7161 does not provide evidence against the null comparison. Importantly, a nonsignificant P-value does not prove that the treatment effects are identical; it indicates that this analysis did not provide strong statistical evidence of a difference under its specified test.
8. Cardiac Function Analysis
The registry also reports a safety-population analysis of baseline LVEF and change in LVEF from baseline at the maximum absolute decrease during treatment. The endpoint was expressed in percentage points of LVEF.
Wilcoxon test of maximum decrease in LVEF
Wilcoxon Rank Sum Test; safety population.
The analysis compared the placebo plus trastuzumab plus docetaxel group with the pertuzumab plus trastuzumab plus docetaxel group among participants with evaluable LVEF assessments.
The Wilcoxon rank-sum test is a nonparametric method for comparing the distributions of an outcome between two independent groups. Here it was used for the maximum decrease in LVEF from baseline.
The reported P-value of 0.7174 does not provide evidence of a difference between the groups under this analysis. The registry does not provide an effect estimate or confidence interval for this analysis, so the magnitude and precision of any between-group difference cannot be quantified from the registry-reported statistical result.
9. Safety Results
The ClinicalTrials.gov record includes serious adverse events by arm as affected participants divided by participants at risk. These counts are presented directly rather than converted into additional rates.
| Group | Serious adverse events affected / at risk |
|---|---|
| Placebo + Trastuzumab + Docetaxel | 116/396 |
| Pertuzumab + Trastuzumab + Docetaxel | 160/408 |
| Crossover From Placebo to Pertuzumab | 10/50 |
The denominators in these registry-reported safety counts differ from the overall randomized enrollment of 808. That distinction matters: the serious-adverse-event counts should not be treated as though 396 and 408 were the original randomized group sizes without further qualification.
10. Statistical Methodology
Kaplan-Meier estimation
The primary PFS endpoint and the other reported time-to-event endpoints are naturally analyzed with survival-analysis methods because follow-up continues until an event occurs or observation is censored. The primary endpoint definition explicitly identifies the Kaplan-Meier method.
Here, di represents the number of events at time ti, while ni is the number at risk immediately before that time.
The advantage of Kaplan-Meier estimation is that participants who have not experienced progression or death by the end of their available observation can still contribute information up to their censoring time.
Stratified log-rank test
The principal PFS analysis used a stratified log-rank test, with stratification by prior treatment status and region. The same general approach was used for the reported stratified overall-survival, investigator-PFS, and time-to-symptom-progression analyses.
A log-rank test compares the observed pattern of events between treatment groups over follow-up. Stratification allows the comparison to account for prespecified strata rather than treating all participants as though they came from one homogeneous risk set.
Cox proportional-hazards model
The reported hazard ratios were based on a Cox proportional-hazards effect measure. The hazard ratio compares the estimated instantaneous event rates between the randomized treatment groups under the model.
For example, an HR of 0.62 corresponds to an estimated 38% lower instantaneous event rate under the model. It does not mean that 38% of patients avoid the event or that each individual experiences the same reduction.
Intention-to-treat analysis
The primary PFS analyses used the ITT population consisting of all randomized participants. An ITT analysis preserves the treatment comparison established by randomization by analyzing participants according to their randomized group rather than redefining groups based on treatment exposure or subsequent events.
This is particularly important for interpreting a randomized treatment effect. Once randomization has occurred, excluding participants after randomization can compromise the balance created by the randomization process.
Cochran-Mantel-Haenszel analysis
The objective-response difference was analyzed using a Mantel-Haenszel method, normalized in the ClinicalTrials.gov record as a Cochran-Mantel-Haenszel test. The analysis was stratified by prior treatment status and region.
The method combines information across strata while accounting for the stratification variables. The reported response-rate difference was defined as the pertuzumab rate minus the placebo rate, and the confidence interval used the Hauck-Anderson method.
Wilcoxon rank-sum test
The maximum decrease in LVEF from baseline was analyzed using the Wilcoxon Rank Sum Test. This nonparametric method compares the relative distributions of a continuous or ordinal outcome between two independent groups without requiring the same normal-distribution assumptions as a conventional two-sample t-test.
11. Interim Analysis and Alpha Spending
Interim monitoring is explicitly supported by the registry's overall-survival analyses. The second interim OS analysis used a pre-defined O'Brien-Fleming stopping boundary for a Lan-DeMets α-spending function.
First interim OS analysis
The pre-defined boundary was HR≤0.603 and p≤0.0012. The reported estimate was HR 0.64 with P = 0.0050.
Second interim OS analysis
The pre-defined boundary was HR≤0.739 and p≤0.0138. The reported estimate was HR 0.66 with P = 0.0008.
The statistical reason for alpha spending is straightforward: if investigators repeatedly inspect accumulating outcome data and apply an ordinary final-analysis threshold at every look, the chance of a false-positive conclusion can exceed the intended type I error rate. A Lan-DeMets spending function provides a flexible way to allocate the available type I error across information times.
The registry-reported second-interim boundary illustrates this principle: the analysis had a prespecified HR and P-value threshold rather than simply applying an unadjusted significance rule at the interim look.
The registry identifies the event-driven final OS analysis as occurring after a total of 385 deaths. It also states that this final analysis was exploratory because the confirmatory OS analysis had already occurred at the second interim analysis.
12. Stratification and Why It Matters
Prior treatment status and region were used as stratification factors in the primary PFS analysis and several secondary time-to-event analyses. Stratification is especially relevant when the randomization scheme or analysis plan anticipates that baseline strata may be associated with the event process.
| Analysis | Stratification | Purpose in the reported analysis |
|---|---|---|
| Primary PFS | Prior treatment status and region | Stratified log-rank comparison and Cox hazard-ratio analysis. |
| Overall survival | Prior treatment status and region | Stratified log-rank comparison and Cox hazard-ratio analysis. |
| Investigator PFS | Prior treatment status and region | Stratified log-rank comparison and Cox hazard-ratio analysis. |
| Objective response | Prior treatment status and region | Cochran-Mantel-Haenszel analysis and stratified response-rate difference. |
A stratified estimate is not simply an average of the separate stratum-specific hazard ratios. It is produced within the statistical framework of the stratified analysis. The key point is that the comparison respects the prespecified strata rather than discarding them from the analysis.
13. Statistical Methods Explained
Why was a log-rank test used for PFS?
PFS is a time-to-event endpoint. Some participants will progress or die during follow-up, while others may remain event-free when their observation ends. The log-rank test is designed to compare the survival distributions of two groups while incorporating the timing of events and censoring.
Why was the primary PFS analysis stratified?
The registry specifies prior treatment status and region as stratification factors. A stratified log-rank test compares treatment groups while accounting for these predefined strata. This is different from simply ignoring the strata and performing one pooled unstratified comparison.
What does a hazard ratio of 0.62 mean?
Under the Cox proportional-hazards framework, 0.62 means that the estimated instantaneous hazard of progression or death was 0.62 times that of the placebo group. Equivalently, it represents an approximately 38% lower estimated instantaneous hazard. It does not mean that 38% of patients benefited or that each patient had exactly a 38% reduction in risk.
Why is the confidence interval important?
The point estimate alone does not communicate statistical precision. The 95% CI of 0.51–0.75 around the primary PFS HR shows the uncertainty associated with estimating the treatment effect. A narrower interval would indicate greater precision; a wider interval would indicate less precision.
Why doesn't the P-value measure effect size?
A P-value measures how compatible the observed data are with a specified null hypothesis under the statistical model. It does not directly quantify the size of the treatment effect. In CLEOPATRA, the HR and its confidence interval describe the relative time-to-event effect, while the P-value addresses evidence against the null comparison.
Why does the interim OS analysis need an alpha-spending boundary?
Because the accumulating OS data were examined more than once, the statistical design needed to account for repeated opportunities to declare efficacy. The ClinicalTrials.gov record identifies an O'Brien-Fleming stopping boundary implemented through a Lan-DeMets alpha-spending function. The boundary protects the overall inferential framework against inflation caused by repeated looks.
Why is the duration-of-response analysis different from ITT PFS?
Duration of response begins only after a participant has achieved an objective response. The analysis therefore conditions on having responded. By contrast, ITT PFS starts at randomization and includes all randomized participants. These endpoints answer different questions and should not be treated as interchangeable measures of treatment effect.
14. Interpreting the Main Statistical Results Together
| Endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Primary IRF PFS, stratified | Hazard ratio | 0.62 | 0.51–0.75 | <0.0001 |
| Primary IRF PFS, unstratified | Hazard ratio | 0.63 | 0.52–0.76 | <0.0001 |
| Overall survival, second interim | Hazard ratio | 0.66 | 0.52–0.84 | 0.0008 |
| Investigator PFS | Hazard ratio | 0.69 | 0.59–0.81 | <0.0001 |
| Objective response | Difference in response rates | 10.83 | 4.2–17.5 | 0.0011 |
| Objective response | Odds ratio | 1.79 | 1.26–2.54 | Not reported |
| Duration of objective response | Hazard ratio | 0.66 | 0.51–0.85 | Not reported |
| Time to symptom progression | Hazard ratio | 0.97 | 0.81–1.16 | 0.7161 |
Several distinct statistical questions are represented in this table. The primary PFS analysis asks whether the time-to-progression-or-death distribution differs between randomized groups. Objective response asks a binary question about tumor response. Duration of response asks how long responses persist among responders. Time to symptom progression asks a separate patient-relevant time-to-event question.
The estimates should therefore not be collapsed into a single overall treatment statistic. A hazard ratio for PFS, an odds ratio for response, and a hazard ratio for duration of response have different estimands and different analysis populations.
15. What the Primary Hazard Ratio Does — and Does Not — Mean
The primary stratified PFS hazard ratio of 0.62 indicates that the estimated instantaneous hazard of progression or death was approximately 38% lower in the pertuzumab arm than in the placebo arm under the reported Cox model.
The HR does not mean that 38% of patients were protected from progression, that every patient experienced the same relative benefit, or that the median PFS was reduced or increased by 38%. It is a model-based relative measure of the event hazard.
The 95% CI of 0.51–0.75 describes uncertainty around the estimated hazard ratio. It does not describe the range of treatment effects across individual patients.
The P-value of <0.0001 addresses the null hypothesis that the PFS survival distributions in the two treatment groups were the same. It is not a measure of clinical magnitude.
Because the effect measure is a Cox proportional-hazards estimate, the interpretation of a single HR is most direct when the proportional-hazards framework is a reasonable representation of the event processes over time. The ClinicalTrials.gov record does not provide a separate proportional-hazards diagnostic.
16. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The randomized ITT comparison of independent-review-facility PFS produced a stratified HR of 0.62 with a 95% CI of 0.51–0.75 and P < 0.0001. The unstratified analysis produced a closely aligned HR of 0.63.
Broader evidence
The registry also reports HRs below 1 for overall survival and investigator-determined PFS, a positive objective-response difference, an objective-response odds ratio above 1, and a duration-of-response HR below 1.
These results address different endpoints and analysis populations. A statistical review should therefore preserve those distinctions rather than treating all estimates as interchangeable evidence of one common effect.
17. Interim Analysis, Confirmatory Testing, and Later Results
The overall-survival results illustrate why the timing and role of an analysis are part of the statistical result itself.
Early efficacy assessment
HR 0.64 with 95% CI 0.47–0.88 and P = 0.0050. The pre-defined O'Brien-Fleming boundary was HR≤0.603 and p≤0.0012.
Confirmatory OS analysis
HR 0.66 with 95% CI 0.52–0.84 and P = 0.0008. The pre-defined boundary was HR≤0.739 and p≤0.0138.
385-death analysis
The registry states that the analysis was planned after a total of 385 deaths and considered exploratory because the confirmatory OS analysis had previously occurred at the second interim analysis.
Later exploratory analysis
HR 0.69 with 95% CI 0.58–0.82 and P < 0.0001. The registry again characterizes this analysis as exploratory.
The important lesson is that a later analysis can be statistically informative without having the same confirmatory status as the prespecified interim analysis. The label "exploratory" describes the role of the analysis in the trial's inferential sequence; it does not mean that the numerical estimate is unusable.
18. Limitations
- Registry-level reporting: this page is restricted to the ClinicalTrials.gov record. The registry does not provide every possible detail of the statistical analysis plan.
- No median survival values reported: the ClinicalTrials.gov record contains hazard ratios, confidence intervals, and P-values for the reported time-to-event analyses but do not provide median PFS or median OS values.
- No subgroup estimates reported: although the overall analyses are stratified by prior treatment status and region, no subgroup-specific treatment-effect estimates are reported here.
- Proportional-hazards assumption: the Cox hazard ratio is model-based, and the ClinicalTrials.gov record does not provide a separate assessment of whether proportional hazards held throughout follow-up.
- Multiple OS looks: repeated interim analyses require alpha spending. The ClinicalTrials.gov record explicitly identify the Lan-DeMets/O'Brien-Fleming framework, and later OS analyses are explicitly described as exploratory.
- Different estimands: ITT PFS, response, duration of response, symptom progression, and LVEF analyses address different statistical questions and use different analysis populations.
- Safety denominators: the registry-reported serious-adverse-event counts use denominators of 396, 408, and 50, rather than the overall randomized enrollment of 808.
- Response selection: duration of objective response is analyzed only among participants who achieved an objective response, so it is not an ITT endpoint.
- Unreported analysis details: the registry does not report a formal method for the posted objective-response odds ratio or duration-of-response hazard ratio in the registry-reported entries.
19. Why This Trial Matters Statistically
CLEOPATRA is a useful statistical teaching case because its registry results combine randomized treatment allocation, a primary time-to-event endpoint, stratified survival analysis, independent review, multiple overall-survival looks, binary response analysis, duration-of-response analysis, symptom progression, and a nonparametric cardiac-function comparison.
| Concept | How it appears in CLEOPATRA |
|---|---|
| Randomization | 808 participants were enrolled in a randomized, parallel phase 3 trial. |
| Triple masking | The registry identifies the trial as triple-masked. |
| ITT analysis | The primary PFS analyses used all randomized participants. |
| Kaplan-Meier estimation | The registered PFS definition identifies the Kaplan-Meier method. |
| Hazard ratio | Primary PFS and several secondary time-to-event analyses use hazard ratios. |
| Confidence interval | Primary and several secondary hazard ratios have two-sided 95% CIs. |
| Log-rank testing | Stratified and unstratified log-rank analyses were reported for PFS. |
| Stratified analysis | Prior treatment status and region were used for the reported stratified analyses. |
| Cochran-Mantel-Haenszel testing | Used for the objective-response comparison. |
| Odds ratio | An OR of 1.79 was reported for objective response. |
| Interim analysis | Multiple OS analyses were reported, including two interim analyses. |
| Alpha spending | Lan-DeMets alpha spending with O'Brien-Fleming stopping boundaries was reported for OS. |
| Nonparametric analysis | Wilcoxon Rank Sum Test was used for maximum LVEF decrease. |
20. Related Tutorials
Learn more about the methods used in this trial:
21. Related Statistical Calculators
22. Sources
- ClinicalTrials.gov: CLEOPATRA, NCT00567190.
- PubMed: PMID 22149875.
- PubMed: PMID 32171426.
- PubMed: PMID 29112701.
- PubMed: PMID 28057664.
- PubMed: PMID 27964843.
Continue through Clinical Biostats
Build deeper understanding of the statistical methods that appear across randomized clinical trials, from survival analysis and hazard ratios to response analysis and interim monitoring.
23. Record Summary
CLEOPATRA provides a compact example of several central principles in clinical-trial statistics. The primary endpoint was a time-to-event outcome assessed by an independent review facility, analyzed in the ITT population using a stratified log-rank test and a Cox proportional-hazards effect measure. The reported primary PFS estimate was HR 0.62 (95% CI 0.51–0.75; P < 0.0001), with a closely aligned unstratified analysis of HR 0.63 (95% CI 0.52–0.76; P < 0.0001).
The secondary results illustrate why clinical-trial statistics cannot be reduced to a single P-value. Overall survival was examined at multiple analysis times under an interim-monitoring framework using Lan-DeMets alpha spending and O'Brien-Fleming stopping boundaries. Objective response was analyzed with both a response-rate difference and an odds ratio. Duration of response used a hazard ratio among responders, while time to symptom progression produced a different hazard-ratio estimate and P-value. LVEF was analyzed separately using a Wilcoxon rank-sum approach in the safety population.
The most informative statistical reading therefore combines the effect estimate, confidence interval, hypothesis test, analysis population, endpoint definition, and timing of the analysis. These elements together explain what the CLEOPATRA registry results establish statistically and what they do not establish.