← Clinical Trials
HR+/HER2- Breast Cancer Phase 3 Time-to-Event Analysis NCT04305496

CAPItello-291: Complete Statistical Analysis of Capivasertib + Fulvestrant in HR+/HER2- Breast Cancer

An independent statistical review of the randomized phase 3 CAPItello-291 trial comparing capivasertib plus fulvestrant with placebo plus fulvestrant in locally advanced (inoperable) or metastatic HR+/HER2- breast cancer, with emphasis on progression-free survival and stratified survival analysis.

Trial start: 2020-04-16  ·  Primary completion: 2023-05-09  ·  Enrollment: 818
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

CAPItello-291 is a randomized, quadruple-masked, parallel phase 3 treatment trial evaluating capivasertib plus fulvestrant versus placebo plus fulvestrant in patients with locally advanced (inoperable) or metastatic breast cancer described in the registry as HR+/HER2-.

818
Enrollment
2 treatment arms
3
Phase
Randomized
0.60
Global PFS HR
95% CI 0.51–0.71
0.51
China PFS HR
95% CI 0.34–0.76
FeatureCAPItello-291
PhasePhase 3
PopulationLocally advanced (inoperable) or metastatic HR+/HER2- breast cancer
DesignRandomized, parallel, quadruple-masked, treatment-purpose trial
AllocationRandomized
Arms2
Primary endpoint typeTime-to-event
Primary endpoints registered8
Statistical method reportedStratified log-rank test
Effect measureHazard ratio
Hypothesis typeSuperiority
StatusActive, not recruiting
Lead sponsorAstraZeneca
ClinicalTrials.govNCT04305496

2. Clinical Question

The central statistical question is whether capivasertib plus fulvestrant is associated with a different time-to-progression-or-death profile than placebo plus fulvestrant in the registered study population.

Population

Patients with locally advanced (inoperable) or metastatic HR+/HER2- breast cancer.

Intervention

Capivasertib plus fulvestrant.

Comparator

Placebo plus fulvestrant.

Primary question

Does the capivasertib-plus-fulvestrant strategy produce a different progression-free survival profile than placebo plus fulvestrant under the prespecified superiority analysis?

3. Trial Design

01
Randomize818 enrolled
02
Two armsCapivasertib or placebo + fulvestrant
03
MaskedQuadruple masking
04
AssessTime-to-event outcomes
05
CompareStratified log-rank analysis
ARM A

Capivasertib + Fulvestrant

  • Fulvestrant
  • Capivasertib
ARM B

Placebo + Fulvestrant

  • Fulvestrant
  • Placebo

The registry describes the allocation as randomized, the design model as parallel, the masking as quadruple, and the primary purpose as treatment. These features establish the basic comparison framework for interpreting the time-to-event results.

Registry cohort accounting: ClinicalTrials.gov reports 818 participants rather than 842 because 24 participants were included in both the global cohort and China cohort. This overlap is explicitly identified in the registry limitations and caveats.

4. Enrollment and Trial Timeline

2020-04-16

Trial start

The registry lists 2020-04-16 as the study start date.

2023-05-09

Primary completion

The registry lists 2023-05-09 as the primary completion date.

Current registry status

Active, not recruiting

The trial is listed as ACTIVE_NOT_RECRUITING.

5. Primary Endpoints

The registry contains eight registered primary endpoints. They are organized around progression-free survival, with results reported for the global cohort and China cohort and for both overall and altered populations. The paired month and percentage entries use the same statistical analysis and effect estimate in the posted registry analyses.

EndpointTime frameAnalysis
Progression Free Survival: Overall Population (Months) in the Global CohortAssessed every 8 weeks for the first 18 months and every 12 weeks thereafter, from randomization to radiological progressionStratified log-rank; HR 0.60 (95% CI 0.51–0.71); P < 0.001
Progression Free Survival: Overall Population (Percentage) in the Global CohortAssessed every 8 weeks for the first 18 months and every 12 weeks thereafter, from randomization to radiological progressionStratified log-rank; HR 0.60 (95% CI 0.51–0.71); P < 0.001
Progression Free Survival: Altered Population (Months) in the Global CohortAssessed every 8 weeks for the first 2 years following objective disease progression or treatment discontinuation and thereafterStratified log-rank; HR 0.50 (95% CI 0.38–0.65); P < 0.001
Progression Free Survival: Altered Population (Percentage) in the Global CohortAssessed every 8 weeks for the first 2 years following objective disease progression or treatment discontinuation and thereafterStratified log-rank; HR 0.50 (95% CI 0.38–0.65); P < 0.001
Progression Free Survival: Overall Population (Months) in the China CohortAssessed every 8 weeks for the first 18 months and every 12 weeks thereafter, from randomization to radiological progressionStratified log-rank; HR 0.51 (95% CI 0.34–0.76); P < 0.001
Progression Free Survival: Overall Population (Percentage) in the China CohortAssessed every 8 weeks for the first 18 months and every 12 weeks thereafter, from randomization to radiological progressionStratified log-rank; HR 0.51 (95% CI 0.34–0.76); P < 0.001
Progression Free Survival: Altered Population (Months) in the China CohortAssessed every 8 weeks for the first 2 years following objective disease progression or treatment discontinuation and thereafterStratified log-rank; HR 0.41 (95% CI 0.19–0.85); P = 0.016
Progression Free Survival: Altered Population (Percentage) in the China CohortAssessed every 8 weeks for the first 2 years following objective disease progression or treatment discontinuation and thereafterStratified log-rank; HR 0.41 (95% CI 0.19–0.85); P = 0.016

The registry defines progression-free survival as the time from randomization until progression per RECIST v1.1, as assessed by the investigator at the local site, or death due to any cause. For the overall global endpoint, the registry additionally states that participants who discontinue treatment prior to progression should continue to be scanned until progression.

6. Statistical Methodology

Stratified log-rank test

The reported formal method for the primary time-to-event analyses is the stratified log-rank test. This is a survival-analysis method designed to compare the timing of events between randomized groups while accounting for prespecified strata when those strata are part of the analysis design.

Conceptual question
Do the observed event-time distributions differ between randomized treatment groups after accounting for the analysis strata?

The test uses the ordering of event and censoring times rather than reducing follow-up to a single binary outcome. That makes it appropriate for a progression-free survival endpoint in which patients can have different lengths of follow-up.

Hazard ratio

The registry reports the hazard ratio as the effect measure for all eight primary endpoint analyses. A hazard ratio compares the estimated instantaneous event rate between groups within a time-to-event framework.

Interpretation of the hazard ratio
HR < 1  →  lower estimated instantaneous event rate in the capivasertib group relative to the comparator

The hazard ratio is a relative time-to-event measure. It is not a probability, a percentage of patients cured, a median survival difference, or an absolute risk difference.

Kaplan-Meier estimation

The registry specifically states that a Kaplan-Meier estimate was used for the altered-population percentage endpoints. Kaplan-Meier estimation is also the standard descriptive framework for presenting progression-free survival over time because it accommodates right-censored observations.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at an event time and ni represents the number at risk immediately before that time.

Intention-to-treat analysis

The China-cohort primary analyses explicitly identify intention-to-treat analysis in the analysis text for the overall and altered populations. The ITT principle analyzes participants according to their randomized treatment assignment and preserves the treatment-comparison framework created by randomization.

Superiority testing

The registered hypothesis type is superiority. This means the statistical question is framed around whether the treatment groups differ, rather than around demonstrating that an experimental treatment is no worse than a prespecified non-inferiority margin.

7. Results: Overall Population in the Global Cohort

The registry reports formal statistical analyses for both the month-based and percentage-based versions of the overall-population global-cohort PFS endpoint. Both entries report the same hazard ratio, confidence interval, and P-value.

Global cohort · Overall population

HR 0.60

95% CI: 0.51–0.71   ·   P < 0.001

Method: stratified log-rank test  ·  Hypothesis: superiority

Registered endpoint formAnalysis populationEffect estimate95% CIP-value
Overall Population (Months)Overall population in the global cohortHR 0.600.51–0.71<0.001
Overall Population (Percentage)Overall population in the Global CohortHR 0.600.51–0.71<0.001
Clinical Biostats interpretation

An HR of 0.60 means that the estimated instantaneous rate of progression or death was 0.60 times the corresponding rate in the comparator under the reported time-to-event analysis. Equivalently, this corresponds to a 40% lower estimated hazard because 1 − 0.60 = 0.40.

The HR does not mean that 40% of participants avoided progression, that every participant had a 40% reduction in risk, or that the absolute probability of progression or death was reduced by 40 percentage points.

The two-sided 95% confidence interval of 0.51–0.71 describes statistical uncertainty around the estimated hazard ratio. It does not describe the range of individual treatment effects across patients.

The P < 0.001 result addresses evidence against the null hypothesis used for the statistical comparison. A P-value does not measure the magnitude of the treatment effect and should not be interpreted as the probability that the treatment effect is real.

Because this is a hazard ratio, interpretation also depends on the time-to-event model and its assumptions. A single HR is a relative summary over follow-up; it is not equivalent to an absolute risk difference at a particular time point.

8. Results: Altered Population in the Global Cohort

The registry reports a separate altered-population PFS analysis for the global cohort. The month-based and percentage-based endpoint entries again report identical formal statistical results.

Global cohort · Altered population

HR 0.50

95% CI: 0.38–0.65   ·   P < 0.001

Method: stratified log-rank test  ·  Hypothesis: superiority

Registered endpoint formAnalysis populationEffect estimate95% CIP-value
Altered Population (Months)Altered population in the Global CohortHR 0.500.38–0.65<0.001
Altered Population (Percentage)Altered population in the Global CohortHR 0.500.38–0.65<0.001
Clinical Biostats interpretation

An HR of 0.50 corresponds to a 50% lower estimated instantaneous hazard of progression or death in the capivasertib-containing comparison, because 1 − 0.50 = 0.50.

The 95% CI of 0.38–0.65 indicates uncertainty around that estimated relative hazard. It does not establish that the true treatment effect for every patient lies within that numerical range.

The P < 0.001 value indicates strong statistical evidence against the null comparison under the reported analysis. It does not say that the treatment has a 99.9% or greater probability of being effective, nor does it quantify clinical importance.

The registry reports this analysis as a stratified log-rank test with a superiority hypothesis. The interpretation therefore remains tied to the time-to-event comparison and its censoring and modeling framework.

9. Results: Overall Population in the China Cohort

The China-cohort analysis was reported using the full analysis set and explicitly identifies the intention-to-treat principle in the analysis text. The registry reports the same statistical result for the month and percentage endpoint forms.

China cohort · Overall population

HR 0.51

95% CI: 0.34–0.76   ·   P < 0.001

Method: stratified log-rank test  ·  Hypothesis: superiority

Registered endpoint formAnalysis populationEffect estimate95% CIP-value
Overall Population (Months)Full analysis set in the China cohortHR 0.510.34–0.76<0.001
Overall Population (Percentage)Full analysis set in the China cohortHR 0.510.34–0.76<0.001
Clinical Biostats interpretation

An HR of 0.51 corresponds to a 49% lower estimated instantaneous hazard relative to the comparator because 1 − 0.51 = 0.49.

The 95% CI of 0.34–0.76 is wider than the corresponding global-cohort overall-population interval of 0.51–0.71. That difference illustrates how estimates from a smaller cohort can carry more statistical uncertainty, although the ClinicalTrials.gov record does not provide the cohort-specific event counts or sample size needed to quantify that precision difference further.

The P < 0.001 value indicates statistical evidence for a difference under the reported superiority analysis. It is not an effect-size measure and does not indicate how large the absolute difference in PFS probability is.

The China-cohort analysis explicitly references intention-to-treat analysis, which is important because treatment assignment remains the basis of the randomized comparison.

10. Results: Altered Population in the China Cohort

The altered-population China-cohort analysis produced the smallest reported hazard ratio among the eight primary endpoint analyses. The registry reports the same result for both the month and percentage versions of the endpoint.

China cohort · Altered population

HR 0.41

95% CI: 0.19–0.85   ·   P = 0.016

Method: stratified log-rank test  ·  Hypothesis: superiority

Registered endpoint formAnalysis populationEffect estimate95% CIP-value
Altered Population (Months)Altered subgroup full analysis set in the China cohortHR 0.410.19–0.850.016
Altered Population (Percentage)Altered subgroup full analysis set in the China cohortHR 0.410.19–0.850.016
Clinical Biostats interpretation

An HR of 0.41 corresponds to a 59% lower estimated instantaneous hazard relative to the comparator because 1 − 0.41 = 0.59.

The confidence interval of 0.19–0.85 is relatively broad around the point estimate. The upper end remains below 1, but the interval itself shows substantial uncertainty about the precise magnitude of the relative effect.

The reported P = 0.016 is evidence against the null hypothesis under the stated analysis. It should not be interpreted as a measure of the size or clinical importance of the treatment effect.

This is a more restricted analysis population than the overall global cohort. It should therefore not be treated as interchangeable with the overall-population estimate or as evidence that the treatment effect is definitively different between these populations. A formal comparison of effects across populations would require an appropriate interaction or heterogeneity analysis, which is not provided in the ClinicalTrials.gov record.

11. Primary Results Summary

The eight posted primary endpoint analyses reduce to four distinct statistical comparisons because each month-based endpoint is paired with a percentage-based endpoint carrying the same estimate, confidence interval, and P-value.

PopulationEndpoint formHR95% CIP-valueMethod
Global · OverallMonths / Percentage0.600.51–0.71<0.001Stratified log-rank
Global · AlteredMonths / Percentage0.500.38–0.65<0.001Stratified log-rank
China · OverallMonths / Percentage0.510.34–0.76<0.001Stratified log-rank
China · AlteredMonths / Percentage0.410.19–0.850.016Stratified log-rank
The registry does not provide median PFS, Kaplan-Meier survival probabilities at specific time points, event counts for each efficacy analysis, or reconstructed Kaplan-Meier curves in the ClinicalTrials.gov record. Those quantities are therefore not inferred or added here.

12. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm for the global and China cohorts. Because 24 participants were included in both cohorts, the cohort counts should not be added together as though they represented mutually exclusive participants.

CohortCapivasertib + FulvestrantPlacebo + Fulvestrant
Global Cohort57 / 355 affected / at risk28 / 350 affected / at risk
China Cohort20 / 71 affected / at risk3 / 62 affected / at risk

Global cohort

The registry reports serious adverse events affecting 57 of 355 participants in the capivasertib group and 28 of 350 in the placebo group.

China cohort

The registry reports serious adverse events affecting 20 of 71 participants in the capivasertib group and 3 of 62 in the placebo group.

Safety interpretation: These figures describe serious adverse events using the affected/at-risk format reported by the registry. They should not be converted into an inferred comparative risk measure or combined across cohorts without the underlying participant-level accounting and the registry's exact safety-analysis definitions.

13. Statistical Methods Explained

Why was a stratified log-rank test used?

The primary endpoints are time-to-event outcomes, so the analysis needs to account for both whether an event occurred and when it occurred. A log-rank test compares the event-time experience of randomized groups across follow-up. The registry specifically reports the stratified form, which incorporates analysis strata into that comparison.

What does an HR of 0.60 mean?

An HR of 0.60 means that the estimated instantaneous event rate in the capivasertib-plus-fulvestrant group was 0.60 times that of the comparator under the reported time-to-event analysis. It can be expressed as a 40% lower estimated hazard, but it should not be translated into a 40% reduction in absolute probability.

Why is the confidence interval important?

A point estimate such as 0.60 is only one estimate from the observed data. The 95% CI of 0.51–0.71 shows the uncertainty around that estimate under the statistical framework. A narrow interval indicates greater numerical precision than a wide interval, although precision and clinical importance are separate questions.

Why does the P-value not measure effect size?

A P-value describes how compatible the observed data are with a specified null hypothesis under the statistical model. It does not tell us how large the treatment effect is. For CAPItello-291, the effect size is communicated primarily through the hazard ratio and its confidence interval.

Why does the China-cohort estimate need separate interpretation?

The China cohort is a distinct analysis population, and the registry explicitly identifies a full analysis set and intention-to-treat analysis for those primary analyses. An HR of 0.51 in the China overall population and 0.60 in the global overall population cannot, by themselves, establish that treatment effect differs between populations. A formal heterogeneity or interaction analysis would be needed for that question.

Why are the month and percentage endpoint entries not eight independent treatment effects?

The registry contains eight primary endpoint entries, but the statistical analyses posted on ClinicalTrials.gov show that the paired month and percentage forms for each population use identical estimates, confidence intervals, P-values, and statistical methods. They therefore represent different registered presentations of the same reported statistical comparison rather than eight distinct numerical treatment effects.

What does the superiority hypothesis imply?

The registered hypothesis type is superiority. The analysis therefore asks whether the randomized groups differ in the relevant time-to-event outcome. This differs from a non-inferiority design, where the central question is whether the treatment remains within a prespecified acceptable margin of the comparator.

14. Confidence Intervals, Censoring, and Time-to-Event Interpretation

Progression-free survival is fundamentally different from a simple binary endpoint because patients can enter follow-up at the time of randomization and contribute information until progression, death, or censoring. The registry definition explicitly identifies progression or death as the event and states that participants who discontinue treatment before progression should continue to be scanned until progression for the overall global endpoint.

Event timing

The endpoint records when progression or death occurs rather than only whether it occurs during an arbitrary fixed window.

Censoring

Participants without an observed event at the relevant follow-up point contribute information up to the point at which they are censored.

Relative effect

The hazard ratio summarizes the relative event rate between treatment groups within the time-to-event framework.

Absolute effect

Absolute PFS probabilities or median PFS would provide a different perspective, but those numerical results are not included in the ClinicalTrials.gov record.

Important distinction: the hazard ratio and confidence interval answer a relative time-to-event question. They do not replace absolute survival probabilities, median event times, or patient-level descriptions of benefit. Those measures should be reported separately when the underlying data are available.

15. Stratification and Intention-to-Treat Analysis

The statistical record identifies stratified log-rank testing as the primary analysis method and explicitly identifies intention-to-treat analysis for the China-cohort primary analyses. These two concepts address different parts of the analysis.

ConceptRole in this trial
RandomizationCreates the treatment-group comparison for the phase 3 parallel design.
Intention-to-treat analysisExplicitly identified in the China-cohort primary analyses.
Stratified log-rank testReported statistical method for the primary time-to-event analyses.
Hazard ratioReported effect measure for all eight primary endpoint analyses.
SuperiorityRegistered hypothesis type for the primary analyses.

The ClinicalTrials.gov record does not identify the specific randomization or analysis stratification factors. Consequently, no particular clinical variables are attributed to the stratification scheme here.

16. What the Hazard Ratios Do — and Do Not — Mean

Relative effect

The global overall-population HR of 0.60 indicates a 40% lower estimated instantaneous hazard of progression or death in the capivasertib comparison relative to placebo plus fulvestrant.

Not an absolute probability

An HR of 0.60 does not mean that 40% of patients avoided progression, that 40% more patients were alive without progression at a particular time, or that every patient experienced the same proportional reduction.

Precision

The corresponding 95% CI of 0.51–0.71 quantifies uncertainty around the estimated relative hazard under the reported statistical framework.

P-value

The P < 0.001 result is evidence from the specified statistical test against its null hypothesis. It is not a measure of treatment magnitude and should not be used as a substitute for the hazard ratio or confidence interval.

17. Comparing the Reported Populations

The registry reports results for both global and China cohorts and distinguishes overall from altered populations. These estimates can be described side by side, but they should not be treated as direct evidence of effect modification without a formal comparison.

PopulationHR95% CIP-valueDescriptive interpretation
Global · Overall0.600.51–0.71<0.001Estimated hazard was 0.60 times the comparator hazard.
Global · Altered0.500.38–0.65<0.001Estimated hazard was 0.50 times the comparator hazard.
China · Overall0.510.34–0.76<0.001Estimated hazard was 0.51 times the comparator hazard.
China · Altered0.410.19–0.850.016Estimated hazard was 0.41 times the comparator hazard.
Subgroup caution: Different point estimates across populations do not automatically demonstrate different treatment effects. A formal claim of heterogeneity requires an appropriate statistical comparison of treatment effects, such as an interaction analysis. No such comparison is included in the ClinicalTrials.gov record.

18. Limitations

19. Why This Trial Matters Statistically

CAPItello-291 is a useful teaching case because its registry record brings together randomized treatment comparison, masking, time-to-event endpoints, stratified survival testing, hazard ratios, confidence intervals, intention-to-treat analysis, and multiple analysis populations.

ConceptHow it appears in CAPItello-291
RandomizationThe trial is registered as randomized with a parallel design and 2 arms.
BlindingThe registry specifies quadruple masking.
Time-to-event endpointAll eight registered primary endpoints are progression-free survival measures.
Stratified log-rank testReported formal statistical method for the posted primary analyses.
Hazard ratioReported effect measure for every primary analysis.
Confidence intervalEvery posted primary analysis includes a two-sided 95% confidence interval.
P-valueEach primary analysis includes a reported P-value.
Intention-to-treatExplicitly identified in the China-cohort primary analyses.
Kaplan-Meier estimationExplicitly identified for the altered-population percentage analyses.
SuperiorityRegistered hypothesis type for the primary analyses.
Cohort overlapThe registry notes that 24 participants were included in both the global and China cohorts.
Safety analysisSerious adverse events are reported by arm for global and China cohorts.

20. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The reported randomized comparisons produced hazard ratios below 1 for all four distinct population-level analyses, with two-sided 95% confidence intervals and P-values reported under stratified log-rank testing.

Clinical interpretation

The registry data support describing the observed progression-free survival comparisons in relative terms. The ClinicalTrials.gov record does not provide enough numerical information to characterize median PFS, absolute PFS differences, or the duration of benefit.

21. A Note on the Eight Primary Endpoint Entries

At first glance, the registry's eight primary endpoints may appear to represent eight separate efficacy questions. The statistical analyses posted on ClinicalTrials.gov show a more structured pattern: four population-level comparisons are each registered twice, once as a month-based outcome and once as a percentage outcome.

PopulationMonth endpointPercentage endpointSame reported analysis?
Global · OverallYesYesYes — HR 0.60; 95% CI 0.51–0.71; P < 0.001
Global · AlteredYesYesYes — HR 0.50; 95% CI 0.38–0.65; P < 0.001
China · OverallYesYesYes — HR 0.51; 95% CI 0.34–0.76; P < 0.001
China · AlteredYesYesYes — HR 0.41; 95% CI 0.19–0.85; P = 0.016

This distinction matters statistically because counting every registry entry as an independent hypothesis test would misrepresent the information reported by the actual statistical analyses.

22. Related Statistical Concepts

Learn more about the methods used in this trial:

23. Related Statistical Calculators

Explore calculators that reinforce the quantitative concepts behind this trial:

24. Sources

Continue through Clinical Biostats

Use the statistical concepts in this trial as a starting point for deeper study of survival analysis, clinical-trial methods, and statistical calculation.

25. Record Summary

CAPItello-291 provides a clear example of randomized time-to-event analysis. The registry describes a phase 3, randomized, parallel, quadruple-masked trial with two treatment arms and eight registered primary endpoints centered on progression-free survival. The posted analyses use stratified log-rank testing and hazard ratios under a superiority framework, with two-sided 95% confidence intervals and reported P-values. The four distinct population-level comparisons have hazard ratios of 0.60, 0.50, 0.51, and 0.41, respectively, with the associated uncertainty and P-values reported above.

The most important statistical lesson is that these numbers must be interpreted as time-to-event treatment-effect estimates, not as absolute probabilities. Confidence intervals describe uncertainty around the estimated relative effects, while P-values address evidence against the corresponding null hypothesis rather than the magnitude of benefit. The global and China cohorts, as well as the overall and altered populations, should be kept analytically distinct, and the registry's 24-participant overlap between the global and China cohorts prevents simple addition of those cohort counts.

Clinical Biostats methodology: A trial-results page should separate registry-reported evidence from statistical interpretation. Where the ClinicalTrials.gov record does not report median event times, absolute survival probabilities, event counts, multiplicity procedures, interim-analysis rules, or missing-data methods, those quantities and methods are not inferred.