This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
CAPItello-291 is a randomized, quadruple-masked, parallel phase 3 treatment trial evaluating capivasertib plus fulvestrant versus placebo plus fulvestrant in patients with locally advanced (inoperable) or metastatic breast cancer described in the registry as HR+/HER2-.
| Feature | CAPItello-291 |
|---|---|
| Phase | Phase 3 |
| Population | Locally advanced (inoperable) or metastatic HR+/HER2- breast cancer |
| Design | Randomized, parallel, quadruple-masked, treatment-purpose trial |
| Allocation | Randomized |
| Arms | 2 |
| Primary endpoint type | Time-to-event |
| Primary endpoints registered | 8 |
| Statistical method reported | Stratified log-rank test |
| Effect measure | Hazard ratio |
| Hypothesis type | Superiority |
| Status | Active, not recruiting |
| Lead sponsor | AstraZeneca |
| ClinicalTrials.gov | NCT04305496 |
2. Clinical Question
The central statistical question is whether capivasertib plus fulvestrant is associated with a different time-to-progression-or-death profile than placebo plus fulvestrant in the registered study population.
Population
Patients with locally advanced (inoperable) or metastatic HR+/HER2- breast cancer.
Intervention
Capivasertib plus fulvestrant.
Comparator
Placebo plus fulvestrant.
Primary question
Does the capivasertib-plus-fulvestrant strategy produce a different progression-free survival profile than placebo plus fulvestrant under the prespecified superiority analysis?
3. Trial Design
Capivasertib + Fulvestrant
- Fulvestrant
- Capivasertib
Placebo + Fulvestrant
- Fulvestrant
- Placebo
The registry describes the allocation as randomized, the design model as parallel, the masking as quadruple, and the primary purpose as treatment. These features establish the basic comparison framework for interpreting the time-to-event results.
4. Enrollment and Trial Timeline
Trial start
The registry lists 2020-04-16 as the study start date.
Primary completion
The registry lists 2023-05-09 as the primary completion date.
Active, not recruiting
The trial is listed as ACTIVE_NOT_RECRUITING.
5. Primary Endpoints
The registry contains eight registered primary endpoints. They are organized around progression-free survival, with results reported for the global cohort and China cohort and for both overall and altered populations. The paired month and percentage entries use the same statistical analysis and effect estimate in the posted registry analyses.
| Endpoint | Time frame | Analysis |
|---|---|---|
| Progression Free Survival: Overall Population (Months) in the Global Cohort | Assessed every 8 weeks for the first 18 months and every 12 weeks thereafter, from randomization to radiological progression | Stratified log-rank; HR 0.60 (95% CI 0.51–0.71); P < 0.001 |
| Progression Free Survival: Overall Population (Percentage) in the Global Cohort | Assessed every 8 weeks for the first 18 months and every 12 weeks thereafter, from randomization to radiological progression | Stratified log-rank; HR 0.60 (95% CI 0.51–0.71); P < 0.001 |
| Progression Free Survival: Altered Population (Months) in the Global Cohort | Assessed every 8 weeks for the first 2 years following objective disease progression or treatment discontinuation and thereafter | Stratified log-rank; HR 0.50 (95% CI 0.38–0.65); P < 0.001 |
| Progression Free Survival: Altered Population (Percentage) in the Global Cohort | Assessed every 8 weeks for the first 2 years following objective disease progression or treatment discontinuation and thereafter | Stratified log-rank; HR 0.50 (95% CI 0.38–0.65); P < 0.001 |
| Progression Free Survival: Overall Population (Months) in the China Cohort | Assessed every 8 weeks for the first 18 months and every 12 weeks thereafter, from randomization to radiological progression | Stratified log-rank; HR 0.51 (95% CI 0.34–0.76); P < 0.001 |
| Progression Free Survival: Overall Population (Percentage) in the China Cohort | Assessed every 8 weeks for the first 18 months and every 12 weeks thereafter, from randomization to radiological progression | Stratified log-rank; HR 0.51 (95% CI 0.34–0.76); P < 0.001 |
| Progression Free Survival: Altered Population (Months) in the China Cohort | Assessed every 8 weeks for the first 2 years following objective disease progression or treatment discontinuation and thereafter | Stratified log-rank; HR 0.41 (95% CI 0.19–0.85); P = 0.016 |
| Progression Free Survival: Altered Population (Percentage) in the China Cohort | Assessed every 8 weeks for the first 2 years following objective disease progression or treatment discontinuation and thereafter | Stratified log-rank; HR 0.41 (95% CI 0.19–0.85); P = 0.016 |
The registry defines progression-free survival as the time from randomization until progression per RECIST v1.1, as assessed by the investigator at the local site, or death due to any cause. For the overall global endpoint, the registry additionally states that participants who discontinue treatment prior to progression should continue to be scanned until progression.
6. Statistical Methodology
Stratified log-rank test
The reported formal method for the primary time-to-event analyses is the stratified log-rank test. This is a survival-analysis method designed to compare the timing of events between randomized groups while accounting for prespecified strata when those strata are part of the analysis design.
The test uses the ordering of event and censoring times rather than reducing follow-up to a single binary outcome. That makes it appropriate for a progression-free survival endpoint in which patients can have different lengths of follow-up.
Hazard ratio
The registry reports the hazard ratio as the effect measure for all eight primary endpoint analyses. A hazard ratio compares the estimated instantaneous event rate between groups within a time-to-event framework.
The hazard ratio is a relative time-to-event measure. It is not a probability, a percentage of patients cured, a median survival difference, or an absolute risk difference.
Kaplan-Meier estimation
The registry specifically states that a Kaplan-Meier estimate was used for the altered-population percentage endpoints. Kaplan-Meier estimation is also the standard descriptive framework for presenting progression-free survival over time because it accommodates right-censored observations.
Here, di represents the number of events at an event time and ni represents the number at risk immediately before that time.
Intention-to-treat analysis
The China-cohort primary analyses explicitly identify intention-to-treat analysis in the analysis text for the overall and altered populations. The ITT principle analyzes participants according to their randomized treatment assignment and preserves the treatment-comparison framework created by randomization.
Superiority testing
The registered hypothesis type is superiority. This means the statistical question is framed around whether the treatment groups differ, rather than around demonstrating that an experimental treatment is no worse than a prespecified non-inferiority margin.
7. Results: Overall Population in the Global Cohort
The registry reports formal statistical analyses for both the month-based and percentage-based versions of the overall-population global-cohort PFS endpoint. Both entries report the same hazard ratio, confidence interval, and P-value.
Global cohort · Overall population
95% CI: 0.51–0.71 · P < 0.001
Method: stratified log-rank test · Hypothesis: superiority
| Registered endpoint form | Analysis population | Effect estimate | 95% CI | P-value |
|---|---|---|---|---|
| Overall Population (Months) | Overall population in the global cohort | HR 0.60 | 0.51–0.71 | <0.001 |
| Overall Population (Percentage) | Overall population in the Global Cohort | HR 0.60 | 0.51–0.71 | <0.001 |
An HR of 0.60 means that the estimated instantaneous rate of progression or death was 0.60 times the corresponding rate in the comparator under the reported time-to-event analysis. Equivalently, this corresponds to a 40% lower estimated hazard because 1 − 0.60 = 0.40.
The HR does not mean that 40% of participants avoided progression, that every participant had a 40% reduction in risk, or that the absolute probability of progression or death was reduced by 40 percentage points.
The two-sided 95% confidence interval of 0.51–0.71 describes statistical uncertainty around the estimated hazard ratio. It does not describe the range of individual treatment effects across patients.
The P < 0.001 result addresses evidence against the null hypothesis used for the statistical comparison. A P-value does not measure the magnitude of the treatment effect and should not be interpreted as the probability that the treatment effect is real.
Because this is a hazard ratio, interpretation also depends on the time-to-event model and its assumptions. A single HR is a relative summary over follow-up; it is not equivalent to an absolute risk difference at a particular time point.
8. Results: Altered Population in the Global Cohort
The registry reports a separate altered-population PFS analysis for the global cohort. The month-based and percentage-based endpoint entries again report identical formal statistical results.
Global cohort · Altered population
95% CI: 0.38–0.65 · P < 0.001
Method: stratified log-rank test · Hypothesis: superiority
| Registered endpoint form | Analysis population | Effect estimate | 95% CI | P-value |
|---|---|---|---|---|
| Altered Population (Months) | Altered population in the Global Cohort | HR 0.50 | 0.38–0.65 | <0.001 |
| Altered Population (Percentage) | Altered population in the Global Cohort | HR 0.50 | 0.38–0.65 | <0.001 |
An HR of 0.50 corresponds to a 50% lower estimated instantaneous hazard of progression or death in the capivasertib-containing comparison, because 1 − 0.50 = 0.50.
The 95% CI of 0.38–0.65 indicates uncertainty around that estimated relative hazard. It does not establish that the true treatment effect for every patient lies within that numerical range.
The P < 0.001 value indicates strong statistical evidence against the null comparison under the reported analysis. It does not say that the treatment has a 99.9% or greater probability of being effective, nor does it quantify clinical importance.
The registry reports this analysis as a stratified log-rank test with a superiority hypothesis. The interpretation therefore remains tied to the time-to-event comparison and its censoring and modeling framework.
9. Results: Overall Population in the China Cohort
The China-cohort analysis was reported using the full analysis set and explicitly identifies the intention-to-treat principle in the analysis text. The registry reports the same statistical result for the month and percentage endpoint forms.
China cohort · Overall population
95% CI: 0.34–0.76 · P < 0.001
Method: stratified log-rank test · Hypothesis: superiority
| Registered endpoint form | Analysis population | Effect estimate | 95% CI | P-value |
|---|---|---|---|---|
| Overall Population (Months) | Full analysis set in the China cohort | HR 0.51 | 0.34–0.76 | <0.001 |
| Overall Population (Percentage) | Full analysis set in the China cohort | HR 0.51 | 0.34–0.76 | <0.001 |
An HR of 0.51 corresponds to a 49% lower estimated instantaneous hazard relative to the comparator because 1 − 0.51 = 0.49.
The 95% CI of 0.34–0.76 is wider than the corresponding global-cohort overall-population interval of 0.51–0.71. That difference illustrates how estimates from a smaller cohort can carry more statistical uncertainty, although the ClinicalTrials.gov record does not provide the cohort-specific event counts or sample size needed to quantify that precision difference further.
The P < 0.001 value indicates statistical evidence for a difference under the reported superiority analysis. It is not an effect-size measure and does not indicate how large the absolute difference in PFS probability is.
The China-cohort analysis explicitly references intention-to-treat analysis, which is important because treatment assignment remains the basis of the randomized comparison.
10. Results: Altered Population in the China Cohort
The altered-population China-cohort analysis produced the smallest reported hazard ratio among the eight primary endpoint analyses. The registry reports the same result for both the month and percentage versions of the endpoint.
China cohort · Altered population
95% CI: 0.19–0.85 · P = 0.016
Method: stratified log-rank test · Hypothesis: superiority
| Registered endpoint form | Analysis population | Effect estimate | 95% CI | P-value |
|---|---|---|---|---|
| Altered Population (Months) | Altered subgroup full analysis set in the China cohort | HR 0.41 | 0.19–0.85 | 0.016 |
| Altered Population (Percentage) | Altered subgroup full analysis set in the China cohort | HR 0.41 | 0.19–0.85 | 0.016 |
An HR of 0.41 corresponds to a 59% lower estimated instantaneous hazard relative to the comparator because 1 − 0.41 = 0.59.
The confidence interval of 0.19–0.85 is relatively broad around the point estimate. The upper end remains below 1, but the interval itself shows substantial uncertainty about the precise magnitude of the relative effect.
The reported P = 0.016 is evidence against the null hypothesis under the stated analysis. It should not be interpreted as a measure of the size or clinical importance of the treatment effect.
This is a more restricted analysis population than the overall global cohort. It should therefore not be treated as interchangeable with the overall-population estimate or as evidence that the treatment effect is definitively different between these populations. A formal comparison of effects across populations would require an appropriate interaction or heterogeneity analysis, which is not provided in the ClinicalTrials.gov record.
11. Primary Results Summary
The eight posted primary endpoint analyses reduce to four distinct statistical comparisons because each month-based endpoint is paired with a percentage-based endpoint carrying the same estimate, confidence interval, and P-value.
| Population | Endpoint form | HR | 95% CI | P-value | Method |
|---|---|---|---|---|---|
| Global · Overall | Months / Percentage | 0.60 | 0.51–0.71 | <0.001 | Stratified log-rank |
| Global · Altered | Months / Percentage | 0.50 | 0.38–0.65 | <0.001 | Stratified log-rank |
| China · Overall | Months / Percentage | 0.51 | 0.34–0.76 | <0.001 | Stratified log-rank |
| China · Altered | Months / Percentage | 0.41 | 0.19–0.85 | 0.016 | Stratified log-rank |
12. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm for the global and China cohorts. Because 24 participants were included in both cohorts, the cohort counts should not be added together as though they represented mutually exclusive participants.
| Cohort | Capivasertib + Fulvestrant | Placebo + Fulvestrant |
|---|---|---|
| Global Cohort | 57 / 355 affected / at risk | 28 / 350 affected / at risk |
| China Cohort | 20 / 71 affected / at risk | 3 / 62 affected / at risk |
Global cohort
The registry reports serious adverse events affecting 57 of 355 participants in the capivasertib group and 28 of 350 in the placebo group.
China cohort
The registry reports serious adverse events affecting 20 of 71 participants in the capivasertib group and 3 of 62 in the placebo group.
13. Statistical Methods Explained
Why was a stratified log-rank test used?
The primary endpoints are time-to-event outcomes, so the analysis needs to account for both whether an event occurred and when it occurred. A log-rank test compares the event-time experience of randomized groups across follow-up. The registry specifically reports the stratified form, which incorporates analysis strata into that comparison.
What does an HR of 0.60 mean?
An HR of 0.60 means that the estimated instantaneous event rate in the capivasertib-plus-fulvestrant group was 0.60 times that of the comparator under the reported time-to-event analysis. It can be expressed as a 40% lower estimated hazard, but it should not be translated into a 40% reduction in absolute probability.
Why is the confidence interval important?
A point estimate such as 0.60 is only one estimate from the observed data. The 95% CI of 0.51–0.71 shows the uncertainty around that estimate under the statistical framework. A narrow interval indicates greater numerical precision than a wide interval, although precision and clinical importance are separate questions.
Why does the P-value not measure effect size?
A P-value describes how compatible the observed data are with a specified null hypothesis under the statistical model. It does not tell us how large the treatment effect is. For CAPItello-291, the effect size is communicated primarily through the hazard ratio and its confidence interval.
Why does the China-cohort estimate need separate interpretation?
The China cohort is a distinct analysis population, and the registry explicitly identifies a full analysis set and intention-to-treat analysis for those primary analyses. An HR of 0.51 in the China overall population and 0.60 in the global overall population cannot, by themselves, establish that treatment effect differs between populations. A formal heterogeneity or interaction analysis would be needed for that question.
Why are the month and percentage endpoint entries not eight independent treatment effects?
The registry contains eight primary endpoint entries, but the statistical analyses posted on ClinicalTrials.gov show that the paired month and percentage forms for each population use identical estimates, confidence intervals, P-values, and statistical methods. They therefore represent different registered presentations of the same reported statistical comparison rather than eight distinct numerical treatment effects.
What does the superiority hypothesis imply?
The registered hypothesis type is superiority. The analysis therefore asks whether the randomized groups differ in the relevant time-to-event outcome. This differs from a non-inferiority design, where the central question is whether the treatment remains within a prespecified acceptable margin of the comparator.
14. Confidence Intervals, Censoring, and Time-to-Event Interpretation
Progression-free survival is fundamentally different from a simple binary endpoint because patients can enter follow-up at the time of randomization and contribute information until progression, death, or censoring. The registry definition explicitly identifies progression or death as the event and states that participants who discontinue treatment before progression should continue to be scanned until progression for the overall global endpoint.
Event timing
The endpoint records when progression or death occurs rather than only whether it occurs during an arbitrary fixed window.
Censoring
Participants without an observed event at the relevant follow-up point contribute information up to the point at which they are censored.
Relative effect
The hazard ratio summarizes the relative event rate between treatment groups within the time-to-event framework.
Absolute effect
Absolute PFS probabilities or median PFS would provide a different perspective, but those numerical results are not included in the ClinicalTrials.gov record.
15. Stratification and Intention-to-Treat Analysis
The statistical record identifies stratified log-rank testing as the primary analysis method and explicitly identifies intention-to-treat analysis for the China-cohort primary analyses. These two concepts address different parts of the analysis.
| Concept | Role in this trial |
|---|---|
| Randomization | Creates the treatment-group comparison for the phase 3 parallel design. |
| Intention-to-treat analysis | Explicitly identified in the China-cohort primary analyses. |
| Stratified log-rank test | Reported statistical method for the primary time-to-event analyses. |
| Hazard ratio | Reported effect measure for all eight primary endpoint analyses. |
| Superiority | Registered hypothesis type for the primary analyses. |
The ClinicalTrials.gov record does not identify the specific randomization or analysis stratification factors. Consequently, no particular clinical variables are attributed to the stratification scheme here.
16. What the Hazard Ratios Do — and Do Not — Mean
The global overall-population HR of 0.60 indicates a 40% lower estimated instantaneous hazard of progression or death in the capivasertib comparison relative to placebo plus fulvestrant.
An HR of 0.60 does not mean that 40% of patients avoided progression, that 40% more patients were alive without progression at a particular time, or that every patient experienced the same proportional reduction.
The corresponding 95% CI of 0.51–0.71 quantifies uncertainty around the estimated relative hazard under the reported statistical framework.
The P < 0.001 result is evidence from the specified statistical test against its null hypothesis. It is not a measure of treatment magnitude and should not be used as a substitute for the hazard ratio or confidence interval.
17. Comparing the Reported Populations
The registry reports results for both global and China cohorts and distinguishes overall from altered populations. These estimates can be described side by side, but they should not be treated as direct evidence of effect modification without a formal comparison.
| Population | HR | 95% CI | P-value | Descriptive interpretation |
|---|---|---|---|---|
| Global · Overall | 0.60 | 0.51–0.71 | <0.001 | Estimated hazard was 0.60 times the comparator hazard. |
| Global · Altered | 0.50 | 0.38–0.65 | <0.001 | Estimated hazard was 0.50 times the comparator hazard. |
| China · Overall | 0.51 | 0.34–0.76 | <0.001 | Estimated hazard was 0.51 times the comparator hazard. |
| China · Altered | 0.41 | 0.19–0.85 | 0.016 | Estimated hazard was 0.41 times the comparator hazard. |
18. Limitations
- Registry-only numerical scope: this page is restricted to the ClinicalTrials.gov record. Additional numerical results from publications are not substituted into missing fields.
- No median PFS reported in the ClinicalTrials.gov record: the ClinicalTrials.gov record provides hazard ratios, confidence intervals, and P-values but do not provide median PFS estimates.
- No time-specific PFS percentages: although some endpoints are registered as percentage outcomes, the statistical analyses posted on ClinicalTrials.gov report hazard ratios rather than specific Kaplan-Meier percentages at named time points.
- No event counts for efficacy analyses: the ClinicalTrials.gov record does not provide the number of progression or death events underlying each hazard ratio.
- Population definitions: the global and China analyses use distinct analysis populations, and the registry notes that 24 participants overlap between those cohorts.
- Altered-population interpretation: the ClinicalTrials.gov record does not provide the detailed eligibility or numerical composition of the altered population beyond its registry label.
- Hazard-ratio interpretation: a hazard ratio is a relative time-to-event measure and should not be treated as an absolute probability or risk difference.
- Proportional-hazards caution: a single hazard ratio provides a compact summary of relative event rates. If the relative hazard changes materially over time, that single summary may not fully describe the treatment-effect pattern.
- Multiplicity: the ClinicalTrials.gov record identifies eight registered primary endpoints and eight posted statistical analyses, but do not provide an alpha-allocation or multiplicity-adjustment strategy. No such strategy is inferred here.
- Interim analysis: the ClinicalTrials.gov record does not report an interim-analysis plan or alpha-spending procedure, so none is described.
- Missing-data methods: the ClinicalTrials.gov record does not identify a formal imputation strategy for missing efficacy observations. No imputation method is inferred.
- Non-inferiority: this is a superiority hypothesis, and the ClinicalTrials.gov record provides no non-inferiority margin. Non-inferiority logic is therefore not applicable to the reported primary hypothesis.
19. Why This Trial Matters Statistically
CAPItello-291 is a useful teaching case because its registry record brings together randomized treatment comparison, masking, time-to-event endpoints, stratified survival testing, hazard ratios, confidence intervals, intention-to-treat analysis, and multiple analysis populations.
| Concept | How it appears in CAPItello-291 |
|---|---|
| Randomization | The trial is registered as randomized with a parallel design and 2 arms. |
| Blinding | The registry specifies quadruple masking. |
| Time-to-event endpoint | All eight registered primary endpoints are progression-free survival measures. |
| Stratified log-rank test | Reported formal statistical method for the posted primary analyses. |
| Hazard ratio | Reported effect measure for every primary analysis. |
| Confidence interval | Every posted primary analysis includes a two-sided 95% confidence interval. |
| P-value | Each primary analysis includes a reported P-value. |
| Intention-to-treat | Explicitly identified in the China-cohort primary analyses. |
| Kaplan-Meier estimation | Explicitly identified for the altered-population percentage analyses. |
| Superiority | Registered hypothesis type for the primary analyses. |
| Cohort overlap | The registry notes that 24 participants were included in both the global and China cohorts. |
| Safety analysis | Serious adverse events are reported by arm for global and China cohorts. |
20. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The reported randomized comparisons produced hazard ratios below 1 for all four distinct population-level analyses, with two-sided 95% confidence intervals and P-values reported under stratified log-rank testing.
Clinical interpretation
The registry data support describing the observed progression-free survival comparisons in relative terms. The ClinicalTrials.gov record does not provide enough numerical information to characterize median PFS, absolute PFS differences, or the duration of benefit.
21. A Note on the Eight Primary Endpoint Entries
At first glance, the registry's eight primary endpoints may appear to represent eight separate efficacy questions. The statistical analyses posted on ClinicalTrials.gov show a more structured pattern: four population-level comparisons are each registered twice, once as a month-based outcome and once as a percentage outcome.
| Population | Month endpoint | Percentage endpoint | Same reported analysis? |
|---|---|---|---|
| Global · Overall | Yes | Yes | Yes — HR 0.60; 95% CI 0.51–0.71; P < 0.001 |
| Global · Altered | Yes | Yes | Yes — HR 0.50; 95% CI 0.38–0.65; P < 0.001 |
| China · Overall | Yes | Yes | Yes — HR 0.51; 95% CI 0.34–0.76; P < 0.001 |
| China · Altered | Yes | Yes | Yes — HR 0.41; 95% CI 0.19–0.85; P = 0.016 |
This distinction matters statistically because counting every registry entry as an independent hypothesis test would misrepresent the information reported by the actual statistical analyses.
22. Related Statistical Concepts
Learn more about the methods used in this trial:
23. Related Statistical Calculators
Explore calculators that reinforce the quantitative concepts behind this trial:
24. Sources
- ClinicalTrials.gov: NCT04305496 — CAPItello-291.
- Linked publication: PubMed record for PMID 39283299.
- Linked publication: PubMed record for PMID 39214106.
- Linked publication: PubMed record for PMID 39159418.
Continue through Clinical Biostats
Use the statistical concepts in this trial as a starting point for deeper study of survival analysis, clinical-trial methods, and statistical calculation.
25. Record Summary
CAPItello-291 provides a clear example of randomized time-to-event analysis. The registry describes a phase 3, randomized, parallel, quadruple-masked trial with two treatment arms and eight registered primary endpoints centered on progression-free survival. The posted analyses use stratified log-rank testing and hazard ratios under a superiority framework, with two-sided 95% confidence intervals and reported P-values. The four distinct population-level comparisons have hazard ratios of 0.60, 0.50, 0.51, and 0.41, respectively, with the associated uncertainty and P-values reported above.
The most important statistical lesson is that these numbers must be interpreted as time-to-event treatment-effect estimates, not as absolute probabilities. Confidence intervals describe uncertainty around the estimated relative effects, while P-values address evidence against the corresponding null hypothesis rather than the magnitude of benefit. The global and China cohorts, as well as the overall and altered populations, should be kept analytically distinct, and the registry's 24-participant overlap between the global and China cohorts prevents simple addition of those cohort counts.