This page separates reported trial results from statistical interpretation. Numerical results on this page are restricted to the ClinicalTrials.gov record for BELLE-3. Where the ClinicalTrials.gov record does not provide an analysis or estimate, no additional result is inferred.
1. Trial at a Glance
BELLE-3 was a randomized, parallel-group, quadruple-masked phase 3 treatment study with 432 enrolled participants. The registered primary endpoint was progression-free survival based on local investigator assessment in the Full Analysis Set, with the primary comparison using a stratified log-rank test.
| Feature | BELLE-3 |
|---|---|
| Trial name | BELLE-3 |
| Phase | Phase 3 |
| Condition | Metastatic Breast Cancer |
| Population | Patients with HR+, HER2-, AI-treated, locally advanced or metastatic breast cancer who progressed on or after mTORi |
| Design | Randomized, parallel-group |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 432 |
| Primary endpoint type | Time-to-event |
| Primary statistical method | Log-rank test; the posted analysis specifies a stratified log-rank test |
| Effect measure | Hazard ratio |
| ClinicalTrials.gov | NCT01633060 |
| Lead sponsor | Novartis Pharmaceuticals |
| Sponsor type | Industry |
2. Clinical Question
The registered clinical question was whether adding BKM120 to fulvestrant, compared with placebo added to fulvestrant, affected progression-free survival in patients with HR+, HER2-, AI-treated, locally advanced or metastatic breast cancer who had progressed on or after mTORi.
Population
Patients with HR+, HER2-, AI-treated, locally advanced or metastatic breast cancer who progressed on or after mTORi.
Intervention
BKM120 with fulvestrant.
Comparator
BKM120 matching placebo with fulvestrant.
Primary question
How did progression-free survival compare between BKM120 100mg plus fulvestrant and placebo plus fulvestrant?
3. Trial Design
The registry describes BELLE-3 as randomized, with a parallel design and quadruple masking. The primary purpose was treatment. These design features are important because the treatment comparison is anchored to randomized assignment rather than to a comparison of patients selected after treatment.
BKM120 + Fulvestrant
- BKM120 100mg
- Fulvestrant
Placebo + Fulvestrant
- BKM120 matching placebo
- Fulvestrant
4. Trial Timeline and Registry Status
Trial start
The registered study start date was 03Oct2012.
Primary efficacy cutoff
The primary efficacy analysis was completed by 23May2016, identified in the registry caveats as the primary PFS analysis cutoff date.
Final safety analysis
The study was later terminated, and the final safety analysis was conducted up to 08Sep2017.
Survival follow-up
One CRF was collecting on 21Sep2017 for survival follow-up, identified in the registry caveat as LPLV.
5. Primary Endpoint
| Endpoint | Registry definition / time frame | Analysis |
|---|---|---|
| Progression Free Survival (PFS) Based on Local Investigator Assessment - Full Analysis Set (FAS) | Every 6 weeks after randomization up to a maximum of 4 years. PFS is defined as the time from date of randomization to the date of first radiologically documented progression or death due to any cause. | Stratified log-rank test at one-sided 2.5% level of significance; hazard ratio reported. |
The registry definition further states that if a patient did not progress or die at the time of the analysis data cut-off or start of new antineoplastic therapy, PFS was censored at the date of the last adequate tumor assessment before the earliest of the cut-off date or the relevant censoring condition described in the registry definition.
6. Statistical Methodology
Full Analysis Set
The primary analysis population was the Full Analysis Set (FAS) based on the Primary Analysis. The statistical analysis record also identifies intention-to-treat analysis as a concept associated with the analysis text.
The distinction is important. A randomized clinical trial generally obtains its strongest protection against allocation bias by maintaining the treatment comparison according to randomized assignment. A time-to-event endpoint can then use the information contributed by participants until progression, death, or censoring under the prespecified endpoint rules.
Stratified log-rank test
The registry reports a Log Rank analysis, while the analysis description specifies comparison of PFS between the two treatment groups using a stratified log-rank test at one-sided 2.5% level of significance.
A log-rank test compares the survival experience of two groups across observed event times. Rather than reducing the analysis to a single time point, it evaluates the ordering of events over follow-up. Stratification can be used when the analysis needs to account for prespecified strata while preserving the randomized comparison.
Hazard ratio
The treatment effect was summarized using a hazard ratio. The reported estimate compares the instantaneous event rate under the BKM120 plus fulvestrant strategy with that under placebo plus fulvestrant, within the framework of the time-to-event analysis.
An HR below 1 indicates a lower estimated instantaneous event rate in the BKM120 group relative to the comparator under the fitted time-to-event framework. It is not an absolute risk difference and does not directly state how many patients avoid progression.
Censoring
PFS is a time-to-event endpoint, so not every participant necessarily has an observed progression or death by the analysis cutoff. The registry definition explicitly incorporates censoring for patients who had not progressed or died under the specified circumstances. This allows partial follow-up information to contribute without treating an unobserved event as though it had occurred.
One-sided significance level
The registry analysis specifies a one-sided 2.5% level of significance. That is an important part of the statistical specification. It should not be silently converted into a conventional two-sided testing framework when interpreting the posted result.
7. Primary PFS Result
The registry reports one formal statistical analysis for the primary endpoint, Progression Free Survival based on local investigator assessment in the Full Analysis Set.
Hazard ratio for progression or death
95% one-sided CI: 0.53–
P < 0.001
Comparison: BKM120 100mg + Fulvestrant vs Placebo + Fulvestrant
| Primary endpoint | BKM120 + Fulvestrant vs Placebo + Fulvestrant |
|---|---|
| Endpoint | Progression Free Survival (PFS) Based on Local Investigator Assessment - Full Analysis Set (FAS) |
| Analysis population | Full Analysis Set (FAS) based on Primary Analysis |
| Method | Stratified log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.67 |
| Confidence interval | 95% one-sided CI: 0.53– |
| P-value | <0.001 |
| Significance framework | One-sided 2.5% level of significance |
An HR of 0.67 means that the estimated instantaneous rate of progression or death in the BKM120 plus fulvestrant group was approximately 67% of the corresponding estimated rate in the placebo plus fulvestrant group under the reported time-to-event analysis. Expressed as a relative hazard comparison, this corresponds to a 33% lower estimated hazard.
That interpretation does not mean that 33% of patients were protected from progression or death, that individual patients experienced exactly a 33% reduction in risk, or that the median PFS differed by a particular number of months. None of those quantities is reported in the ClinicalTrials.gov record.
The reported confidence interval is one-sided and has a lower bound of 0.53 as reported in the registry. It therefore should not be read as though it were an ordinary two-sided interval extending from 0.53 to another numerical upper bound. The ClinicalTrials.gov record does not provide that upper bound.
The P-value <0.001 addresses the strength of evidence against the statistical null framework under the specified one-sided testing procedure. A P-value does not measure the magnitude of the treatment effect, and it should not be interpreted as the probability that the treatment effect is real.
Finally, the analysis is based on time-to-event data and censoring. A hazard ratio is most naturally interpreted as a model-based relative event-rate measure over follow-up; it is not interchangeable with a risk ratio or an absolute difference in the probability of progression by a particular date. The ClinicalTrials.gov record also do not state a proportional-hazards assumption assessment, so no additional claim about that assumption is made here.
8. What the PFS Hazard Ratio Tells Us
The primary estimate is easiest to understand by separating three different questions: relative event rate, absolute event probability, and time until an event.
Relative event rate
HR 0.67 indicates a lower estimated instantaneous rate of progression or death in the BKM120 plus fulvestrant group within the reported analysis framework.
Absolute probability
The hazard ratio does not tell us the absolute probability that a particular patient will progress or die by a specified time.
Median PFS
The ClinicalTrials.gov record does not provide a median PFS for either treatment group, so no median-time comparison is presented.
Individual benefit
The HR is a population-level treatment comparison. It does not imply that every participant experienced the same proportional reduction in event hazard.
A useful statistical reading combines the HR of 0.67 with its one-sided confidence interval, the reported one-sided 2.5% testing level, the Full Analysis Set, the censoring rules, and the stratified log-rank method. Reporting only the HR would omit important information about how the estimate was obtained and how it should be interpreted.
9. Secondary Endpoints and Posted Analyses
The ClinicalTrials.gov record reports 11 outcome measures, but the ClinicalTrials.gov record contains only 1 statistical analysis, corresponding to the primary PFS endpoint.
| Registry information | Reported in the ClinicalTrials.gov record |
|---|---|
| Outcome measures posted | 11 |
| Statistical analyses posted | 1 |
| Primary-endpoint analyses | 1 |
| Primary analyses with estimate + CI | 1 |
| Formal secondary statistical comparisons reported | None |
Accordingly, this page does not manufacture secondary-endpoint estimates. For an endpoint of the same general time-to-event type, a Kaplan-Meier analysis with an appropriate between-group test and an effect estimate such as a hazard ratio would commonly be considered, but the ClinicalTrials.gov record does not establish that such a formal analysis was posted for any secondary endpoint in BELLE-3.
10. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm using affected patients over patients at risk.
| Safety group | Serious adverse events affected / at risk |
|---|---|
| BKM120 100 mg + Fulvestrant | 74 / 288 |
| Placebo + Fulvestrant | 26 / 140 |
| All Patients | 100 / 428 |
These figures are descriptive safety counts. They should not be substituted for the primary efficacy analysis, and the ClinicalTrials.gov record does not provide a formal between-group statistical test for serious adverse events.
11. Statistical Methods Explained
Why was a time-to-event endpoint used for PFS?
PFS records not only whether progression or death occurred, but also when the first qualifying event occurred. This matters because participants can have different lengths of follow-up. A time-to-event framework can incorporate those different observation periods through censoring rather than requiring every participant to have an observed event.
What does an HR of 0.67 mean?
Within the reported analysis framework, an HR of 0.67 means the estimated instantaneous rate of progression or death was 0.67 times that in the comparator group. The complementary interpretation is a 33% lower estimated hazard. It is not a statement that 33% of patients avoided an event or that individual patients each had a 33% reduction in risk.
Why use a log-rank test?
The log-rank test is designed for comparing time-to-event distributions between groups. Rather than evaluating only one fixed follow-up time, it uses information across observed event times. BELLE-3's registry analysis specifically identifies a stratified log-rank comparison.
Why does stratification matter?
A stratified log-rank test evaluates the treatment comparison within the framework of prespecified strata rather than treating all event times as if they necessarily arose from one homogeneous stratum. The ClinicalTrials.gov record identifies the method as stratified but do not list the specific stratification factors, so none are inferred here.
What does the one-sided confidence interval mean?
The posted analysis gives a 95% one-sided confidence interval with lower bound 0.53. A one-sided interval provides a boundary in one direction rather than the two numerical endpoints ordinarily shown for a two-sided interval. The missing upper boundary should not be reconstructed from the ClinicalTrials.gov record.
Why doesn't P < 0.001 measure the size of the treatment effect?
The P-value describes how incompatible the observed data are with the null framework under the specified testing procedure. It depends on both the observed effect and the amount of information in the analysis. Effect size is instead communicated by the hazard ratio and its uncertainty. Thus, HR 0.67 and P < 0.001 answer different statistical questions.
Why does the Full Analysis Set matter?
The registry identifies the Full Analysis Set as the population for the primary analysis and associates intention-to-treat analysis with the statistical analysis text. Keeping the analysis population tied to the prespecified randomized framework helps preserve the comparability created by randomization. It also prevents a post-randomization selection rule from silently changing the treatment comparison.
12. Intention-to-Treat Analysis and Randomization
Randomization is the design feature that establishes the initial comparability of treatment groups in expectation. Once participants are randomized, analyzing them according to that assignment preserves that randomized contrast even when subsequent clinical experience differs.
The BELLE-3 statistical analysis explicitly identifies intention-to-treat analysis as a concept in the analysis text and identifies the Full Analysis Set as the primary analysis population. This is particularly relevant for PFS because participants can discontinue therapy, initiate other therapy, or become censored during follow-up.
| Concept | Role in BELLE-3 |
|---|---|
| Randomization | Registry-designated allocation method |
| Parallel design | Two treatment groups followed in parallel |
| Full Analysis Set | Primary PFS analysis population |
| Intention-to-treat concept | Identified in the posted statistical analysis text |
| Time-to-event endpoint | Primary PFS outcome |
| Stratified log-rank | Formal primary comparison |
13. Censoring and Follow-Up
Censoring is central to interpreting the BELLE-3 PFS analysis. The registry defines PFS from randomization to the first radiologically documented progression or death due to any cause. Participants who have not experienced one of those events by the relevant analysis point can contribute follow-up information up to the specified censoring time.
This is different from treating a censored participant as though the event never occurred. Instead, the statistical method recognizes that the participant was observed for a particular period without an observed qualifying event. The Kaplan-Meier framework is designed to use that partial information.
Event
First radiologically documented progression or death due to any cause, according to the registered PFS definition.
Assessment schedule
Every 6 weeks after randomization, up to a maximum of 4 years.
Censoring
The registry specifies censoring at the date of the last adequate tumor assessment under the stated circumstances.
Analysis
Stratified log-rank testing compares the time-to-event experience between the randomized groups.
14. Why a Kaplan-Meier Framework Fits This Endpoint
The ClinicalTrials.gov record identifies PFS as a time-to-event endpoint and the statistical method as a log-rank test. These methods are naturally associated with Kaplan-Meier estimation because the analysis must accommodate variable follow-up and censoring.
Here, di represents the number of events at an observed event time and ni represents the number at risk immediately before that time. The registry-reported BELLE-3 data do not contain the event-by-event information needed to reconstruct a Kaplan-Meier curve.
15. Confidence Intervals and Precision
The primary analysis reports an HR of 0.67 and a 95% one-sided confidence interval with lower bound 0.53. The confidence interval provides information about uncertainty around the estimated treatment effect under the analysis framework.
A confidence interval and a P-value provide related but different information. The interval communicates uncertainty around the effect estimate, while the P-value addresses the specified hypothesis-testing framework. Neither one describes the range of treatment effects that individual patients will experience.
The registry-reported interval is one-sided, which is especially important here. Because only the lower bound is provided in the trial data, the page reports 0.53– rather than inventing an upper bound. The correct statistical response to incomplete reporting is to preserve the reported information rather than reconstruct a number that is not present.
16. Multiplicity, Interim Analysis, and Other Design Topics
The registry-reported BELLE-3 data do not provide a multiplicity strategy, endpoint hierarchy, interim-analysis procedure, alpha-spending method, or statistical power calculation. They also do not provide a non-inferiority margin, factorial design, crossover specification, missing-data imputation method, or Bayesian method.
| Design topic | What the ClinicalTrials.gov record establishes |
|---|---|
| Multiplicity | Not specified in the ClinicalTrials.gov record |
| Interim analysis | Not specified in the ClinicalTrials.gov record |
| Alpha spending | Not specified in the ClinicalTrials.gov record |
| Non-inferiority margin | Not specified; the posted analysis does not identify a non-inferiority design |
| Crossover | Not specified in the ClinicalTrials.gov record |
| Factorial design | Not present; the registry identifies a parallel design |
| Missing-data / imputation method | Not specified in the ClinicalTrials.gov record |
| Bayesian methods | Not specified in the ClinicalTrials.gov record |
| Stratification factors | Not specified in the ClinicalTrials.gov record |
This distinction is methodologically important. The absence of a registry-reported detail is not evidence that the underlying protocol lacked the procedure; it means only that the information available for this page does not establish it.
17. Interpreting the One-Sided Testing Framework
The primary comparison was specified as a stratified log-rank test at one-sided 2.5% level of significance. This tells us how the formal test was framed, but the ClinicalTrials.gov record lists the hypothesis type as Other / not stated.
What is known
The analysis used a one-sided 2.5% significance level and produced P < 0.001.
What is not known
The ClinicalTrials.gov record does not state the formal null and alternative hypotheses in words.
Why this matters
A one-sided test assigns its rejection region to one direction, so the direction of the prespecified hypothesis matters.
What should not be inferred
The listed hypothesis type should not be silently converted into a superiority, equivalence, or non-inferiority label beyond what the registry explicitly states.
18. Safety and Efficacy Answer Different Questions
The primary PFS analysis and the serious-adverse-event counts represent different statistical tasks. The PFS analysis asks whether time to progression or death differs between randomized groups. The safety data describe the number of patients affected by serious adverse events among the reported at-risk populations.
| Evidence type | BELLE-3 information reported | Statistical interpretation |
|---|---|---|
| Efficacy | PFS HR 0.67; 95% one-sided CI lower bound 0.53; P < 0.001 | Formal time-to-event comparison |
| Safety | 74/288 vs 26/140 serious adverse events | Descriptive affected/at-risk counts |
| Analysis population | FAS for primary PFS analysis | Primary efficacy analysis framework |
| Safety denominator | 428 across all patients | Reported safety at-risk population |
A statistically persuasive efficacy result does not itself quantify the frequency or severity of adverse events, just as a safety count does not provide evidence about the magnitude of the PFS treatment effect. Keeping these analyses separate is essential for clear clinical-trial interpretation.
19. Registry Status and Data Cutoffs
BELLE-3 is listed as TERMINATED. The registry caveat is particularly important because it distinguishes the primary efficacy analysis from later safety and survival follow-up activities.
23May2016
The primary PFS analysis cutoff date.
08Sep2017
Final safety analysis conducted up to this date.
21Sep2017
One CRF was collecting for survival follow-up, identified as LPLV in the registry caveat.
These dates demonstrate why a clinical-trial record should be read as a sequence of analysis activities rather than as a single undifferentiated dataset. The PFS result belongs to the primary efficacy analysis cutoff, while the serious-adverse-event information is associated with a later safety analysis.
20. What This Analysis Does Not Establish
- No median PFS: the ClinicalTrials.gov record does not provide median PFS for either treatment group.
- No absolute PFS rates: the ClinicalTrials.gov record does not provide PFS probabilities at specific follow-up times.
- No secondary estimates: although 11 outcome measures are posted, only one formal statistical analysis is reported.
- No subgroup estimates: the ClinicalTrials.gov record contains no subgroup-specific hazard ratios or confidence intervals.
- No baseline table: detailed baseline demographic or disease-characteristic values are not reported.
- No reconstructed survival curve: the reported HR and P-value are insufficient to recreate an authentic Kaplan-Meier curve.
- No formal safety comparison: serious adverse events are reported as affected/at-risk counts without a registry-reported inferential comparison.
- No multiplicity conclusion: the ClinicalTrials.gov record does not establish how multiple outcome measures were handled statistically.
21. Important Limitations and Interpretation Issues
- One-sided confidence interval: the registry-reported 95% CI is explicitly one-sided, and only its lower bound of 0.53 is provided. It should not be converted into a two-sided interval.
- Hazard ratio interpretation: HR 0.67 is a relative time-to-event measure, not an absolute risk reduction or a statement about the percentage of patients who benefit.
- Proportional hazards: the ClinicalTrials.gov record does not report an assessment of the proportional-hazards assumption. A single HR should therefore not be treated as a complete description of how treatment effects behave at every point in follow-up.
- Censoring: PFS interpretation depends on the registered censoring rules and the assumption that the resulting censoring mechanism is appropriate for the analysis.
- Analysis population: the primary result is tied to the Full Analysis Set. The analysis population should not be silently substituted with the overall enrollment count.
- Safety denominator: the serious-adverse-event analysis reports 428 patients at risk across all patients, compared with 432 enrolled. The ClinicalTrials.gov record does not explain this difference.
- Incomplete secondary reporting: 11 outcome measures are posted, but only one formal statistical analysis is reported in the ClinicalTrials.gov record used for this page.
- Study status: the trial is listed as terminated, and the primary efficacy cutoff differs from the later safety-analysis period.
22. Why This Trial Matters Statistically
BELLE-3 is a useful teaching case because it concentrates several fundamental clinical-trial concepts in a relatively compact statistical record: randomization, masking, a time-to-event primary endpoint, censoring, a Full Analysis Set, a stratified log-rank test, a hazard ratio, a one-sided significance framework, and the separation of efficacy and safety populations.
| Concept | How it appears in BELLE-3 |
|---|---|
| Randomization | Registry-designated randomized allocation |
| Parallel design | Two treatment arms followed in parallel |
| Blinding | Quadruple masking |
| Time-to-event endpoint | Primary PFS endpoint |
| Kaplan-Meier framework | Natural framework for the reported time-to-event endpoint and log-rank comparison |
| Log-rank test | Primary statistical comparison |
| Stratification | Primary analysis specifies a stratified log-rank test |
| Hazard ratio | Primary effect measure; estimate 0.67 |
| Confidence interval | 95% one-sided CI with lower bound 0.53 |
| P-value | <0.001 under the reported testing framework |
| Intention-to-treat concept | Identified in the statistical analysis text |
| Safety analysis | Serious adverse events reported by affected/at-risk counts |
| Data-cutoff interpretation | Primary efficacy cutoff separated from later safety and survival follow-up |
23. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Sources
- ClinicalTrials.gov: BELLE-3, NCT01633060.
- PubMed: Publication associated with BELLE-3, PMID 29223745.
Continue through the Clinical Biostats statistical pathway
Connect this trial's time-to-event endpoint and treatment-effect measure to deeper tutorials and statistical tools for clinical-trial analysis.
26. Record Summary
BELLE-3 is a randomized phase 3, parallel-group, quadruple-masked trial with 432 enrolled participants evaluating BKM120 100mg plus fulvestrant against placebo plus fulvestrant in patients with HR+, HER2-, AI-treated, locally advanced or metastatic breast cancer who progressed on or after mTORi. Its registered primary endpoint was progression-free survival based on local investigator assessment in the Full Analysis Set, assessed every 6 weeks after randomization up to a maximum of 4 years.
The posted primary analysis used a stratified log-rank test at a one-sided 2.5% significance level and reported a hazard ratio of 0.67, with a 95% one-sided confidence interval lower bound of 0.53 and P < 0.001. Statistically, the HR indicates a 33% lower estimated hazard of progression or death under the reported comparison. It does not provide an absolute risk difference, a median PFS, or an individual-level probability of benefit.
The registry record also reports serious adverse events of 74/288 for BKM120 100 mg plus fulvestrant and 26/140 for placebo plus fulvestrant, with 100/428 across all patients. The study's primary efficacy cutoff was 23May2016, while the later final safety analysis extended through 08Sep2017, with one CRF collecting on 21Sep2017 for survival follow-up.