← Clinical Trials
Metastatic Breast Cancer Phase 3 Time-to-Event Analysis NCT01633060

BELLE-3: Complete Statistical Analysis of BKM120 With Fulvestrant in Metastatic Breast Cancer

An independent statistical review of the randomized phase 3 BELLE-3 trial evaluating BKM120 with fulvestrant versus placebo with fulvestrant in patients with HR+, HER2-, AI-treated, locally advanced or metastatic breast cancer who progressed on or after mTORi.

Trial start: 03Oct2012  ·  Primary completion: 23May2016  ·  Status: Terminated
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results on this page are restricted to the ClinicalTrials.gov record for BELLE-3. Where the ClinicalTrials.gov record does not provide an analysis or estimate, no additional result is inferred.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

BELLE-3 was a randomized, parallel-group, quadruple-masked phase 3 treatment study with 432 enrolled participants. The registered primary endpoint was progression-free survival based on local investigator assessment in the Full Analysis Set, with the primary comparison using a stratified log-rank test.

432
Enrollment
ClinicalTrials.gov record
2
Arms
Parallel design
0.67
PFS HR
95% one-sided CI 0.53–
<0.001
P-value
Primary PFS analysis
FeatureBELLE-3
Trial nameBELLE-3
PhasePhase 3
ConditionMetastatic Breast Cancer
PopulationPatients with HR+, HER2-, AI-treated, locally advanced or metastatic breast cancer who progressed on or after mTORi
DesignRandomized, parallel-group
MaskingQuadruple
Primary purposeTreatment
Enrollment432
Primary endpoint typeTime-to-event
Primary statistical methodLog-rank test; the posted analysis specifies a stratified log-rank test
Effect measureHazard ratio
ClinicalTrials.govNCT01633060
Lead sponsorNovartis Pharmaceuticals
Sponsor typeIndustry

2. Clinical Question

The registered clinical question was whether adding BKM120 to fulvestrant, compared with placebo added to fulvestrant, affected progression-free survival in patients with HR+, HER2-, AI-treated, locally advanced or metastatic breast cancer who had progressed on or after mTORi.

Population

Patients with HR+, HER2-, AI-treated, locally advanced or metastatic breast cancer who progressed on or after mTORi.

Intervention

BKM120 with fulvestrant.

Comparator

BKM120 matching placebo with fulvestrant.

Primary question

How did progression-free survival compare between BKM120 100mg plus fulvestrant and placebo plus fulvestrant?

3. Trial Design

01
Enroll432 participants
02
Randomize2 parallel arms
03
TreatBKM120 or matching placebo with fulvestrant
04
AssessPFS every 6 weeks
05
AnalyzeStratified log-rank test

The registry describes BELLE-3 as randomized, with a parallel design and quadruple masking. The primary purpose was treatment. These design features are important because the treatment comparison is anchored to randomized assignment rather than to a comparison of patients selected after treatment.

ARM A

BKM120 + Fulvestrant

  • BKM120 100mg
  • Fulvestrant
ARM B

Placebo + Fulvestrant

  • BKM120 matching placebo
  • Fulvestrant
Masking matters statistically. The registry classifies the trial as quadruple-masked. The ClinicalTrials.gov record does not specify the four masked parties, so this page does not assign particular roles to the masking designation. The important methodological point is that masking can reduce opportunities for knowledge of treatment assignment to influence trial conduct and assessment.

4. Trial Timeline and Registry Status

03Oct2012

Trial start

The registered study start date was 03Oct2012.

23May2016

Primary efficacy cutoff

The primary efficacy analysis was completed by 23May2016, identified in the registry caveats as the primary PFS analysis cutoff date.

08Sep2017

Final safety analysis

The study was later terminated, and the final safety analysis was conducted up to 08Sep2017.

21Sep2017

Survival follow-up

One CRF was collecting on 21Sep2017 for survival follow-up, identified in the registry caveat as LPLV.

Why the dates matter: the primary PFS analysis and the later safety/follow-up activities occurred at different times. A statistical analysis should therefore keep the primary efficacy cutoff separate from the later safety analysis rather than treating all registry information as if it came from one simultaneous database lock.

5. Primary Endpoint

EndpointRegistry definition / time frameAnalysis
Progression Free Survival (PFS) Based on Local Investigator Assessment - Full Analysis Set (FAS) Every 6 weeks after randomization up to a maximum of 4 years. PFS is defined as the time from date of randomization to the date of first radiologically documented progression or death due to any cause. Stratified log-rank test at one-sided 2.5% level of significance; hazard ratio reported.

The registry definition further states that if a patient did not progress or die at the time of the analysis data cut-off or start of new antineoplastic therapy, PFS was censored at the date of the last adequate tumor assessment before the earliest of the cut-off date or the relevant censoring condition described in the registry definition.

6. Statistical Methodology

Full Analysis Set

The primary analysis population was the Full Analysis Set (FAS) based on the Primary Analysis. The statistical analysis record also identifies intention-to-treat analysis as a concept associated with the analysis text.

The distinction is important. A randomized clinical trial generally obtains its strongest protection against allocation bias by maintaining the treatment comparison according to randomized assignment. A time-to-event endpoint can then use the information contributed by participants until progression, death, or censoring under the prespecified endpoint rules.

Stratified log-rank test

The registry reports a Log Rank analysis, while the analysis description specifies comparison of PFS between the two treatment groups using a stratified log-rank test at one-sided 2.5% level of significance.

A log-rank test compares the survival experience of two groups across observed event times. Rather than reducing the analysis to a single time point, it evaluates the ordering of events over follow-up. Stratification can be used when the analysis needs to account for prespecified strata while preserving the randomized comparison.

Hazard ratio

The treatment effect was summarized using a hazard ratio. The reported estimate compares the instantaneous event rate under the BKM120 plus fulvestrant strategy with that under placebo plus fulvestrant, within the framework of the time-to-event analysis.

Conceptual interpretation
HR = estimated hazard in BKM120 + fulvestrant ÷ estimated hazard in placebo + fulvestrant

An HR below 1 indicates a lower estimated instantaneous event rate in the BKM120 group relative to the comparator under the fitted time-to-event framework. It is not an absolute risk difference and does not directly state how many patients avoid progression.

Censoring

PFS is a time-to-event endpoint, so not every participant necessarily has an observed progression or death by the analysis cutoff. The registry definition explicitly incorporates censoring for patients who had not progressed or died under the specified circumstances. This allows partial follow-up information to contribute without treating an unobserved event as though it had occurred.

One-sided significance level

The registry analysis specifies a one-sided 2.5% level of significance. That is an important part of the statistical specification. It should not be silently converted into a conventional two-sided testing framework when interpreting the posted result.

7. Primary PFS Result

The registry reports one formal statistical analysis for the primary endpoint, Progression Free Survival based on local investigator assessment in the Full Analysis Set.

Hazard ratio for progression or death

0.67

95% one-sided CI: 0.53–

P < 0.001

Comparison: BKM120 100mg + Fulvestrant vs Placebo + Fulvestrant

Primary endpointBKM120 + Fulvestrant vs Placebo + Fulvestrant
EndpointProgression Free Survival (PFS) Based on Local Investigator Assessment - Full Analysis Set (FAS)
Analysis populationFull Analysis Set (FAS) based on Primary Analysis
MethodStratified log-rank test
Effect measureHazard ratio
Estimate0.67
Confidence interval95% one-sided CI: 0.53–
P-value<0.001
Significance frameworkOne-sided 2.5% level of significance
Clinical Biostats interpretation

An HR of 0.67 means that the estimated instantaneous rate of progression or death in the BKM120 plus fulvestrant group was approximately 67% of the corresponding estimated rate in the placebo plus fulvestrant group under the reported time-to-event analysis. Expressed as a relative hazard comparison, this corresponds to a 33% lower estimated hazard.

That interpretation does not mean that 33% of patients were protected from progression or death, that individual patients experienced exactly a 33% reduction in risk, or that the median PFS differed by a particular number of months. None of those quantities is reported in the ClinicalTrials.gov record.

The reported confidence interval is one-sided and has a lower bound of 0.53 as reported in the registry. It therefore should not be read as though it were an ordinary two-sided interval extending from 0.53 to another numerical upper bound. The ClinicalTrials.gov record does not provide that upper bound.

The P-value <0.001 addresses the strength of evidence against the statistical null framework under the specified one-sided testing procedure. A P-value does not measure the magnitude of the treatment effect, and it should not be interpreted as the probability that the treatment effect is real.

Finally, the analysis is based on time-to-event data and censoring. A hazard ratio is most naturally interpreted as a model-based relative event-rate measure over follow-up; it is not interchangeable with a risk ratio or an absolute difference in the probability of progression by a particular date. The ClinicalTrials.gov record also do not state a proportional-hazards assumption assessment, so no additional claim about that assumption is made here.

8. What the PFS Hazard Ratio Tells Us

The primary estimate is easiest to understand by separating three different questions: relative event rate, absolute event probability, and time until an event.

Relative event rate

HR 0.67 indicates a lower estimated instantaneous rate of progression or death in the BKM120 plus fulvestrant group within the reported analysis framework.

Absolute probability

The hazard ratio does not tell us the absolute probability that a particular patient will progress or die by a specified time.

Median PFS

The ClinicalTrials.gov record does not provide a median PFS for either treatment group, so no median-time comparison is presented.

Individual benefit

The HR is a population-level treatment comparison. It does not imply that every participant experienced the same proportional reduction in event hazard.

Why the estimate is not the whole result
Effect estimate + uncertainty interval + testing framework + analysis population

A useful statistical reading combines the HR of 0.67 with its one-sided confidence interval, the reported one-sided 2.5% testing level, the Full Analysis Set, the censoring rules, and the stratified log-rank method. Reporting only the HR would omit important information about how the estimate was obtained and how it should be interpreted.

9. Secondary Endpoints and Posted Analyses

The ClinicalTrials.gov record reports 11 outcome measures, but the ClinicalTrials.gov record contains only 1 statistical analysis, corresponding to the primary PFS endpoint.

Registry informationReported in the ClinicalTrials.gov record
Outcome measures posted11
Statistical analyses posted1
Primary-endpoint analyses1
Primary analyses with estimate + CI1
Formal secondary statistical comparisons reportedNone

Accordingly, this page does not manufacture secondary-endpoint estimates. For an endpoint of the same general time-to-event type, a Kaplan-Meier analysis with an appropriate between-group test and an effect estimate such as a hazard ratio would commonly be considered, but the ClinicalTrials.gov record does not establish that such a formal analysis was posted for any secondary endpoint in BELLE-3.

10. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm using affected patients over patients at risk.

Safety groupSerious adverse events affected / at risk
BKM120 100 mg + Fulvestrant74 / 288
Placebo + Fulvestrant26 / 140
All Patients100 / 428

These figures are descriptive safety counts. They should not be substituted for the primary efficacy analysis, and the ClinicalTrials.gov record does not provide a formal between-group statistical test for serious adverse events.

Denominator caution: the serious-adverse-event table uses an at-risk total of 428 patients, whereas the overall trial enrollment is 432. The ClinicalTrials.gov record does not explain the difference, so this page reports both figures exactly as provided rather than assuming why the safety denominator differs from enrollment.

11. Statistical Methods Explained

Why was a time-to-event endpoint used for PFS?

PFS records not only whether progression or death occurred, but also when the first qualifying event occurred. This matters because participants can have different lengths of follow-up. A time-to-event framework can incorporate those different observation periods through censoring rather than requiring every participant to have an observed event.

What does an HR of 0.67 mean?

Within the reported analysis framework, an HR of 0.67 means the estimated instantaneous rate of progression or death was 0.67 times that in the comparator group. The complementary interpretation is a 33% lower estimated hazard. It is not a statement that 33% of patients avoided an event or that individual patients each had a 33% reduction in risk.

Why use a log-rank test?

The log-rank test is designed for comparing time-to-event distributions between groups. Rather than evaluating only one fixed follow-up time, it uses information across observed event times. BELLE-3's registry analysis specifically identifies a stratified log-rank comparison.

Why does stratification matter?

A stratified log-rank test evaluates the treatment comparison within the framework of prespecified strata rather than treating all event times as if they necessarily arose from one homogeneous stratum. The ClinicalTrials.gov record identifies the method as stratified but do not list the specific stratification factors, so none are inferred here.

What does the one-sided confidence interval mean?

The posted analysis gives a 95% one-sided confidence interval with lower bound 0.53. A one-sided interval provides a boundary in one direction rather than the two numerical endpoints ordinarily shown for a two-sided interval. The missing upper boundary should not be reconstructed from the ClinicalTrials.gov record.

Why doesn't P < 0.001 measure the size of the treatment effect?

The P-value describes how incompatible the observed data are with the null framework under the specified testing procedure. It depends on both the observed effect and the amount of information in the analysis. Effect size is instead communicated by the hazard ratio and its uncertainty. Thus, HR 0.67 and P < 0.001 answer different statistical questions.

Why does the Full Analysis Set matter?

The registry identifies the Full Analysis Set as the population for the primary analysis and associates intention-to-treat analysis with the statistical analysis text. Keeping the analysis population tied to the prespecified randomized framework helps preserve the comparability created by randomization. It also prevents a post-randomization selection rule from silently changing the treatment comparison.

12. Intention-to-Treat Analysis and Randomization

Randomization is the design feature that establishes the initial comparability of treatment groups in expectation. Once participants are randomized, analyzing them according to that assignment preserves that randomized contrast even when subsequent clinical experience differs.

The BELLE-3 statistical analysis explicitly identifies intention-to-treat analysis as a concept in the analysis text and identifies the Full Analysis Set as the primary analysis population. This is particularly relevant for PFS because participants can discontinue therapy, initiate other therapy, or become censored during follow-up.

ConceptRole in BELLE-3
RandomizationRegistry-designated allocation method
Parallel designTwo treatment groups followed in parallel
Full Analysis SetPrimary PFS analysis population
Intention-to-treat conceptIdentified in the posted statistical analysis text
Time-to-event endpointPrimary PFS outcome
Stratified log-rankFormal primary comparison

13. Censoring and Follow-Up

Censoring is central to interpreting the BELLE-3 PFS analysis. The registry defines PFS from randomization to the first radiologically documented progression or death due to any cause. Participants who have not experienced one of those events by the relevant analysis point can contribute follow-up information up to the specified censoring time.

This is different from treating a censored participant as though the event never occurred. Instead, the statistical method recognizes that the participant was observed for a particular period without an observed qualifying event. The Kaplan-Meier framework is designed to use that partial information.

Event

First radiologically documented progression or death due to any cause, according to the registered PFS definition.

Assessment schedule

Every 6 weeks after randomization, up to a maximum of 4 years.

Censoring

The registry specifies censoring at the date of the last adequate tumor assessment under the stated circumstances.

Analysis

Stratified log-rank testing compares the time-to-event experience between the randomized groups.

14. Why a Kaplan-Meier Framework Fits This Endpoint

The ClinicalTrials.gov record identifies PFS as a time-to-event endpoint and the statistical method as a log-rank test. These methods are naturally associated with Kaplan-Meier estimation because the analysis must accommodate variable follow-up and censoring.

Conceptual Kaplan-Meier estimator
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at an observed event time and ni represents the number at risk immediately before that time. The registry-reported BELLE-3 data do not contain the event-by-event information needed to reconstruct a Kaplan-Meier curve.

Educational note: this page does not draw a fabricated Kaplan-Meier curve from the reported hazard ratio and P-value. A valid reconstruction requires the underlying event and censoring information or sufficiently detailed digitized source data.

15. Confidence Intervals and Precision

The primary analysis reports an HR of 0.67 and a 95% one-sided confidence interval with lower bound 0.53. The confidence interval provides information about uncertainty around the estimated treatment effect under the analysis framework.

Precision is not the same as significance

A confidence interval and a P-value provide related but different information. The interval communicates uncertainty around the effect estimate, while the P-value addresses the specified hypothesis-testing framework. Neither one describes the range of treatment effects that individual patients will experience.

The registry-reported interval is one-sided, which is especially important here. Because only the lower bound is provided in the trial data, the page reports 0.53– rather than inventing an upper bound. The correct statistical response to incomplete reporting is to preserve the reported information rather than reconstruct a number that is not present.

16. Multiplicity, Interim Analysis, and Other Design Topics

The registry-reported BELLE-3 data do not provide a multiplicity strategy, endpoint hierarchy, interim-analysis procedure, alpha-spending method, or statistical power calculation. They also do not provide a non-inferiority margin, factorial design, crossover specification, missing-data imputation method, or Bayesian method.

Design topicWhat the ClinicalTrials.gov record establishes
MultiplicityNot specified in the ClinicalTrials.gov record
Interim analysisNot specified in the ClinicalTrials.gov record
Alpha spendingNot specified in the ClinicalTrials.gov record
Non-inferiority marginNot specified; the posted analysis does not identify a non-inferiority design
CrossoverNot specified in the ClinicalTrials.gov record
Factorial designNot present; the registry identifies a parallel design
Missing-data / imputation methodNot specified in the ClinicalTrials.gov record
Bayesian methodsNot specified in the ClinicalTrials.gov record
Stratification factorsNot specified in the ClinicalTrials.gov record

This distinction is methodologically important. The absence of a registry-reported detail is not evidence that the underlying protocol lacked the procedure; it means only that the information available for this page does not establish it.

17. Interpreting the One-Sided Testing Framework

The primary comparison was specified as a stratified log-rank test at one-sided 2.5% level of significance. This tells us how the formal test was framed, but the ClinicalTrials.gov record lists the hypothesis type as Other / not stated.

What is known

The analysis used a one-sided 2.5% significance level and produced P < 0.001.

What is not known

The ClinicalTrials.gov record does not state the formal null and alternative hypotheses in words.

Why this matters

A one-sided test assigns its rejection region to one direction, so the direction of the prespecified hypothesis matters.

What should not be inferred

The listed hypothesis type should not be silently converted into a superiority, equivalence, or non-inferiority label beyond what the registry explicitly states.

18. Safety and Efficacy Answer Different Questions

The primary PFS analysis and the serious-adverse-event counts represent different statistical tasks. The PFS analysis asks whether time to progression or death differs between randomized groups. The safety data describe the number of patients affected by serious adverse events among the reported at-risk populations.

Evidence typeBELLE-3 information reportedStatistical interpretation
EfficacyPFS HR 0.67; 95% one-sided CI lower bound 0.53; P < 0.001Formal time-to-event comparison
Safety74/288 vs 26/140 serious adverse eventsDescriptive affected/at-risk counts
Analysis populationFAS for primary PFS analysisPrimary efficacy analysis framework
Safety denominator428 across all patientsReported safety at-risk population

A statistically persuasive efficacy result does not itself quantify the frequency or severity of adverse events, just as a safety count does not provide evidence about the magnitude of the PFS treatment effect. Keeping these analyses separate is essential for clear clinical-trial interpretation.

19. Registry Status and Data Cutoffs

BELLE-3 is listed as TERMINATED. The registry caveat is particularly important because it distinguishes the primary efficacy analysis from later safety and survival follow-up activities.

Primary efficacy

23May2016

The primary PFS analysis cutoff date.

Later safety

08Sep2017

Final safety analysis conducted up to this date.

Survival follow-up

21Sep2017

One CRF was collecting for survival follow-up, identified as LPLV in the registry caveat.

These dates demonstrate why a clinical-trial record should be read as a sequence of analysis activities rather than as a single undifferentiated dataset. The PFS result belongs to the primary efficacy analysis cutoff, while the serious-adverse-event information is associated with a later safety analysis.

20. What This Analysis Does Not Establish

Methodological principle: missing information should remain missing. A complete statistical analysis is not one that fills every section with numbers; it is one that clearly distinguishes reported quantities, valid statistical interpretation, and information that the available record does not establish.

21. Important Limitations and Interpretation Issues

22. Why This Trial Matters Statistically

BELLE-3 is a useful teaching case because it concentrates several fundamental clinical-trial concepts in a relatively compact statistical record: randomization, masking, a time-to-event primary endpoint, censoring, a Full Analysis Set, a stratified log-rank test, a hazard ratio, a one-sided significance framework, and the separation of efficacy and safety populations.

ConceptHow it appears in BELLE-3
RandomizationRegistry-designated randomized allocation
Parallel designTwo treatment arms followed in parallel
BlindingQuadruple masking
Time-to-event endpointPrimary PFS endpoint
Kaplan-Meier frameworkNatural framework for the reported time-to-event endpoint and log-rank comparison
Log-rank testPrimary statistical comparison
StratificationPrimary analysis specifies a stratified log-rank test
Hazard ratioPrimary effect measure; estimate 0.67
Confidence interval95% one-sided CI with lower bound 0.53
P-value<0.001 under the reported testing framework
Intention-to-treat conceptIdentified in the statistical analysis text
Safety analysisSerious adverse events reported by affected/at-risk counts
Data-cutoff interpretationPrimary efficacy cutoff separated from later safety and survival follow-up

23. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue through the Clinical Biostats statistical pathway

Connect this trial's time-to-event endpoint and treatment-effect measure to deeper tutorials and statistical tools for clinical-trial analysis.

26. Record Summary

BELLE-3 is a randomized phase 3, parallel-group, quadruple-masked trial with 432 enrolled participants evaluating BKM120 100mg plus fulvestrant against placebo plus fulvestrant in patients with HR+, HER2-, AI-treated, locally advanced or metastatic breast cancer who progressed on or after mTORi. Its registered primary endpoint was progression-free survival based on local investigator assessment in the Full Analysis Set, assessed every 6 weeks after randomization up to a maximum of 4 years.

The posted primary analysis used a stratified log-rank test at a one-sided 2.5% significance level and reported a hazard ratio of 0.67, with a 95% one-sided confidence interval lower bound of 0.53 and P < 0.001. Statistically, the HR indicates a 33% lower estimated hazard of progression or death under the reported comparison. It does not provide an absolute risk difference, a median PFS, or an individual-level probability of benefit.

The registry record also reports serious adverse events of 74/288 for BKM120 100 mg plus fulvestrant and 26/140 for placebo plus fulvestrant, with 100/428 across all patients. The study's primary efficacy cutoff was 23May2016, while the later final safety analysis extended through 08Sep2017, with one CRF collecting on 21Sep2017 for survival follow-up.

Clinical Biostats methodology: The purpose of a trial-results page is not simply to repeat a registry record. It is to explain how the reported endpoint, analysis population, statistical test, effect measure, confidence interval, P-value, censoring framework, and safety information fit together—while preserving the distinction between documented results and statistical interpretation.