← Clinical Trials
Breast Cancer Phase 3 Supportive Care NCT00556374

ABCSG-18: Complete Statistical Analysis of Denosumab in Breast Cancer

An independent statistical review of the randomized phase 3 ABCSG-18 trial evaluating denosumab in patients with breast cancer receiving aromatase inhibitor therapy, with emphasis on clinical fractures, bone mineral density, vertebral fractures, disease-free survival, bone metastases-free survival, overall survival, and the statistical methods used to analyze them.

Trial status: Completed  ·  Enrollment: 3420  ·  Primary completion: 07 October 2014
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the statistical analyses posted for ABCSG-18 in the ClinicalTrials.gov record. Where the registry does not report a formal statistical method or p-value, that distinction is preserved.

Registry context: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

ABCSG-18 was a randomized, parallel-group phase 3 trial evaluating denosumab versus placebo in patients with breast cancer receiving non-steroidal aromatase inhibitor therapy. The registered primary endpoint was time to first clinical fracture.

3420
Enrolled
Phase 3 trial
4
Registered arms
Parallel design
0.504
Primary HR
95% CI 0.39–0.65
<0.0001
Primary P-value
Superiority analysis
FeatureABCSG-18
PhasePhase 3
ConditionBreast Cancer
Brief titleStudy to Determine Treatment Effects of Denosumab in Patients With Breast Cancer Receiving Aromatase Inhibitor Therapy
DesignRandomized, parallel-group
MaskingQuadruple
Primary purposeSupportive care
Enrollment3420
Primary endpointTime to First Clinical Fracture
Primary endpoint typeTime-to-event
Lead sponsorAmgen
ClinicalTrials.govNCT00556374

2. Clinical Question

The primary statistical question was whether denosumab, administered subcutaneously at a dose of 60 mg every 6 months, would reduce the rate of first clinical fracture compared with placebo in patients with non-metastatic breast cancer receiving aromatase inhibitor therapy.

Population

Patients with breast cancer receiving non-steroidal aromatase inhibitor therapy, as represented in the randomized trial.

Intervention

Denosumab, a biological intervention, administered subcutaneously at 60 mg every 6 months according to the registered efficacy hypothesis.

Comparator

Placebo.

Primary question

Does denosumab reduce the rate of first on-study clinical fracture relative to placebo?

3. Trial Design

01
Randomize3420 participants
02
Parallel groupsDenosumab vs placebo efficacy comparison
03
Blinded treatmentQuadruple masking
04
Follow-upClinical fractures and bone outcomes
05
Time-to-eventPrimary fracture analysis
EFFICACY COMPARISON · PLACEBO

Placebo

  • Placebo comparator
  • Compared with denosumab for the primary clinical-fracture endpoint
  • Also used as the reference group for the registered BMD and vertebral-fracture analyses
EFFICACY COMPARISON · DENOSUMAB

Denosumab

  • Biological intervention
  • 60 mg every 6 months in the registered efficacy hypothesis
  • Compared with placebo for clinical fractures, BMD, vertebral fractures, DFS, BMFS, and OS

The registry profile lists 4 arms and several interventions, including placebo, denosumab, non-steroidal aromatase inhibitor therapy, zoledronic acid, and standard of care. The posted statistical analyses in the ClinicalTrials.gov record compare placebo versus denosumab.

4. Trial Timing and Registry Structure

Start
18 December 2006
Primary completion
07 October 2014
Primary analysis cutoff
26 March 2014
DFS data cutoff
15 September 2015

The primary fracture analysis and later secondary endpoints use different analysis windows. That distinction is statistically important: estimates from different data cutoffs should not be treated as though they arose from the same information set.

5. Primary Endpoint

EndpointRegistered definition and time frameAnalysis
Time to First Clinical Fracture From randomization until the primary analysis cut-off date of 26 March 2014; maximum time on main study at the cut-off was as registered. Time to first on-study clinical fracture was defined as the number of days from randomization to the date of the x-ray confirming the clinical fracture. A clinical fracture is any clinically evident fracture with associated symptoms and confirmed by x-ray. Participants who died or withdrew without experiencing a clinical fracture were censored at the date of last contact or study termination as registered. Cox proportional-hazards model

The primary analysis used the Full Analysis Set: all randomized participants. The treatment comparison was placebo versus denosumab, with intention-to-treat and stratified-analysis concepts identified in the registry analysis text.

6. Primary Result: Time to First Clinical Fracture

Hazard ratio for first clinical fracture

0.504

95% CI: 0.39–0.65   ·   P < 0.0001

Two-sided 95% confidence interval · Full Analysis Set · Superiority hypothesis

The registry reports a hazard ratio of 0.504 for denosumab versus placebo. Because the hazard ratio is below 1, the estimated average event rate over the analyzed follow-up was lower in the denosumab group under the fitted Cox model. A simple descriptive translation is that 0.504 corresponds to an estimated hazard approximately 49.6% lower than the placebo hazard.

Clinical Biostats interpretation

What the estimate means: The HR of 0.504 is a relative time-to-event measure from a Cox proportional-hazards model. It compares the modeled hazard of experiencing a first clinical fracture between the randomized treatment groups.

What it does not mean: It does not mean that 49.6% of patients avoided a fracture, that each individual had exactly a 49.6% reduction in probability, or that the absolute difference in fracture risk was 49.6 percentage points.

What the confidence interval says: The two-sided 95% CI of 0.39–0.65 quantifies uncertainty around the estimated hazard ratio under the model and sampling framework. It does not describe the range of effects for individual patients.

Why the p-value is different: The p-value of <0.0001 addresses evidence against the null hypothesis in the prespecified statistical framework. It does not measure the size of the treatment effect; the HR and its confidence interval describe magnitude and precision.

Important model caution: Interpretation of a Cox hazard ratio depends on the proportional-hazards framework. The registry identifies a Cox proportional-hazards model, but the ClinicalTrials.gov record does not provide a separate assessment of that assumption.

Primary analysis structure
HR = 0.504   |   95% CI = 0.39–0.65   |   P < 0.0001

The comparison is placebo versus denosumab, analyzed in all randomized participants using a Cox proportional-hazards model. The registry analysis text identifies intention-to-treat and stratified analysis concepts.

7. Secondary Endpoint Results: Bone Mineral Density

Three secondary endpoints evaluated percent change from baseline in bone mineral density at Month 36 at pre-selected sites. All three analyses used ANCOVA and adjusted for baseline value and the randomization stratification factors.

EndpointEstimate: Difference from Placebo95% CIP-value
Total lumbar spine BMD, Month 3610.029.04–11.01<0.0001
Total hip BMD, Month 367.926.87–8.97<0.0001
Femoral neck BMD, Month 366.515.62–7.39<0.0001

Total Lumbar Spine BMD

Adjusted difference from placebo

10.02

95% CI: 9.04–11.01   ·   P < 0.0001

Total Hip BMD

Adjusted difference from placebo

7.92

95% CI: 6.87–8.97   ·   P < 0.0001

Femoral Neck BMD

Adjusted difference from placebo

6.51

95% CI: 5.62–7.39   ·   P < 0.0001

These are adjusted differences in percent change from baseline, not hazard ratios or odds ratios. The estimates therefore answer a different statistical question from the primary fracture analysis.

Why ANCOVA is appropriate here: The registered BMD analyses compare a continuous change-from-baseline outcome at Month 36. The model included treatment group as the independent variable and adjusted for baseline BMD and randomization stratification factors. Baseline adjustment can improve precision by accounting for prognostic variation that existed before treatment.

8. Secondary Endpoint Results: Vertebral Fractures

The registry reports two binary vertebral-fracture outcomes at 36 months. Both were analyzed with logistic regression in the Vertebral Fracture Analysis Set, defined as participants with a baseline and at least one post-baseline vertebral x-ray prior to or at Month 36.

EndpointOdds ratio95% CIP-value
Participants with new vertebral fractures0.530.33–0.850.0088
Participants with new or worsening vertebral fractures0.540.34–0.840.0070
How to interpret the odds ratios

An odds ratio of 0.53 compares the odds of a new vertebral fracture between the denosumab and placebo groups in the analyzed population. An odds ratio below 1 indicates lower estimated odds in the denosumab group.

Likewise, an odds ratio of 0.54 indicates lower estimated odds for the new-or-worsening vertebral-fracture endpoint. Neither odds ratio is itself a risk ratio or an absolute probability difference.

The confidence intervals describe uncertainty around the estimated odds ratios. The p-values address statistical evidence under the corresponding superiority tests; they do not quantify the magnitude of the effect.

Population matters: These logistic-regression analyses did not use the same analysis-set definition as the primary clinical-fracture analysis. The primary analysis used all randomized participants, whereas the vertebral-fracture analyses used participants with the required radiographic assessments. Comparisons across endpoints should therefore account for these different analysis populations.

9. Secondary Endpoint Results: Disease-Free Survival

Disease-free survival was analyzed from randomization until the DFS data cut-off date of 15 September 2015. The Full Analysis Set included all randomized participants.

Disease-free survival hazard ratio

0.816

95% CI: 0.66–1.00   ·   P = 0.0515

Two-sided 95% confidence interval · Cox proportional-hazards model

The registry states that the analysis was stratified by the randomization strata: hospital type, use of aromatase inhibitor, and baseline lumbar spine BMD. Treatment was fitted as a covariate in the stratified Cox model.

Clinical Biostats interpretation

Estimate: The HR of 0.816 indicates a lower estimated average event rate and longer disease-free time for denosumab relative to placebo under the reported Cox model.

Precision: The 95% CI of 0.66–1.00 is substantially wider relative to the primary fracture estimate and reaches 1.00 at its upper boundary. This indicates greater uncertainty about the precise magnitude of the relative effect.

P-value: The reported P = 0.0515 is close to, but above, the conventional 0.05 threshold. More importantly, a p-value should not be treated as a continuous measure of effect size. The HR and confidence interval provide the effect estimate and its precision.

Multiplicity: The registry describes later BMFS and OS analyses as conditional on the DFS outcome because of a hierarchical testing strategy. That structure must be considered when interpreting the later endpoints.

10. Secondary Endpoint Results: Bone Metastases-free Survival

Bone Metastases-free Survival (BMFS) was evaluated from randomization until the end of the main study, with a maximum time on the main study of 152 months. The analysis population was the Full Analysis Set: all randomized participants.

Bone metastases-free survival hazard ratio

0.808

95% CI: 0.654–0.997

Two-sided 95% confidence interval · P-value not reported

The registry does not report a formal statistical method for BMFS in the ClinicalTrials.gov record. It reports the hazard ratio and confidence interval but states that the analysis was conditional on the outcome of DFS because of a hierarchical testing strategy; therefore p-value not reported.

Clinical Biostats interpretation

The HR of 0.808 indicates a lower estimated average event rate and longer bone metastases-free time for denosumab relative to placebo under a hazard-ratio interpretation. The 95% CI of 0.654–0.997 describes uncertainty around that estimate.

The absence of a reported p-value is not a missing number that should be reconstructed from the confidence interval. The registry explicitly states that the BMFS analysis was conditional on DFS within a hierarchical testing strategy and therefore did not report a p-value.

This is an important example of why a confidence interval and a p-value are not interchangeable pieces of output. The registry provides an effect estimate and interval while deliberately withholding a formal p-value for this endpoint.

11. Secondary Endpoint Results: Overall Survival

Overall survival was evaluated from randomization until the end of the main study, with a maximum duration of the main study of 152 months. The analysis population was the Full Analysis Set: all randomized participants.

Overall survival hazard ratio

0.802

95% CI: 0.635–1.013

Two-sided 95% confidence interval · P-value not reported

The registry does not report a statistical method for this OS analysis in the ClinicalTrials.gov record. It reports the hazard ratio and its 95% confidence interval, and states that OS was analyzed conditionally on DFS because of the hierarchical testing strategy; therefore no p-value was reported.

Clinical Biostats interpretation

The HR of 0.802 corresponds to an estimated hazard about 19.8% lower for denosumab relative to placebo under a hazard-ratio interpretation. It does not mean that 19.8% more participants survived, nor does it describe an absolute survival difference.

The 95% CI of 0.635–1.013 expresses uncertainty around the estimated relative hazard. Because the interval extends above 1.0, the ClinicalTrials.gov record does not establish a precise relative hazard estimate below 1.0 at the 95% confidence level.

The registry does not report a p-value for this endpoint, and one should not be manufactured from the confidence interval. The hierarchical testing note is part of the statistical interpretation of the result.

12. Statistical Methodology

Cox proportional-hazards model

The primary endpoint and disease-free survival were analyzed using Cox proportional-hazards models. The Cox model estimates a relative hazard associated with treatment while allowing the baseline hazard to remain unspecified.

Conceptual Cox model
h(t | X) = h0(t) exp(βX)

For a binary treatment indicator, exp(β) represents the model-based hazard ratio comparing treatment with the reference group, conditional on the model structure.

For the primary fracture endpoint, the registry identifies the analysis population as all randomized participants and identifies intention-to-treat and stratified analysis concepts. For DFS, the registry explicitly states that the model was stratified by hospital type, aromatase-inhibitor use, and baseline lumbar-spine BMD.

Kaplan-Meier estimation

A time-to-event endpoint such as time to first clinical fracture is naturally represented through survival-function methods. Kaplan-Meier estimation is commonly used to describe the probability of remaining event-free over time while accounting for right censoring.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

Here di is the number of events at time ti and ni is the number at risk immediately before that time.

The registry-reported ABCSG-18 statistical-analysis records identify the primary method as a Cox proportional-hazards model rather than separately reporting Kaplan-Meier output. Kaplan-Meier estimation is nevertheless an important complementary way to understand the time-to-event framework.

ANCOVA

The three Month 36 BMD analyses used ANCOVA. The registry specifies that the model included treatment group as the independent variable and adjusted for the baseline value and the randomization stratification factors.

Why baseline adjustment matters
Adjusted Month 36 outcome = treatment effect + baseline adjustment + stratification adjustment + residual variation

The exact fitted model is more detailed than this schematic representation, but the important principle is that baseline BMD is incorporated rather than ignored.

Logistic regression

The two vertebral-fracture endpoints were binary outcomes and were analyzed using logistic regression. The models included treatment group as the independent variable and were stratified by the randomization stratification factors.

Odds-ratio interpretation
OR = odds in denosumab group ÷ odds in placebo group

An OR below 1 indicates lower estimated odds in the denosumab group. The OR should not automatically be described as a relative risk.

Intention-to-treat analysis

The primary fracture analysis was performed in the Full Analysis Set consisting of all randomized participants, and the analysis text identifies intention-to-treat analysis as a concept. An ITT framework preserves the randomized comparison rather than redefining groups according to subsequent treatment exposure.

Stratified analysis

Stratification accounts for prespecified randomization strata in the treatment-effect model. For DFS, the registry explicitly identifies hospital type, use of aromatase inhibitor, and baseline lumbar spine BMD as the randomization strata used in the Cox analysis.

13. Statistical Methods Explained

Why was ANCOVA used for BMD?

BMD at Month 36 is a quantitative outcome, making a linear-model framework appropriate. ANCOVA additionally adjusts for baseline BMD and the randomization stratification factors. This can improve precision because participants can differ in their baseline BMD even after randomization.

What does a hazard ratio of 0.504 mean?

It means that the fitted Cox model estimated a hazard in the denosumab group equal to 0.504 times the placebo hazard. Equivalently, 1 − 0.504 = 0.496, so the estimate corresponds to approximately a 49.6% lower hazard. It does not mean a 49.6% absolute reduction in the number of fractures.

Why is a hazard ratio different from an odds ratio?

The primary endpoint is a time-to-event outcome, so its effect is expressed as a hazard ratio. The vertebral-fracture outcomes are binary endpoints at 36 months, so their effects are expressed as odds ratios from logistic regression. A hazard ratio incorporates event timing, whereas an odds ratio compares odds of the binary outcome in the analyzed population.

Why does the DFS result need careful interpretation?

The DFS HR is 0.816 with a 95% CI of 0.66–1.00 and P = 0.0515. The estimate is below 1, but the interval reaches 1.00 and the p-value is above 0.05. The appropriate lesson is not to convert the p-value into an effect-size judgment: the HR and confidence interval should be read together.

Why was no p-value reported for BMFS and OS?

The registry explicitly states that both analyses were conditional on the DFS outcome because of a hierarchical testing strategy. It therefore reports the hazard ratios and confidence intervals but does not report p-values. This is a feature of the prespecified testing hierarchy, not an invitation to calculate a substitute p-value.

Why does the analysis population matter?

The primary fracture analysis uses all randomized participants, while the BMD and vertebral-fracture analyses use endpoint-specific analysis sets requiring evaluable measurements. An estimate from one analysis population should not be interpreted as though it came from another population simply because both compare denosumab with placebo.

Why should a confidence interval be read with the estimate?

The point estimate summarizes the fitted treatment effect, while the confidence interval communicates uncertainty around that estimate. For example, the primary HR of 0.504 has a 95% CI of 0.39–0.65, whereas the OS HR of 0.802 has a 95% CI of 0.635–1.013. These intervals provide important information that the point estimates alone cannot convey.

14. Multiplicity and Hierarchical Testing

The ClinicalTrials.gov record explicitly support a hierarchical testing strategy for later time-to-event outcomes. BMFS and OS were conditional on the outcome of DFS, and the registry states that p-values were therefore not reported for those analyses.

EndpointRoleReported statistical information
Time to First Clinical FracturePrimaryHR 0.504; 95% CI 0.39–0.65; P < 0.0001
DFSSecondaryHR 0.816; 95% CI 0.66–1.00; P = 0.0515
BMFSSecondaryHR 0.808; 95% CI 0.654–0.997; p-value not reported because analysis was conditional on DFS
OSSecondaryHR 0.802; 95% CI 0.635–1.013; p-value not reported because analysis was conditional on DFS
Why the hierarchy matters: Multiple endpoints create multiple opportunities for statistical significance unless the analysis plan controls how hypotheses are tested. ABCSG-18's registry record specifically documents a hierarchical relationship for BMFS and OS following DFS. The later endpoints therefore should not be interpreted as though each had an independent, unadjusted confirmatory p-value.

15. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment phase and arm. These counts are presented as affected participants over participants at risk.

Study phase / groupSerious adverse events affected / at risk
Double-blind Phase: Placebo515 / 1690
Double-blind Phase: Denosumab521 / 1709
Open-label Phase: Placebo/Denosumab40 / 245
Open-label Phase: Denosumab/Denosumab0 / 1

These are serious-adverse-event counts, not estimates of the primary fracture treatment effect. They should also not be combined across phases or groups without an analysis specification defining the appropriate denominator and exposure framework.

Safety interpretation: The double-blind phase reports 515 affected participants among 1690 at risk in the placebo group and 521 among 1709 in the denosumab group. The registry also separately reports open-label-phase groups. Because the ClinicalTrials.gov record does not provide a formal statistical comparison for these safety figures, the page reports the counts without constructing an unregistered hypothesis test.

16. Interpreting the Different Effect Measures

EndpointEffect measureEstimateWhat it compares
Time to First Clinical FractureHazard ratio0.504Modeled time-to-event hazard
Lumbar spine BMDDifference from placebo10.02Adjusted difference in percent change from baseline
Total hip BMDDifference from placebo7.92Adjusted difference in percent change from baseline
Femoral neck BMDDifference from placebo6.51Adjusted difference in percent change from baseline
New vertebral fracturesOdds ratio0.53Odds of a binary fracture outcome
New or worsening vertebral fracturesOdds ratio0.54Odds of a binary fracture outcome
DFSHazard ratio0.816Modeled time-to-event hazard
BMFSHazard ratio0.808Modeled time-to-event hazard
OSHazard ratio0.802Modeled time-to-event hazard

This distinction is central to reading the trial correctly. A hazard ratio, odds ratio, and adjusted continuous-outcome difference are not interchangeable. Each is tied to a different endpoint structure and therefore answers a different statistical question.

17. Censoring and Time-to-Event Analysis

The primary endpoint is explicitly defined as time from randomization to the x-ray-confirmed clinical fracture. Participants who died or withdrew without a clinical fracture were censored at the date of last contact or study termination as registered.

Censoring allows participants who have not experienced the event by their last observed follow-up to contribute the information available up to that point. This is one reason time-to-event methods are preferable to simply counting fractures without accounting for unequal follow-up.

Event information

A participant contributes an event time when the clinical fracture is confirmed by x-ray.

Censored information

A participant without an observed fracture contributes follow-up until the registered censoring point.

Model output

The Cox model summarizes the relative hazard between denosumab and placebo.

Interpretation caution

Censoring assumptions and the proportional-hazards framework affect how a Cox hazard ratio should be interpreted.

18. Primary Analysis vs Later Outcomes

One of the most useful statistical lessons from ABCSG-18 is that the trial's endpoints do not all represent the same analysis question.

FeaturePrimary fracture analysisLater secondary analyses
EndpointTime to First Clinical FractureBMD, vertebral fractures, DFS, BMFS, OS
Primary cutoff26 March 2014Endpoint-specific later time frames
Analysis populationAll randomized participantsEndpoint-specific populations for BMD and vertebral fractures; all randomized for DFS, BMFS, and OS
Statistical methodCox proportional-hazards modelANCOVA, logistic regression, Cox model, or not reported depending on endpoint
Multiplicity structurePrimary superiority hypothesisBMFS and OS explicitly conditional on DFS through hierarchical testing

Keeping these distinctions visible prevents a common analytical error: treating every reported number from a clinical trial as though it came from one homogeneous statistical analysis.

19. Why the Primary Result Is a Time-to-Event Result

The primary endpoint is not simply whether a participant ever fractured. It is when the first clinical fracture occurred, measured from randomization. That time dimension carries information that a binary endpoint would discard.

For example, two participants could both eventually experience a fracture but have very different follow-up times before the event. Conversely, a participant who has not fractured by the end of observation provides useful information through their event-free follow-up, even though their observation is censored rather than an event.

The Cox model is therefore well matched to the registered endpoint structure. Its hazard ratio provides a relative comparison of event rates over time while accounting for censoring under the model assumptions.

20. Why This Trial Matters Statistically

ABCSG-18 is a useful teaching case because it connects several core clinical-trial methods within one randomized study. The primary endpoint illustrates survival analysis, while the secondary outcomes demonstrate why endpoint type determines the statistical model.

ConceptHow it appears in ABCSG-18
Randomization3420 participants enrolled in a randomized phase 3 parallel-group study
BlindingQuadruple masking
Intention-to-treatPrimary analysis in all randomized participants with ITT identified in the analysis text
Time-to-event endpointTime to First Clinical Fracture
Cox modelPrimary fracture analysis and DFS
Hazard ratioPrimary fracture, DFS, BMFS, and OS effect measure
ANCOVAMonth 36 lumbar spine, total hip, and femoral neck BMD analyses
Covariate adjustmentBMD analyses adjusted for baseline value and randomization stratification factors
Logistic regressionNew and new-or-worsening vertebral-fracture analyses
Odds ratioEffect measure for the binary vertebral-fracture endpoints
Stratified analysisUsed in the fracture, BMD, vertebral-fracture, and DFS analysis descriptions
Confidence intervalsReported for every posted statistical analysis in the ClinicalTrials.gov record
Hierarchical testingBMFS and OS conditional on DFS, with p-values not reported

21. Important Limitations and Interpretation Issues

22. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The primary randomized comparison produced a Cox-model HR of 0.504 with a two-sided 95% CI of 0.39–0.65 and P < 0.0001. Secondary analyses produced adjusted BMD differences, vertebral-fracture odds ratios, and later time-to-event hazard ratios with endpoint-specific uncertainty.

Clinical interpretation

The ClinicalTrials.gov record describes differences in clinical fracture, bone mineral density, vertebral fractures, disease-free survival, bone metastases-free survival, and overall survival. Clinical interpretation should consider the endpoint definition, absolute clinical context, analysis population, follow-up period, and safety information rather than relying on a single statistical measure.

23. A Closer Look at the Confidence Intervals

The confidence intervals in ABCSG-18 illustrate why statistical interpretation should not stop at a point estimate.

EndpointEstimate95% CIInterpretive focus
Clinical fractureHR 0.5040.39–0.65Precision around the primary time-to-event effect
Lumbar spine BMD10.029.04–11.01Precision around adjusted BMD difference
Total hip BMD7.926.87–8.97Precision around adjusted BMD difference
Femoral neck BMD6.515.62–7.39Precision around adjusted BMD difference
New vertebral fractureOR 0.530.33–0.85Precision around binary-outcome odds ratio
New or worsening vertebral fractureOR 0.540.34–0.84Precision around binary-outcome odds ratio
DFSHR 0.8160.66–1.00Interval reaches 1.00
BMFSHR 0.8080.654–0.997Interval remains below 1.00
OSHR 0.8020.635–1.013Interval extends above 1.00

The widths and locations of these intervals also show why different endpoints should not be summarized simply as "positive" or "negative." The primary fracture estimate is relatively precise, while the later OS estimate has an interval that extends above 1.0.

24. Clinical Biostats Synthesis

The statistical story of ABCSG-18 is most clearly understood by following the endpoint hierarchy and the analysis method assigned to each type of data.

Primary endpoint
Clinical fracture → Cox model → HR 0.504 → 95% CI 0.39–0.65 → P < 0.0001

The primary endpoint is a time-to-event comparison in all randomized participants.

Bone outcomes
BMD → ANCOVA → differences 10.02, 7.92, 6.51

These are adjusted differences in percent change from baseline at Month 36, not time-to-event effects.

Vertebral fractures
Binary fracture outcomes → logistic regression → OR 0.53 and 0.54

These odds ratios describe binary outcomes at 36 months in the specified vertebral-fracture analysis set.

Later survival outcomes
DFS → HR 0.816   ·   BMFS → HR 0.808   ·   OS → HR 0.802

The registry explicitly places BMFS and OS within a hierarchical testing structure conditional on DFS and does not report p-values for those two endpoints.

25. Sources

26. Related Tutorials

Learn more about the methods used in this trial:

27. Related Calculators

Continue through the Clinical Biostats statistical learning pathway

Connect the endpoints and methods in ABCSG-18 to deeper tutorials and statistical calculators for clinical-trial analysis.

28. Limitations of the Supplied Registry Data

The ClinicalTrials.gov record contains nine posted statistical analyses and complete results for the primary endpoint, several secondary endpoints, and serious adverse events by arm. They do not provide every descriptive component that might appear in a full clinical-study report.

Reported here

Trial design, enrollment, endpoint definitions, analysis populations, statistical methods where reported, effect estimates, confidence intervals, p-values where reported, hierarchical-testing notes, and serious-adverse-event counts.

Not inferred

Unreported baseline characteristics, median event times, subgroup estimates, formal BMFS/OS methods, or additional safety comparisons are not reconstructed from outside information.

Clinical Biostats methodology: A trial-results page should distinguish what the registry actually reports from what a statistician can explain about the analysis. This page therefore preserves endpoint definitions, analysis populations, effect measures, confidence intervals, and testing structure rather than filling gaps with inferred results.