This page separates reported trial results from statistical interpretation. Numerical results, endpoint definitions, analysis populations, and reported statistical methods are taken only from the ClinicalTrials.gov trial data posted on ClinicalTrials.gov for ABCSG-16.
1. Trial at a Glance
ABCSG-16 was a randomized, parallel-group, phase 3 trial evaluating whether 5 years of additional anastrozole was more effective than 2 years of additional anastrozole after 5 years of adjuvant endocrine therapy in terms of disease-free survival.
| Feature | ABCSG-16 |
|---|---|
| Trial name | ABCSG-16 |
| Brief title | Secondary Adjuvant Long Term Study With Arimidex |
| Phase | Phase 3 |
| Condition | Breast Cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 3484 |
| Arms | 2 |
| Intervention | Anastrozole |
| Primary endpoint type | Time-to-event |
| Primary analysis method | Log-rank test |
| Effect measure | Hazard ratio |
| Hypothesis type | Superiority |
| Lead sponsor | AstraZeneca |
| Trial status | COMPLETED |
| Start | 2004-03-01 |
| Primary completion | 2017-06-30 |
2. Clinical Question
The central question was whether 5 years of additional anastrozole was more effective than 2 years of additional anastrozole after 5 years of adjuvant endocrine therapy in terms of disease-free survival.
Population
The trial enrolled 3484 patients with breast cancer into the randomized phase 3 study.
Intervention
Arm B received anastrozole at 1 mg per day for 5 years.
Comparator
Arm A received anastrozole at 1 mg per day for 2 years.
Primary question
Does 5 years of additional anastrozole produce a different disease-free survival experience from 2 years of additional anastrozole?
3. Trial Design
Anastrozole for 2 years
- Anastrozole
- 1 mg per day
- 2 years of additional treatment
Anastrozole for 5 years
- Anastrozole
- 1 mg per day
- 5 years of additional treatment
The registry identifies the design as randomized, parallel, unmasked, and intended for treatment. The two groups therefore represent different durations of the same intervention rather than a drug-versus-placebo comparison.
4. Endpoints
| Endpoint | Registry definition / time frame | Type |
|---|---|---|
| Disease-free Survival After Prolonged Endocrine Treatment | DFS was defined as the time from two years after randomization to the earliest occurrence of loco-regional recurrence, distant recurrence, contralateral new breast cancer, second cancer or death from any cause, assessed up to a maximum of 8.5 years | Time-to-event |
| Overall Survival After Prolonged Endocrine Treatment | Overall survival was defined as the time from two years after randomization to death due to any cause, assessed up to a maximum of 8.5 years | Time-to-event |
| Time to First Clinical Fracture | Time to first clinical fracture was defined as time to first clinical fracture, in the period from 2 years until 5 years after randomization for each patient. | Time-to-event |
| Time to Secondary Carcinoma | Risk of secondary carcinoma was defined as the time from two years after randomization to first occurrence of new secondary cancer without new breast cancer (local or contralateral), assessed up to a maximum of 8.5 years | Time-to-event |
| Time to Contralateral Breast Cancer | Risk of contralateral breast cancer was defined as the time from two years after randomization to first occurrence of new contralateral breast cancer, assessed up to a maximum of 8.5 years | Time-to-event |
5. Statistical Methodology
Primary analysis: log-rank test
The registered primary endpoint was a time-to-event outcome, and the posted primary analysis used a log-rank test. The log-rank test compares the survival experience of two groups across observed event times while accounting for the fact that some patients may be censored before experiencing the event.
The registry analysis notes explicitly define the null hypothesis as no difference between the two treatment arms in the probability of a DFS event.
Hazard ratio
The reported effect measure for every posted statistical analysis was the hazard ratio (HR). A hazard ratio compares the instantaneous event rates between two groups over the analyzed follow-up. An HR below 1 indicates a lower estimated event hazard for the first-listed group relative to the second-listed group; an HR above 1 indicates a higher estimated event hazard.
Values below or above 1 describe the direction and relative magnitude of the estimated hazard comparison. The HR is not itself an absolute probability, a median survival time, or the percentage of patients who benefit.
Per-protocol analysis
The primary DFS analysis was conducted in the Per-Protocol (PP) Population. The registry defines this population as all randomized patients who signed the informed consent, complied with the inclusion and exclusion criteria, and had at least one follow-up examination after the specified treatment period in the registry-reported definition.
This is an important feature of the statistical interpretation. Randomization establishes the initial treatment allocation, but a per-protocol analysis focuses on patients meeting the protocol-defined criteria for inclusion in that analysis. Consequently, the estimate should be interpreted as the result reported for that specified PP population rather than automatically as an all-randomized-patient estimate.
Secondary endpoint analyses
All four posted secondary analyses also used the log-rank test with hazard ratios and two-sided 95% confidence intervals. Each analysis had an endpoint-specific analysis population defined in the registry.
| Endpoint | Analysis population | Method | Effect measure |
|---|---|---|---|
| Overall Survival After Prolonged Endocrine Treatment | All randomized patients who signed the informed consent and who did not die during the first two years | Log-rank | Hazard ratio |
| Time to First Clinical Fracture | All randomized patients who signed the informed consent and who had no fractures during the first two years | Log-rank | Hazard ratio |
| Time to Secondary Carcinoma | All randomized patients who signed the informed consent and had no secondary carcinoma, excluding contralateral mammacarcinoma, during the first two years | Log-rank | Hazard ratio |
| Time to Contralateral Breast Cancer | All randomized patients who signed the informed consent and had no contralateral mammacarcinoma during the first two years | Log-rank | Hazard ratio |
6. Statistical Methods Explained
Why was a log-rank test used?
Because the primary endpoint was disease-free survival, the outcome is not simply whether an event occurred. The timing of the event matters, and some participants may not have experienced the event by the end of their available follow-up. A log-rank test is designed for comparing time-to-event distributions between groups while incorporating censored observations.
What does a hazard ratio of 0.93 mean?
The primary DFS estimate of 0.93 indicates that the estimated hazard in Arm A relative to Arm B was 0.93. Equivalently, the estimated hazard in Arm A was approximately 7% lower than in Arm B, based on the reported HR. This does not mean that 7% of patients avoided a DFS event, nor does it represent a 7-percentage-point difference in disease-free survival.
What does the 95% confidence interval tell us?
The primary 95% confidence interval was 0.79 to 1.11. It describes the statistical uncertainty surrounding the estimated hazard ratio under the analysis framework. Because the interval includes 1, the data are compatible with both a lower and a higher hazard for Arm A relative to Arm B within the interval.
Why is the p-value not a measure of effect size?
The primary p-value was 0.425. A p-value quantifies how unusual the observed data would be under the specified null hypothesis; it does not measure the magnitude of the treatment effect. The HR and its confidence interval provide the effect estimate and its precision, while the p-value addresses evidence against the null hypothesis.
Why does the per-protocol population matter?
The primary analysis was reported in the PP population rather than simply being described as an all-randomized analysis. Patients had to satisfy the registry's protocol-defined criteria for inclusion in that population. This means the reported HR describes the treatment comparison in that analysis population and should not automatically be treated as though it were calculated from every randomized patient.
How should an HR above 1 be interpreted?
For the secondary endpoint of time to first clinical fracture, the reported HR was 1.35 for Arm A versus Arm B. An HR above 1 indicates a higher estimated event hazard in Arm A relative to Arm B. It does not by itself state the absolute probability of fracture, the number of additional fractures, or the clinical importance of the difference.
7. Primary Results: Disease-free Survival
The primary endpoint was Disease-free Survival After Prolonged Endocrine Treatment. The registry defines DFS as the time from two years after randomization to the earliest occurrence of the specified disease-free-survival event in its endpoint definition. The analysis compared Arm A, anastrozole for 2 years, with Arm B, anastrozole for 5 years.
Primary DFS hazard ratio
95% CI: 0.79–1.11 · P = 0.425
Two-sided 95% confidence interval; log-rank analysis; Per-Protocol population.
| Primary endpoint | Arm A | Arm B | Effect estimate | P-value |
|---|---|---|---|---|
| Disease-free Survival After Prolonged Endocrine Treatment | Anastrozole for 2 years | Anastrozole for 5 years | HR 0.93 (95% CI 0.79–1.11) | 0.425 |
The reported HR of 0.93 means that the estimated hazard of a DFS event in Arm A was approximately 7% lower than the estimated hazard in Arm B under the reported analysis. The direction of the point estimate therefore favors Arm A numerically, but the estimate is close to the no-difference value of 1.
The HR does not mean that 7% fewer patients experienced a DFS event, and it does not describe an absolute difference in disease-free survival. It is a relative time-to-event measure.
The 95% CI of 0.79–1.11 indicates substantial uncertainty around the point estimate and includes 1. Thus, the confidence interval is compatible with a lower hazard as well as a higher hazard for Arm A relative to Arm B.
The p-value of 0.425 is a test of the specified null hypothesis; it is not a measure of how large or clinically important the treatment difference is. Interpretation also needs to respect the fact that the analysis was performed in the registry-defined per-protocol population and used a time-to-event comparison.
8. Secondary Results: Overall Survival
Overall survival was defined in the registry as the time from two years after randomization to death due to any cause, with the registry-reported time-frame field ending in an incomplete phrase. The analysis population included all randomized patients who signed informed consent and who did not die during the first two years.
Overall survival hazard ratio
95% CI: 0.83–1.25 · P = 0.867
Two-sided 95% confidence interval; log-rank analysis.
The HR of 1.02 indicates that the estimated hazard of death in Arm A was approximately 2% higher than in Arm B under the reported analysis. The point estimate is very close to 1.
The 95% CI of 0.83–1.25 includes 1, so the reported interval encompasses both a lower and a higher hazard for Arm A relative to Arm B. The confidence interval should be read as an expression of statistical precision, not as a range of individual patient outcomes.
The p-value of 0.867 addresses the null hypothesis specified for the comparison. It does not quantify the magnitude of the observed HR and should not be interpreted as the probability that the null hypothesis is true.
9. Secondary Results: Time to First Clinical Fracture
Time to first clinical fracture was defined as time to first clinical fracture in the period from 2 years until 5 years. The analysis included all randomized patients who signed informed consent and who had no fractures during the first two years.
Time to first clinical fracture
95% CI: 1.00–1.84 · P = 0.052
Two-sided 95% confidence interval; log-rank analysis.
The HR of 1.35 indicates a higher estimated hazard of first clinical fracture in Arm A relative to Arm B during the analyzed period. Expressed descriptively, the estimated hazard in Arm A was approximately 35% higher than in Arm B.
The estimate is a relative hazard measure rather than an absolute fracture probability. It does not say that 35% of patients fractured, nor that the absolute fracture risk differed by 35 percentage points.
The 95% CI of 1.00–1.84 reaches the no-difference value of 1. The p-value of 0.052 is close to, but above, 0.05. It should not be treated as a graded measure of effect size. The HR and confidence interval provide the more direct description of the estimated treatment comparison and its uncertainty.
10. Secondary Results: Time to Secondary Carcinoma
Risk of secondary carcinoma was defined in the registry as the time from two years after randomization to first occurrence of the specified new secondary carcinoma event. The analysis excluded patients with secondary carcinoma, excluding contralateral mammacarcinoma, during the first two years.
Time to secondary carcinoma
95% CI: 0.81–1.38 · P = 0.678
Two-sided 95% confidence interval; log-rank analysis.
The HR of 1.06 indicates a slightly higher estimated hazard of the specified secondary carcinoma event in Arm A relative to Arm B. The point estimate is close to 1, so the estimated relative difference is small.
The 95% CI of 0.81–1.38 includes 1 and extends in both directions around the no-difference value. This indicates uncertainty that is materially larger than the small point estimate alone might suggest.
The p-value of 0.678 tests the specified null hypothesis and does not measure clinical effect size. The result should be interpreted using the HR, its confidence interval, and the endpoint-specific analysis population together.
11. Secondary Results: Time to Contralateral Breast Cancer
Risk of contralateral breast cancer was defined in the registry as the time from two years after randomization to first occurrence of the specified contralateral breast cancer event. The analysis included patients who had no contralateral mammacarcinoma during the first two years.
Time to contralateral breast cancer
95% CI: 0.75–1.77 · P = 0.531
Two-sided 95% confidence interval; log-rank analysis.
The HR of 1.15 indicates a higher estimated hazard of contralateral breast cancer in Arm A relative to Arm B, with the point estimate corresponding to an approximately 15% higher estimated hazard.
This does not mean that the absolute probability of contralateral breast cancer was 15% higher. The hazard ratio is a relative time-to-event measure.
The 95% CI of 0.75–1.77 includes 1 and is relatively broad compared with the point estimate. The p-value of 0.531 is a test statistic under the specified null hypothesis, not a measure of the size or practical importance of the observed HR.
12. Consolidated Statistical Results
| Endpoint | Role | Analysis population | HR | 95% CI | P-value |
|---|---|---|---|---|---|
| Disease-free Survival After Prolonged Endocrine Treatment | Primary | Per-Protocol | 0.93 | 0.79–1.11 | 0.425 |
| Overall Survival After Prolonged Endocrine Treatment | Secondary | All randomized patients meeting endpoint-specific criteria | 1.02 | 0.83–1.25 | 0.867 |
| Time to First Clinical Fracture | Secondary | All randomized patients meeting endpoint-specific criteria | 1.35 | 1.00–1.84 | 0.052 |
| Time to Secondary Carcinoma | Secondary | All randomized patients meeting endpoint-specific criteria | 1.06 | 0.81–1.38 | 0.678 |
| Time to Contralateral Breast Cancer | Secondary | All randomized patients meeting endpoint-specific criteria | 1.15 | 0.75–1.77 | 0.531 |
All five posted statistical analyses use the same broad analytical framework: time-to-event outcomes compared with a log-rank test and expressed using hazard ratios with two-sided 95% confidence intervals. The primary analysis is distinguished by its per-protocol population and its status as the prespecified primary endpoint.
13. Reading the Hazard Ratios Across Endpoints
Viewed together, the point estimates lie on both sides of the HR = 1 reference point. The estimates themselves should not be compared solely by whether their p-values are above or below a threshold. Each endpoint has its own definition, analysis population, censoring structure, and clinical interpretation.
14. Why Confidence Intervals Matter
The DFS HR was 0.93, with a 95% CI of 0.79–1.11. The interval shows that the point estimate should not be interpreted as a precise estimate of a 7% hazard reduction. Values on both sides of 1 remain compatible with the reported interval.
The clinical-fracture HR was 1.35, with a 95% CI of 1.00–1.84. The interval conveys both the direction of the point estimate and the uncertainty around it. The lower bound reaches 1.00, so the no-difference value remains at the boundary of the reported interval.
Confidence intervals are especially useful when several related endpoints are reported because they prevent the analysis from being reduced to a binary p-value classification. An HR of 1.35 with a 95% CI of 1.00–1.84 communicates considerably more statistical information than the p-value of 0.052 alone.
15. Censoring and Time-to-Event Interpretation
All five posted statistical analyses are classified as time-to-event analyses. In this setting, patients do not necessarily contribute the same amount of observable follow-up. A participant can contribute information until an event occurs or until the observation period ends or otherwise becomes unavailable under the analysis rules.
This is why a time-to-event analysis differs from simply comparing the percentage of patients with events. The log-rank test uses the ordering of event times and the numbers at risk at those times. The hazard ratio likewise describes a relative event-hazard comparison rather than a simple difference in cumulative proportions.
The reported ABCSG-16 analyses use this framework for DFS, overall survival, clinical fracture, secondary carcinoma, and contralateral breast cancer.
16. Analysis Populations and Eligibility for Secondary Endpoints
A notable feature of the registry record is that the secondary analyses do not all use exactly the same population definition. The overall-survival analysis excludes patients who died during the first two years. The clinical-fracture analysis excludes patients with fractures during the first two years. The secondary-carcinoma and contralateral-breast-cancer analyses likewise define populations according to the absence of the relevant event during the first two years.
| Endpoint | Population restriction reported in registry |
|---|---|
| DFS | Per-Protocol population defined by protocol compliance and follow-up criteria. |
| Overall survival | No death during the first two years. |
| Clinical fracture | No fracture during the first two years. |
| Secondary carcinoma | No secondary carcinoma, excluding contralateral mammacarcinoma, during the first two years. |
| Contralateral breast cancer | No contralateral mammacarcinoma during the first two years. |
These distinctions matter because a hazard ratio is always conditional on the population and time origin used for the analysis. The registry specifies two years after randomization as the starting point for the primary DFS definition and for the registry-reported secondary endpoint definitions.
17. Statistical Features Supported by the Registry
Randomization
The trial allocation is identified as randomized, providing the design basis for comparing the two treatment-duration strategies.
Time-to-event analysis
The primary endpoint and all four posted secondary analyses are classified as time-to-event outcomes.
Log-rank testing
The reported statistical method for the primary and secondary analyses is the log-rank test.
Hazard ratios
Every posted statistical analysis reports a hazard ratio as its effect measure.
18. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm as affected patients over patients at risk.
| Arm | Serious adverse events | Affected / at risk |
|---|---|---|
| Arm A: Anastrozol - 1 mg Per Day for 2 Y | Serious adverse events | 452 / 1705 |
| Arm B: Anastrozol - 1 mg Per Day for 5 Y | Serious adverse events | 687 / 1710 |
The ClinicalTrials.gov record reports affected and at-risk counts rather than a formal statistical comparison of serious adverse-event rates. The numbers therefore describe the reported safety burden by arm but do not establish a hypothesis-tested difference between the treatment durations.
19. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The primary DFS analysis reported HR 0.93 with a two-sided 95% CI of 0.79–1.11 and P = 0.425 in the per-protocol population. The interval includes the no-difference value of 1.
Clinical interpretation
The registry question concerns the comparative effect of 2 versus 5 years of additional anastrozole after 5 years of adjuvant endocrine therapy. The statistical evidence should be interpreted across DFS, overall survival, fracture, secondary carcinoma, and contralateral breast cancer rather than from one number alone.
The distinction is important because statistical evidence does not automatically answer every clinical question. The HR describes a relative time-to-event comparison; the confidence interval describes statistical uncertainty; and the p-value evaluates the specified null hypothesis. None of these alone describes an individual's absolute outcome.
20. Important Limitations and Interpretation Issues
- Per-protocol primary analysis: the primary DFS estimate was reported in the registry-defined PP population, so it should not automatically be interpreted as an all-randomized analysis.
- Endpoint-specific populations: the secondary analyses use different eligibility restrictions based on events occurring during the first two years.
- Hazard-ratio interpretation: HRs are relative time-to-event measures and should not be interpreted as absolute risk differences or percentage-point differences.
- Confidence intervals: intervals that include 1 indicate that the no-difference value remains compatible with the reported statistical uncertainty.
- Multiple endpoints: five statistical analyses are posted, but the ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure. The secondary p-values should therefore be interpreted as the reported endpoint-specific tests rather than assuming an unreported familywise error-control strategy.
- Proportional-hazards assumptions: the ClinicalTrials.gov record does not report an assessment of the proportional-hazards assumption. A single HR should therefore be understood as the reported relative hazard summary, not as proof that hazards were proportional throughout follow-up.
- Safety comparisons: serious adverse-event counts are reported by arm, but no formal comparative safety analysis is reported.
21. Why This Trial Matters Statistically
ABCSG-16 is a useful teaching case because it illustrates how a randomized clinical trial can compare different durations of the same therapy rather than comparing an active intervention with an unrelated control. The statistical question is then whether extending treatment changes the time-to-event experience.
| Concept | How it appears in ABCSG-16 |
|---|---|
| Randomization | Patients were allocated in a randomized, parallel-group phase 3 design. |
| Time-to-event endpoints | DFS, overall survival, clinical fracture, secondary carcinoma, and contralateral breast cancer were analyzed as time-to-event outcomes. |
| Log-rank test | Used as the reported method for the primary and secondary statistical analyses. |
| Hazard ratio | Used as the reported effect measure for all five statistical analyses. |
| Confidence interval | Each reported HR includes a two-sided 95% confidence interval. |
| Per-protocol analysis | The primary DFS analysis used the registry-defined PP population. |
| Endpoint-specific populations | Secondary analyses apply restrictions based on events during the first two years. |
| Safety reporting | Serious adverse events are reported as affected patients over patients at risk by arm. |
22. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
23. Related Statistical Calculators
24. A Practical Framework for Reading ABCSG-16
Step 1 · Identify the estimand
The primary question concerns disease-free survival after two years following randomization and compares different durations of additional anastrozole.
Step 2 · Identify the population
The primary result is explicitly reported in the per-protocol population, while the secondary endpoints have their own registry-defined restrictions.
Step 3 · Read the HR
Determine whether the HR is below, equal to, or above 1 and describe its direction without converting it into an absolute risk difference.
Step 4 · Read the CI
Assess the width of the 95% CI and whether it contains 1. This adds information about precision that a point estimate alone cannot provide.
25. Sources
- ClinicalTrials.gov: NCT00295620 — ABCSG-16.
- PubMed: PubMed record for PMID 34320285.
Continue through the Clinical Biostats statistical pathway
Use the related tutorials and calculators to explore the survival-analysis concepts represented in ABCSG-16, including hazard ratios, confidence intervals, log-rank testing, and time-to-event endpoints.
26. Record Summary
ABCSG-16 provides a focused example of randomized time-to-event analysis in which the intervention is duration of additional anastrozole. The primary DFS analysis used a log-rank test and reported a hazard ratio of 0.93 with a two-sided 95% CI of 0.79–1.11 and P = 0.425 in the registry-defined per-protocol population. Four secondary time-to-event analyses used the same broad statistical framework, with HRs of 1.02 for overall survival, 1.35 for time to first clinical fracture, 1.06 for time to secondary carcinoma, and 1.15 for time to contralateral breast cancer.
The statistical lesson is broader than any individual p-value. A sound reading combines the analysis population, time origin, endpoint definition, hazard ratio, confidence interval, log-rank test, and endpoint-specific eligibility criteria. The serious-adverse-event counts provide a separate safety perspective and should not be conflated with the efficacy analyses.