← Clinical Trials
Breast Adenocarcinoma Phase 3 Time-to-Event NCT00310180

TAILORx: Complete Statistical Analysis of Hormone Therapy in Node-Negative Breast Cancer

An independent statistical analysis of the randomized phase 3 TAILORx trial, focusing on 5-year disease-free survival, the Cox proportional-hazards model, hazard ratios, confidence intervals, and the registry-described noninferiority framework.

Start: 2006-04-07  ·  Primary completion: 2018-03-02  ·  Status: Active, not recruiting
Scope of this record

This page separates reported trial results from statistical interpretation. Trial-specific numerical results and methodological details are restricted to the ClinicalTrials.gov record for NCT00310180. This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

TAILORx is a randomized, parallel, unmasked phase 3 treatment trial in women who have undergone surgery for node-negative breast cancer. The ClinicalTrials.gov record reports 10,273 participants enrolled and four arms, with 5-year disease-free survival as the single registered primary endpoint.

10,273
Enrollment
ClinicalTrials.gov record
4
Arms
Parallel design
1
Primary endpoint
5-year DFS
1.08
DFS HR
95% CI 0.94–1.24
FeatureTAILORx
TrialTAILORx
NCT IDNCT00310180
PhasePhase 3
Therapeutic areaOncology
ConditionsBreast Adenocarcinoma; Hormone Receptor Positive; Stage IA Breast Cancer AJCC v7; Stage IB Breast Cancer AJCC v7; Stage IIA Breast Cancer AJCC v6 and v7; Stage IIB Breast Cancer AJCC v6 and v7; Stage IIIB Breast Cancer AJCC v7
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment10,273
Lead sponsorNational Cancer Institute (NCI)
Sponsor typeNIH
StatusActive, not recruiting

2. Clinical Question

The registry title describes TAILORx as a study of hormone therapy with or without combination chemotherapy in women who have undergone surgery for node-negative breast cancer. The primary statistical comparison reported in the posted analysis is Arm B versus Arm C for 5-year disease-free survival.

Population

Women who have undergone surgery for node-negative breast cancer, within the breast-cancer and hormone-receptor-positive conditions listed in the registry record.

Intervention

The ClinicalTrials.gov record describes the overall intervention set as including anastrozole, exemestane, letrozole, tamoxifen citrate, laboratory biomarker analysis, quality-of-life assessment, and radiation therapy. The ClinicalTrials.gov record does not provide enough arm-level detail to assign these components individually to Arm B or Arm C.

Comparator

The primary posted analysis compares Arm B vs Arm C. The ClinicalTrials.gov record does not provide arm names or arm-level treatment descriptions beyond those labels.

Primary question

For the registered 5-year disease-free survival endpoint, what is the relative time-to-event difference between Arm B and Arm C, and does the observed result meet the trial's registry-described noninferiority framework?

3. Trial Design

01
Enroll 10,273 participants
02
Randomize Four treatment arms
03
Follow DFS assessments
04
Analyze Cox regression
05
Estimate 5-year DFS
Allocation
Randomized allocation.
Design model
Parallel.
Masking
None.
Primary purpose
Treatment.
Phase
Phase 3.
Results status
Results posted; seven outcome measures and two statistical analyses are reported in the ClinicalTrials.gov record.
Registry scope: The ClinicalTrials.gov record identifies four arms and reports the primary statistical comparison as Arm B versus Arm C. It does not provide enough arm-level treatment information to reconstruct all four regimens without introducing facts from outside the ClinicalTrials.gov record.

4. Endpoints

EndpointRegistry definition / time frameRole
5-year Disease-free Survival Time frame: Assessed every 6 months within 5 years from registration and then annually up to 20 years, DFS rate estimated at 5 years.

Definition: Disease-free survival (DFS) is defined to be time from randomization to first event, where the first event is any of ipsilateral breast tumor recurrence, local recurrence, regional recurrence, distant recurrence, contralateral second primary invasive cancer, second primary non-breast invasive cancer (excluding non-melanoma skin cancers), or death without evidence of recurrence.
Primary
5-year Disease-free Survival by Age and Recurrence Score Groups Assessed every 6 months within 5 years from registration and then annually up to 20 years, DFS rate estimated at 5 years. Secondary

The registry classifies the primary endpoint as a time-to-event outcome and reports the outcome unit as percentage of participants. The registry-reported definition establishes a composite event endpoint: the clock begins at randomization and ends at the first qualifying recurrence, second primary cancer, or death without evidence of recurrence.

5. Statistical Methodology

Analysis population

The posted primary analysis includes all eligible patients who had on-study data and follow-up data. This is the analysis population specified in the registry statistical-analysis record.

Cox proportional-hazards model

The registry reports Regression, Cox as the analysis method, normalized to a Cox proportional-hazards model. The model produces a hazard ratio comparing the instantaneous event rates between Arm B and Arm C over the analyzed follow-up.

Reported primary analysis
Arm B vs Arm C  →  HR = 1.08  ·  95% CI 0.94–1.24  ·  P = 0.13

The confidence interval is two-sided and has a confidence level of 95%. The reported hypothesis type is listed as Superiority, while the analysis notes separately describe the study as using a noninferiority design.

Why survival analysis is appropriate

DFS is not simply a binary endpoint assessed at one fixed time. The registered definition describes the time from randomization to the first qualifying event, with assessments continuing beyond the 5-year DFS estimate. Time-to-event methods preserve information about when events occur and can account for participants who have follow-up without an observed event during the analysis period.

Kaplan-Meier estimation

For a time-to-event endpoint such as DFS, Kaplan-Meier estimation is the standard descriptive framework for estimating the event-free survival function over time. It accommodates right-censored observations by allowing participants without an observed event to contribute information through their available follow-up.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents events at an event time and ni represents participants at risk immediately before that time. The ClinicalTrials.gov record does not provide the underlying event and censoring records needed to reconstruct the TAILORx Kaplan-Meier curve.

6. The Noninferiority Framework

The most important statistical nuance in the registry-reported TAILORx analysis is that the registry simultaneously labels the posted hypothesis type as Superiority and describes the study in the analysis notes as a noninferiority design.

Registry-described noninferiority logic: the registry-reported analysis notes state that the noninferiority question is formulated using a conventional superiority null hypothesis of equal DFS on the two arms, with the null framed as Arm B not being inferior to Arm C. The alternative is that Arm B has substantially worse DFS than Arm C, specified by a hazard ratio for B vs C of 1.322.

This is different from the simpler question, "Is the hazard ratio different from 1?" In an ordinary superiority analysis, 1 is the central null value. In a noninferiority framework, the relevant clinical boundary is the prespecified margin. Here, the registry-reported analysis notes identify 1.322 as the hazard-ratio boundary for the alternative in which Arm B is substantially worse than Arm C.

ComponentRegistry-the ClinicalTrials.gov record
Primary endpoint5-year Disease-free Survival
ComparisonArm B vs Arm C
Effect measureHazard ratio
Observed HR1.08
95% CI0.94–1.24
P-value0.13
Hypothesis fieldSuperiority
Design description in analysis notesNoninferiority design
Specified hazard-ratio threshold1.322

The numerical relationship between the confidence interval and the stated margin is therefore important. The upper confidence-limit value of 1.24 is below the registry-reported threshold of 1.322. That comparison addresses the registry-described noninferiority logic; it is not the same statistical question as testing whether the hazard ratio differs from 1.

7. Primary Result: 5-year Disease-free Survival

The posted primary statistical analysis compares Arm B with Arm C for 5-year disease-free survival. The analysis population consists of all eligible patients who had on-study data and follow-up data.

Hazard ratio for disease-free survival

1.08

95% CI: 0.94–1.24   ·   P = 0.13

Arm B vs Arm C  ·  Cox proportional-hazards model

Primary endpointArm B vs Arm CAnalysis
5-year Disease-free Survival HR 1.08 (95% CI 0.94–1.24); P = 0.13 Cox proportional-hazards model
Clinical Biostats interpretation

An HR of 1.08 means that the fitted Cox model estimates the instantaneous rate of a qualifying DFS event in Arm B to be about 1.08 times the corresponding rate in Arm C over the analyzed follow-up. Expressed descriptively, that is an estimated hazard about 8% higher for Arm B than Arm C under the model.

The HR does not mean that 8% more participants experienced an event. It is not an absolute risk difference, an 8-percentage-point difference in 5-year DFS, or a statement about the outcome of every individual participant.

The 95% confidence interval of 0.94–1.24 describes statistical uncertainty around the estimated hazard ratio under the model and sampling framework. It does not describe the range of individual patient outcomes.

The P = 0.13 value is a measure of compatibility with the particular statistical hypothesis being tested under the analysis framework. It is not a measure of effect size, clinical importance, or the probability that the treatment effect is true.

For the registry-described noninferiority question, the relevant comparison is with the specified hazard-ratio threshold of 1.322, not merely with 1. The reported upper confidence limit of 1.24 is below that threshold.

Finally, a Cox hazard ratio relies on the proportional-hazards modeling framework. The ClinicalTrials.gov record does not provide information allowing an independent assessment of whether that assumption held throughout follow-up.

Important distinction: The registry's posted analysis lists the hypothesis type as "Superiority," but its analysis notes explicitly describe the study as a noninferiority design and identify 1.322 as the hazard-ratio threshold. Those two pieces of registry metadata should not be silently collapsed into a generic superiority interpretation.

8. Secondary Result: 5-year Disease-free Survival by Age and Recurrence Score Groups

The second posted statistical analysis evaluates 5-year Disease-free Survival by Age and Recurrence Score Groups. It compares Arm B versus Arm C and uses Cox proportional-hazards analysis.

Treatment interaction test

P = 0.004

Treatment interaction test across 9 age-by-Recurrence Score subsets

FeatureRegistry-reported analysis
Endpoint5-year Disease-free Survival by Age and Recurrence Score Groups
ComparisonArm B vs Arm C
MethodCox proportional-hazards model
Analysis populationAll eligible patients who had on-study data and follow-up data
Interaction structure9 age-by-Recurrence Score subsets
Age groups≤50; 51-65; 66-75
Recurrence Score groups0-10; 11-25; >25
Treatment interaction P-value0.004
Clinical Biostats interpretation

The reported P = 0.004 is associated with the treatment interaction test across the 9 age-by-Recurrence Score subsets. It therefore addresses whether the treatment comparison varies across the specified combinations of age and Recurrence Score; it is not itself a hazard ratio and does not quantify the magnitude of treatment effect in any individual subgroup.

The registry-reported analysis does not provide subgroup-specific hazard-ratio estimates or confidence intervals. Accordingly, the interaction result should not be converted into a claim about which individual age or Recurrence Score category has the largest treatment effect.

This distinction is fundamental in subgroup analysis: an interaction test asks a different question from testing each subgroup separately. A subgroup-specific P-value, if one were available, would not by itself establish that treatment effects differ between subgroups.

9. Statistical Methods Explained

Why is a Cox proportional-hazards model appropriate for DFS?

The primary endpoint is explicitly a time-to-event outcome. DFS measures time from randomization to the first qualifying event, so participants can have different lengths of follow-up and can be censored without an observed event. A Cox model is designed to compare event hazards while retaining the timing information contained in a survival endpoint.

What does an HR of 1.08 mean?

An HR of 1.08 for Arm B versus Arm C means that the fitted model estimates an instantaneous event rate 1.08 times that of Arm C. It does not mean that 8% more people had a DFS event, nor does it directly provide the difference in 5-year DFS percentages.

Why does the confidence interval matter?

The 95% CI of 0.94–1.24 shows the statistical uncertainty around the estimated HR of 1.08. The interval is especially important in this trial because the registry-described noninferiority framework identifies a specific upper boundary of 1.322. The upper confidence limit of 1.24 remains below that boundary.

Why is the noninferiority margin different from 1?

A hazard ratio of 1 represents equal estimated hazards. A noninferiority margin represents the largest deterioration that the study is designed to rule out under its prespecified framework. In the registry-reported TAILORx analysis notes, that boundary is expressed as a hazard ratio of 1.322 for Arm B versus Arm C.

What does P = 0.13 tell us?

A P-value measures evidence relative to the statistical hypothesis being evaluated; it does not measure the size of the treatment effect. The primary estimate is HR 1.08 with a 95% CI of 0.94–1.24. For the registry-described noninferiority question, the confidence interval's position relative to 1.322 is therefore a central part of interpretation.

What does the P = 0.004 interaction test tell us?

It tests the treatment interaction across the 9 specified age-by-Recurrence Score subsets. It does not tell us the hazard ratio within any particular subset. Because the ClinicalTrials.gov record does not provide those subgroup estimates and confidence intervals, the interaction result should be interpreted as evidence concerning treatment-effect variation across the specified grouping structure rather than as a quantitative effect estimate.

Why does the analysis population matter?

The posted analyses include all eligible patients who had on-study data and follow-up data. That definition determines who contributes to the reported estimates. It is therefore important not to substitute a different population definition when interpreting the hazard ratio or interaction test.

10. Understanding the Primary Result in Context

Relative effect

The primary model estimates HR 1.08 for Arm B versus Arm C. This is a relative time-to-event measure, not an absolute DFS percentage.

Precision

The 95% CI is 0.94–1.24. Precision should be judged from the interval rather than from the point estimate alone.

Hypothesis framework

The registry labels the hypothesis type as Superiority, while the analysis notes describe a noninferiority design with a 1.322 hazard-ratio threshold.

Time-to-event endpoint

DFS is defined from randomization to the first qualifying recurrence, second primary cancer, or death without evidence of recurrence.

One useful way to read the result is to separate three questions that can otherwise become conflated:

  1. What is the estimated relative effect? The reported HR is 1.08.
  2. How precise is that estimate? The two-sided 95% CI is 0.94–1.24.
  3. What hypothesis is the study actually evaluating? The registry analysis notes describe a noninferiority framework with a specified HR threshold of 1.322, even though the hypothesis-type field says "Superiority."

This separation is particularly important in noninferiority trials because a result can be interpreted differently depending on whether the question is equality, superiority, or whether the treatment's possible disadvantage remains below a prespecified clinically relevant boundary.

11. Age and Recurrence Score Interaction

The secondary analysis divides the population into nine combinations formed from three age groups and three Recurrence Score groups.

Recurrence Score 0-10Recurrence Score 11-25Recurrence Score >25
Age ≤50Subset 1Subset 2Subset 3
Age 51-65Subset 4Subset 5Subset 6
Age 66-75Subset 7Subset 8Subset 9

The registry states that a treatment interaction test was performed for these nine age-by-Recurrence Score subsets in patients randomized to Arms B and C, using Cox proportional-hazards analysis. The reported P-value is 0.004.

Interpretation caution: the ClinicalTrials.gov record does not provide nine separate hazard ratios or confidence intervals. The appropriate conclusion from the posted result is therefore limited to the reported interaction analysis. It would be inappropriate to infer the direction or magnitude of treatment effects in individual subsets from the interaction P-value alone.

12. What the Hazard Ratio Does — and Does Not — Mean

Statistical interpretation

The primary HR of 1.08 is a model-based comparison of the event hazard for Arm B relative to Arm C. An HR above 1 indicates a higher estimated instantaneous event rate for the numerator arm under the fitted model.

It does not mean that Arm B had 8% fewer or more disease-free participants at 5 years. It also does not represent an 8-percentage-point difference in DFS.

Why the confidence interval matters

The 95% CI of 0.94–1.24 places the uncertainty around the estimated HR. Both values below and above 1 are contained in the interval, while the entire interval remains below the registry-specified noninferiority threshold of 1.322.

Why the p-value is not the effect size

The P-value of 0.13 is tied to the hypothesis tested in the posted analysis. It does not quantify how large the treatment difference is. The HR and its confidence interval provide the effect estimate and its precision; the P-value addresses evidence under the relevant hypothesis.

13. Planned Follow-up and Time Frame

The registered primary endpoint specifies a long follow-up structure: DFS is assessed every 6 months within 5 years from registration and then annually up to 20 years, with the DFS rate estimated at 5 years.

2006-04-07

Trial start

The registry profile lists April 7, 2006 as the study start date.

5-year endpoint

Primary DFS assessment

The primary endpoint specifies DFS assessment every 6 months within 5 years, with the DFS rate estimated at 5 years.

Up to 20 years

Extended DFS follow-up

The registered assessment schedule continues annually after the first 5 years up to 20 years.

2018-03-02

Primary completion

The registry profile lists March 2, 2018 as the primary completion date.

14. Results Coverage in the Registry

The ClinicalTrials.gov record reports 7 outcome measures and 2 statistical analyses. One of those statistical analyses is for the registered primary endpoint, and one is for the secondary age-by-Recurrence Score analysis.

Posted itemInformation available in the ClinicalTrials.gov record
Primary endpoint analysis5-year Disease-free Survival; Arm B vs Arm C; Cox proportional-hazards model; HR 1.08; 95% CI 0.94–1.24; P = 0.13
Secondary analysis5-year Disease-free Survival by Age and Recurrence Score Groups; Arm B vs Arm C; Cox proportional-hazards model; treatment interaction P = 0.004
Other posted outcomesThe record reports seven outcome measures in total, but the ClinicalTrials.gov record does not provide additional numerical analyses for them.

Accordingly, this page does not introduce median survival, subgroup hazard ratios, baseline characteristics, event counts, adverse-event rates, Kaplan-Meier estimates, or other numerical results that are not contained in the ClinicalTrials.gov record.

15. Safety and Other Outcomes

the ClinicalTrials.gov record lists quality-of-life assessment, laboratory biomarker analysis, radiation therapy, and several drug interventions among the study interventions. However, the ClinicalTrials.gov record does not provide serious adverse-event counts by arm or other numerical safety results.

Data boundary: Because the registry-reported TAILORx data do not contain serious adverse-event counts by treatment arm, this page does not construct a safety table or infer safety comparisons from the intervention list. The presence of an intervention in the registry is not itself a quantitative safety result.

16. Design Topics Not Reported in the Supplied Analysis

TopicWhat can be stated from the ClinicalTrials.gov record
Noninferiority marginA hazard-ratio threshold of 1.322 is specified in the primary analysis notes.
CrossoverNo crossover information is reported.
Factorial designNo factorial design is identified; the design model is parallel.
MultiplicityNo multiplicity-adjustment procedure is reported in the posted analysis data.
Interim analysisNo interim-analysis procedure is reported in the posted analysis data.
Missing-data / imputationNo missing-data or imputation method is reported.
StratificationNo stratification factors are reported in the ClinicalTrials.gov record.
Bayesian methodsNo Bayesian method is reported; the posted method is a Cox proportional-hazards model.

These omissions are important because design features such as interim monitoring, multiplicity control, censoring rules, and missing-data handling can affect the interpretation of a clinical-trial analysis. They should not be reconstructed from assumptions about how similar trials are commonly analyzed.

17. Limitations and Interpretation Issues

18. Why This Trial Matters Statistically

TAILORx is a useful statistical teaching case because the ClinicalTrials.gov record illustrates how a randomized clinical trial can combine a time-to-event endpoint with a Cox model and a noninferiority framework. It also shows why the choice of hypothesis matters: the same hazard ratio can be viewed differently depending on whether the analysis asks about superiority, equality, or whether an unfavorable effect remains below a prespecified noninferiority boundary.

ConceptHow it appears in TAILORx
RandomizationThe trial uses randomized allocation.
Parallel designThe registered design model is parallel.
Time-to-event endpoint5-year DFS is defined as time from randomization to the first qualifying event.
Kaplan-Meier estimationAppropriate descriptive framework for a time-to-event DFS endpoint.
Cox proportional-hazards modelPosted primary and secondary analyses use Cox regression.
Hazard ratioThe primary analysis reports HR 1.08 for Arm B vs Arm C.
Confidence intervalThe primary HR has a two-sided 95% CI of 0.94–1.24.
Noninferiority logicThe primary analysis notes specify a hazard-ratio threshold of 1.322.
Interaction testingA treatment interaction test was performed across 9 age-by-Recurrence Score subsets.
Subgroup analysisAge groups and Recurrence Score groups form the secondary analysis structure.

19. Primary Analysis: A Step-by-Step Statistical Reading

Step 1 · Define the event

DFS begins at randomization and ends at the first qualifying recurrence, second primary cancer, or death without evidence of recurrence.

Step 2 · Preserve event timing

Because the endpoint is time-to-event, participants contribute information according to their observed follow-up rather than only as event/no-event observations.

Step 3 · Fit the Cox model

The registry reports Cox regression for the Arm B versus Arm C comparison.

Step 4 · Interpret the HR

The fitted HR is 1.08, describing the estimated relative event hazard for Arm B compared with Arm C.

The next step is to distinguish the ordinary reference value of 1 from the noninferiority boundary. An HR of 1 would represent equal estimated hazards. The registry's noninferiority analysis instead identifies 1.322 as the relevant threshold for a substantially worse DFS hazard in Arm B.

Primary statistical comparison
HR = 1.08   |   95% CI = 0.94–1.24   |   Noninferiority threshold = 1.322

The reported upper confidence limit, 1.24, is below the specified threshold of 1.322. This is the central numerical relationship for interpreting the posted analysis under the registry-described noninferiority framework.

20. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

21. Related Statistical Calculators

22. Sources

Continue through the Clinical Biostats statistical pathway

Use the related tutorials and calculators to examine the survival-analysis concepts illustrated by this trial.

23. Record Summary

TAILORx provides a useful example of how a randomized phase 3 trial can be analyzed when the primary outcome is a time-to-event endpoint and the principal effect measure is a hazard ratio. The registry analysis reports an HR of 1.08 for Arm B versus Arm C, with a two-sided 95% CI of 0.94–1.24 and P = 0.13. The same registry analysis notes describe the study as using a noninferiority design and specify a hazard-ratio threshold of 1.322. The upper confidence limit of 1.24 is below that threshold, making the relationship between the confidence interval and the noninferiority boundary central to interpretation.

The secondary analysis adds a different statistical question: whether the treatment comparison varies across nine age-by-Recurrence Score subsets. Its reported interaction P-value is 0.004, but the ClinicalTrials.gov record does not provide the corresponding subgroup hazard ratios or confidence intervals. The appropriate statistical reading is therefore to distinguish the interaction test from the magnitude of treatment effect in any individual subgroup.

Clinical Biostats methodology: A trial-results page should distinguish the reported statistical estimate from the hypothesis it addresses. For TAILORx, that means reading the Cox hazard ratio, its confidence interval, the P-value, and the registry-described noninferiority threshold together rather than treating any single number as the complete statistical conclusion.