This page separates reported trial results from statistical interpretation. Trial-specific numerical results and methodological details are restricted to the ClinicalTrials.gov record for NCT00310180. This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
TAILORx is a randomized, parallel, unmasked phase 3 treatment trial in women who have undergone surgery for node-negative breast cancer. The ClinicalTrials.gov record reports 10,273 participants enrolled and four arms, with 5-year disease-free survival as the single registered primary endpoint.
| Feature | TAILORx |
|---|---|
| Trial | TAILORx |
| NCT ID | NCT00310180 |
| Phase | Phase 3 |
| Therapeutic area | Oncology |
| Conditions | Breast Adenocarcinoma; Hormone Receptor Positive; Stage IA Breast Cancer AJCC v7; Stage IB Breast Cancer AJCC v7; Stage IIA Breast Cancer AJCC v6 and v7; Stage IIB Breast Cancer AJCC v6 and v7; Stage IIIB Breast Cancer AJCC v7 |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 10,273 |
| Lead sponsor | National Cancer Institute (NCI) |
| Sponsor type | NIH |
| Status | Active, not recruiting |
2. Clinical Question
The registry title describes TAILORx as a study of hormone therapy with or without combination chemotherapy in women who have undergone surgery for node-negative breast cancer. The primary statistical comparison reported in the posted analysis is Arm B versus Arm C for 5-year disease-free survival.
Population
Women who have undergone surgery for node-negative breast cancer, within the breast-cancer and hormone-receptor-positive conditions listed in the registry record.
Intervention
The ClinicalTrials.gov record describes the overall intervention set as including anastrozole, exemestane, letrozole, tamoxifen citrate, laboratory biomarker analysis, quality-of-life assessment, and radiation therapy. The ClinicalTrials.gov record does not provide enough arm-level detail to assign these components individually to Arm B or Arm C.
Comparator
The primary posted analysis compares Arm B vs Arm C. The ClinicalTrials.gov record does not provide arm names or arm-level treatment descriptions beyond those labels.
Primary question
For the registered 5-year disease-free survival endpoint, what is the relative time-to-event difference between Arm B and Arm C, and does the observed result meet the trial's registry-described noninferiority framework?
3. Trial Design
4. Endpoints
| Endpoint | Registry definition / time frame | Role |
|---|---|---|
| 5-year Disease-free Survival |
Time frame: Assessed every 6 months within 5 years from registration and then annually up to 20 years, DFS rate estimated at 5 years. Definition: Disease-free survival (DFS) is defined to be time from randomization to first event, where the first event is any of ipsilateral breast tumor recurrence, local recurrence, regional recurrence, distant recurrence, contralateral second primary invasive cancer, second primary non-breast invasive cancer (excluding non-melanoma skin cancers), or death without evidence of recurrence. |
Primary |
| 5-year Disease-free Survival by Age and Recurrence Score Groups | Assessed every 6 months within 5 years from registration and then annually up to 20 years, DFS rate estimated at 5 years. | Secondary |
The registry classifies the primary endpoint as a time-to-event outcome and reports the outcome unit as percentage of participants. The registry-reported definition establishes a composite event endpoint: the clock begins at randomization and ends at the first qualifying recurrence, second primary cancer, or death without evidence of recurrence.
5. Statistical Methodology
Analysis population
The posted primary analysis includes all eligible patients who had on-study data and follow-up data. This is the analysis population specified in the registry statistical-analysis record.
Cox proportional-hazards model
The registry reports Regression, Cox as the analysis method, normalized to a Cox proportional-hazards model. The model produces a hazard ratio comparing the instantaneous event rates between Arm B and Arm C over the analyzed follow-up.
The confidence interval is two-sided and has a confidence level of 95%. The reported hypothesis type is listed as Superiority, while the analysis notes separately describe the study as using a noninferiority design.
Why survival analysis is appropriate
DFS is not simply a binary endpoint assessed at one fixed time. The registered definition describes the time from randomization to the first qualifying event, with assessments continuing beyond the 5-year DFS estimate. Time-to-event methods preserve information about when events occur and can account for participants who have follow-up without an observed event during the analysis period.
Kaplan-Meier estimation
For a time-to-event endpoint such as DFS, Kaplan-Meier estimation is the standard descriptive framework for estimating the event-free survival function over time. It accommodates right-censored observations by allowing participants without an observed event to contribute information through their available follow-up.
Here, di represents events at an event time and ni represents participants at risk immediately before that time. The ClinicalTrials.gov record does not provide the underlying event and censoring records needed to reconstruct the TAILORx Kaplan-Meier curve.
6. The Noninferiority Framework
The most important statistical nuance in the registry-reported TAILORx analysis is that the registry simultaneously labels the posted hypothesis type as Superiority and describes the study in the analysis notes as a noninferiority design.
This is different from the simpler question, "Is the hazard ratio different from 1?" In an ordinary superiority analysis, 1 is the central null value. In a noninferiority framework, the relevant clinical boundary is the prespecified margin. Here, the registry-reported analysis notes identify 1.322 as the hazard-ratio boundary for the alternative in which Arm B is substantially worse than Arm C.
| Component | Registry-the ClinicalTrials.gov record |
|---|---|
| Primary endpoint | 5-year Disease-free Survival |
| Comparison | Arm B vs Arm C |
| Effect measure | Hazard ratio |
| Observed HR | 1.08 |
| 95% CI | 0.94–1.24 |
| P-value | 0.13 |
| Hypothesis field | Superiority |
| Design description in analysis notes | Noninferiority design |
| Specified hazard-ratio threshold | 1.322 |
The numerical relationship between the confidence interval and the stated margin is therefore important. The upper confidence-limit value of 1.24 is below the registry-reported threshold of 1.322. That comparison addresses the registry-described noninferiority logic; it is not the same statistical question as testing whether the hazard ratio differs from 1.
7. Primary Result: 5-year Disease-free Survival
The posted primary statistical analysis compares Arm B with Arm C for 5-year disease-free survival. The analysis population consists of all eligible patients who had on-study data and follow-up data.
Hazard ratio for disease-free survival
95% CI: 0.94–1.24 · P = 0.13
Arm B vs Arm C · Cox proportional-hazards model
| Primary endpoint | Arm B vs Arm C | Analysis |
|---|---|---|
| 5-year Disease-free Survival | HR 1.08 (95% CI 0.94–1.24); P = 0.13 | Cox proportional-hazards model |
An HR of 1.08 means that the fitted Cox model estimates the instantaneous rate of a qualifying DFS event in Arm B to be about 1.08 times the corresponding rate in Arm C over the analyzed follow-up. Expressed descriptively, that is an estimated hazard about 8% higher for Arm B than Arm C under the model.
The HR does not mean that 8% more participants experienced an event. It is not an absolute risk difference, an 8-percentage-point difference in 5-year DFS, or a statement about the outcome of every individual participant.
The 95% confidence interval of 0.94–1.24 describes statistical uncertainty around the estimated hazard ratio under the model and sampling framework. It does not describe the range of individual patient outcomes.
The P = 0.13 value is a measure of compatibility with the particular statistical hypothesis being tested under the analysis framework. It is not a measure of effect size, clinical importance, or the probability that the treatment effect is true.
For the registry-described noninferiority question, the relevant comparison is with the specified hazard-ratio threshold of 1.322, not merely with 1. The reported upper confidence limit of 1.24 is below that threshold.
Finally, a Cox hazard ratio relies on the proportional-hazards modeling framework. The ClinicalTrials.gov record does not provide information allowing an independent assessment of whether that assumption held throughout follow-up.
8. Secondary Result: 5-year Disease-free Survival by Age and Recurrence Score Groups
The second posted statistical analysis evaluates 5-year Disease-free Survival by Age and Recurrence Score Groups. It compares Arm B versus Arm C and uses Cox proportional-hazards analysis.
Treatment interaction test
Treatment interaction test across 9 age-by-Recurrence Score subsets
| Feature | Registry-reported analysis |
|---|---|
| Endpoint | 5-year Disease-free Survival by Age and Recurrence Score Groups |
| Comparison | Arm B vs Arm C |
| Method | Cox proportional-hazards model |
| Analysis population | All eligible patients who had on-study data and follow-up data |
| Interaction structure | 9 age-by-Recurrence Score subsets |
| Age groups | ≤50; 51-65; 66-75 |
| Recurrence Score groups | 0-10; 11-25; >25 |
| Treatment interaction P-value | 0.004 |
The reported P = 0.004 is associated with the treatment interaction test across the 9 age-by-Recurrence Score subsets. It therefore addresses whether the treatment comparison varies across the specified combinations of age and Recurrence Score; it is not itself a hazard ratio and does not quantify the magnitude of treatment effect in any individual subgroup.
The registry-reported analysis does not provide subgroup-specific hazard-ratio estimates or confidence intervals. Accordingly, the interaction result should not be converted into a claim about which individual age or Recurrence Score category has the largest treatment effect.
This distinction is fundamental in subgroup analysis: an interaction test asks a different question from testing each subgroup separately. A subgroup-specific P-value, if one were available, would not by itself establish that treatment effects differ between subgroups.
9. Statistical Methods Explained
Why is a Cox proportional-hazards model appropriate for DFS?
The primary endpoint is explicitly a time-to-event outcome. DFS measures time from randomization to the first qualifying event, so participants can have different lengths of follow-up and can be censored without an observed event. A Cox model is designed to compare event hazards while retaining the timing information contained in a survival endpoint.
What does an HR of 1.08 mean?
An HR of 1.08 for Arm B versus Arm C means that the fitted model estimates an instantaneous event rate 1.08 times that of Arm C. It does not mean that 8% more people had a DFS event, nor does it directly provide the difference in 5-year DFS percentages.
Why does the confidence interval matter?
The 95% CI of 0.94–1.24 shows the statistical uncertainty around the estimated HR of 1.08. The interval is especially important in this trial because the registry-described noninferiority framework identifies a specific upper boundary of 1.322. The upper confidence limit of 1.24 remains below that boundary.
Why is the noninferiority margin different from 1?
A hazard ratio of 1 represents equal estimated hazards. A noninferiority margin represents the largest deterioration that the study is designed to rule out under its prespecified framework. In the registry-reported TAILORx analysis notes, that boundary is expressed as a hazard ratio of 1.322 for Arm B versus Arm C.
What does P = 0.13 tell us?
A P-value measures evidence relative to the statistical hypothesis being evaluated; it does not measure the size of the treatment effect. The primary estimate is HR 1.08 with a 95% CI of 0.94–1.24. For the registry-described noninferiority question, the confidence interval's position relative to 1.322 is therefore a central part of interpretation.
What does the P = 0.004 interaction test tell us?
It tests the treatment interaction across the 9 specified age-by-Recurrence Score subsets. It does not tell us the hazard ratio within any particular subset. Because the ClinicalTrials.gov record does not provide those subgroup estimates and confidence intervals, the interaction result should be interpreted as evidence concerning treatment-effect variation across the specified grouping structure rather than as a quantitative effect estimate.
Why does the analysis population matter?
The posted analyses include all eligible patients who had on-study data and follow-up data. That definition determines who contributes to the reported estimates. It is therefore important not to substitute a different population definition when interpreting the hazard ratio or interaction test.
10. Understanding the Primary Result in Context
Relative effect
The primary model estimates HR 1.08 for Arm B versus Arm C. This is a relative time-to-event measure, not an absolute DFS percentage.
Precision
The 95% CI is 0.94–1.24. Precision should be judged from the interval rather than from the point estimate alone.
Hypothesis framework
The registry labels the hypothesis type as Superiority, while the analysis notes describe a noninferiority design with a 1.322 hazard-ratio threshold.
Time-to-event endpoint
DFS is defined from randomization to the first qualifying recurrence, second primary cancer, or death without evidence of recurrence.
One useful way to read the result is to separate three questions that can otherwise become conflated:
- What is the estimated relative effect? The reported HR is 1.08.
- How precise is that estimate? The two-sided 95% CI is 0.94–1.24.
- What hypothesis is the study actually evaluating? The registry analysis notes describe a noninferiority framework with a specified HR threshold of 1.322, even though the hypothesis-type field says "Superiority."
This separation is particularly important in noninferiority trials because a result can be interpreted differently depending on whether the question is equality, superiority, or whether the treatment's possible disadvantage remains below a prespecified clinically relevant boundary.
11. Age and Recurrence Score Interaction
The secondary analysis divides the population into nine combinations formed from three age groups and three Recurrence Score groups.
| Recurrence Score 0-10 | Recurrence Score 11-25 | Recurrence Score >25 | |
|---|---|---|---|
| Age ≤50 | Subset 1 | Subset 2 | Subset 3 |
| Age 51-65 | Subset 4 | Subset 5 | Subset 6 |
| Age 66-75 | Subset 7 | Subset 8 | Subset 9 |
The registry states that a treatment interaction test was performed for these nine age-by-Recurrence Score subsets in patients randomized to Arms B and C, using Cox proportional-hazards analysis. The reported P-value is 0.004.
12. What the Hazard Ratio Does — and Does Not — Mean
The primary HR of 1.08 is a model-based comparison of the event hazard for Arm B relative to Arm C. An HR above 1 indicates a higher estimated instantaneous event rate for the numerator arm under the fitted model.
It does not mean that Arm B had 8% fewer or more disease-free participants at 5 years. It also does not represent an 8-percentage-point difference in DFS.
The 95% CI of 0.94–1.24 places the uncertainty around the estimated HR. Both values below and above 1 are contained in the interval, while the entire interval remains below the registry-specified noninferiority threshold of 1.322.
The P-value of 0.13 is tied to the hypothesis tested in the posted analysis. It does not quantify how large the treatment difference is. The HR and its confidence interval provide the effect estimate and its precision; the P-value addresses evidence under the relevant hypothesis.
13. Planned Follow-up and Time Frame
The registered primary endpoint specifies a long follow-up structure: DFS is assessed every 6 months within 5 years from registration and then annually up to 20 years, with the DFS rate estimated at 5 years.
Trial start
The registry profile lists April 7, 2006 as the study start date.
Primary DFS assessment
The primary endpoint specifies DFS assessment every 6 months within 5 years, with the DFS rate estimated at 5 years.
Extended DFS follow-up
The registered assessment schedule continues annually after the first 5 years up to 20 years.
Primary completion
The registry profile lists March 2, 2018 as the primary completion date.
14. Results Coverage in the Registry
The ClinicalTrials.gov record reports 7 outcome measures and 2 statistical analyses. One of those statistical analyses is for the registered primary endpoint, and one is for the secondary age-by-Recurrence Score analysis.
| Posted item | Information available in the ClinicalTrials.gov record |
|---|---|
| Primary endpoint analysis | 5-year Disease-free Survival; Arm B vs Arm C; Cox proportional-hazards model; HR 1.08; 95% CI 0.94–1.24; P = 0.13 |
| Secondary analysis | 5-year Disease-free Survival by Age and Recurrence Score Groups; Arm B vs Arm C; Cox proportional-hazards model; treatment interaction P = 0.004 |
| Other posted outcomes | The record reports seven outcome measures in total, but the ClinicalTrials.gov record does not provide additional numerical analyses for them. |
Accordingly, this page does not introduce median survival, subgroup hazard ratios, baseline characteristics, event counts, adverse-event rates, Kaplan-Meier estimates, or other numerical results that are not contained in the ClinicalTrials.gov record.
15. Safety and Other Outcomes
the ClinicalTrials.gov record lists quality-of-life assessment, laboratory biomarker analysis, radiation therapy, and several drug interventions among the study interventions. However, the ClinicalTrials.gov record does not provide serious adverse-event counts by arm or other numerical safety results.
16. Design Topics Not Reported in the Supplied Analysis
| Topic | What can be stated from the ClinicalTrials.gov record |
|---|---|
| Noninferiority margin | A hazard-ratio threshold of 1.322 is specified in the primary analysis notes. |
| Crossover | No crossover information is reported. |
| Factorial design | No factorial design is identified; the design model is parallel. |
| Multiplicity | No multiplicity-adjustment procedure is reported in the posted analysis data. |
| Interim analysis | No interim-analysis procedure is reported in the posted analysis data. |
| Missing-data / imputation | No missing-data or imputation method is reported. |
| Stratification | No stratification factors are reported in the ClinicalTrials.gov record. |
| Bayesian methods | No Bayesian method is reported; the posted method is a Cox proportional-hazards model. |
These omissions are important because design features such as interim monitoring, multiplicity control, censoring rules, and missing-data handling can affect the interpretation of a clinical-trial analysis. They should not be reconstructed from assumptions about how similar trials are commonly analyzed.
17. Limitations and Interpretation Issues
- Registry-level detail: the ClinicalTrials.gov record provides the primary Cox result and one interaction analysis but do not provide the complete statistical-analysis plan.
- Arm-level treatment detail: the primary statistical comparison is identified as Arm B versus Arm C, but the ClinicalTrials.gov record does not provide sufficient arm-level descriptions to map the complete intervention list onto those arms.
- Noninferiority terminology: the registry's hypothesis-type field says "Superiority," while the analysis notes describe a noninferiority design. Both pieces of information should be retained rather than silently choosing one label.
- Hazard-ratio interpretation: the HR is a relative time-to-event measure and should not be interpreted as an absolute difference in 5-year DFS.
- Proportional-hazards assumption: the Cox model relies on its proportional-hazards framework. The ClinicalTrials.gov record does not provide diagnostics for that assumption.
- Subgroup interpretation: the secondary analysis provides an interaction P-value but no subgroup-specific HRs or confidence intervals. The individual subgroup effects therefore cannot be reconstructed from the ClinicalTrials.gov record.
- Analysis population: the posted analysis includes all eligible patients with on-study and follow-up data. Results should not be attributed to a different analysis population.
- Longitudinal follow-up: the registry specifies assessment through 20 years, but the ClinicalTrials.gov record does not provide numerical long-term DFS results beyond the posted primary analysis.
- Safety: quantitative serious-adverse-event results by arm are not included in the ClinicalTrials.gov record.
18. Why This Trial Matters Statistically
TAILORx is a useful statistical teaching case because the ClinicalTrials.gov record illustrates how a randomized clinical trial can combine a time-to-event endpoint with a Cox model and a noninferiority framework. It also shows why the choice of hypothesis matters: the same hazard ratio can be viewed differently depending on whether the analysis asks about superiority, equality, or whether an unfavorable effect remains below a prespecified noninferiority boundary.
| Concept | How it appears in TAILORx |
|---|---|
| Randomization | The trial uses randomized allocation. |
| Parallel design | The registered design model is parallel. |
| Time-to-event endpoint | 5-year DFS is defined as time from randomization to the first qualifying event. |
| Kaplan-Meier estimation | Appropriate descriptive framework for a time-to-event DFS endpoint. |
| Cox proportional-hazards model | Posted primary and secondary analyses use Cox regression. |
| Hazard ratio | The primary analysis reports HR 1.08 for Arm B vs Arm C. |
| Confidence interval | The primary HR has a two-sided 95% CI of 0.94–1.24. |
| Noninferiority logic | The primary analysis notes specify a hazard-ratio threshold of 1.322. |
| Interaction testing | A treatment interaction test was performed across 9 age-by-Recurrence Score subsets. |
| Subgroup analysis | Age groups and Recurrence Score groups form the secondary analysis structure. |
19. Primary Analysis: A Step-by-Step Statistical Reading
Step 1 · Define the event
DFS begins at randomization and ends at the first qualifying recurrence, second primary cancer, or death without evidence of recurrence.
Step 2 · Preserve event timing
Because the endpoint is time-to-event, participants contribute information according to their observed follow-up rather than only as event/no-event observations.
Step 3 · Fit the Cox model
The registry reports Cox regression for the Arm B versus Arm C comparison.
Step 4 · Interpret the HR
The fitted HR is 1.08, describing the estimated relative event hazard for Arm B compared with Arm C.
The next step is to distinguish the ordinary reference value of 1 from the noninferiority boundary. An HR of 1 would represent equal estimated hazards. The registry's noninferiority analysis instead identifies 1.322 as the relevant threshold for a substantially worse DFS hazard in Arm B.
The reported upper confidence limit, 1.24, is below the specified threshold of 1.322. This is the central numerical relationship for interpreting the posted analysis under the registry-described noninferiority framework.
20. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
21. Related Statistical Calculators
22. Sources
- ClinicalTrials.gov: TAILORx, NCT00310180.
- PubMed: PMID 37043210.
- PubMed: PMID 35175284.
- PubMed: PMID 34137783.
- PubMed: PMID 34090143.
- PubMed: PMID 33626896.
Continue through the Clinical Biostats statistical pathway
Use the related tutorials and calculators to examine the survival-analysis concepts illustrated by this trial.
23. Record Summary
TAILORx provides a useful example of how a randomized phase 3 trial can be analyzed when the primary outcome is a time-to-event endpoint and the principal effect measure is a hazard ratio. The registry analysis reports an HR of 1.08 for Arm B versus Arm C, with a two-sided 95% CI of 0.94–1.24 and P = 0.13. The same registry analysis notes describe the study as using a noninferiority design and specify a hazard-ratio threshold of 1.322. The upper confidence limit of 1.24 is below that threshold, making the relationship between the confidence interval and the noninferiority boundary central to interpretation.
The secondary analysis adds a different statistical question: whether the treatment comparison varies across nine age-by-Recurrence Score subsets. Its reported interaction P-value is 0.004, but the ClinicalTrials.gov record does not provide the corresponding subgroup hazard ratios or confidence intervals. The appropriate statistical reading is therefore to distinguish the interaction test from the magnitude of treatment effect in any individual subgroup.