This page separates reported trial results from statistical interpretation. Numerical results on this page are restricted to the ClinicalTrials.gov trial data posted on ClinicalTrials.gov for KATHERINE. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
KATHERINE was a randomized, parallel, phase 3 oncology trial evaluating trastuzumab emtansine versus trastuzumab as adjuvant therapy in patients with HER2-positive breast cancer who had residual tumor in the breast or axillary lymph nodes following preoperative therapy. The registry reports one primary time-to-event endpoint and a series of secondary and other pre-specified time-to-event analyses.
| Feature | KATHERINE |
|---|---|
| Trial name | KATHERINE |
| NCT identifier | NCT01772472 |
| Phase | Phase 3 |
| Condition | Breast Cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 1486 |
| Interventions | Trastuzumab; trastuzumab emtansine |
| Primary endpoint | Invasive Disease-free Survival (IDFS) Rate at 3 Years |
| Primary endpoint type | Time-to-event |
| Primary analysis | Log-rank test with hazard ratio as the effect measure |
| Hypothesis | Superiority |
| Results posted | Yes |
| Statistical analyses posted | 15 |
| Lead sponsor | Hoffmann-La Roche |
| Sponsor type | Industry |
2. Clinical Question
The clinical question was whether trastuzumab emtansine, compared with trastuzumab, improved invasive disease-free survival when used as adjuvant therapy in patients with HER2-positive breast cancer who had residual tumor in the breast or axillary lymph nodes following preoperative therapy.
Population
Patients with HER2-positive breast cancer who had residual tumor in the breast or axillary lymph nodes following preoperative therapy.
Intervention
Trastuzumab emtansine.
Comparator
Trastuzumab.
Primary question
Does trastuzumab emtansine improve the time-to-event endpoint of invasive disease-free survival relative to trastuzumab?
3. Trial Design
Trastuzumab
- Trastuzumab
- Comparator treatment in the randomized trial
Trastuzumab emtansine
- Trastuzumab emtansine
- Intervention treatment in the randomized trial
The registry identifies randomization, rather than treatment actually received, as the basis for the primary efficacy analysis. This distinction becomes important when interpreting the reported hazard ratio because randomization is what establishes the treatment groups for the principal comparative analysis.
4. Endpoints
Primary endpoint
| Endpoint | Time frame | Type | Analysis |
|---|---|---|---|
| Invasive Disease-free Survival (IDFS) Rate at 3 Years | At Year 3 | Time-to-event | Log-rank test; hazard ratio |
The registry definition states that an IDFS event was defined as the first occurrence of specified events, beginning with ipsilateral invasive breast tumor recurrence and ipsilateral local-regional invasive breast cancer recurrence involving the axilla, regional lymph nodes, chest wall and/or skin of the ipsilateral breast); distant recurrence (i.e., evidence of breast cancer in any anatomic site-other than the 2 above-mentioned sites - that has either been histologically confirmed or clinically diagnosed as recurrent invasive breast cancer); contralateral invasive breast cancer; death attributable to any cause including breast cancer, non-breast cancer or unknown cause . 3-year IDFS rate in ITT population was estimated using Kaplan Meier (KM) method and the percentage of participants who were event-free 3 years after randomization was estimated.
Secondary and other pre-specified endpoints with posted analyses
| Endpoint | Time frame | Role |
|---|---|---|
| IDFS Including Second Primary Non-breast Cancer (SPNBC) Rate at 3 Years | At Year 3 | Secondary |
| IDFS Including SPNBC Rate at 7 Years | At Year 7 | Secondary |
| IDFS Including SPNBC Rate at 8 Years | At Year 8 | Secondary |
| Disease-free Survival (DFS) Rate at 3 Years | At Year 3 | Secondary |
| DFS Rate at 7 Years | At Year 7 | Secondary |
| DFS Rate at 8 Years | At Year 8 | Secondary |
| Overall Survival (OS) Rate at 5 Years | At Year 5 | Secondary |
| OS Rate at 7 Years | At Year 7 | Secondary |
| OS Rate at 8 Years | At Year 8 | Secondary |
| Distant Recurrence-free Interval (DRFI) Rate at 3 Years | At Year 3 | Secondary |
| DRFI Rate at 7 Years | At Year 7 | Secondary |
| DRFI Rate at 8 Years | At Year 8 | Secondary |
| IDFS Rate at 7 Years | At Year 7 | Other pre-specified |
| IDFS Rate at 8 Years | At Year 8 | Other pre-specified |
All of the posted analyses in the ClinicalTrials.gov record use the same basic comparative framework: the ITT population, comparison of trastuzumab with trastuzumab emtansine, a log-rank test, and a hazard ratio as the effect measure. That consistency makes the registry results useful for examining how one randomized time-to-event framework was applied across several related endpoints and time points.
5. Statistical Methodology
Intention-to-treat analysis
The registry defines the ITT population as including all participants who were randomized to the study regardless of whether they received any study treatment. The reported primary and secondary efficacy analyses use this population.
The statistical importance of ITT analysis is that treatment assignment remains the defining group variable after randomization. The analysis therefore preserves the treatment comparison created by the randomized design rather than redefining groups according to subsequent treatment exposure.
Time-to-event analysis
The primary endpoint and all of the statistical analyses posted on ClinicalTrials.gov are classified as time-to-event endpoints. Unlike a simple binary outcome observed at a single visit, a time-to-event analysis uses information about when an event occurs and can accommodate participants whose event status is not observed during the available follow-up.
The registry reports the treatment effect using a hazard ratio and compares the groups using a log-rank test.
Log-rank test
The reported statistical method is a log-rank test. The log-rank procedure compares the survival experience of two groups across the observed follow-up rather than reducing the analysis to a single fixed-time proportion.
Conceptually, at each observed event time the method compares the events that occurred with the numbers of participants at risk in the two groups. The evidence is then accumulated across follow-up. This makes the method particularly appropriate for randomized comparisons in which the outcome is naturally expressed as time to an event.
Hazard ratio
The registry reports the treatment effect as a hazard ratio. A hazard ratio compares the instantaneous event rates between treatment groups within the time-to-event framework.
For the analyses posted on ClinicalTrials.gov, the groups are reported as trastuzumab versus trastuzumab emtansine. Thus an HR below 1 corresponds to a lower estimated hazard for trastuzumab relative to trastuzumab emtansine under that ordering.
That direction matters. The numerical value of a hazard ratio cannot be interpreted correctly without knowing which treatment is in the numerator. In this page, the registry's group comparison is explicitly Trastuzumab vs Trastuzumab Emtansine.
Confidence intervals
The primary analysis reports a 95% confidence interval of 0.44 to 0.66 around the hazard ratio of 0.54. The secondary analyses likewise provide 95% confidence intervals around their hazard-ratio estimates.
A confidence interval communicates statistical precision around the estimated effect under the analysis framework. It does not describe the range of outcomes that individual patients will experience, and it is not a probability statement that the true hazard ratio lies within the particular interval.
6. Primary Result: Invasive Disease-free Survival at 3 Years
The registry reports a formal statistical analysis of the primary endpoint, Invasive Disease-free Survival (IDFS) Rate at 3 Years, using a log-rank test in the ITT population. The effect measure is a hazard ratio and the hypothesis type is superiority.
Hazard ratio for invasive disease-free survival
95% CI: 0.44–0.66 · P < 0.0001
Comparison: Trastuzumab vs Trastuzumab Emtansine · ITT population
| Primary endpoint | Analysis population | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|---|
| IDFS Rate at 3 Years | ITT | Log-rank test | Hazard ratio | 0.54 | 0.44–0.66 | <0.0001 |
The reported HR of 0.54 means that, with trastuzumab as the numerator group in the reported comparison, the estimated instantaneous IDFS event hazard was 54% of that estimated for trastuzumab emtansine under the fitted time-to-event comparison. Expressed as a simple relative complement, this corresponds to a 46% lower estimated hazard for trastuzumab relative to trastuzumab emtansine.
It does not mean that 46% of participants avoided an event, that 46% of participants benefited, or that each participant experienced a 46% reduction in personal risk. A hazard ratio is a relative time-to-event measure, not an absolute risk difference.
The 95% CI of 0.44–0.66 indicates the statistical uncertainty surrounding the estimated hazard ratio. Its entire range is below 1, so the reported interval is consistent with a lower estimated hazard for the numerator group throughout the interval. The width of the interval also matters: it communicates that the estimate is not an exact quantity.
The p-value of <0.0001 addresses evidence against the relevant null hypothesis under the stated statistical test. It does not measure the size of the treatment effect, the probability that the treatment is beneficial, or the clinical importance of the result.
Because the endpoint is time-to-event and the effect measure is a hazard ratio, interpretation should remain tied to the survival-analysis framework. The registry extract in the ClinicalTrials.gov record does not provide additional model details about proportional-hazards diagnostics, so no stronger claim about that assumption is made here.
7. Secondary Results: IDFS Including Second Primary Non-breast Cancer
Three-year analysis
IDFS Including SPNBC at 3 Years
95% CI: 0.46–0.69 · P < .0001
ITT population · Log-rank test · Superiority
The HR of 0.57 indicates a lower estimated event hazard for trastuzumab relative to trastuzumab emtansine in the reported comparison, corresponding to a 43% lower estimated hazard on the relative hazard scale.
The 95% CI of 0.46–0.69 describes uncertainty around that estimate and remains below 1 throughout the reported interval. The p-value of <.0001 is evidence from the reported log-rank analysis against the relevant null hypothesis; it should not be interpreted as a measure of effect magnitude.
Seven- and eight-year analyses
| Endpoint | Time frame | HR | 95% CI | P-value |
|---|---|---|---|---|
| IDFS Including SPNBC Rate at 7 Years | At Year 7 | 0.57 | 0.46–0.69 | <0.0001 |
| IDFS Including SPNBC Rate at 8 Years | At Year 8 | 0.57 | 0.46–0.69 | <.0001 |
The ClinicalTrials.gov record reports the same hazard-ratio estimate and confidence interval at Years 7 and 8 as at Year 3 for this endpoint. These entries should be read as the registry's posted analyses at those specified time frames; they should not be interpreted as proof that the underlying survival curves, event counts, or follow-up distributions are identical.
8. Secondary Results: Disease-free Survival
| Endpoint | Time frame | HR | 95% CI | P-value |
|---|---|---|---|---|
| Disease-free Survival (DFS) Rate at 3 Years | At Year 3 | 0.58 | 0.48–0.70 | <0.0001 |
| DFS Rate at 7 Years | At Year 7 | 0.58 | 0.48–0.70 | <0.0001 |
| DFS Rate at 8 Years | At Year 8 | 0.58 | 0.48–0.70 | <.0001 |
Each DFS analysis was performed in the ITT population using a log-rank test, with hazard ratio as the effect measure and a superiority hypothesis. The registry-reported estimates are below 1, indicating a lower estimated hazard for trastuzumab in the reported trastuzumab-versus-trastuzumab-emtansine ordering.
The HR of 0.58 corresponds to a 42% lower estimated hazard on the relative hazard scale for the numerator group. The 95% CI of 0.48–0.70 communicates uncertainty around that estimate. The p-values of <0.0001 or <.0001, depending on the registry entry, quantify statistical evidence under the reported test rather than the magnitude of the observed effect.
9. Secondary Results: Overall Survival
| Endpoint | Time frame | HR | 95% CI | P-value |
|---|---|---|---|---|
| Overall Survival (OS) Rate at 5 Years | At Year 5 | 0.70 | 0.53–0.91 | 0.0082 |
| OS Rate at 7 Years | At Year 7 | 0.70 | 0.53–0.91 | 0.0082 |
| OS Rate at 8 Years | At Year 8 | 0.70 | 0.53–0.91 | 0.0082 |
Overall survival hazard ratio
95% CI: 0.53–0.91 · P = 0.0082
Reported at Years 5, 7, and 8 · ITT population
An HR of 0.70 means the estimated instantaneous death hazard for trastuzumab was 70% of that for trastuzumab emtansine in the reported comparison. On the relative hazard scale, that corresponds to a 30% lower estimated hazard for the numerator group.
The 95% CI of 0.53–0.91 gives the reported uncertainty interval around the HR. It is narrower than a point estimate alone can communicate and remains below 1 across the interval. The p-value of 0.0082 provides evidence under the reported log-rank analysis but does not say that the effect size is 0.0082, nor does it quantify clinical importance.
OS is a time-to-death endpoint, so the hazard ratio should not be confused with a difference in mortality percentages at a particular time point. The ClinicalTrials.gov record does not provide the corresponding absolute survival percentages, median survival times, event counts, or Kaplan-Meier estimates needed for those additional interpretations.
10. Secondary Results: Distant Recurrence-free Interval
| Endpoint | Time frame | HR | 95% CI | P-value |
|---|---|---|---|---|
| Distant Recurrence-free Interval (DRFI) Rate at 3 Years | At Year 3 | 0.60 | 0.47–0.76 | <0.0001 |
| DRFI Rate at 7 Years | At Year 7 | 0.60 | 0.47–0.76 | <0.0001 |
| DRFI Rate at 8 Years | At Year 8 | 0.60 | 0.47–0.76 | <.0001 |
The DRFI analyses again use the ITT population and log-rank testing. The HR of 0.60 corresponds to a 40% lower estimated hazard for the numerator group on the relative hazard scale. The 95% CI of 0.47–0.76 describes uncertainty around that estimate.
Although several KATHERINE endpoints use the same log-rank and hazard-ratio framework, the endpoints are not interchangeable. IDFS, DFS, OS, and DRFI represent different clinical event definitions. A hazard ratio from one endpoint should therefore be interpreted for that endpoint rather than transferred to another outcome.
11. Other Pre-specified IDFS Results
| Endpoint | Time frame | HR | 95% CI | P-value | Role |
|---|---|---|---|---|---|
| IDFS Rate at 7 Years | At Year 7 | 0.54 | 0.44–0.66 | <0.0001 | Other pre-specified |
| IDFS Rate at 8 Years | At Year 8 | 0.54 | 0.44–0.66 | <.0001 | Other pre-specified |
The registry reports the same HR and 95% confidence interval for IDFS at Years 7 and 8 as for the primary IDFS analysis at Year 3. These results reinforce the importance of distinguishing the endpoint definition from the time frame named in the outcome measure: the registry labels the primary endpoint at Year 3, while the other pre-specified entries are separately labeled at Years 7 and 8.
12. Results Across the Time Frames
This visual is a direct representation of the hazard-ratio estimates reported in the ClinicalTrials.gov record; it is not a reconstruction of Kaplan-Meier curves. The estimates span from 0.54 for IDFS to 0.70 for OS, with the other reported endpoint families between those values.
The most important statistical point is that these estimates should not be treated as five measurements of the same outcome. Each endpoint has its own event definition and time frame. The fact that several estimates are numerically similar does not establish equivalence among the endpoints.
13. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm as affected participants over participants at risk.
| Safety measure | Trastuzumab | Trastuzumab emtansine |
|---|---|---|
| Serious adverse events | 58 / 720 | 94 / 740 |
Trastuzumab
58 of 720 participants were reported as affected by serious adverse events.
Trastuzumab emtansine
94 of 740 participants were reported as affected by serious adverse events.
These figures are presented as reported counts over the corresponding numbers at risk. They should not be converted into a new derived percentage for this page because the data rule for this record is to report numbers exactly as reported in the registry. The safety data also represent a different analytical question from the efficacy hazard ratios: serious adverse events describe an adverse-event outcome, whereas the reported efficacy estimates describe time-to-event treatment comparisons.
14. Statistical Methods Explained
Why was a log-rank test used?
The primary endpoint is a time-to-event outcome, and the registry reports the log-rank test as the analysis method. A log-rank test is designed to compare the event-time distributions of two groups across follow-up while incorporating the timing of observed events and the information contributed by participants who remain event-free during follow-up.
What does the primary hazard ratio of 0.54 mean?
Because the reported comparison is trastuzumab versus trastuzumab emtansine, an HR of 0.54 means that the estimated instantaneous event hazard for the trastuzumab group was 54% of that for the trastuzumab emtansine group under the reported analysis. The corresponding relative complement is a 46% lower estimated hazard. It does not mean a 46% absolute reduction in events.
Why is the analysis population important?
The registry explicitly states that the ITT population includes all participants randomized to the study regardless of whether they received any study treatment. This keeps participants attached to their randomized treatment assignment and preserves the principal comparison created by randomization.
What does the 95% confidence interval tell us?
For the primary HR of 0.54, the 95% CI is 0.44–0.66. The interval describes statistical uncertainty around the estimated hazard ratio under the analysis framework. A confidence interval is more informative than the point estimate alone because it shows the range of effect estimates compatible with the statistical uncertainty represented by the interval.
Why doesn't the p-value measure effect size?
The primary p-value is <0.0001. A p-value quantifies evidence against a null hypothesis under the specified statistical model and test. It is affected by both the magnitude of the observed effect and the amount of information in the data, so it cannot be interpreted as a direct measure of how large or clinically important the treatment effect is.
Why is this a time-to-event analysis rather than a simple percentage comparison?
The registry classifies the primary endpoint as time-to-event. Time-to-event methods preserve information about when events occur rather than treating every participant as if they had exactly the same follow-up duration. The log-rank test and hazard ratio are therefore aligned with the structure of the endpoint.
Why should different endpoints not be collapsed into one result?
IDFS, IDFS including SPNBC, DFS, OS, and DRFI are separately defined outcomes in the registry. Even when they use the same statistical method, their events and clinical meanings differ. A hazard ratio of 0.54 for IDFS and a hazard ratio of 0.70 for OS are therefore separate statistical findings, not interchangeable estimates of one underlying quantity.
15. Understanding the Primary Confidence Interval
The point estimate is 0.54, while the interval communicates the statistical precision around that estimate. The interval should be read as a range describing uncertainty about the estimated treatment effect, not as a range of individual patient outcomes.
The interval is particularly useful because it shows that the primary estimate is not being presented as an exact property of the trial population. Sampling variability means that a future study conducted under the same general design could produce a different estimate. The confidence interval gives a quantitative indication of that uncertainty for the observed analysis.
It is also important that the confidence interval is attached to the hazard ratio. It therefore does not directly answer questions about absolute event rates, absolute risk differences, number needed to treat, or survival probabilities at a particular time. Those require corresponding absolute outcome estimates, which are not reported in the ClinicalTrials.gov record used for this page.
16. Reading the P-values Correctly
| Endpoint family | Reported P-value | What it addresses |
|---|---|---|
| Primary IDFS at 3 Years | <0.0001 | Evidence from the reported log-rank comparison under a superiority hypothesis |
| IDFS including SPNBC | <.0001 or <0.0001 | Evidence from the reported log-rank comparison |
| DFS | <.0001 or <0.0001 | Evidence from the reported log-rank comparison |
| OS | 0.0082 | Evidence from the reported log-rank comparison |
| DRFI | <.0001 or <0.0001 | Evidence from the reported log-rank comparison |
The repeated presence of small p-values does not mean that all endpoints have the same effect size. For example, the OS HR is 0.70, whereas the primary IDFS HR is 0.54. The p-values answer a different question from the hazard ratios.
Hazard ratio → magnitude and direction of relative effect.
Confidence interval → uncertainty and precision around the effect estimate.
P-value → evidence against the null hypothesis under the specified test.
17. Randomization and the ITT Principle
Randomization is the central design feature that allows the trial to compare the two treatment strategies while reducing systematic differences in treatment assignment. The registry identifies the study as randomized and parallel, with two interventions: trastuzumab and trastuzumab emtansine.
The ITT definition posted on ClinicalTrials.gov for the statistical analyses is explicit: all participants randomized to the study are included regardless of whether they received any study treatment. The same ITT framework is reported for the primary endpoint and the registry-reported secondary analyses.
What ITT preserves
The treatment comparison generated by randomization, because participants remain classified according to their randomized assignment.
What ITT does not guarantee
It does not guarantee equal numbers, equal follow-up, or equal event experience between randomized groups.
For time-to-event outcomes, ITT also means that the statistical analysis is not restricted only to participants who completed treatment. That distinction helps prevent post-randomization treatment behavior from redefining the principal efficacy comparison.
18. Why the Hazard Ratio Direction Matters
The analyses posted on ClinicalTrials.gov consistently specify the groups compared as Trastuzumab vs Trastuzumab Emtansine. This ordering is essential when interpreting every HR on the page.
| Endpoint family | HR | Relative interpretation for trastuzumab vs trastuzumab emtansine |
|---|---|---|
| IDFS | 0.54 | 46% lower estimated hazard |
| IDFS including SPNBC | 0.57 | 43% lower estimated hazard |
| DFS | 0.58 | 42% lower estimated hazard |
| OS | 0.70 | 30% lower estimated hazard |
| DRFI | 0.60 | 40% lower estimated hazard |
These are simple interpretations of the reported HRs using 1 − HR. They do not transform the hazard ratios into absolute risk reductions. The distinction is important because a relative hazard reduction and an absolute probability difference are fundamentally different measures.
19. Interpreting the Repeated Long-term Estimates
The ClinicalTrials.gov record contains several endpoint families measured at multiple time points. Some estimates remain numerically identical across those time points. For example, the IDFS including SPNBC HR is reported as 0.57 at Years 3, 7, and 8, with a 95% CI of 0.46–0.69. The DFS HR is likewise reported as 0.58 at Years 3, 7, and 8.
These repeated numbers should be handled carefully. A repeated registry estimate is a reported statistical result; it should not be reverse-engineered into an assumption about the underlying event process. Without individual participant event and censoring information, one cannot reconstruct the Kaplan-Meier curves or independently verify how the estimates were generated.
20. Limitations
- Registry-level detail: the ClinicalTrials.gov extract provides the primary and selected secondary statistical analyses, but it does not provide every component of a full statistical analysis plan.
- No absolute outcome estimates reported: the data contain hazard ratios and confidence intervals but do not provide the corresponding absolute IDFS, DFS, OS, or DRFI rates by arm.
- No median event times reported: median survival or disease-free times are not included in the ClinicalTrials.gov record and therefore are not reported here.
- No Kaplan-Meier curves reported: the numerical hazard ratios do not provide enough information to reconstruct the underlying survival curves.
- Hazard-ratio interpretation: a hazard ratio is a relative time-to-event measure and should not be substituted for absolute risk measures.
- Proportional-hazards considerations: the registry extract does not provide additional diagnostics or model details that would permit a fuller assessment of the assumptions behind a hazard-ratio interpretation.
- Safety comparison: serious adverse events are reported as affected/at-risk counts, but the ClinicalTrials.gov record does not include a formal comparative statistical test for those counts.
- Multiplicity: the ClinicalTrials.gov record identifies a primary endpoint and multiple secondary and other pre-specified analyses, but do not provide enough information to reconstruct a complete multiplicity-control strategy. The page therefore does not assign confirmatory status to individual secondary p-values beyond their registry role.
- Missing-data and imputation methods: the ClinicalTrials.gov record does not describe a missing-data or imputation strategy, so no such method is attributed to the trial here.
21. Why This Trial Matters Statistically
KATHERINE is a useful teaching example because its registry record connects a randomized phase 3 design to a coherent family of time-to-event analyses. The primary endpoint, secondary endpoints, and other pre-specified outcomes are all analyzed within an ITT framework using the log-rank test and hazard ratio.
| Statistical concept | How it appears in KATHERINE |
|---|---|
| Randomization | The study uses randomized allocation with two parallel treatment groups. |
| Intention-to-treat analysis | The reported efficacy population includes all randomized participants regardless of whether they received study treatment. |
| Time-to-event endpoint | The primary endpoint is classified as time-to-event. |
| Log-rank test | The reported comparative method for the primary and registry-reported secondary analyses. |
| Hazard ratio | The reported effect measure across the statistical analyses reported. |
| Confidence interval | 95% intervals quantify uncertainty around the reported hazard ratios. |
| Superiority testing | The primary analysis is explicitly classified under a superiority hypothesis. |
| Multiple endpoints | The registry reports a primary endpoint plus multiple secondary and other pre-specified outcomes. |
| Safety analysis | Serious adverse events are reported by randomized treatment arm as affected/at-risk counts. |
The statistical lesson is broader than the individual HR of 0.54. A clinical trial result becomes interpretable only when the endpoint, analysis population, comparison direction, statistical method, effect measure, confidence interval, and hypothesis framework are read together.
22. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The primary randomized comparison reports an HR of 0.54 with a 95% CI of 0.44–0.66 and P < 0.0001 using a log-rank analysis in the ITT population. The registry classifies the hypothesis as superiority.
Clinical interpretation
The statistical result concerns invasive disease-free survival as defined by the registry. It should be interpreted alongside the endpoint definition, follow-up time, absolute outcome measures when available, and safety information rather than as a standalone measure of overall treatment value.
The distinction is deliberate. Statistical significance describes evidence under a specified model and test. Clinical interpretation requires understanding what outcome was measured, how large the estimated effect was, how precisely it was estimated, and what other evidence is relevant to the treatment comparison.
23. A Compact Statistical Reading of KATHERINE
| Question | Answer from the ClinicalTrials.gov record |
|---|---|
| What was randomized? | Participants were randomized between trastuzumab and trastuzumab emtansine in a parallel phase 3 study. |
| What was the primary endpoint? | IDFS Rate at 3 Years, a time-to-event endpoint. |
| Who was analyzed? | The ITT population, including all randomized participants regardless of whether they received study treatment. |
| How was the primary comparison performed? | Log-rank test. |
| What was the effect measure? | Hazard ratio. |
| What was the primary HR? | 0.54. |
| What was the primary 95% CI? | 0.44–0.66. |
| What was the primary P-value? | <0.0001. |
| What was the hypothesis type? | Superiority. |
| What safety information is reported? | Serious adverse events: 58/720 for trastuzumab and 94/740 for trastuzumab emtansine. |
24. Related Tutorials
Learn more about the methods used in this trial:
25. Related Calculators
26. Sources
- ClinicalTrials.gov: NCT01772472 — KATHERINE.
- Linked publication: PubMed record for PMID 30516102.
- Linked publication: PubMed record for PMID 39813643.
Continue with the statistical methods
Explore the core survival-analysis concepts that connect randomized trials with hazard ratios, confidence intervals, log-rank tests, and time-to-event endpoints.
27. Record Summary
KATHERINE is a randomized phase 3 trial with 1486 enrolled participants and two parallel treatment groups, trastuzumab and trastuzumab emtansine. Its primary endpoint is a time-to-event measure, Invasive Disease-free Survival (IDFS) Rate at 3 Years, analyzed in the ITT population using a log-rank test and a hazard ratio under a superiority hypothesis.
The primary reported HR is 0.54 with a 95% CI of 0.44–0.66 and P < 0.0001. The same general statistical framework is used for the registry-reported secondary analyses of IDFS including SPNBC, DFS, OS, and DRFI, as well as the other pre-specified IDFS analyses. The reported HRs range from 0.54 to 0.70 across these endpoint families.
The statistical interpretation depends on keeping several concepts distinct: the hazard ratio describes relative event hazard, the confidence interval describes uncertainty around that estimate, and the p-value addresses evidence under the specified hypothesis test. None of these quantities alone describes absolute clinical risk or the experience of an individual participant.
The registry also supplies serious adverse-event counts of 58/720 for trastuzumab and 94/740 for trastuzumab emtansine. These safety figures should be considered separately from the efficacy hazard ratios because they represent a different outcome and are not accompanied in the ClinicalTrials.gov record by a formal comparative statistical analysis.