← Clinical Trials
HER2-Positive Breast Cancer Phase 3 Completed NCT03523585

DESTINY-Breast02: Complete Statistical Analysis of Trastuzumab Deruxtecan in HER2-Positive Breast Cancer

An independent statistical review of the randomized phase 3 DESTINY-Breast02 trial evaluating trastuzumab deruxtecan versus treatment of investigator's/physician's choice in participants with HER2-positive, unresectable and/or metastatic breast cancer previously treated with trastuzumab emtansine.

Trial start: 2018-08-01  ·  Primary completion: 2022-06-30  ·  Enrollment: 608
Scope of this record

This page separates reported trial results from statistical interpretation. Trial facts and numerical results on this page are restricted to the ClinicalTrials.gov record data. This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

DESTINY-Breast02 was a randomized, open-label, phase 3 parallel-group trial evaluating trastuzumab deruxtecan in participants with HER2-positive, unresectable and/or metastatic breast cancer previously treated with trastuzumab emtansine. The registry reports 608 participants, three arms, one registered primary time-to-event endpoint, and five posted statistical analyses.

608
Enrollment
Registered participants
3
Arms
Parallel-group design
0.3589
Primary PFS HR
95% CI 0.2840–0.4535
<0.000001
Primary PFS P-value
Two-sided log-rank test
FeatureDESTINY-Breast02
PhasePhase 3
ConditionBreast Cancer
PopulationParticipants with HER2-positive, unresectable and/or metastatic breast cancer previously treated with trastuzumab emtansine
DesignRandomized, parallel-group, open-label
AllocationRandomized
Primary purposeTreatment
Enrollment608
Arms3
Primary endpoint typeTime-to-event
Primary endpointProgression-Free Survival (PFS) Based on Blinded Independent Central Review (BICR)
Hypothesis typeSuperiority
Results postedYes
Lead sponsorDaiichi Sankyo
StatusCOMPLETED

2. Clinical Question

The trial addresses whether trastuzumab deruxtecan improves progression-free survival compared with treatment of investigator's/physician's choice in participants with HER2-positive, unresectable and/or metastatic breast cancer who had previously been treated with trastuzumab emtansine.

Population

Participants with HER2-positive, unresectable and/or metastatic breast cancer previously treated with trastuzumab emtansine.

Intervention

Trastuzumab deruxtecan (T-DXd).

Comparator

Treatment of investigator's/physician's choice (TPC), represented in the registry by trastuzumab+capecitabine and lapatinib+capecitabine treatment arms.

Primary question

Does T-DXd produce a superior progression-free survival outcome compared with TPC when PFS is assessed by blinded independent central review?

3. Trial Design

01
Randomize 608 participants
02
3 Arms Randomized parallel groups
03
Treatment T-DXd or TPC
04
Assessment PFS and other outcomes
05
Analysis Time-to-event and categorical methods
Allocation
Randomized
Design model
Parallel
Masking
None
Primary purpose
Treatment
Phase
Phase 3
Enrollment
608
ARM · T-DXd

Trastuzumab deruxtecan

  • Intervention: trastuzumab deruxtecan (drug)
  • Compared with treatment of investigator's/physician's choice
COMPARATOR · TPC

Treatment of investigator's/physician's choice

  • Trastuzumab (drug)
  • Capecitabine (drug)
  • Lapatinib (drug)

The registry identifies three arms overall. The reported efficacy comparisons presented in the statistical analyses are between trastuzumab deruxtecan and treatment of investigator's/physician's choice rather than a separate three-way statistical comparison.

4. Trial Timeline

2018-08-01

Trial start

The registered trial start date was 2018-08-01.

2022-06-30

Primary completion

The registered primary completion date was 2022-06-30.

30 Jun 2022

Data cutoff for posted efficacy analyses

The Full Analysis Set for the reported PFS, OS, investigator-assessed PFS, and ORR analyses used a data cut-off date of 30 Jun 2022.

5. Endpoints

The registry reports one primary endpoint and four additional statistical analyses. The primary endpoint is a time-to-event measure based on blinded independent central review.

RoleEndpointTime frameType
Primary Progression-Free Survival (PFS) Based on Blinded Independent Central Review (BICR) in Participants With HER2-positive, Unresectable and/or Metastatic Breast Cancer Participants Previously Treated With Trastuzumab Emtansine Baseline up to 46 months postdose Time-to-event
Secondary Overall Survival (OS) in Participants With HER2-positive, Unresectable and/or Metastatic Breast Cancer Participants Previously Treated With Trastuzumab Emtansine Baseline up to 46 months postdose Time-to-event
Secondary Progression-Free Survival (PFS) Based on Investigator Assessment in Participants With HER2-positive, Unresectable and/or Metastatic Breast Cancer Participants Previously Treated With Trastuzumab Emtansine Up to 46 months Time-to-event
Secondary Percentage of Participants With Objective Response Rate (ORR) in Participants With HER2-positive, Unresectable and/or Metastatic Breast Cancer Participants Previously Treated With Trastuzumab Emtansine Baseline up to 46 months postdose Binary

Primary endpoint definition

Progression-free survival by BICR was defined as the time from the date of randomization to the earlier of the dates of the first objective documentation of disease progression, as per RECIST v1.1, or death due to any cause.

Why the definition matters: PFS combines two possible event types—objective disease progression and death from any cause—and therefore requires time-to-event methods rather than a simple comparison of percentages at a fixed visit.

6. Analysis Populations and Stratification

The posted statistical analyses state that PFS, OS, investigator-assessed PFS, and ORR were assessed in the Full Analysis Set at the data cut-off date of 30 Jun 2022. The analysis text also identifies an intention-to-treat framework and stratified analysis.

Analysis featureRegistry-supported detail
Primary PFS populationFull Analysis Set
Secondary OS populationFull Analysis Set
Investigator-assessed PFS populationFull Analysis Set
ORR populationFull Analysis Set
Data cut-off30 Jun 2022
Analysis frameworkIntention-to-treat analysis identified in the analysis text
Stratification factorsHormone receptor status; prior treatment with pertuzumab; history of visceral disease, as defined by IXRS

The same three factors were used in the stratified log-rank testing and in the stratified Cox proportional-hazards model: hormone receptor status, prior treatment with pertuzumab, and history of visceral disease, as defined by the IXRS.

7. Statistical Methodology

Kaplan-Meier estimation

For a time-to-event endpoint such as PFS or OS, Kaplan-Meier estimation is used to describe the estimated event-free probability over time while accounting for participants whose event time is censored. The ClinicalTrials.gov record identifies PFS and OS as time-to-event endpoints and identify the log-rank test and hazard ratio as the principal comparative measures.

Conceptual Kaplan-Meier form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at event time ti, while ni is the number at risk immediately before that time.

Stratified log-rank test

The reported primary PFS analysis used a log-rank test. The registry states that the P-value was based on a log-rank test stratified by hormone receptor status, prior treatment with pertuzumab, and history of visceral disease, as defined by the IXRS.

Stratification is important because the randomized comparison can be evaluated within the levels of prespecified factors rather than treating the trial population as completely homogeneous with respect to those factors.

Stratified Cox proportional-hazards model

The reported hazard ratio was based on a stratified Cox proportional-hazards model using the same three stratification factors. In general, a Cox hazard ratio summarizes the relative instantaneous event rate associated with treatment under the model.

Interpretation of a hazard ratio
HR < 1  →  lower estimated instantaneous event rate in the T-DXd group

The hazard ratio is a relative time-to-event measure. It is not a probability, not a median, not an absolute risk difference, and not the percentage of patients who benefit.

Cochran-Mantel-Haenszel test

The two posted ORR analyses used the Cochran-Mantel-Haenszel test, one for BICR assessment and one for investigator assessment. This is a categorical-data method that can compare treatment groups while accounting for stratification.

Superiority framework

The registry identifies the hypothesis type for the posted analyses as superiority. This differs conceptually from a non-inferiority framework: the objective here is to establish evidence that the treatment comparison favors the intervention rather than merely showing that it remains within a prespecified acceptable loss of efficacy.

8. Primary Result: BICR-Assessed Progression-Free Survival

The primary endpoint was PFS based on blinded independent central review. The Full Analysis Set was evaluated at the 30 Jun 2022 data cut-off. The comparison was T-DXd versus treatment of investigator's/physician's choice.

Hazard ratio for progression or death

0.3589

95% CI: 0.2840–0.4535   ·   P < 0.000001

Two-sided stratified log-rank test; hazard ratio from a stratified Cox proportional-hazards model.

Clinical Biostats interpretation

What the estimate means: A hazard ratio of 0.3589 means that, under the fitted stratified Cox model, the estimated instantaneous rate of the PFS event—objective disease progression or death—was about 35.89% of that in the TPC group. Equivalently, 1 − 0.3589 = 0.6411, so the estimated hazard was approximately 64.11% lower under the model.

What it does not mean: it does not mean that 64.11% of patients avoided progression, that every patient experienced a 64.11% reduction, or that median PFS was reduced or increased by that percentage. A hazard ratio is a relative time-to-event measure, not an absolute probability.

What the confidence interval says: the two-sided 95% confidence interval of 0.2840–0.4535 describes the statistical uncertainty around the estimated hazard ratio under the analysis framework. It is not the range of individual patient outcomes.

Why the P-value is different: the P-value of <0.000001 addresses the evidence against the null hypothesis under the specified test. It does not measure the magnitude of the treatment effect. Effect magnitude is described by the hazard ratio and its confidence interval.

Important cautions: the hazard ratio comes from a stratified Cox model, so its interpretation depends on the model framework and the proportional-hazards interpretation. The analysis also incorporates censoring inherent to time-to-event data. Because the P-value comes from a stratified log-rank test, it should not be interpreted independently of the prespecified stratification factors.

What can and cannot be concluded from the primary result

The reported result provides a treatment comparison for BICR-assessed PFS using the Full Analysis Set, a stratified log-rank test, and a stratified Cox model. The confidence interval lies below 1.0, and the registry identifies the hypothesis as superiority. The ClinicalTrials.gov record does not provide median PFS, Kaplan-Meier survival probabilities at specified time points, or the number of PFS events, so those quantities are not reported here.

Educational note: a valid Kaplan-Meier curve cannot be reconstructed from the reported hazard ratio, confidence interval, and P-value alone. Underlying event and censoring information would be required for a quantitative reconstruction.

9. Secondary Result: Overall Survival

Overall survival was assessed in the Full Analysis Set at the 30 Jun 2022 data cut-off. The reported comparison was T-DXd versus treatment of investigator's/physician's choice.

Hazard ratio for overall survival

0.6575

95% CI: 0.5023–0.8605   ·   P = 0.0021

Two-sided stratified log-rank test; hazard ratio from a stratified Cox proportional-hazards model.

Clinical Biostats interpretation

What the estimate means: A hazard ratio of 0.6575 indicates that the estimated instantaneous rate of death in the T-DXd group was about 65.75% of that in the TPC group under the fitted stratified Cox model. The corresponding relative reduction in estimated hazard is approximately 34.25%.

What it does not mean: the estimate is not a 34.25% absolute improvement in survival, and it does not mean that 34.25% of participants lived longer. It summarizes a relative time-to-event comparison under the model.

Precision: the 95% confidence interval of 0.5023–0.8605 gives the uncertainty interval for the estimated hazard ratio. It does not describe the distribution of individual survival times.

P-value: the reported P = 0.0021 quantifies evidence against the null hypothesis under the two-sided stratified log-rank framework. It does not quantify the size or clinical importance of the effect.

Cautions: OS is affected by the complete treatment pathway after randomization, including events and censoring over follow-up. The ClinicalTrials.gov record does not report median OS, survival probabilities at fixed times, or subsequent-treatment patterns, so those measures cannot be added to this analysis.

10. Secondary Result: Investigator-Assessed Progression-Free Survival

PFS was also evaluated by investigator assessment. The Full Analysis Set was analyzed at the 30 Jun 2022 data cut-off using a stratified log-rank test and a stratified Cox proportional-hazards model.

Hazard ratio for investigator-assessed PFS

0.2828

95% CI: 0.2270–0.3524   ·   P < 0.000001

Two-sided stratified log-rank test; hazard ratio from a stratified Cox proportional-hazards model.

Clinical Biostats interpretation

What the estimate means: A hazard ratio of 0.2828 corresponds to an estimated instantaneous rate of progression or death that was approximately 28.28% of the rate in the TPC group under the stratified Cox model. The corresponding relative reduction in estimated hazard is approximately 71.72%.

What it does not mean: it does not mean that 71.72% of participants were progression-free, nor that each participant experienced the same relative reduction in event risk.

Precision: the 95% confidence interval of 0.2270–0.3524 describes uncertainty around the hazard ratio. The ClinicalTrials.gov record does not report median investigator-assessed PFS or event counts.

P-value: P < 0.000001 is evidence against the null under the specified stratified log-rank test; it should not be treated as a measure of effect magnitude.

Assessment distinction: this endpoint differs from the primary BICR-assessed PFS endpoint because the disease assessment source is investigator assessment rather than blinded independent central review. The two analyses therefore answer closely related but not identical measurement questions.

11. Secondary Results: Objective Response Rate

The registry contains two statistical analyses for objective response rate: one based on blinded independent central review and one based on investigator assessment. Both compare T-DXd with treatment of investigator's/physician's choice and both use the Full Analysis Set at the 30 Jun 2022 data cut-off.

ORR analysisAssessmentMethodP-valueHypothesis
Objective Response Rate BICR assessment Cochran-Mantel-Haenszel test <0.0001 Superiority
Objective Response Rate Investigator assessment Cochran-Mantel-Haenszel test <0.0001 Superiority

The ClinicalTrials.gov record does not provide the ORR estimates, response counts, confidence intervals, or response percentages by treatment group. Accordingly, this page reports only the posted statistical comparisons and does not construct an ORR estimate from the P-value.

Statistical distinction: a categorical endpoint such as ORR is analyzed differently from PFS or OS. The Cochran-Mantel-Haenszel procedure evaluates the categorical treatment comparison while accounting for the stratification structure. A P-value alone cannot tell the reader how large the difference in response rates was.

12. Safety: Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by arm as affected participants divided by participants at risk. These counts should be interpreted as safety-event counts rather than efficacy outcomes.

ArmSerious adverse eventsInterpretation of reported count
Trastuzumab Deruxtecan (T-DXd)103/404103 affected participants among 404 at risk
Trastuzumab + Capecitabine19/8719 affected participants among 87 at risk
Lapatinib + Capecitabine27/10827 affected participants among 108 at risk

The denominators used for the serious-adverse-event reporting differ from the overall enrollment of 608. Therefore, these safety counts should not be interpreted as if 404, 87, and 108 were the original randomized arm sizes unless the registry explicitly establishes that correspondence. The ClinicalTrials.gov record identifies them as affected/at-risk counts.

Safety interpretation: the serious-adverse-event ClinicalTrials.gov record support reporting the affected and at-risk counts by arm. They do not provide confidence intervals, comparative P-values, exposure-adjusted incidence rates, or a formal statistical comparison of serious adverse events. Those quantities should not be inferred.

13. Statistical Methods Explained

Why was a log-rank test used for the primary PFS endpoint?

PFS is a time-to-event endpoint, so participants can experience an event at different times and some participants may be censored before experiencing progression or death. The log-rank test is designed to compare event-time distributions between randomized groups while incorporating the timing of events and censoring rather than reducing the endpoint to a single fixed-time proportion.

Why was the log-rank test stratified?

The registry specifies stratification by hormone receptor status, prior treatment with pertuzumab, and history of visceral disease as defined by the IXRS. Stratification allows the treatment comparison to account for these prespecified factors in the time-to-event test rather than ignoring their role in the randomized design.

What does a hazard ratio of 0.3589 mean?

For the primary BICR-assessed PFS analysis, 0.3589 is the estimated ratio of the instantaneous PFS-event rates between T-DXd and TPC under the stratified Cox model. A value below 1 indicates a lower estimated event hazard in the T-DXd group. It is not a percentage of patients who benefit and is not a direct statement about median PFS.

Why is the confidence interval important?

The 95% confidence interval shows the statistical precision of the estimated hazard ratio. For the primary PFS result, the interval is 0.2840–0.4535. Reporting the interval alongside 0.3589 communicates more information than the point estimate alone because it shows the range of values compatible with the statistical uncertainty represented by the model and sample.

Why does the P-value not measure effect size?

The P-value depends on both the observed treatment difference and the amount of statistical information available. A very small P-value can accompany either a large or a relatively small effect when the sample and event information are sufficient. The hazard ratio and confidence interval are therefore needed to describe the magnitude and precision of the time-to-event effect.

Why is BICR assessment different from investigator assessment?

The primary PFS endpoint uses blinded independent central review, while a secondary PFS analysis uses investigator assessment. Independent central review can provide a standardized external assessment of radiographic progression, whereas investigator assessment reflects the treating site's assessment. Comparing the two analyses can therefore provide information about the robustness of the time-to-event finding, but they are not literally the same endpoint measurement process.

Why is ORR analyzed with the Cochran-Mantel-Haenszel test?

ORR is a binary endpoint: participants are classified according to whether they meet the response definition. The Cochran-Mantel-Haenszel test is appropriate for comparing categorical outcomes while accounting for stratification. This differs from the log-rank and Cox approaches used for endpoints where the timing of events is part of the outcome.

14. The Role of the Full Analysis Set and Intention-to-Treat Analysis

The registry states that the efficacy analyses were conducted in the Full Analysis Set and identifies intention-to-treat analysis among the concepts present in the analysis text. The central statistical advantage of an intention-to-treat framework is that treatment groups remain anchored to randomized assignment rather than being redefined according to treatment received or later events.

Why randomization matters

Randomization establishes the treatment groups before outcomes occur, supporting a causal comparison when the trial is properly conducted and analyzed according to the randomized assignment.

Why analysis population matters

The interpretation of an effect estimate depends on which participants are included. The registry explicitly identifies the Full Analysis Set for the posted efficacy analyses.

The ClinicalTrials.gov record does not define the Full Analysis Set in greater detail, so no additional inclusion or exclusion rules are added here.

15. Stratification in This Trial

The posted PFS and OS analyses used three stratification factors: hormone receptor status, prior treatment with pertuzumab, and history of visceral disease, as defined by the IXRS.

Stratification factorHow it enters the analysis
Hormone receptor statusIncluded in the stratified log-rank comparison and stratified Cox model
Prior treatment with pertuzumabIncluded in the stratified log-rank comparison and stratified Cox model
History of visceral diseaseIncluded in the stratified log-rank comparison and stratified Cox model

Stratification should not be confused with a subgroup analysis. A stratification factor is incorporated into the primary comparison framework; a subgroup analysis instead asks how the treatment effect behaves within a particular subset. The ClinicalTrials.gov record does not report separate subgroup hazard ratios for these factors.

16. Multiplicity and Interpretation of Multiple Endpoints

The registry contains one registered primary endpoint and four posted statistical analyses: the primary BICR-assessed PFS analysis, secondary OS, secondary investigator-assessed PFS, and two ORR analyses. The ClinicalTrials.gov record identifies all posted analyses as superiority tests, but they do not provide a multiplicity-adjustment scheme or an alpha-allocation hierarchy.

Endpoint / analysisRoleReported methodReported P-value
BICR-assessed PFSPrimaryStratified log-rank<0.000001
OSSecondaryStratified log-rank0.0021
Investigator-assessed PFSSecondaryStratified log-rank<0.000001
ORR, BICR assessmentSecondaryCochran-Mantel-Haenszel<0.0001
ORR, investigator assessmentSecondaryCochran-Mantel-Haenszel<0.0001
Multiplicity caution: the ClinicalTrials.gov record does not specify an alpha-spending procedure, hierarchical testing sequence, or other multiplicity adjustment for the secondary endpoints. Therefore, the individual reported P-values should be presented as reported rather than assuming an unstated familywise-error strategy.

17. What the Primary Hazard Ratio Does — and Does Not — Mean

Effect size

The primary PFS hazard ratio of 0.3589 is a relative measure of the instantaneous rate of progression or death under the stratified Cox model. Expressed as a simple complement, 1 − 0.3589 corresponds to an approximately 64.11% lower estimated hazard for T-DXd relative to TPC.

Not an absolute benefit

The hazard ratio does not tell us how many additional months an individual patient will remain progression-free, what proportion of patients will avoid progression, or what the absolute difference in survival probability is at a particular time. Those questions require absolute time-to-event estimates, which are not reported in the ClinicalTrials.gov record.

Confidence interval

The 95% CI of 0.2840–0.4535 provides a measure of uncertainty around the primary hazard-ratio estimate. A narrower interval would indicate greater statistical precision, while a wider interval would indicate more uncertainty. The interval should not be interpreted as a range containing the effects for individual patients.

P-value

The P-value of <0.000001 describes the evidence against the null hypothesis under the specified two-sided stratified log-rank test. It is not a probability that the null hypothesis is true, and it is not a measure of the size of the treatment effect.

18. Limitations

19. Why This Trial Matters Statistically

DESTINY-Breast02 provides a useful teaching example because the same randomized comparison illustrates several distinct statistical ideas: time-to-event analysis, stratification, hazard ratios, central versus investigator assessment, categorical-response analysis, and the distinction between a primary endpoint and secondary analyses.

ConceptHow it appears in DESTINY-Breast02
RandomizationRandomized phase 3 parallel-group trial with 608 enrolled participants
Time-to-event analysisPFS and OS were analyzed as time-to-event outcomes
Kaplan-Meier estimationAppropriate framework for describing PFS and OS over follow-up
Log-rank testingUsed for the reported PFS and OS comparisons
Stratified analysisHormone receptor status, prior pertuzumab treatment, and history of visceral disease were incorporated into the reported stratified analyses
Cox modelUsed to estimate hazard ratios for PFS and OS
Hazard ratioPrimary PFS HR 0.3589; OS HR 0.6575; investigator-assessed PFS HR 0.2828
Confidence interval95% two-sided intervals reported for the three time-to-event hazard ratios
Cochran-Mantel-Haenszel testUsed for the two posted ORR analyses
Central versus investigator assessmentBoth BICR and investigator-assessed PFS were reported
Superiority testingThe posted analyses are identified with a superiority hypothesis type
Safety analysisSerious adverse events are reported by arm using affected/at-risk counts

20. Related Tutorials

Learn more about the methods used in this trial:

21. Related Statistical Calculators

22. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical concepts behind randomized time-to-event trials, stratified comparisons, hazard ratios, categorical endpoints, and confidence intervals.

23. Record Summary

DESTINY-Breast02 illustrates how a randomized phase 3 trial can use a time-to-event primary endpoint to compare treatment groups while accounting for censoring and prespecified stratification. The primary BICR-assessed PFS analysis used a stratified log-rank test and a stratified Cox proportional-hazards model, producing a hazard ratio of 0.3589 with a two-sided 95% confidence interval of 0.2840–0.4535 and a P-value of <0.000001. Secondary analyses extended the statistical framework to OS, investigator-assessed PFS, and ORR, while the ClinicalTrials.gov record provides serious-adverse-event counts by arm.

The statistical story is therefore broader than a single P-value. Proper interpretation requires distinguishing the primary endpoint from secondary endpoints, BICR from investigator assessment, relative hazard from absolute event probability, confidence intervals from P-values, and efficacy analysis from safety reporting.

Clinical Biostats methodology: A trial-results page should not merely repeat a registry record. The goal is to explain how the endpoint definitions, randomized design, analysis populations, stratification, statistical tests, effect measures, uncertainty, and limitations fit together while keeping reported evidence separate from statistical interpretation.