← Clinical Trials
HER2-Positive Breast Cancer Phase 3 Time-to-Event Analysis NCT03975647

HER2CLIMB-02: Complete Statistical Analysis of Tucatinib in HER2-Positive Breast Cancer

An independent statistical analysis of the randomized phase 3 HER2CLIMB-02 trial evaluating tucatinib versus placebo in combination with ado-trastuzumab emtansine (T-DM1) for patients with advanced or metastatic HER2-positive breast cancer.

Trial start: 2019-10-02  ·  Primary completion: 2023-06-29  ·  Enrollment: 466
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

HER2CLIMB-02 was a randomized, parallel, quadruple-masked phase 3 trial evaluating tucatinib versus placebo, each in combination with ado-trastuzumab emtansine (T-DM1), in patients with advanced or metastatic HER2-positive breast cancer. The registered primary endpoint was investigator-assessed progression-free survival (PFS) by RECIST v1.1.

466
Enrolled
2 treatment arms
3
Phase
Randomized phase 3
0.759
Primary PFS HR
95% CI 0.607–0.950
0.0163
P-value
Log-rank analysis
FeatureHER2CLIMB-02
Trial nameHER2CLIMB-02
ClinicalTrials.gov identifierNCT03975647
PhasePhase 3
ConditionHER2-positive breast cancer
PopulationPatients with advanced or metastatic HER2-positive breast cancer
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment466
Primary endpointInvestigator-assessed PFS according to RECIST v1.1
Lead sponsorSeagen, a wholly owned subsidiary of Pfizer
Trial statusActive, not recruiting

2. Clinical Question

The central question was whether adding tucatinib to ado-trastuzumab emtansine could improve investigator-assessed progression-free survival compared with ado-trastuzumab emtansine plus placebo in patients with advanced or metastatic HER2-positive breast cancer.

Population

Patients with advanced or metastatic HER2-positive breast cancer.

Intervention

Tucatinib in combination with ado-trastuzumab emtansine (T-DM1).

Comparator

Placebo in combination with ado-trastuzumab emtansine (T-DM1).

Primary question

Does tucatinib plus T-DM1 produce a different progression-free survival experience from placebo plus T-DM1?

3. Trial Design

01
Randomize 466 participants
02
Parallel arms 2 treatment groups
03
Quadruple masking Registry design designation
04
Assess PFS Investigator RECIST v1.1
05
Compare Log-rank and Cox methods
TREATMENT ARM

Tucatinib + T-DM1

  • Tucatinib
  • Ado-trastuzumab emtansine (T-DM1)
  • Randomized comparison against placebo + T-DM1
COMPARATOR ARM

Placebo + T-DM1

  • Placebo
  • Ado-trastuzumab emtansine (T-DM1)
  • Randomized comparison against tucatinib + T-DM1
Design structure. The registry describes HER2CLIMB-02 as randomized, parallel, quadruple-masked, and treatment-purpose phase 3 research. The ClinicalTrials.gov record does not report a crossover design, factorial structure, non-inferiority margin, or Bayesian analysis.

4. Randomization and Analysis Population

The primary efficacy analysis used the intention-to-treat (ITT) analysis set. The registry defines this population as all participants who were randomized on or before the date of last patient in (LPI), in the global study.

Analysis featureRegistry information
Primary efficacy populationIntention-to-treat analysis set
ITT definitionAll participants randomized on or before the date of LPI in the global study
Primary comparisonTucatinib + ado-trastuzumab emtansine vs placebo + ado-trastuzumab emtansine
Primary endpoint typeTime-to-event
Primary analysis methodLog-rank test
Hazard-ratio modelCox proportional hazards model

The ITT framework is important because treatment groups are compared according to randomized assignment rather than being redefined according to subsequent treatment exposure. This preserves the treatment contrast created by randomization for the efficacy analysis.

5. Primary Endpoint

EndpointRegistry definition / time frameAnalysis
Progression-Free Survival (PFS) as Per Response Evaluation Criteria in Solid Tumors (RECIST) Version (v)1.1 Based on Investigator Assessment Time from the date of randomization to the investigator assessment of disease progression (PD) as per RECIST v1.1 or death from any cause, whichever occurred first. The registry time frame is from randomization to PD or death from any cause or censoring, whichever occurred first. Log-rank test; hazard ratio calculated from Cox proportional hazards model

The registry-reported endpoint definition specifies progression under RECIST v1.1 as at least a 20% increase in the sum of diameters of target lesions, taking as reference the smallest sum on study, with the registry definition continuing beyond the registry-reported text. The endpoint is therefore a time-to-event measure rather than a simple proportion of patients who progressed.

Why PFS is a time-to-event endpoint
Time origin = randomization  →  event = progression or death

Participants who have not experienced progression or death by the relevant follow-up point can contribute censored observations. This allows the analysis to incorporate both the timing of events and incomplete follow-up rather than reducing the endpoint to a simple yes/no status at one arbitrary date.

6. Primary PFS Result

The posted primary analysis compared tucatinib plus T-DM1 with placebo plus T-DM1 in the ITT population using a log-rank test. The hazard ratio was calculated from a Cox proportional hazards model.

Hazard ratio for progression or death

0.759

95% CI: 0.607–0.950   ·   P = 0.0163

Analysis: log-rank test; hazard ratio from Cox proportional hazards model.

Clinical Biostats interpretation

An HR of 0.759 means that, under the fitted proportional-hazards model, the estimated instantaneous rate of progression or death in the tucatinib + T-DM1 group was approximately 24.1% lower than in the placebo + T-DM1 group. That is a relative hazard interpretation, not a statement that 24.1% of patients avoided progression or death.

The HR does not mean that every patient experienced a 24.1% reduction in risk, nor does it give the absolute difference in the probability of progression or death at a particular time point. Absolute risks and survival probabilities require the underlying time-to-event distribution.

The two-sided 95% confidence interval of 0.607–0.950 describes statistical uncertainty around the estimated hazard ratio under the model and analysis framework. Because the interval lies below 1, the reported estimate is below the null value of equal hazards throughout that interval.

The P-value of 0.0163 addresses the compatibility of the observed test statistic with the null hypothesis under the specified testing framework. It does not measure the size of the treatment effect. Effect size is described by the HR, while precision is described by the confidence interval.

The Cox interpretation also depends on the proportional-hazards model used to obtain the HR. A single HR summarizes a relative event-rate contrast over follow-up and does not by itself describe the complete shape of the two survival curves.

Reading the primary result together

ComponentReported valueWhat it tells us
Point estimateHR 0.759Estimated relative event-rate contrast between the randomized groups
95% CI0.607–0.950Precision and uncertainty around the HR estimate
P-value0.0163Evidence from the reported log-rank test under its specified testing framework
Analysis populationITTParticipants analyzed according to randomized treatment assignment
Model for HRCox proportional hazardsModel used to calculate the reported hazard ratio
Educational note: the ClinicalTrials.gov record provides a hazard ratio, confidence interval, and P-value, but do not provide the underlying event and censoring records needed to reconstruct a Kaplan-Meier curve. This page therefore does not fabricate a survival curve from summary statistics.

7. Covariates in the Cox Model

The registry analysis notes state that the hazard ratio was calculated from a Cox proportional hazards model. The registry-reported analysis notes identify the following comparison variables in the model:

Model variableComparison stated in the registry analysis notes
Line of treatment for metastatic diseaseFirst versus other
Hormone receptor statusNegative versus positive
Presence or history of brain metastasesYes versus no
ECOG status at randomization0,1

These variables are part of the reported Cox-model analysis notes. They should not be interpreted as evidence that each variable independently modifies the treatment effect. The ClinicalTrials.gov record does not report interaction tests for these factors.

8. Secondary Endpoint: PFS in Participants With Brain Metastases at Baseline

A secondary time-to-event analysis evaluated progression-free survival in participants with brain metastases at baseline, using investigator assessment according to RECIST v1.1. The comparison remained tucatinib + T-DM1 versus placebo + T-DM1 and used a log-rank test in the ITT framework.

Hazard ratio for PFS in participants with brain metastases at baseline

0.639

95% CI: 0.459–0.891   ·   P = 0.0078

Analysis: log-rank test; hazard ratio from Cox proportional hazards model.

Clinical Biostats interpretation

An HR of 0.639 corresponds to an estimated instantaneous progression-or-death rate approximately 36.1% lower for tucatinib + T-DM1 than for placebo + T-DM1 under the fitted model.

The 95% CI of 0.459–0.891 describes the uncertainty around this subgroup estimate. It does not mean that individual patients' effects must fall within this numerical range.

The P-value of 0.0078 is evidence from the reported log-rank analysis under its specified framework. It is not a measure of how large the treatment effect is, and it should not be used by itself to compare the subgroup result with the overall PFS result.

Because this is a subgroup defined by baseline brain metastases, the estimate is based on a more restricted population than the primary ITT analysis. A difference between a subgroup estimate and the overall estimate does not, by itself, establish treatment-effect heterogeneity. A formal interaction analysis would be needed to support such a claim, and the ClinicalTrials.gov record does not report one.

Secondary analysis componentReported information
EndpointPFS as per RECIST v1.1 in participants with brain metastases at baseline
PopulationITT analysis set; subgroup with brain metastases at baseline
ComparisonTucatinib + T-DM1 vs placebo + T-DM1
MethodLog-rank test
Effect measureHazard ratio
Estimate0.639
95% CI0.459–0.891
P-value0.0078

9. Secondary Endpoint: Objective Response Rate

Objective response rate (ORR) as per RECIST v1.1 based on investigator assessment was another posted secondary outcome. Unlike PFS, ORR is a binary endpoint rather than a time-to-event endpoint.

EndpointComparisonMethodP-value
Objective Response Rate (ORR) as Per RECIST v1.1 Based on Investigator Assessment Tucatinib + T-DM1 vs placebo + T-DM1 Cochran-Mantel-Haenszel test 0.2055
Clinical Biostats interpretation

The registry reports a Cochran-Mantel-Haenszel P-value of 0.2055 for the ORR comparison. The ClinicalTrials.gov record does not provide the treatment-group response percentages, response counts, confidence interval, or an effect estimate for ORR.

Therefore, this page does not manufacture an ORR difference or an odds ratio. The reported information establishes the analysis method and P-value, but not a numerical magnitude of the between-group response difference.

The use of the Cochran-Mantel-Haenszel test is consistent with a categorical comparison in which the analysis accounts for the factors specified in the analysis framework. The ClinicalTrials.gov record identifies the method but do not provide enough information to reproduce the calculation.

10. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized arm. These counts are presented as affected participants divided by participants at risk.

Safety measurePlacebo + T-DM1Tucatinib + T-DM1
Serious adverse events 52 / 233 70 / 231
Interpretation of the safety counts: the ClinicalTrials.gov record reports affected participants and participants at risk, but do not provide a statistical comparison, confidence interval, P-value, event-level definition, or exposure-adjusted rate for serious adverse events. The counts therefore should be described as reported rather than converted into an inferential treatment comparison.

Safety and efficacy answer different statistical questions. The primary PFS analysis asks whether the time-to-progression-or-death experience differs between randomized groups. Serious adverse-event counts describe an important safety dimension but do not provide a direct measure of PFS benefit or harm.

11. Statistical Methodology

Log-rank test

The primary PFS analysis used a log-rank test. The log-rank framework is designed for comparing time-to-event distributions between groups while accounting for the timing of events and censored observations.

Conceptual comparison
H0: survival experience is the same between randomized groups

The reported P-value comes from the statistical comparison of the observed event-time experience between the randomized treatment groups. It is distinct from the magnitude of the hazard ratio.

Cox proportional hazards model

The registry states that the hazard ratio was calculated from a Cox proportional hazards model. The Cox model provides a relative hazard measure while allowing the baseline hazard to remain unspecified.

Interpretation of the hazard ratio
HR < 1  →  lower estimated instantaneous event rate in the treatment group

For HER2CLIMB-02, the reported primary HR of 0.759 is a model-based relative measure of progression or death. It is not a probability, an absolute risk difference, or the percentage of patients who benefit.

Intention-to-treat analysis

The primary analysis used the ITT analysis set. In a randomized trial, this approach maintains the randomized treatment assignment as the basis for efficacy comparison. It helps preserve the causal contrast generated by randomization even when post-randomization events occur.

Cochran-Mantel-Haenszel test

The secondary ORR analysis used the Cochran-Mantel-Haenszel test. This is a method for categorical data that can combine information across strata rather than simply treating the data as one unstratified two-group table.

Kaplan-Meier estimation

Kaplan-Meier estimation is the standard descriptive framework for displaying time-to-event distributions such as PFS. The ClinicalTrials.gov record identifies Kaplan-Meier estimation as part of the trial's statistical learning pathway, but the posted primary ClinicalTrials.gov record does not provide a Kaplan-Meier estimate, median PFS, or survival probabilities. Those values are therefore not reported on this page.

12. Statistical Methods Explained

Why was a log-rank test used for the primary endpoint?

PFS is a time-to-event endpoint. Patients can experience progression or death at different times, while some observations can be censored. A log-rank test is designed to compare the event-time distributions between randomized groups rather than reducing follow-up to a single binary outcome.

What does an HR of 0.759 mean?

The HR summarizes the relative instantaneous event rate under the Cox model. An HR of 0.759 corresponds to an estimated 24.1% lower instantaneous rate of progression or death in the tucatinib + T-DM1 group relative to the placebo + T-DM1 group. It does not mean that 24.1% of patients avoided an event.

Why is the confidence interval important?

The point estimate alone does not show how precisely the treatment effect has been estimated. The 95% CI of 0.607–0.950 shows the uncertainty surrounding the primary HR within the reported statistical framework. A confidence interval is therefore more informative than the point estimate alone when assessing precision.

Why does the P-value not measure effect size?

The P-value of 0.0163 describes the evidence from the reported log-rank test against its null hypothesis. It is affected by the magnitude of the observed difference, variability, information available, and sample size. The HR is the effect-size measure; the confidence interval describes its precision.

Why is the brain-metastases analysis not automatically a test of treatment heterogeneity?

The brain-metastases subgroup has its own HR of 0.639, but comparing that HR with the overall HR is not the same as formally testing whether the treatment effect differs according to brain-metastasis status. A formal interaction analysis would be needed to establish effect modification, and the ClinicalTrials.gov record does not report such an interaction test.

Why are PFS and ORR analyzed differently?

PFS records the timing of progression or death and is therefore a time-to-event endpoint. ORR is a binary response outcome. The registry consequently reports a log-rank test for PFS and a Cochran-Mantel-Haenszel test for ORR.

13. Confidence Intervals and Statistical Precision

The primary PFS estimate illustrates why a clinical-trial result should not be summarized by a P-value alone.

AnalysisEstimate95% CIP-value
Primary PFSHR 0.7590.607–0.9500.0163
PFS with baseline brain metastasesHR 0.6390.459–0.8910.0078
ORRNot reported in the ClinicalTrials.gov recordNot reported in the ClinicalTrials.gov record0.2055

The two PFS analyses both have confidence intervals entirely below 1. However, the subgroup estimate has a wider interval relative to the overall primary analysis, reflecting the greater uncertainty associated with analyzing a restricted population. The registry does not provide the underlying event counts needed to further quantify that precision.

A practical reading rule

Estimate: What is the estimated treatment effect?

Confidence interval: How precisely has that effect been estimated?

P-value: How does the observed test statistic compare with the specified null hypothesis under the statistical testing framework?

These three quantities answer related but different questions and should not be treated as interchangeable.

14. Multiplicity and Hypothesis Testing

The ClinicalTrials.gov record identifies the primary endpoint as PFS and report three statistical analyses: one primary PFS analysis, one secondary PFS analysis in participants with brain metastases at baseline, and one secondary ORR analysis.

AnalysisRoleReported methodReported P-value
Investigator-assessed PFSPrimaryLog-rank test0.0163
PFS in participants with brain metastases at baselineSecondaryLog-rank test0.0078
Objective response rateSecondaryCochran-Mantel-Haenszel test0.2055
Multiplicity caution: the ClinicalTrials.gov record does not state a multiplicity-adjustment procedure, alpha allocation, hierarchical testing strategy, or formal hypothesis hierarchy beyond identifying the endpoint roles. The individual reported P-values should therefore not be assigned an unreported multiplicity interpretation.

Similarly, the ClinicalTrials.gov record does not report an interim-analysis strategy or alpha-spending method. Those design elements should not be inferred simply because the study was a phase 3 randomized trial.

15. Censoring and the Proportional-Hazards Assumption

Time-to-event analysis requires careful handling of participants whose event has not been observed by the end of their available follow-up. The registered PFS endpoint explicitly includes a censoring date in its time frame. This is why PFS analysis cannot be reproduced from treatment-group sample sizes alone.

The Cox proportional-hazards model adds another layer of interpretation. Its hazard-ratio summary is based on a proportional-hazards framework. If the relative hazard changes materially over time, a single HR can compress a more complicated pattern into one number.

Censoring

Censored observations contribute information up to the point at which their event status is no longer observed under the analysis framework.

Proportional hazards

The Cox HR provides a concise relative event-rate summary under its proportional-hazards modeling framework.

The ClinicalTrials.gov record does not report a formal diagnostic for the proportional-hazards assumption, so no conclusion about that assumption is made here.

16. Missing Data and Imputation

The ClinicalTrials.gov record does not report a specific missing-data or imputation method for the posted analyses. For that reason, this page does not attribute a particular imputation procedure to HER2CLIMB-02.

For a time-to-event endpoint, censoring is not equivalent to ordinary missing-data imputation. A participant without an observed event at the relevant follow-up point can contribute information through the censoring time without assigning an artificial progression date.

Registry-data boundary: the absence of a registry-reported imputation method should not be interpreted as evidence that no missing-data procedures existed in the underlying protocol or statistical analysis plan. It means only that such a method is not reported in the ClinicalTrials.gov record.

17. Interim Analysis, Non-Inferiority, Crossover, and Bayesian Methods

The ClinicalTrials.gov record supports a conventional superiority-style time-to-event interpretation of the posted HRs, but it does not state a formal hypothesis type beyond Other / not stated.

Design topicWhat the ClinicalTrials.gov record supports
Non-inferiority marginNot reported in the ClinicalTrials.gov record
CrossoverNot reported in the ClinicalTrials.gov record
Factorial designNot reported; design model is parallel
Interim analysisNot reported in the ClinicalTrials.gov record
Alpha-spending procedureNot reported in the ClinicalTrials.gov record
Bayesian analysisNot reported in the ClinicalTrials.gov record
Multiplicity adjustmentNot reported in the ClinicalTrials.gov record

This distinction matters because statistical methods should be reconstructed from documented trial information rather than inferred from the treatment effect. A phase 3 trial can incorporate many different approaches to interim monitoring, multiplicity, or hypothesis testing; the ClinicalTrials.gov record does not establish those features for HER2CLIMB-02.

18. Trial Timeline

2019-10-02

Trial start

The registry lists October 2, 2019 as the study start date.

2023-06-29

Primary completion

The registry lists June 29, 2023 as the primary completion date.

Current registry status

Active, not recruiting

the ClinicalTrials.gov record lists the study as ACTIVE_NOT_RECRUITING.

19. What the Primary Hazard Ratio Does — and Does Not — Mean

Effect size

The primary HR of 0.759 indicates an estimated relative reduction in the instantaneous rate of progression or death of approximately 24.1% for tucatinib + T-DM1 relative to placebo + T-DM1 under the fitted Cox model.

This is not a statement that 24.1% of participants benefited, that 24.1% fewer participants eventually progressed, or that every participant experienced the same proportional reduction.

Precision

The 95% CI of 0.607–0.950 communicates uncertainty around the estimated HR. It is not a range of individual patient outcomes and should not be interpreted as saying that the true effect for each patient lies within that interval.

Statistical evidence

The reported two-sided P-value of 0.0163 comes from the log-rank analysis. It is not a measure of clinical importance or treatment-effect magnitude. The HR and its confidence interval provide the principal quantitative description of the reported relative treatment effect.

20. Primary and Secondary Analyses in Context

Primary PFS

The primary endpoint produced an HR of 0.759 with a 95% CI of 0.607–0.950 and a reported P-value of 0.0163.

Brain-metastases PFS

The baseline brain-metastases subgroup produced an HR of 0.639 with a 95% CI of 0.459–0.891 and a reported P-value of 0.0078.

ORR

The registry reports a Cochran-Mantel-Haenszel P-value of 0.2055 but does not provide an ORR effect estimate in the ClinicalTrials.gov record.

Safety

Serious adverse events were reported as 52/233 in the placebo + T-DM1 arm and 70/231 in the tucatinib + T-DM1 arm.

These analyses should not be collapsed into a single numerical "trial score." PFS, ORR, and safety describe different dimensions of the randomized comparison, and the registry supplies different levels of quantitative detail for each.

21. Important Limitations and Interpretation Issues

22. Why This Trial Matters Statistically

HER2CLIMB-02 is a useful statistical teaching case because it connects a randomized phase 3 design with a time-to-event primary endpoint, an ITT analysis population, a log-rank comparison, a Cox-model hazard ratio, a baseline brain-metastases subgroup analysis, and a categorical ORR analysis using the Cochran-Mantel-Haenszel method.

ConceptHow it appears in HER2CLIMB-02
RandomizationRandomized phase 3 parallel-group design
BlindingQuadruple masking
ITT analysisPrimary PFS analysis used the ITT analysis set
Time-to-event endpointInvestigator-assessed PFS by RECIST v1.1
Log-rank testReported method for primary PFS and brain-metastases PFS analyses
Hazard ratioPrimary PFS HR 0.759; brain-metastases PFS HR 0.639
Cox modelUsed to calculate the reported hazard ratios
Confidence interval95% two-sided intervals reported for the PFS HRs
Categorical analysisORR analyzed using the Cochran-Mantel-Haenszel test
Subgroup analysisPFS in participants with brain metastases at baseline
Safety analysisSerious adverse events reported by randomized arm

23. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical methods behind randomized trials, time-to-event endpoints, categorical comparisons, confidence intervals, and treatment-effect measures.

26. Record Summary

HER2CLIMB-02 provides a compact example of modern randomized time-to-event analysis. The trial enrolled 466 participants in a randomized, parallel, quadruple-masked phase 3 design comparing tucatinib + T-DM1 with placebo + T-DM1. Its primary endpoint was investigator-assessed PFS according to RECIST v1.1, analyzed in the ITT population using a log-rank test, with the hazard ratio calculated from a Cox proportional hazards model.

The primary PFS analysis reported an HR of 0.759 with a 95% CI of 0.607–0.950 and a P-value of 0.0163. A secondary analysis among participants with brain metastases at baseline reported an HR of 0.639 with a 95% CI of 0.459–0.891 and a P-value of 0.0078. A secondary ORR analysis used the Cochran-Mantel-Haenszel test and reported a P-value of 0.2055, without an ORR effect estimate in the ClinicalTrials.gov record.

The statistical interpretation should therefore focus on the distinction between effect size (hazard ratio), precision (confidence interval), and evidence under a testing framework (P-value). It should also keep the primary ITT analysis separate from the baseline brain-metastases subgroup and avoid inferring unreported design features such as multiplicity procedures, interim monitoring, crossover, or Bayesian methods.

Clinical Biostats methodology: A trial-results page should distinguish reported registry evidence from statistical interpretation. When the registry supplies an estimate and confidence interval, those values are reported exactly; when it does not supply a numerical effect measure, the analysis does not manufacture one.