← Clinical Trials
Breast Cancer Phase 3 Completed NCT00829166

EMILIA: Complete Statistical Analysis of Trastuzumab Emtansine in HER2-Positive Breast Cancer

An independent statistical review of the randomized phase 3 EMILIA trial comparing trastuzumab emtansine with lapatinib plus capecitabine in participants with HER2-positive locally advanced or metastatic breast cancer.

Trial start: 2009-02  ·  Primary completion: 2012-07  ·  Enrollment: 991
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the statistical analyses contained in the ClinicalTrials.gov record. This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

EMILIA was a randomized, parallel-group, open-label phase 3 trial evaluating trastuzumab emtansine versus lapatinib plus capecitabine in participants with HER2-positive locally advanced or metastatic breast cancer.

991
Enrolled
Randomized trial
2
Arms
Parallel design
0.650
PFS HR
95% CI 0.549–0.771
0.749
Final OS HR
95% CI 0.639–0.877
FeatureEMILIA
PhasePhase 3
ConditionBreast Cancer
PopulationParticipants with HER2-positive locally advanced or metastatic breast cancer
DesignRandomized, parallel
MaskingNone
Primary purposeTreatment
Enrollment991
Arms2
Primary endpoint typesBinary; Time-to-event
Results postedYes
Outcome measures posted17
Statistical analyses posted8
Lead sponsorHoffmann-La Roche
Sponsor typeIndustry
ClinicalTrials.govNCT00829166

2. Clinical Question

The central statistical question was whether treatment assignment to trastuzumab emtansine produced better time-to-event and tumor-response outcomes than lapatinib plus capecitabine in participants with HER2-positive locally advanced or metastatic breast cancer.

Population

Participants with HER2-positive locally advanced or metastatic breast cancer.

Intervention

Trastuzumab emtansine.

Comparator

Lapatinib plus capecitabine.

Primary question

Does trastuzumab emtansine improve the registered efficacy endpoints compared with lapatinib plus capecitabine?

3. Trial Design

01
Randomize991 participants
02
Two armsParallel allocation
03
TreatmentStudy interventions
04
AssessProgression, survival, response
05
AnalyzeITT and stratified methods
ARM 1

Trastuzumab emtansine

  • Trastuzumab emtansine
ARM 2

Lapatinib + capecitabine

  • Lapatinib
  • Capecitabine

The trial was randomized, used a parallel design, and had no masking. These design characteristics matter statistically because randomization establishes the basis for comparing outcomes according to assigned treatment, while the absence of masking can be relevant when outcomes or treatment decisions have subjective components.

4. Trial Timeline

2009-02

Trial start

The EMILIA trial began in February 2009.

14 Jan 2012

Primary efficacy cutoff for PFS

The registered PFS and several response-related endpoints use a data cutoff of 14 January 2012, with a time frame from randomization of up to 2 years, 11 months.

31 Jul 2012

Second interim OS analysis

The confirmatory second interim analysis of overall survival uses a data cutoff of 31 July 2012, with a time frame from randomization of up to 3 years, 5 months.

2012-07

Primary completion

The registered primary completion date was July 2012.

31 Dec 2014

Final OS analysis

The final overall-survival endpoint uses a data cutoff of 31 December 2014, with a time frame from randomization of up to 5 years, 11 months.

5. Analysis Populations and Endpoint Structure

The posted primary and secondary statistical analyses identify the intention-to-treat population as the analysis population. The registry defines this population as including all randomized participants on the basis of the treatment assigned at randomization.

Analysis population / restrictionRole in the analyses posted on ClinicalTrials.gov
ITT populationUsed for the posted primary and secondary efficacy analyses; participants are analyzed according to randomized treatment assignment.
Measurable disease at baselineRequired for the posted objective-response analysis and clinical-benefit analysis.
Female participants with baseline assessment and at least 1 follow-up assessmentRestriction identified for the posted time-to-symptom-progression analysis.

The ITT principle is particularly important for a randomized trial. Once participants are randomized, analyzing them according to the assigned group preserves the comparison created by randomization. It does not require every participant to remain on treatment or to complete every assessment.

6. Primary Endpoints

The registry lists 8 primary endpoints. They span binary and time-to-event outcome types. Three of those primary endpoints have formal statistical analyses in the ClinicalTrials.gov record: PFS as assessed by an independent review committee, overall survival at the second interim analysis, and overall survival at the final analysis.

Registered primary endpointTime frameFormal analysis in the ClinicalTrials.gov record
Percentage of Participants With PD or Death as Assessed by an Independent Review Committee (IRC) From the date of randomization through the data cut-off date of 14 Jan 2012 (up to 2 years, 11 months) No formal statistical analysis posted in the ClinicalTrials.gov record
Progression-free Survival (PFS) as Assessed by an IRC (Co-primary Endpoint) From the date of randomization through the data cut-off date of 14 Jan 2012 (up to 2 years, 11 months) Log-rank test; Cox regression for HR
Percentage of Participants Who Died: Second Interim Analysis From the date of randomization through the data cut-off date of 31 Jul 2012 (up to 3 years, 5 months) No formal statistical analysis posted in the ClinicalTrials.gov record
Overall Survival: Second Interim Analysis (Co-primary Endpoint) From the date of randomization through the data cut-off date of 31 Jul 2012 (up to 3 years, 5 months) Log-rank test; Cox regression for HR
Percentage of Participants Who Died: Final Analysis From the date of randomization through the data cut-off date of 31 Dec 2014 (up to 5 years, 11 months) No formal statistical analysis posted in the ClinicalTrials.gov record
Overall Survival: Final Analysis From the date of randomization through the data cut-off date of 31 Dec 2014 (up to 5 years, 11 months) Log-rank test; Cox regression for HR
Percentage of Participants Who Were Alive at Year 1 Year 1 No formal statistical analysis posted in the ClinicalTrials.gov record
Percentage of Participants Who Were Alive at Year 2 Year 2 No formal statistical analysis posted in the ClinicalTrials.gov record

Registry definitions

PFS as assessed by an IRC: Tumor response was assessed by an IRC according to modified RECIST. Measurable lesions were identified as target lesions at baseline, with up to 5 target lesions per organ and 10 in total. The registry describes the sum of the longest diameter for target lesions as the baseline SLD.
Overall survival: OS was defined as the time from the date of randomization to the date of death from any cause. The median duration of OS was estimated using the Kaplan-Meier method, and the 95% CI was computed using the method of Brookmeyer and Crowley.

7. Primary Result: Progression-Free Survival

The first formal primary analysis in the registry-reported statistical record concerns progression-free survival as assessed by an independent review committee. The analysis used the ITT population and compared trastuzumab emtansine with lapatinib plus capecitabine.

Hazard ratio for progression or death

0.650

95% CI: 0.549–0.771   ·   P < 0.0001

Method: stratified log-rank test; HR estimated by Cox regression

CharacteristicReported value
EndpointProgression-free Survival (PFS) as Assessed by an IRC (Co-primary Endpoint)
Analysis populationITT
ComparisonTrastuzumab emtansine vs lapatinib + capecitabine
MethodLog-rank test
Effect measureHazard ratio
Estimate0.650
95% CI0.549–0.771
P-value<0.0001
Hypothesis typeSuperiority
StratificationRegion of enrollment; number of prior chemotherapeutic regimens (0-1 or >1); visceral/non-visceral disease
Clinical Biostats interpretation

The estimated HR of 0.650 means that, under the Cox model used for this analysis, the estimated instantaneous hazard of progression or death in the trastuzumab emtansine group was 65.0% of that in the lapatinib-plus-capecitabine group. Expressed as a simple derived comparison, this corresponds to an estimated 35.0% lower hazard under the model.

The HR does not mean that 35.0% of participants avoided progression, that each individual had exactly a 35.0% reduction in risk, or that the absolute probability of progression or death was reduced by 35.0 percentage points.

The 95% CI of 0.549–0.771 describes uncertainty around the estimated relative hazard under the statistical model. It does not describe the range of outcomes that individual patients might experience.

The P-value of <0.0001 addresses evidence against the null hypothesis in the specified testing framework. A P-value does not measure the magnitude of the treatment effect; the HR and its confidence interval are needed to describe magnitude and precision.

The analysis was stratified by region, number of prior chemotherapeutic regimens, and visceral versus non-visceral disease. Because the HR came from a Cox regression model, its interpretation also depends on the model's assumptions, including the proportional-hazards framework. Censoring and the timing of progression or death are therefore integral to the analysis.

8. Primary Result: Overall Survival — Second Interim Analysis

The second formal primary analysis was the registered co-primary overall-survival endpoint at the second interim analysis. The registry states that this second interim analysis was deemed to be the confirmatory analysis.

Hazard ratio for death

0.682

95% CI: 0.548–0.849   ·   P = 0.0006

Method: stratified log-rank test; HR estimated by Cox regression

CharacteristicReported value
EndpointOverall Survival: Second Interim Analysis (Co-primary Endpoint)
DefinitionTime from randomization to death from any cause
Analysis populationITT
ComparisonTrastuzumab emtansine vs lapatinib + capecitabine
MethodLog-rank test
Effect measureHazard ratio
Estimate0.682
95% CI0.548–0.849
P-value0.0006
Hypothesis typeSuperiority
Analysis designationSecond interim analysis; deemed confirmatory
Clinical Biostats interpretation

The estimated HR of 0.682 means that the modeled instantaneous hazard of death in the trastuzumab emtansine group was 68.2% of the corresponding hazard in the comparator group. As a simple derived statement, this is an estimated 31.8% lower hazard under the fitted model.

This is a relative time-to-event measure, not an absolute mortality reduction. It does not mean that 31.8% of participants were prevented from dying, nor does it specify how many additional participants were alive at a particular time point.

The 95% CI of 0.548–0.849 quantifies uncertainty around the estimated HR. Its width reflects the precision of the estimated relative treatment effect; it does not represent individual-level variability.

The P-value of 0.0006 is evidence against the relevant null hypothesis under the trial's stated testing framework. It should not be interpreted as a probability that the treatment effect is real, and it does not indicate the clinical magnitude of the effect.

The analysis was stratified by region of enrollment, number of prior chemotherapeutic regimens, and visceral/non-visceral disease. The Cox-derived HR therefore carries the assumptions of the time-to-event model and should be interpreted together with the confidence interval and the trial's censoring structure.

9. Primary Result: Overall Survival — Final Analysis

The final overall-survival analysis used the same randomized comparison and ITT population but a later data cutoff of 31 December 2014. The registry explicitly describes the final analysis as descriptive.

Final-analysis hazard ratio for death

0.749

95% CI: 0.639–0.877   ·   P = 0.0003

Method: stratified log-rank test; HR estimated by Cox regression

CharacteristicReported value
EndpointOverall Survival: Final Analysis
DefinitionTime from randomization to death from any cause
Analysis populationITT
ComparisonTrastuzumab emtansine vs lapatinib + capecitabine
MethodLog-rank test
Effect measureHazard ratio
Estimate0.749
95% CI0.639–0.877
P-value0.0003
Hypothesis typeSuperiority
Analysis designationFinal analysis; descriptive
Clinical Biostats interpretation

The final-analysis HR of 0.749 corresponds to an estimated instantaneous hazard of death equal to 74.9% of the comparator hazard under the Cox model. As a derived relative statement, that is approximately a 25.1% lower estimated hazard for trastuzumab emtansine.

The estimate does not mean that 25.1% of participants survived, that 25.1% of deaths were prevented, or that every participant experienced the same proportional reduction in hazard.

The 95% CI of 0.639–0.877 gives the uncertainty around the HR estimate. The interval is narrower than the corresponding second-interim-analysis interpretation would be if additional information increased precision, but the numerical width should always be considered in relation to the underlying event and censoring information rather than treated as a measure of clinical importance by itself.

The P-value of 0.0003 provides evidence against the relevant null hypothesis in the specified analysis framework. It is not an effect-size statistic. The HR and CI describe the relative effect; absolute survival probabilities would answer a different question.

The registry identifies this final analysis as descriptive. That designation is important: the presence of a nominal P-value does not by itself transform a descriptive later analysis into a newly independent confirmatory hypothesis test.

Educational note: the ClinicalTrials.gov record provides hazard ratios and confidence intervals but do not provide enough event-time information to reconstruct a valid Kaplan-Meier curve. A survival curve should not be fabricated from a single HR and confidence interval.

10. Primary Endpoints Without Formal Statistical Comparisons in the Supplied Record

Five of the eight registered primary endpoints are binary or time-specific survival summaries for which the ClinicalTrials.gov record does not provide a formal comparison. The registry does state what these endpoints measure.

EndpointRegistry definition / time frameTypical statistical approach
Percentage of Participants With PD or Death as Assessed by an IRC PD was assessed by an IRC using modified RECIST; from randomization through 14 Jan 2012, up to 2 years, 11 months. A binary comparison could use a stratified categorical-data method such as a Cochran-Mantel-Haenszel test when stratification factors are part of the prespecified analysis.
Percentage of Participants Who Died: Second Interim Analysis Percentage who died from any cause; second interim analysis deemed confirmatory; through 31 Jul 2012, up to 3 years, 5 months. A fixed-time mortality percentage can be compared using a categorical-data framework, while the underlying OS endpoint is more appropriately analyzed as a time-to-event outcome.
Percentage of Participants Who Died: Final Analysis Percentage who died from any cause; final analysis described as descriptive; through 31 Dec 2014, up to 5 years, 11 months. A descriptive fixed-time percentage can be reported directly; formal OS inference is generally based on time-to-event methods.
Percentage of Participants Who Were Alive at Year 1 Percentage alive 1 year after starting treatment; final analysis. A time-specific survival estimate is typically obtained from a survival-analysis framework such as Kaplan-Meier estimation.
Percentage of Participants Who Were Alive at Year 2 Percentage alive 2 years after starting treatment; final analysis. A time-specific survival estimate is typically obtained from a survival-analysis framework such as Kaplan-Meier estimation.

The ClinicalTrials.gov record does not provide arm-specific numerical estimates or formal comparisons for these five endpoints. Accordingly, this page does not infer values from the corresponding HR analyses.

11. Secondary Endpoint Results

The ClinicalTrials.gov record contains five secondary statistical analyses. These cover investigator-assessed PFS, objective response, clinical benefit, time to treatment failure, and time to symptom progression.

Investigator-Assessed PFS

Hazard ratio

0.658

95% CI: 0.560–0.774   ·   P < 0.0001

Trastuzumab emtansine vs lapatinib + capecitabine

This analysis used the ITT population and a stratified log-rank test, with the HR estimated by Cox regression. The analysis was stratified by region of enrollment, number of prior chemotherapeutic regimens, and visceral/non-visceral disease.

Objective Response as Assessed by an IRC

Difference in objective response rates

12.7

95% CI: 6.0–19.4   ·   P = 0.0002

Difference defined as trastuzumab emtansine minus lapatinib + capecitabine

The objective-response analysis used the ITT population, with only participants with measurable disease at baseline included in the analysis. The statistical method was the Cochran-Mantel-Haenszel test, reported as a Mantel-Haenszel chi-squared test in the registry, with stratification by region, prior chemotherapeutic regimens, and visceral/non-visceral disease. The 95% CI for the difference was computed using an approximate normal method.

Difference in objective response rate
Estimated difference
12.7 percentage points

The bar is a visual representation of the reported difference and is not a reconstruction of the underlying response rates.

Clinical Benefit as Assessed by an IRC

Difference in clinical benefit rate

14

95% CI: 7.0–20.9

Difference defined as trastuzumab emtansine minus lapatinib + capecitabine

The registry identifies a Wald / z-test framework for this analysis. The 95% CI for the difference in clinical benefit rate was computed using the normal approximation method. The analysis used the ITT population and the registry states that only participants with measurable disease at baseline were included.

Time to Treatment Failure

Hazard ratio

0.703

95% CI: 0.602–0.820   ·   P < 0.0001

Method: stratified log-rank test; HR estimated by Cox regression

The HR of 0.703 corresponds to an estimated instantaneous treatment-failure hazard equal to 70.3% of the comparator hazard under the fitted model, or a derived estimated 29.7% lower hazard.

Time to Symptom Progression

Hazard ratio

0.796

95% CI: 0.667–0.951   ·   P = 0.0121

Method: stratified log-rank test; HR estimated by Cox regression

The analysis used the ITT population, with the registry specifying a restriction to female participants with a baseline assessment and at least 1 follow-up assessment. The analysis was stratified by region, number of prior chemotherapeutic regimens, and visceral/non-visceral disease.

Secondary endpointEffect estimate95% CIP-valueMethod
PFS as assessed by the investigatorHR 0.6580.560–0.774<0.0001Log-rank; Cox regression
Objective response as assessed by an IRCDifference 12.76.0–19.40.0002Cochran-Mantel-Haenszel
Clinical benefit as assessed by an IRCDifference 147.0–20.9Not provided in registry-reported analysisWald / z-test framework
Time to treatment failureHR 0.7030.602–0.820<0.0001Log-rank; Cox regression
Time to symptom progressionHR 0.7960.667–0.9510.0121Log-rank; Cox regression

12. Statistical Methodology

Kaplan-Meier estimation

The registry specifies Kaplan-Meier estimation for overall survival. Kaplan-Meier estimation is designed for time-to-event data in which some participants may not experience the event during observed follow-up. Those observations can be right-censored and still contribute information until their censoring time.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

where di is the number of events at time ti and ni is the number at risk immediately before that time.

For EMILIA, the registered OS definition is the time from randomization to death from any cause. The registry also states that the median OS was estimated using Kaplan-Meier and that its 95% CI was computed using the Brookmeyer and Crowley method.

Stratified log-rank test

The formal time-to-event comparisons in the analyses posted on ClinicalTrials.gov used the log-rank test. The analyses were stratified by region of enrollment, number of prior chemotherapeutic regimens, and visceral versus non-visceral disease.

Stratification allows the comparison to account for prespecified categories while preserving the time-to-event structure of the data. Rather than simply comparing the percentage of events at one arbitrary time, the log-rank framework uses the ordering of event times across follow-up.

Cox regression and hazard ratios

The registry states that the HR was estimated using Cox regression. A hazard ratio is a relative measure of the instantaneous event hazard under the fitted model.

Interpretation of the hazard ratio
HR < 1  →  lower estimated instantaneous event hazard in the trastuzumab emtansine group

An HR is not the same as a relative risk, an absolute risk difference, a median-survival ratio, or the proportion of participants who benefit.

Intention-to-treat analysis

The posted analyses define the ITT population as all randomized participants analyzed on the basis of treatment assigned at randomization. This is the natural analysis population for preserving the treatment comparison created by randomization.

Cochran-Mantel-Haenszel analysis

The objective-response endpoint used a Mantel-Haenszel chi-squared test, normalized here as a Cochran-Mantel-Haenszel test. This method provides a way to compare categorical outcomes while accounting for stratification variables.

Wald / z-test framework

The clinical-benefit analysis is identified in the ClinicalTrials.gov record with a Wald / z-test framework. In general, a Wald-type statistic compares an estimated effect with its estimated standard error, producing a standardized quantity for inference under the relevant large-sample approximation.

13. Statistical Methods Explained

Why was the analysis stratified?

The time-to-event analyses were stratified by region of enrollment, number of prior chemotherapeutic regimens, and visceral versus non-visceral disease. Stratification can improve the validity and efficiency of a treatment comparison when these factors are incorporated into the trial's design and analysis plan. It also avoids treating the observed distribution of these factors as if it had arisen independently of the randomized design.

What does an HR of 0.650 mean?

An HR of 0.650 means that the fitted model estimates the instantaneous hazard in the trastuzumab emtansine group at 65.0% of the comparator hazard. The simple arithmetic interpretation is a 35.0% lower estimated hazard. It does not mean a 35.0% absolute reduction in the probability of progression or death.

Why use an ITT population?

ITT analysis keeps participants in the group to which they were randomized. This maintains the comparison generated by randomization and avoids redefining treatment groups according to events that occurred after randomization.

What is the difference between an HR and a response-rate difference?

The HR describes a relative comparison of event hazards over time. The objective-response difference of 12.7 is a difference between two categorical response rates, defined here as trastuzumab emtansine minus lapatinib plus capecitabine. The two measures therefore summarize different aspects of treatment effect.

Why does the confidence interval matter?

A point estimate such as HR 0.749 is only one estimate from the observed data. The corresponding 95% CI of 0.639–0.877 communicates uncertainty around that estimate under the model and inferential framework. It does not give a range in which every individual patient's treatment effect must lie.

Why doesn't the P-value measure the size of the effect?

The P-value addresses the evidence against a specified null hypothesis. Its magnitude depends on both the observed data and the amount of information available. Effect size should therefore be read from the HR or difference, while precision is assessed using the confidence interval.

Why is a time-to-event endpoint different from a binary endpoint?

A binary endpoint records whether an event occurred within a defined framework. A time-to-event endpoint retains information about when the event occurred and can accommodate censoring. This is why PFS, OS, time to treatment failure, and time to symptom progression were analyzed with survival-analysis methods in the ClinicalTrials.gov record.

14. Multiplicity, Interim Analysis, and Hypothesis Type

The ClinicalTrials.gov record identifies superiority as the hypothesis type for all eight posted statistical analyses. Three primary endpoint analyses are formally reported: PFS as assessed by an IRC, OS at the second interim analysis, and final OS.

FeatureWhat the ClinicalTrials.gov record establishes
Hypothesis typeSuperiority
Primary endpoints8 registered primary endpoints
Primary endpoint analyses3 formal statistical analyses in the ClinicalTrials.gov record
Second interim OS analysisReported as the confirmatory analysis
Final OS analysisReported as descriptive
Multiplicity adjustment detailsNot provided in the ClinicalTrials.gov record
Alpha-spending detailsNot provided in the ClinicalTrials.gov record
Non-inferiority marginNot applicable to the registry-reported superiority analyses
Bayesian methodsNot identified in the statistical analyses posted on ClinicalTrials.gov
Factorial designNot used; the design model is parallel

The distinction between the second interim analysis and final analysis is statistically important. The registry explicitly states that the second interim OS analysis was deemed confirmatory, whereas the final analysis is described as descriptive. A later P-value should not automatically be interpreted as if it were associated with a newly defined confirmatory testing budget.

15. Crossover, Missing Data, and Other Design Issues

Crossover

The ClinicalTrials.gov record does not identify a crossover procedure or crossover rate. No crossover adjustment is therefore applied or inferred in this analysis.

Missing data

The ClinicalTrials.gov record does not specify a missing-data or imputation method for the posted analyses.

Non-inferiority margin

No non-inferiority hypothesis or margin is identified. The posted analyses use superiority hypotheses.

Bayesian methods

No Bayesian method is identified among the statistical analyses posted on ClinicalTrials.gov.

These omissions matter because statistical interpretation should follow the documented analysis rather than fill gaps with assumptions. For example, a time-to-event analysis can be sensitive to censoring rules, while a categorical analysis can depend on how participants with unavailable assessments are handled. The ClinicalTrials.gov record does not provide enough information to reconstruct those details.

16. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm. The denominators are presented as affected participants divided by participants at risk.

Safety groupSerious adverse eventsAffected / at risk
Trastuzumab emtansineSerious adverse events92 / 490
Lapatinib + capecitabineSerious adverse events99 / 488
Lapatinib + capecitabine / Trastuzumab emtansineSerious adverse events19 / 136
Serious adverse events by registry-reported safety group
Trastuzumab emtansine
92 / 490
Lapatinib + capecitabine
99 / 488
Lapatinib + capecitabine / trastuzumab emtansine
19 / 136

The third safety group is explicitly identified in the ClinicalTrials.gov record as Lapatinib + Capecitabine/ Trastuzumab Em with 19 affected participants among 136 at risk. Because the ClinicalTrials.gov record does not further define this group's construction, these figures should not be silently combined with either of the two main randomized treatment groups.

17. Understanding the Primary PFS and OS Results Together

The primary efficacy analyses show a consistent direction of estimated treatment effect across two different time-to-event endpoints and two OS analyses.

EndpointHR95% CIP-valueInterpretation of HR
IRC-assessed PFS0.6500.549–0.771<0.000135.0% lower estimated hazard
OS, second interim0.6820.548–0.8490.000631.8% lower estimated hazard
OS, final analysis0.7490.639–0.8770.000325.1% lower estimated hazard

The three estimates should not be treated as interchangeable. PFS measures progression or death, whereas OS measures death from any cause. The second-interim and final OS analyses also correspond to different data cutoffs and different analysis designations in the registry.

How to read the pattern

The HR estimates are all below 1, meaning that each fitted comparison estimates a lower instantaneous event hazard for trastuzumab emtansine relative to lapatinib plus capecitabine. The numerical HR changes across analyses do not imply that a treatment effect "decayed" by a specific percentage, because each endpoint and analysis has its own information set and follow-up structure.

18. Time-to-Event Endpoints: What Is Actually Being Compared?

EMILIA contains several time-to-event endpoints: IRC-assessed PFS, investigator-assessed PFS, overall survival, time to treatment failure, and time to symptom progression. These endpoints differ in what constitutes an event.

EndpointEvent conceptStatistical structure
IRC-assessed PFSProgression or death, with progression assessed by an IRCTime-to-event; stratified log-rank; Cox HR
Investigator-assessed PFSProgression or death assessed by investigatorsTime-to-event; stratified log-rank; Cox HR
Overall survivalDeath from any causeTime-to-event; Kaplan-Meier; stratified log-rank; Cox HR
Time to treatment failureTreatment-failure endpoint as defined by the trial registryTime-to-event; stratified log-rank; Cox HR
Time to symptom progressionSymptom progression endpointTime-to-event; stratified log-rank; Cox HR

Using multiple time-to-event endpoints can provide complementary information, but it also creates a distinction between the statistical evidence for a particular endpoint and the broader clinical story. The HR for PFS should not be interpreted as if it were an OS HR, and an endpoint's P-value does not establish that every other endpoint must show the same magnitude of effect.

19. Clinical Biostats Interpretation of the Confidence Intervals

PFS precision

The IRC-assessed PFS HR was 0.650 with a 95% CI of 0.549–0.771. The interval provides the uncertainty range around the estimated relative hazard under the specified analysis model. It is more informative than the point estimate alone because it shows how precisely the treatment effect was estimated.

Second-interim OS precision

The second-interim OS HR was 0.682 with a 95% CI of 0.548–0.849. The interval remains below 1, while still showing uncertainty around the magnitude of the relative effect.

Final OS precision

The final OS HR was 0.749 with a 95% CI of 0.639–0.877. The estimate remains below 1, but the final-analysis point estimate differs from the second-interim estimate. Such differences are expected when additional follow-up and events change the information available for estimation.

Confidence interval caution: A 95% CI is not a statement that there is a 95% probability that the true treatment effect lies inside this particular interval. It is an interval generated by a statistical procedure with stated long-run coverage properties under the relevant assumptions.

20. Why the Objective-Response Analysis Uses a Different Method

Objective response is a categorical outcome rather than a time-to-event endpoint. The registry-reported analysis therefore uses a Cochran-Mantel-Haenszel test rather than a log-rank test.

Difference in response rates
Difference = Response ratetrastuzumab emtansine − Response ratelapatinib + capecitabine

The registry-reported estimate is 12.7, with a two-sided 95% CI of 6.0–19.4.

The difference is an absolute effect measure on the percentage-point scale. That makes it fundamentally different from an HR. The response analysis also has a specific analysis restriction: only participants with measurable disease at baseline were included.

The confidence interval was computed using an approximate normal method. This is another reason to distinguish the response analysis from the survival analyses: different outcome structures lead naturally to different estimators, test statistics, and confidence intervals.

21. Limitations and Interpretation Issues

22. Why This Trial Matters Statistically

EMILIA is a useful teaching case because its registry record connects several core clinical-trial methods within one randomized phase 3 study. It includes independent-review PFS, investigator-assessed PFS, overall survival at an interim analysis and final analysis, categorical response outcomes, stratified time-to-event methods, and ITT analysis.

ConceptHow it appears in EMILIA
RandomizationRandomized, parallel-group phase 3 design
ITT analysisPrimary and secondary efficacy analyses use the ITT population
Kaplan-Meier estimationRegistered method for estimating median OS
Hazard ratioPrimary and secondary time-to-event effect measure
Confidence intervalsReported around primary HRs and categorical differences
Log-rank testPrimary and secondary time-to-event comparisons
Cox regressionHR estimation for the time-to-event analyses
Stratified analysisRegion, prior chemotherapy regimen count, and visceral/non-visceral disease
Cochran-Mantel-Haenszel testIRC-assessed objective-response comparison
Wald / z-testClinical-benefit analysis framework
Interim analysisSecond interim OS analysis designated confirmatory
Descriptive final analysisFinal OS analysis explicitly described as descriptive

23. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

24. Related Statistical Calculators

Apply the same statistical concepts with related Clinical Biostats calculators:

25. Sources

Continue through the Clinical Biostats statistical library

Explore the underlying methods used in randomized clinical trials, from survival analysis and hazard ratios to categorical-data tests and confidence intervals.

26. Record Summary

EMILIA provides a compact example of how randomized clinical-trial evidence can be analyzed across several endpoint structures. The ClinicalTrials.gov record includes a randomized phase 3 parallel design, an ITT analysis population, stratified log-rank comparisons, Cox-derived hazard ratios, Kaplan-Meier estimation for overall survival, a Cochran-Mantel-Haenszel analysis for objective response, and a Wald / z-test framework for clinical benefit.

The three formal primary analyses show HR estimates of 0.650 for IRC-assessed PFS, 0.682 for OS at the second interim analysis, and 0.749 for final OS. Their corresponding 95% confidence intervals and P-values describe statistical uncertainty and evidence within their respective analysis frameworks. The secondary analyses extend the same statistical story to investigator-assessed PFS, objective response, clinical benefit, treatment failure, and symptom progression.

Clinical Biostats methodology: The most useful interpretation of a clinical trial combines the treatment effect estimate, its confidence interval, the endpoint definition, the analysis population, the statistical method, the data cutoff, and the design features that determine how the result should be understood. A P-value alone is not a measure of treatment effect, and a hazard ratio should not be interpreted as an absolute risk difference.