This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the statistical analyses contained in the ClinicalTrials.gov record. This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
EMILIA was a randomized, parallel-group, open-label phase 3 trial evaluating trastuzumab emtansine versus lapatinib plus capecitabine in participants with HER2-positive locally advanced or metastatic breast cancer.
| Feature | EMILIA |
|---|---|
| Phase | Phase 3 |
| Condition | Breast Cancer |
| Population | Participants with HER2-positive locally advanced or metastatic breast cancer |
| Design | Randomized, parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 991 |
| Arms | 2 |
| Primary endpoint types | Binary; Time-to-event |
| Results posted | Yes |
| Outcome measures posted | 17 |
| Statistical analyses posted | 8 |
| Lead sponsor | Hoffmann-La Roche |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT00829166 |
2. Clinical Question
The central statistical question was whether treatment assignment to trastuzumab emtansine produced better time-to-event and tumor-response outcomes than lapatinib plus capecitabine in participants with HER2-positive locally advanced or metastatic breast cancer.
Population
Participants with HER2-positive locally advanced or metastatic breast cancer.
Intervention
Trastuzumab emtansine.
Comparator
Lapatinib plus capecitabine.
Primary question
Does trastuzumab emtansine improve the registered efficacy endpoints compared with lapatinib plus capecitabine?
3. Trial Design
Trastuzumab emtansine
- Trastuzumab emtansine
Lapatinib + capecitabine
- Lapatinib
- Capecitabine
The trial was randomized, used a parallel design, and had no masking. These design characteristics matter statistically because randomization establishes the basis for comparing outcomes according to assigned treatment, while the absence of masking can be relevant when outcomes or treatment decisions have subjective components.
4. Trial Timeline
Trial start
The EMILIA trial began in February 2009.
Primary efficacy cutoff for PFS
The registered PFS and several response-related endpoints use a data cutoff of 14 January 2012, with a time frame from randomization of up to 2 years, 11 months.
Second interim OS analysis
The confirmatory second interim analysis of overall survival uses a data cutoff of 31 July 2012, with a time frame from randomization of up to 3 years, 5 months.
Primary completion
The registered primary completion date was July 2012.
Final OS analysis
The final overall-survival endpoint uses a data cutoff of 31 December 2014, with a time frame from randomization of up to 5 years, 11 months.
5. Analysis Populations and Endpoint Structure
The posted primary and secondary statistical analyses identify the intention-to-treat population as the analysis population. The registry defines this population as including all randomized participants on the basis of the treatment assigned at randomization.
| Analysis population / restriction | Role in the analyses posted on ClinicalTrials.gov |
|---|---|
| ITT population | Used for the posted primary and secondary efficacy analyses; participants are analyzed according to randomized treatment assignment. |
| Measurable disease at baseline | Required for the posted objective-response analysis and clinical-benefit analysis. |
| Female participants with baseline assessment and at least 1 follow-up assessment | Restriction identified for the posted time-to-symptom-progression analysis. |
The ITT principle is particularly important for a randomized trial. Once participants are randomized, analyzing them according to the assigned group preserves the comparison created by randomization. It does not require every participant to remain on treatment or to complete every assessment.
6. Primary Endpoints
The registry lists 8 primary endpoints. They span binary and time-to-event outcome types. Three of those primary endpoints have formal statistical analyses in the ClinicalTrials.gov record: PFS as assessed by an independent review committee, overall survival at the second interim analysis, and overall survival at the final analysis.
| Registered primary endpoint | Time frame | Formal analysis in the ClinicalTrials.gov record |
|---|---|---|
| Percentage of Participants With PD or Death as Assessed by an Independent Review Committee (IRC) | From the date of randomization through the data cut-off date of 14 Jan 2012 (up to 2 years, 11 months) | No formal statistical analysis posted in the ClinicalTrials.gov record |
| Progression-free Survival (PFS) as Assessed by an IRC (Co-primary Endpoint) | From the date of randomization through the data cut-off date of 14 Jan 2012 (up to 2 years, 11 months) | Log-rank test; Cox regression for HR |
| Percentage of Participants Who Died: Second Interim Analysis | From the date of randomization through the data cut-off date of 31 Jul 2012 (up to 3 years, 5 months) | No formal statistical analysis posted in the ClinicalTrials.gov record |
| Overall Survival: Second Interim Analysis (Co-primary Endpoint) | From the date of randomization through the data cut-off date of 31 Jul 2012 (up to 3 years, 5 months) | Log-rank test; Cox regression for HR |
| Percentage of Participants Who Died: Final Analysis | From the date of randomization through the data cut-off date of 31 Dec 2014 (up to 5 years, 11 months) | No formal statistical analysis posted in the ClinicalTrials.gov record |
| Overall Survival: Final Analysis | From the date of randomization through the data cut-off date of 31 Dec 2014 (up to 5 years, 11 months) | Log-rank test; Cox regression for HR |
| Percentage of Participants Who Were Alive at Year 1 | Year 1 | No formal statistical analysis posted in the ClinicalTrials.gov record |
| Percentage of Participants Who Were Alive at Year 2 | Year 2 | No formal statistical analysis posted in the ClinicalTrials.gov record |
Registry definitions
7. Primary Result: Progression-Free Survival
The first formal primary analysis in the registry-reported statistical record concerns progression-free survival as assessed by an independent review committee. The analysis used the ITT population and compared trastuzumab emtansine with lapatinib plus capecitabine.
Hazard ratio for progression or death
95% CI: 0.549–0.771 · P < 0.0001
Method: stratified log-rank test; HR estimated by Cox regression
| Characteristic | Reported value |
|---|---|
| Endpoint | Progression-free Survival (PFS) as Assessed by an IRC (Co-primary Endpoint) |
| Analysis population | ITT |
| Comparison | Trastuzumab emtansine vs lapatinib + capecitabine |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.650 |
| 95% CI | 0.549–0.771 |
| P-value | <0.0001 |
| Hypothesis type | Superiority |
| Stratification | Region of enrollment; number of prior chemotherapeutic regimens (0-1 or >1); visceral/non-visceral disease |
The estimated HR of 0.650 means that, under the Cox model used for this analysis, the estimated instantaneous hazard of progression or death in the trastuzumab emtansine group was 65.0% of that in the lapatinib-plus-capecitabine group. Expressed as a simple derived comparison, this corresponds to an estimated 35.0% lower hazard under the model.
The HR does not mean that 35.0% of participants avoided progression, that each individual had exactly a 35.0% reduction in risk, or that the absolute probability of progression or death was reduced by 35.0 percentage points.
The 95% CI of 0.549–0.771 describes uncertainty around the estimated relative hazard under the statistical model. It does not describe the range of outcomes that individual patients might experience.
The P-value of <0.0001 addresses evidence against the null hypothesis in the specified testing framework. A P-value does not measure the magnitude of the treatment effect; the HR and its confidence interval are needed to describe magnitude and precision.
The analysis was stratified by region, number of prior chemotherapeutic regimens, and visceral versus non-visceral disease. Because the HR came from a Cox regression model, its interpretation also depends on the model's assumptions, including the proportional-hazards framework. Censoring and the timing of progression or death are therefore integral to the analysis.
8. Primary Result: Overall Survival — Second Interim Analysis
The second formal primary analysis was the registered co-primary overall-survival endpoint at the second interim analysis. The registry states that this second interim analysis was deemed to be the confirmatory analysis.
Hazard ratio for death
95% CI: 0.548–0.849 · P = 0.0006
Method: stratified log-rank test; HR estimated by Cox regression
| Characteristic | Reported value |
|---|---|
| Endpoint | Overall Survival: Second Interim Analysis (Co-primary Endpoint) |
| Definition | Time from randomization to death from any cause |
| Analysis population | ITT |
| Comparison | Trastuzumab emtansine vs lapatinib + capecitabine |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.682 |
| 95% CI | 0.548–0.849 |
| P-value | 0.0006 |
| Hypothesis type | Superiority |
| Analysis designation | Second interim analysis; deemed confirmatory |
The estimated HR of 0.682 means that the modeled instantaneous hazard of death in the trastuzumab emtansine group was 68.2% of the corresponding hazard in the comparator group. As a simple derived statement, this is an estimated 31.8% lower hazard under the fitted model.
This is a relative time-to-event measure, not an absolute mortality reduction. It does not mean that 31.8% of participants were prevented from dying, nor does it specify how many additional participants were alive at a particular time point.
The 95% CI of 0.548–0.849 quantifies uncertainty around the estimated HR. Its width reflects the precision of the estimated relative treatment effect; it does not represent individual-level variability.
The P-value of 0.0006 is evidence against the relevant null hypothesis under the trial's stated testing framework. It should not be interpreted as a probability that the treatment effect is real, and it does not indicate the clinical magnitude of the effect.
The analysis was stratified by region of enrollment, number of prior chemotherapeutic regimens, and visceral/non-visceral disease. The Cox-derived HR therefore carries the assumptions of the time-to-event model and should be interpreted together with the confidence interval and the trial's censoring structure.
9. Primary Result: Overall Survival — Final Analysis
The final overall-survival analysis used the same randomized comparison and ITT population but a later data cutoff of 31 December 2014. The registry explicitly describes the final analysis as descriptive.
Final-analysis hazard ratio for death
95% CI: 0.639–0.877 · P = 0.0003
Method: stratified log-rank test; HR estimated by Cox regression
| Characteristic | Reported value |
|---|---|
| Endpoint | Overall Survival: Final Analysis |
| Definition | Time from randomization to death from any cause |
| Analysis population | ITT |
| Comparison | Trastuzumab emtansine vs lapatinib + capecitabine |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.749 |
| 95% CI | 0.639–0.877 |
| P-value | 0.0003 |
| Hypothesis type | Superiority |
| Analysis designation | Final analysis; descriptive |
The final-analysis HR of 0.749 corresponds to an estimated instantaneous hazard of death equal to 74.9% of the comparator hazard under the Cox model. As a derived relative statement, that is approximately a 25.1% lower estimated hazard for trastuzumab emtansine.
The estimate does not mean that 25.1% of participants survived, that 25.1% of deaths were prevented, or that every participant experienced the same proportional reduction in hazard.
The 95% CI of 0.639–0.877 gives the uncertainty around the HR estimate. The interval is narrower than the corresponding second-interim-analysis interpretation would be if additional information increased precision, but the numerical width should always be considered in relation to the underlying event and censoring information rather than treated as a measure of clinical importance by itself.
The P-value of 0.0003 provides evidence against the relevant null hypothesis in the specified analysis framework. It is not an effect-size statistic. The HR and CI describe the relative effect; absolute survival probabilities would answer a different question.
The registry identifies this final analysis as descriptive. That designation is important: the presence of a nominal P-value does not by itself transform a descriptive later analysis into a newly independent confirmatory hypothesis test.
10. Primary Endpoints Without Formal Statistical Comparisons in the Supplied Record
Five of the eight registered primary endpoints are binary or time-specific survival summaries for which the ClinicalTrials.gov record does not provide a formal comparison. The registry does state what these endpoints measure.
| Endpoint | Registry definition / time frame | Typical statistical approach |
|---|---|---|
| Percentage of Participants With PD or Death as Assessed by an IRC | PD was assessed by an IRC using modified RECIST; from randomization through 14 Jan 2012, up to 2 years, 11 months. | A binary comparison could use a stratified categorical-data method such as a Cochran-Mantel-Haenszel test when stratification factors are part of the prespecified analysis. |
| Percentage of Participants Who Died: Second Interim Analysis | Percentage who died from any cause; second interim analysis deemed confirmatory; through 31 Jul 2012, up to 3 years, 5 months. | A fixed-time mortality percentage can be compared using a categorical-data framework, while the underlying OS endpoint is more appropriately analyzed as a time-to-event outcome. |
| Percentage of Participants Who Died: Final Analysis | Percentage who died from any cause; final analysis described as descriptive; through 31 Dec 2014, up to 5 years, 11 months. | A descriptive fixed-time percentage can be reported directly; formal OS inference is generally based on time-to-event methods. |
| Percentage of Participants Who Were Alive at Year 1 | Percentage alive 1 year after starting treatment; final analysis. | A time-specific survival estimate is typically obtained from a survival-analysis framework such as Kaplan-Meier estimation. |
| Percentage of Participants Who Were Alive at Year 2 | Percentage alive 2 years after starting treatment; final analysis. | A time-specific survival estimate is typically obtained from a survival-analysis framework such as Kaplan-Meier estimation. |
The ClinicalTrials.gov record does not provide arm-specific numerical estimates or formal comparisons for these five endpoints. Accordingly, this page does not infer values from the corresponding HR analyses.
11. Secondary Endpoint Results
The ClinicalTrials.gov record contains five secondary statistical analyses. These cover investigator-assessed PFS, objective response, clinical benefit, time to treatment failure, and time to symptom progression.
Investigator-Assessed PFS
Hazard ratio
95% CI: 0.560–0.774 · P < 0.0001
Trastuzumab emtansine vs lapatinib + capecitabine
This analysis used the ITT population and a stratified log-rank test, with the HR estimated by Cox regression. The analysis was stratified by region of enrollment, number of prior chemotherapeutic regimens, and visceral/non-visceral disease.
Objective Response as Assessed by an IRC
Difference in objective response rates
95% CI: 6.0–19.4 · P = 0.0002
Difference defined as trastuzumab emtansine minus lapatinib + capecitabine
The objective-response analysis used the ITT population, with only participants with measurable disease at baseline included in the analysis. The statistical method was the Cochran-Mantel-Haenszel test, reported as a Mantel-Haenszel chi-squared test in the registry, with stratification by region, prior chemotherapeutic regimens, and visceral/non-visceral disease. The 95% CI for the difference was computed using an approximate normal method.
Clinical Benefit as Assessed by an IRC
Difference in clinical benefit rate
95% CI: 7.0–20.9
Difference defined as trastuzumab emtansine minus lapatinib + capecitabine
The registry identifies a Wald / z-test framework for this analysis. The 95% CI for the difference in clinical benefit rate was computed using the normal approximation method. The analysis used the ITT population and the registry states that only participants with measurable disease at baseline were included.
Time to Treatment Failure
Hazard ratio
95% CI: 0.602–0.820 · P < 0.0001
Method: stratified log-rank test; HR estimated by Cox regression
The HR of 0.703 corresponds to an estimated instantaneous treatment-failure hazard equal to 70.3% of the comparator hazard under the fitted model, or a derived estimated 29.7% lower hazard.
Time to Symptom Progression
Hazard ratio
95% CI: 0.667–0.951 · P = 0.0121
Method: stratified log-rank test; HR estimated by Cox regression
The analysis used the ITT population, with the registry specifying a restriction to female participants with a baseline assessment and at least 1 follow-up assessment. The analysis was stratified by region, number of prior chemotherapeutic regimens, and visceral/non-visceral disease.
| Secondary endpoint | Effect estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| PFS as assessed by the investigator | HR 0.658 | 0.560–0.774 | <0.0001 | Log-rank; Cox regression |
| Objective response as assessed by an IRC | Difference 12.7 | 6.0–19.4 | 0.0002 | Cochran-Mantel-Haenszel |
| Clinical benefit as assessed by an IRC | Difference 14 | 7.0–20.9 | Not provided in registry-reported analysis | Wald / z-test framework |
| Time to treatment failure | HR 0.703 | 0.602–0.820 | <0.0001 | Log-rank; Cox regression |
| Time to symptom progression | HR 0.796 | 0.667–0.951 | 0.0121 | Log-rank; Cox regression |
12. Statistical Methodology
Kaplan-Meier estimation
The registry specifies Kaplan-Meier estimation for overall survival. Kaplan-Meier estimation is designed for time-to-event data in which some participants may not experience the event during observed follow-up. Those observations can be right-censored and still contribute information until their censoring time.
where di is the number of events at time ti and ni is the number at risk immediately before that time.
For EMILIA, the registered OS definition is the time from randomization to death from any cause. The registry also states that the median OS was estimated using Kaplan-Meier and that its 95% CI was computed using the Brookmeyer and Crowley method.
Stratified log-rank test
The formal time-to-event comparisons in the analyses posted on ClinicalTrials.gov used the log-rank test. The analyses were stratified by region of enrollment, number of prior chemotherapeutic regimens, and visceral versus non-visceral disease.
Stratification allows the comparison to account for prespecified categories while preserving the time-to-event structure of the data. Rather than simply comparing the percentage of events at one arbitrary time, the log-rank framework uses the ordering of event times across follow-up.
Cox regression and hazard ratios
The registry states that the HR was estimated using Cox regression. A hazard ratio is a relative measure of the instantaneous event hazard under the fitted model.
An HR is not the same as a relative risk, an absolute risk difference, a median-survival ratio, or the proportion of participants who benefit.
Intention-to-treat analysis
The posted analyses define the ITT population as all randomized participants analyzed on the basis of treatment assigned at randomization. This is the natural analysis population for preserving the treatment comparison created by randomization.
Cochran-Mantel-Haenszel analysis
The objective-response endpoint used a Mantel-Haenszel chi-squared test, normalized here as a Cochran-Mantel-Haenszel test. This method provides a way to compare categorical outcomes while accounting for stratification variables.
Wald / z-test framework
The clinical-benefit analysis is identified in the ClinicalTrials.gov record with a Wald / z-test framework. In general, a Wald-type statistic compares an estimated effect with its estimated standard error, producing a standardized quantity for inference under the relevant large-sample approximation.
13. Statistical Methods Explained
Why was the analysis stratified?
The time-to-event analyses were stratified by region of enrollment, number of prior chemotherapeutic regimens, and visceral versus non-visceral disease. Stratification can improve the validity and efficiency of a treatment comparison when these factors are incorporated into the trial's design and analysis plan. It also avoids treating the observed distribution of these factors as if it had arisen independently of the randomized design.
What does an HR of 0.650 mean?
An HR of 0.650 means that the fitted model estimates the instantaneous hazard in the trastuzumab emtansine group at 65.0% of the comparator hazard. The simple arithmetic interpretation is a 35.0% lower estimated hazard. It does not mean a 35.0% absolute reduction in the probability of progression or death.
Why use an ITT population?
ITT analysis keeps participants in the group to which they were randomized. This maintains the comparison generated by randomization and avoids redefining treatment groups according to events that occurred after randomization.
What is the difference between an HR and a response-rate difference?
The HR describes a relative comparison of event hazards over time. The objective-response difference of 12.7 is a difference between two categorical response rates, defined here as trastuzumab emtansine minus lapatinib plus capecitabine. The two measures therefore summarize different aspects of treatment effect.
Why does the confidence interval matter?
A point estimate such as HR 0.749 is only one estimate from the observed data. The corresponding 95% CI of 0.639–0.877 communicates uncertainty around that estimate under the model and inferential framework. It does not give a range in which every individual patient's treatment effect must lie.
Why doesn't the P-value measure the size of the effect?
The P-value addresses the evidence against a specified null hypothesis. Its magnitude depends on both the observed data and the amount of information available. Effect size should therefore be read from the HR or difference, while precision is assessed using the confidence interval.
Why is a time-to-event endpoint different from a binary endpoint?
A binary endpoint records whether an event occurred within a defined framework. A time-to-event endpoint retains information about when the event occurred and can accommodate censoring. This is why PFS, OS, time to treatment failure, and time to symptom progression were analyzed with survival-analysis methods in the ClinicalTrials.gov record.
14. Multiplicity, Interim Analysis, and Hypothesis Type
The ClinicalTrials.gov record identifies superiority as the hypothesis type for all eight posted statistical analyses. Three primary endpoint analyses are formally reported: PFS as assessed by an IRC, OS at the second interim analysis, and final OS.
| Feature | What the ClinicalTrials.gov record establishes |
|---|---|
| Hypothesis type | Superiority |
| Primary endpoints | 8 registered primary endpoints |
| Primary endpoint analyses | 3 formal statistical analyses in the ClinicalTrials.gov record |
| Second interim OS analysis | Reported as the confirmatory analysis |
| Final OS analysis | Reported as descriptive |
| Multiplicity adjustment details | Not provided in the ClinicalTrials.gov record |
| Alpha-spending details | Not provided in the ClinicalTrials.gov record |
| Non-inferiority margin | Not applicable to the registry-reported superiority analyses |
| Bayesian methods | Not identified in the statistical analyses posted on ClinicalTrials.gov |
| Factorial design | Not used; the design model is parallel |
The distinction between the second interim analysis and final analysis is statistically important. The registry explicitly states that the second interim OS analysis was deemed confirmatory, whereas the final analysis is described as descriptive. A later P-value should not automatically be interpreted as if it were associated with a newly defined confirmatory testing budget.
15. Crossover, Missing Data, and Other Design Issues
Crossover
The ClinicalTrials.gov record does not identify a crossover procedure or crossover rate. No crossover adjustment is therefore applied or inferred in this analysis.
Missing data
The ClinicalTrials.gov record does not specify a missing-data or imputation method for the posted analyses.
Non-inferiority margin
No non-inferiority hypothesis or margin is identified. The posted analyses use superiority hypotheses.
Bayesian methods
No Bayesian method is identified among the statistical analyses posted on ClinicalTrials.gov.
These omissions matter because statistical interpretation should follow the documented analysis rather than fill gaps with assumptions. For example, a time-to-event analysis can be sensitive to censoring rules, while a categorical analysis can depend on how participants with unavailable assessments are handled. The ClinicalTrials.gov record does not provide enough information to reconstruct those details.
16. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm. The denominators are presented as affected participants divided by participants at risk.
| Safety group | Serious adverse events | Affected / at risk |
|---|---|---|
| Trastuzumab emtansine | Serious adverse events | 92 / 490 |
| Lapatinib + capecitabine | Serious adverse events | 99 / 488 |
| Lapatinib + capecitabine / Trastuzumab emtansine | Serious adverse events | 19 / 136 |
The third safety group is explicitly identified in the ClinicalTrials.gov record as Lapatinib + Capecitabine/ Trastuzumab Em with 19 affected participants among 136 at risk. Because the ClinicalTrials.gov record does not further define this group's construction, these figures should not be silently combined with either of the two main randomized treatment groups.
17. Understanding the Primary PFS and OS Results Together
The primary efficacy analyses show a consistent direction of estimated treatment effect across two different time-to-event endpoints and two OS analyses.
| Endpoint | HR | 95% CI | P-value | Interpretation of HR |
|---|---|---|---|---|
| IRC-assessed PFS | 0.650 | 0.549–0.771 | <0.0001 | 35.0% lower estimated hazard |
| OS, second interim | 0.682 | 0.548–0.849 | 0.0006 | 31.8% lower estimated hazard |
| OS, final analysis | 0.749 | 0.639–0.877 | 0.0003 | 25.1% lower estimated hazard |
The three estimates should not be treated as interchangeable. PFS measures progression or death, whereas OS measures death from any cause. The second-interim and final OS analyses also correspond to different data cutoffs and different analysis designations in the registry.
The HR estimates are all below 1, meaning that each fitted comparison estimates a lower instantaneous event hazard for trastuzumab emtansine relative to lapatinib plus capecitabine. The numerical HR changes across analyses do not imply that a treatment effect "decayed" by a specific percentage, because each endpoint and analysis has its own information set and follow-up structure.
18. Time-to-Event Endpoints: What Is Actually Being Compared?
EMILIA contains several time-to-event endpoints: IRC-assessed PFS, investigator-assessed PFS, overall survival, time to treatment failure, and time to symptom progression. These endpoints differ in what constitutes an event.
| Endpoint | Event concept | Statistical structure |
|---|---|---|
| IRC-assessed PFS | Progression or death, with progression assessed by an IRC | Time-to-event; stratified log-rank; Cox HR |
| Investigator-assessed PFS | Progression or death assessed by investigators | Time-to-event; stratified log-rank; Cox HR |
| Overall survival | Death from any cause | Time-to-event; Kaplan-Meier; stratified log-rank; Cox HR |
| Time to treatment failure | Treatment-failure endpoint as defined by the trial registry | Time-to-event; stratified log-rank; Cox HR |
| Time to symptom progression | Symptom progression endpoint | Time-to-event; stratified log-rank; Cox HR |
Using multiple time-to-event endpoints can provide complementary information, but it also creates a distinction between the statistical evidence for a particular endpoint and the broader clinical story. The HR for PFS should not be interpreted as if it were an OS HR, and an endpoint's P-value does not establish that every other endpoint must show the same magnitude of effect.
19. Clinical Biostats Interpretation of the Confidence Intervals
The IRC-assessed PFS HR was 0.650 with a 95% CI of 0.549–0.771. The interval provides the uncertainty range around the estimated relative hazard under the specified analysis model. It is more informative than the point estimate alone because it shows how precisely the treatment effect was estimated.
The second-interim OS HR was 0.682 with a 95% CI of 0.548–0.849. The interval remains below 1, while still showing uncertainty around the magnitude of the relative effect.
The final OS HR was 0.749 with a 95% CI of 0.639–0.877. The estimate remains below 1, but the final-analysis point estimate differs from the second-interim estimate. Such differences are expected when additional follow-up and events change the information available for estimation.
20. Why the Objective-Response Analysis Uses a Different Method
Objective response is a categorical outcome rather than a time-to-event endpoint. The registry-reported analysis therefore uses a Cochran-Mantel-Haenszel test rather than a log-rank test.
The registry-reported estimate is 12.7, with a two-sided 95% CI of 6.0–19.4.
The difference is an absolute effect measure on the percentage-point scale. That makes it fundamentally different from an HR. The response analysis also has a specific analysis restriction: only participants with measurable disease at baseline were included.
The confidence interval was computed using an approximate normal method. This is another reason to distinguish the response analysis from the survival analyses: different outcome structures lead naturally to different estimators, test statistics, and confidence intervals.
21. Limitations and Interpretation Issues
- Registry-level detail: The ClinicalTrials.gov record contains eight formal statistical analyses but do not provide every implementation detail that would normally be available in a complete statistical analysis plan.
- Endpoint heterogeneity: PFS, OS, response, clinical benefit, treatment failure, and symptom progression measure different outcomes and should not be collapsed into a single statistic.
- Different data cutoffs: The PFS analysis uses 14 Jan 2012, the second-interim OS analysis uses 31 Jul 2012, and the final OS analysis uses 31 Dec 2014. Estimates from these analyses therefore represent different information sets.
- Final analysis designation: The registry describes the final OS analysis as descriptive, while the second interim analysis was deemed confirmatory. This distinction affects how the corresponding P-value should be interpreted.
- Proportional-hazards assumption: Cox HRs summarize relative hazards under a proportional-hazards model. A single HR may be less informative if the relative hazard changes substantially over time.
- Censoring: Time-to-event analyses depend on how participants are followed and censored. The ClinicalTrials.gov record does not provide the complete censoring rules.
- Response population: Objective response and clinical benefit analyses were restricted to participants with measurable disease at baseline, unlike the general ITT definition.
- Symptom-progression population: The registry-reported analysis specifies female participants with a baseline assessment and at least 1 follow-up assessment, creating a different analysis restriction from the general ITT description.
- Multiplicity: The ClinicalTrials.gov record identifies multiple primary endpoints and secondary analyses but do not provide the complete multiplicity-control procedure. Individual P-values should therefore be interpreted in their documented testing context rather than assumed to represent independent hypotheses.
- Safety-group interpretation: The third reported serious-adverse-event group has a different denominator and is labeled as a lapatinib-plus-capecitabine / trastuzumab emtansine group. It should not be assumed to represent a third randomized arm.
22. Why This Trial Matters Statistically
EMILIA is a useful teaching case because its registry record connects several core clinical-trial methods within one randomized phase 3 study. It includes independent-review PFS, investigator-assessed PFS, overall survival at an interim analysis and final analysis, categorical response outcomes, stratified time-to-event methods, and ITT analysis.
| Concept | How it appears in EMILIA |
|---|---|
| Randomization | Randomized, parallel-group phase 3 design |
| ITT analysis | Primary and secondary efficacy analyses use the ITT population |
| Kaplan-Meier estimation | Registered method for estimating median OS |
| Hazard ratio | Primary and secondary time-to-event effect measure |
| Confidence intervals | Reported around primary HRs and categorical differences |
| Log-rank test | Primary and secondary time-to-event comparisons |
| Cox regression | HR estimation for the time-to-event analyses |
| Stratified analysis | Region, prior chemotherapy regimen count, and visceral/non-visceral disease |
| Cochran-Mantel-Haenszel test | IRC-assessed objective-response comparison |
| Wald / z-test | Clinical-benefit analysis framework |
| Interim analysis | Second interim OS analysis designated confirmatory |
| Descriptive final analysis | Final OS analysis explicitly described as descriptive |
23. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
24. Related Statistical Calculators
Apply the same statistical concepts with related Clinical Biostats calculators:
25. Sources
- ClinicalTrials.gov: NCT00829166 — EMILIA.
- PubMed: PMID 28526536.
- PubMed: PMID 23020162.
- PubMed: PMID 22437872.
Continue through the Clinical Biostats statistical library
Explore the underlying methods used in randomized clinical trials, from survival analysis and hazard ratios to categorical-data tests and confidence intervals.
26. Record Summary
EMILIA provides a compact example of how randomized clinical-trial evidence can be analyzed across several endpoint structures. The ClinicalTrials.gov record includes a randomized phase 3 parallel design, an ITT analysis population, stratified log-rank comparisons, Cox-derived hazard ratios, Kaplan-Meier estimation for overall survival, a Cochran-Mantel-Haenszel analysis for objective response, and a Wald / z-test framework for clinical benefit.
The three formal primary analyses show HR estimates of 0.650 for IRC-assessed PFS, 0.682 for OS at the second interim analysis, and 0.749 for final OS. Their corresponding 95% confidence intervals and P-values describe statistical uncertainty and evidence within their respective analysis frameworks. The secondary analyses extend the same statistical story to investigator-assessed PFS, objective response, clinical benefit, treatment failure, and symptom progression.