This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the registry-reported EMBER-3 trial data.
1. Trial at a Glance
EMBER-3 was a randomized, parallel, open-label phase 3 trial evaluating three treatment strategies in participants with ER+, HER2- advanced breast cancer. The registry reports 874 enrolled participants, three arms, three registered primary endpoints, and three posted formal primary statistical analyses.
| Feature | EMBER-3 |
|---|---|
| Trial | EMBER-3 |
| NCT identifier | NCT04975308 |
| Phase | Phase 3 |
| Therapeutic area | Oncology |
| Condition | Breast Neoplasms; Neoplasm Metastasis |
| Population | Participants with ER+, HER2- advanced breast cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 874 |
| Arms | 3 |
| Trial status | Active, not recruiting |
| Start date | October 4, 2021 |
| Primary completion | June 24, 2024 |
| Lead sponsor | Eli Lilly and Company |
| Sponsor type | Industry |
2. Clinical Question
The statistical questions in EMBER-3 are organized around randomized comparisons of investigator-assessed progression-free survival. The registry reports one primary PFS comparison of imlunestrant versus investigator's choice of endocrine therapy, a second primary PFS comparison of imlunestrant plus abemaciclib versus imlunestrant, and a third primary PFS analysis restricted to the ESR1-mutation detected population comparing imlunestrant with investigator's choice of endocrine therapy.
Population
Participants with ER+, HER2- advanced breast cancer, with the ESR1-mutation detected population defined for the third primary analysis as participants in arms A or B with ESR1 mutations detected at baseline.
Intervention
Imlunestrant in Arm A, and imlunestrant plus abemaciclib in Arm C for the combination comparison.
Comparator
Investigator's choice of endocrine therapy in Arm B for the Arm A versus Arm B comparisons, and imlunestrant in Arm A for the Arm C versus Arm A comparison.
Primary questions
How does investigator-assessed PFS compare between the randomized treatment strategies, including the prespecified ESR1-mutation detected population?
3. Trial Design
Imlunestrant
- Imlunestrant
- Randomized comparison with investigator's choice of endocrine therapy
- Randomized comparison with imlunestrant plus abemaciclib
Investigator's Choice of Endocrine Therapy
- Investigator's choice of endocrine therapy
- Registered intervention includes exemestane
- Registered intervention includes fulvestrant
Imlunestrant + Abemaciclib
- Imlunestrant
- Abemaciclib
- Compared with imlunestrant in the primary Arm C versus Arm A analysis
The absence of masking is an important design characteristic when interpreting an investigator-assessed endpoint. PFS is a time-to-event outcome, but its event definition includes investigator-assessed disease progression using RECIST version 1.1 criteria. Because treatment assignment was not masked, assessment processes can be a potential source of bias even though the randomized treatment comparison remains the central basis for causal inference.
4. Primary Endpoints
| Registered primary endpoint | Time frame | Type | Formal analysis |
|---|---|---|---|
| Investigator-assessed Progression Free Survival (PFS) (Between Arm A and Arm B) | Randomization to the date of first documented progression of disease or death from any cause (up to 28 months) | Time-to-event | Log-rank test; hazard ratio |
| Investigator-assessed PFS (Between Arm C and Arm A) | Randomization to the date of first documented progression of disease or death from any cause (up to 26 months) | Time-to-event | Log-rank test; hazard ratio |
| Investigator-assessed PFS in the Estrogen Receptor 1 (ESR1)-Mutation Detected Population (Between Arm A and Arm B) | Randomization to the date of first documented progression of disease or death from any cause (up to 28 months) | Time-to-event | Log-rank test; hazard ratio |
Registry endpoint definition
For each of the three registered primary endpoints, PFS was defined as the time from randomization to the date of first documented progression of disease or death from any cause in the absence of disease progression, using Response Evaluation Criteria in Solid Tumors (RECIST) version 1.1 criteria, as assessed by investigator. The registry definition states that progressive disease was defined as at least a 20% increase in the sum of the diameters of target lesions, with reference to the registry definition.
5. Statistical Methodology
Log-rank test
All three posted primary analyses used the log-rank test. The log-rank test is designed to compare time-to-event distributions between randomized groups while incorporating the timing of events and accommodating right censoring.
The registry identifies the hypothesis type for each primary analysis as superiority. The reported effect measure is a hazard ratio, accompanied by a two-sided 95% confidence interval.
Hazard ratio
The effect measure reported for all three primary analyses is the hazard ratio (HR). The hazard ratio compares the estimated instantaneous event rates between the randomized groups over the analyzed follow-up.
For the Arm C versus Arm A analysis, for example, an HR of 0.569 means the estimated hazard of progression or death was 0.569 times that of Arm A under the reported analysis. It does not mean that 56.9% of patients avoided progression, nor does it directly provide an absolute difference in PFS.
Kaplan-Meier estimation
Although the ClinicalTrials.gov record identifies the formal comparison method as a log-rank test rather than separately reporting Kaplan-Meier estimates, PFS is a time-to-event endpoint for which Kaplan-Meier estimation is the standard descriptive framework. Kaplan-Meier curves estimate the probability of remaining event-free over time while retaining information from censored participants up to their censoring times.
Here, di represents the number of events at event time ti, while ni represents the number at risk immediately before that time.
Censoring
The ClinicalTrials.gov record explicitly identify censored participants. In the Arm A versus Arm B analysis, 94 participants in Arm A and 77 in Arm B were censored. In the Arm C versus Arm A analysis, 99 participants in Arm C and 64 in Arm A were censored. In the ESR1-mutation detected analysis, 29 participants in Arm A and 16 in Arm B were censored.
Censoring is not equivalent to a treatment failure. A censored participant contributes information to the analysis until the time at which the participant is censored. The validity of standard survival analysis depends on assumptions about the relationship between censoring and the event process.
Analysis populations
The first primary analysis used all participants randomly assigned to either Arm A or Arm B, including censored participants. The second used all participants randomly assigned to either Arm A or Arm C concurrently, including censored participants. The ESR1 analysis used participants randomly assigned to Arm A or Arm B who had ESR1 mutations detected at baseline, again including censored participants.
| Primary comparison | Analysis population | Censored observations reported |
|---|---|---|
| Arm A vs Arm B | All participants randomly assigned to Arm A or Arm B, including censored | Arm A = 94; Arm B = 77 |
| Arm C vs Arm A | All participants randomly assigned to Arm A or Arm C concurrently, including censored | Arm C = 99; Arm A = 64 |
| ESR1-mutation detected: Arm A vs Arm B | All participants randomly assigned to Arm A or Arm B with ESR1 mutations detected at baseline, including censored | Arm A = 29; Arm B = 16 |
6. Primary Results
ClinicalTrials.gov reports three formal primary statistical analyses. Each uses a log-rank test and reports a hazard ratio with a two-sided 95% confidence interval. The following results are presented exactly as reported in the ClinicalTrials.gov record.
6.1 Investigator-Assessed PFS: Arm A vs Arm B
The first primary comparison evaluated investigator-assessed PFS between imlunestrant and investigator's choice of endocrine therapy.
Hazard ratio for progression or death
95% CI: 0.724–1.039 · P = 0.1158
Comparison: Arm A, imlunestrant vs Arm B, investigator's choice of endocrine therapy
| Analysis feature | Reported result |
|---|---|
| Endpoint | Investigator-assessed Progression Free Survival (PFS) (Between Arm A and Arm B) |
| Time frame | Randomization to the date of first documented progression of disease or death from any cause (up to 28 months) |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.867 |
| 95% CI | 0.724–1.039 |
| P-value | 0.1158 |
| Hypothesis | Superiority |
The estimated hazard ratio of 0.867 is below 1, so the point estimate corresponds to an estimated hazard of progression or death that is approximately 86.7% of the hazard in the investigator's-choice group under the reported analysis. Expressed as a relative difference in the estimated hazard, this corresponds to approximately a 13.3% lower estimated hazard for Arm A.
That interpretation applies to the estimated hazard; it does not mean that 13.3% of patients benefited, that PFS was 13.3% longer, or that the probability of progression was reduced by exactly 13.3% for every participant.
The 95% confidence interval, 0.724–1.039, describes uncertainty around the estimated hazard ratio. Because the interval extends across 1, the data are compatible with both a lower and a higher hazard under the statistical model and sampling framework. The p-value of 0.1158 addresses the statistical evidence against the specified null hypothesis; it is not a measure of the magnitude or clinical importance of the estimated effect.
Because PFS is a time-to-event endpoint, the interpretation also depends on appropriate handling of censoring and on the assumptions underlying the hazard-ratio framework. The ClinicalTrials.gov record does not provide a separate test of proportional hazards, so the HR should not be interpreted as an assumption-free description of the entire PFS experience.
6.2 Investigator-Assessed PFS: Arm C vs Arm A
The second primary comparison evaluated whether adding abemaciclib to imlunestrant changed investigator-assessed PFS relative to imlunestrant alone.
Hazard ratio for progression or death
95% CI: 0.441–0.733 · P < 0.0001
Comparison: Arm C, imlunestrant + abemaciclib vs Arm A, imlunestrant
| Analysis feature | Reported result |
|---|---|
| Endpoint | Investigator-assessed PFS (Between Arm C and Arm A) |
| Time frame | Randomization to the date of first documented progression of disease or death from any cause (up to 26 months) |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.569 |
| 95% CI | 0.441–0.733 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
The hazard ratio of 0.569 indicates that the estimated instantaneous hazard of progression or death in Arm C was 0.569 times that in Arm A under the reported analysis. Put another way, the point estimate corresponds to an approximately 43.1% lower estimated hazard for Arm C relative to Arm A.
This is a relative time-to-event measure. It does not mean that 43.1% of participants avoided progression, that individual patients experienced a 43.1% reduction in their personal risk, or that the median PFS differed by 43.1%.
The 95% confidence interval of 0.441–0.733 gives a range of values compatible with the estimated effect under the analysis framework. The entire interval is below 1, which is consistent with a lower estimated hazard for Arm C across the interval of uncertainty represented by the reported confidence interval.
The p-value of <0.0001 quantifies the statistical evidence against the null hypothesis used for the comparison. It does not tell us that the treatment effect is large, clinically important, or certain. Effect size and precision are communicated by the HR and confidence interval; the p-value addresses evidence against the null.
As with any hazard-ratio analysis, interpretation requires attention to censoring and the proportional-hazards assumption. The ClinicalTrials.gov record does not report a separate assessment of proportional hazards.
6.3 Investigator-Assessed PFS in the ESR1-Mutation Detected Population
The third primary analysis restricted the Arm A versus Arm B comparison to participants who had ESR1 mutations detected at baseline.
Hazard ratio for progression or death
95% CI: 0.464–0.821 · P = 0.0008
Comparison: Arm A, imlunestrant vs Arm B, investigator's choice of endocrine therapy
| Analysis feature | Reported result |
|---|---|
| Endpoint | Investigator-assessed PFS in the Estrogen Receptor 1 (ESR1)-Mutation Detected Population (Between Arm A and Arm B) |
| Time frame | Randomization to the date of first documented progression of disease or death from any cause (up to 28 months) |
| Population | Participants randomly assigned to Arm A or Arm B with ESR1 mutations detected at baseline |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.617 |
| 95% CI | 0.464–0.821 |
| P-value | 0.0008 |
| Hypothesis | Superiority |
The hazard ratio of 0.617 means that the estimated instantaneous hazard of progression or death in Arm A was 0.617 times that in Arm B for the ESR1-mutation detected population under the reported analysis. The point estimate therefore corresponds to approximately a 38.3% lower estimated hazard for Arm A relative to Arm B.
The estimate does not mean that 38.3% of participants were protected from progression, nor does it establish an absolute PFS improvement of 38.3%. It is a relative time-to-event measure.
The 95% confidence interval of 0.464–0.821 does not cross 1, indicating that the range of effects represented by this interval is on the lower-hazard side of the null value. The interval is also substantially wider than a point estimate alone can communicate, illustrating why precision should always accompany a hazard ratio.
The p-value of 0.0008 measures the evidence against the relevant null hypothesis; it does not measure the size of the treatment effect. The HR and its confidence interval provide the effect-size information.
This is a population-restricted primary analysis. It should not automatically be interpreted as proof that ESR1 mutation status modifies treatment effect unless a formal interaction analysis or other appropriate comparison of treatment effects supports such a conclusion. The ClinicalTrials.gov record does not report an interaction test.
7. Comparing the Three Primary Analyses
| Primary endpoint | Comparison | HR | 95% CI | P-value |
|---|---|---|---|---|
| Investigator-assessed PFS | Arm A vs Arm B | 0.867 | 0.724–1.039 | 0.1158 |
| Investigator-assessed PFS | Arm C vs Arm A | 0.569 | 0.441–0.733 | <0.0001 |
| Investigator-assessed PFS, ESR1-mutation detected population | Arm A vs Arm B | 0.617 | 0.464–0.821 | 0.0008 |
The three analyses answer different questions and should not be collapsed into a single overall treatment effect. The first asks about imlunestrant versus investigator's choice of endocrine therapy in the broader randomized comparison. The second asks about adding abemaciclib to imlunestrant. The third asks the Arm A versus Arm B question within the baseline ESR1-mutation detected population.
There is also an important distinction between the point estimates and the evidence reported by the confidence intervals and p-values. The first estimate is below 1 but has a 95% confidence interval extending above 1 and a p-value of 0.1158. The second and third confidence intervals remain below 1, with p-values of <0.0001 and 0.0008, respectively. These are descriptions of the reported statistical results rather than a ranking of the treatment strategies.
8. Secondary Endpoints and Other Posted Outcomes
The registry profile posted on ClinicalTrials.gov for EMBER-3 reports 21 outcome measures and three posted statistical analyses. However, the ClinicalTrials.gov record identifies only the three formal primary statistical analyses and do not provide numerical results for additional secondary outcome measures.
For a time-to-event secondary endpoint, an appropriate analysis would ordinarily use a survival-analysis framework such as Kaplan-Meier estimation together with a log-rank comparison and a hazard-ratio model when prespecified by the statistical analysis plan. The exact method, effect estimate, confidence interval, and multiplicity treatment should be taken from the corresponding registered analysis rather than inferred from the primary PFS results.
9. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm as affected participants over participants at risk. These entries should be reproduced as reported rather than converted into additional safety statistics.
| Arm | Serious adverse events affected / at risk |
|---|---|
| Arm A: Imlunestrant | 46/327 |
| Arm B: Investigator's Choice of Endocrine Therapy | 6/32 |
| Arm B: Investigator's Choice of Endocrine Therapy | 49/292 |
| Arm C: Imlunestrant + Abemaciclib | 42/208 |
The two Arm B entries are preserved separately because the ClinicalTrials.gov record contains them as separate affected/at-risk records. They should not be combined into a new denominator or interpreted as a single safety estimate without the underlying registry context that explains the two entries.
10. Statistical Methods Explained
Why was a log-rank test used?
PFS is a time-to-event endpoint. A simple comparison of the proportion of participants who progressed would discard information about when progression occurred and how long censored participants were observed. The log-rank test instead compares the event experience over follow-up while accommodating right censoring.
What does an HR of 0.569 mean?
An HR of 0.569 means that, under the reported time-to-event analysis, the estimated instantaneous hazard of progression or death in Arm C was 0.569 times that in Arm A. The corresponding point estimate can be described as approximately a 43.1% lower estimated hazard. It does not mean that 43.1% of patients avoided progression or that PFS increased by exactly 43.1%.
Why does the confidence interval matter?
A hazard ratio is an estimate, not a certainty. The 95% confidence interval describes the uncertainty around that estimate under the statistical model and sampling framework. For the Arm C versus Arm A analysis, the interval is 0.441–0.733; for the Arm A versus Arm B analysis, it is 0.724–1.039. These intervals communicate much more information about precision than the point estimates alone.
Why does a p-value not measure effect size?
A p-value addresses the strength of evidence against a null hypothesis under the specified statistical model. It does not tell us how large the treatment effect is. A hazard ratio and its confidence interval are needed to describe the magnitude and precision of the estimated relative effect.
Why is the ESR1 analysis not automatically an interaction test?
The ESR1-mutation detected analysis is a restricted population comparison. A treatment effect estimated within one biomarker-defined group does not by itself establish that the treatment effect differs from the effect in another group. To demonstrate effect modification, the appropriate question is whether the treatment-by-biomarker interaction is supported by a formal statistical test or equivalent prespecified heterogeneity analysis. No such interaction result is included in the ClinicalTrials.gov record.
What does censoring mean in these PFS analyses?
A censored participant is not automatically classified as having had a favorable or unfavorable outcome. Instead, the participant contributes observed follow-up information until the censoring time. The survival analysis then uses the available event and censoring information to estimate the time-to-event distribution. The validity of that approach depends on assumptions concerning the censoring mechanism.
Why is the absence of masking statistically relevant?
The registry describes EMBER-3 as having no masking. Randomization can balance measured and unmeasured prognostic factors in expectation, but lack of masking can still matter for outcomes involving investigator assessment. Because the primary PFS endpoint includes investigator-assessed progression under RECIST version 1.1, the possibility of assessment-related bias is a relevant limitation when interpreting the reported hazard ratios.
11. Multiplicity and the Three Primary Analyses
EMBER-3 has three registered primary endpoints and three posted primary analyses. That structure is statistically important because multiple confirmatory questions can create a multiplicity problem if each hypothesis is tested independently at the same nominal significance level.
| Primary question | Reported hypothesis type | Reported analysis | P-value |
|---|---|---|---|
| Arm A vs Arm B PFS | Superiority | Log-rank test; HR 0.867 | 0.1158 |
| Arm C vs Arm A PFS | Superiority | Log-rank test; HR 0.569 | <0.0001 |
| ESR1-mutation detected Arm A vs Arm B PFS | Superiority | Log-rank test; HR 0.617 | 0.0008 |
The ClinicalTrials.gov record identifies these as primary analyses but do not provide an alpha-allocation strategy, gatekeeping procedure, hierarchical testing sequence, or other multiplicity-adjustment details. Therefore, this page does not infer one.
12. Non-Inferiority, Bayesian Methods, Crossover, and Interim Analysis
The registry-reported EMBER-3 trial data identify all three primary hypotheses as superiority. No non-inferiority margin is reported, so no non-inferiority interpretation is appropriate for these primary analyses.
The registry-reported statistical-method fields identify the log-rank test as the method and the hazard ratio as the effect measure. No Bayesian method is reported. No crossover analysis is reported in the ClinicalTrials.gov record. No interim-analysis or alpha-spending procedure is reported in the ClinicalTrials.gov record.
| Design topic | What the ClinicalTrials.gov record supports |
|---|---|
| Hypothesis type | Superiority |
| Non-inferiority margin | Not reported in the ClinicalTrials.gov record |
| Bayesian analysis | Not reported in the ClinicalTrials.gov record |
| Crossover | Not reported in the ClinicalTrials.gov record |
| Interim analysis | Not reported in the ClinicalTrials.gov record |
| Alpha spending | Not reported in the ClinicalTrials.gov record |
| Missing-data imputation | Not reported in the ClinicalTrials.gov record |
These omissions should not be interpreted as evidence that the underlying protocol or statistical analysis plan contains no such procedures. They mean only that the ClinicalTrials.gov record does not provide them, so they are not reconstructed here.
13. Randomization and Analysis Populations
Randomization is central to the causal interpretation of the primary comparisons. The registry describes EMBER-3 as randomized and parallel, with no masking. The primary analyses were defined around participants randomly assigned to the relevant treatment groups.
Arm A vs Arm B
All participants randomly assigned to either Arm A or Arm B, including censored participants, were included in the reported analysis population.
Arm C vs Arm A
All participants randomly assigned to either Arm A or Arm C concurrently, including censored participants, were included in the reported analysis population.
ESR1 analysis
The analysis was restricted to participants randomly assigned to Arm A or Arm B who had ESR1 mutations detected at baseline.
Why randomization matters
Random assignment provides the basis for comparing outcomes across treatment strategies without relying solely on adjustment for measured baseline characteristics.
The ESR1 analysis illustrates an important distinction: randomization remains relevant within the selected population, but restricting an analysis to a biomarker-defined subgroup generally reduces the available information and can increase statistical uncertainty.
14. Interpretation of Hazard Ratios
The reported HR of 0.867 corresponds to a point estimate below 1, but the 95% CI of 0.724–1.039 crosses 1. The appropriate statistical description is therefore the complete estimate, interval, and p-value—not simply the direction of the point estimate.
The reported HR of 0.569 indicates a lower estimated hazard in Arm C relative to Arm A. The 95% CI of 0.441–0.733 remains below 1, providing the reported precision around that estimate.
The reported HR of 0.617 indicates a lower estimated hazard in Arm A relative to Arm B within the ESR1-mutation detected population. Its 95% CI of 0.464–0.821 remains below 1.
These three interpretations should remain tied to their respective populations and comparisons. A hazard ratio is not a universal property of a drug independent of the comparison group, endpoint definition, follow-up period, or analysis population.
15. Confidence Intervals and Statistical Precision
Confidence intervals are especially useful for distinguishing a point estimate from the range of uncertainty surrounding it.
| Comparison | HR | 95% CI width and position relative to 1 | Statistical interpretation |
|---|---|---|---|
| Arm A vs Arm B | 0.867 | 0.724–1.039; crosses 1 | The interval includes the null hazard ratio. |
| Arm C vs Arm A | 0.569 | 0.441–0.733; below 1 | The reported interval is entirely below the null hazard ratio. |
| ESR1 Arm A vs Arm B | 0.617 | 0.464–0.821; below 1 | The reported interval is entirely below the null hazard ratio. |
The confidence interval does not describe the range of treatment effects experienced by individual patients. It describes uncertainty around the estimated population-level effect under the statistical framework used for the analysis.
16. P-Values in Context
The three reported p-values are 0.1158, <0.0001, and 0.0008. They should be interpreted in the context of the corresponding comparison, endpoint, analysis population, and multiplicity structure.
A p-value is calculated under a null hypothesis and reflects the compatibility of the observed data, or more extreme data, with that null under the statistical model. It does not directly provide the probability that the null hypothesis is true.
Similarly, a very small p-value does not imply a large treatment effect. EMBER-3 illustrates why the HR and confidence interval should be read alongside the p-value rather than replaced by it.
17. Missing Data and Imputation
The ClinicalTrials.gov record does not report a missing-data or imputation strategy for the primary PFS analyses. Because PFS is a time-to-event endpoint, missing follow-up and censoring are handled differently from missing continuous or binary measurements.
For a standard survival analysis, the key issue is whether censoring can be treated as sufficiently independent of the future event process conditional on the information used by the analysis. The ClinicalTrials.gov record identifies censored observations but do not provide enough information to evaluate the censoring mechanism or any sensitivity analyses around it.
18. Trial Timeline
Trial start
The EMBER-3 trial began on October 4, 2021 according to the registry profile.
Primary completion
the ClinicalTrials.gov record lists June 24, 2024 as the primary completion date.
Active, not recruiting
The ClinicalTrials.gov profile classifies EMBER-3 as active, not recruiting.
19. Why This Trial Matters Statistically
EMBER-3 is a useful teaching case because the ClinicalTrials.gov record bring together several central ideas in clinical-trial biostatistics: randomized comparisons, time-to-event endpoints, censoring, log-rank testing, hazard ratios, confidence intervals, p-values, and biomarker-defined analysis populations.
| Statistical concept | How it appears in EMBER-3 |
|---|---|
| Randomization | The trial is randomized and uses a parallel design. |
| Three-arm design | The registry describes imlunestrant, investigator's choice of endocrine therapy, and imlunestrant plus abemaciclib. |
| Time-to-event endpoint | All three primary endpoints are investigator-assessed PFS analyses. |
| Log-rank test | The reported formal method for all three primary statistical analyses. |
| Hazard ratio | The reported effect measure for all three primary analyses. |
| Confidence interval | Each primary analysis reports a two-sided 95% CI. |
| P-value | Each primary analysis reports a p-value. |
| Censoring | Censored participant counts are explicitly reported for each primary comparison. |
| Biomarker-defined population | The third primary analysis is restricted to participants with ESR1 mutations detected at baseline. |
| Multiplicity | Three primary analyses create a setting in which the error-control strategy matters. |
| Open-label assessment | The registry describes the trial as having no masking, relevant to investigator-assessed PFS. |
20. Important Limitations and Interpretation Issues
- Three primary analyses: the ClinicalTrials.gov record reports three primary endpoints and three formal analyses, but do not provide the multiplicity-control procedure. Nominal p-values therefore should not be assigned an unreported familywise-error interpretation.
- Investigator assessment: PFS is investigator-assessed. Because the trial is unmasked, assessment-related bias is a relevant consideration when interpreting progression endpoints.
- Hazard-ratio assumptions: a single HR summarizes a relative time-to-event effect under a model. The ClinicalTrials.gov record does not report a separate assessment of proportional hazards.
- Censoring: the primary analyses include censored observations, but the ClinicalTrials.gov record does not provide sufficient information to assess the censoring mechanism or sensitivity to alternative censoring assumptions.
- Biomarker restriction: the ESR1-mutation detected analysis uses a more restricted population than the broader Arm A versus Arm B analysis, so its precision and interpretation are population-specific.
- Missing-data methods: no imputation or missing-data sensitivity strategy is reported in the ClinicalTrials.gov record.
- Unreported design details: the ClinicalTrials.gov record does not report an interim-analysis strategy, alpha spending, Bayesian methods, crossover, or a non-inferiority margin.
- Safety denominator context: the registry-reported serious-adverse-event data contain two separate Arm B affected/at-risk entries. They are reproduced without combining them or inferring an additional safety denominator.
- Secondary results: the registry profile reports 21 outcome measures, but numerical secondary statistical analyses were not included in the ClinicalTrials.gov recordset. This page therefore does not construct secondary results from the primary analyses.
21. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The three reported primary PFS analyses used log-rank tests and hazard ratios. The Arm A versus Arm B estimate was 0.867 with a 95% CI of 0.724–1.039; the Arm C versus Arm A estimate was 0.569 with a 95% CI of 0.441–0.733; and the ESR1-mutation detected Arm A versus Arm B estimate was 0.617 with a 95% CI of 0.464–0.821.
Clinical interpretation
The statistical results describe differences in the time-to-progression-or-death endpoint between the specified randomized groups. They do not, by themselves, describe every dimension of benefit, toxicity, patient experience, or long-term outcome.
A statistically estimated hazard ratio should therefore be treated as one component of the evidence. Absolute event probabilities, median event times, safety, quality of life, subsequent therapy, and other clinical outcomes can add information when those data are available. The registry-reported EMBER-3 dataset does not provide those additional numerical efficacy measures, so they are not inferred here.
22. What the Three Primary Hazard Ratios Do — and Do Not — Mean
The point estimate indicates a lower estimated hazard in Arm A, corresponding to approximately a 13.3% lower estimated hazard relative to Arm B. The 95% CI of 0.724–1.039 includes 1, so the uncertainty around the estimate is important to the interpretation.
The point estimate corresponds to approximately a 43.1% lower estimated hazard in Arm C relative to Arm A. The 95% CI of 0.441–0.733 remains below 1. This does not mean that every participant experienced a 43.1% reduction in risk.
The point estimate corresponds to approximately a 38.3% lower estimated hazard in Arm A relative to Arm B among participants with ESR1 mutations detected at baseline. The result is specific to that analysis population and should not automatically be interpreted as evidence of biomarker interaction.
The distinction between relative hazard and
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Sources
- ClinicalTrials.gov: EMBER-3, NCT04975308.
- Linked publication: PubMed record, PMID 39660834.
Continue learning from the statistical methods
The EMBER-3 results illustrate how randomized treatment comparisons, time-to-event endpoints, hazard ratios, confidence intervals, log-rank testing, censoring, and biomarker-defined populations fit together in clinical-trial analysis.
26. Record Summary
EMBER-3 provides a useful example of a randomized phase 3 trial with a three-arm parallel design and multiple primary time-to-event questions. The registry reports three formal investigator-assessed PFS analyses, all using log-rank testing with hazard ratios and two-sided 95% confidence intervals. The Arm A versus Arm B comparison produced an HR of 0.867 (95% CI 0.724–1.039; P = 0.1158), the Arm C versus Arm A comparison produced an HR of 0.569 (95% CI 0.441–0.733; P < 0.0001), and the ESR1-mutation detected Arm A versus Arm B analysis produced an HR of 0.617 (95% CI 0.464–0.821; P = 0.0008).
The statistical story is therefore not simply a collection of p-values. It includes the randomized comparisons, the precise definition of PFS, censoring, the selected analysis populations, the interpretation of hazard ratios, the uncertainty represented by confidence intervals, and the multiplicity implications of having three primary analyses. The open-label design and investigator assessment are also relevant to interpretation.