← Clinical Trials
ER+, HER2- Advanced Breast Cancer Phase 3 Randomized NCT04975308

EMBER-3: Complete Statistical Analysis of Imlunestrant in ER+, HER2- Advanced Breast Cancer

An independent statistical analysis of the randomized phase 3 EMBER-3 trial evaluating imlunestrant, investigator's choice of endocrine therapy, and imlunestrant plus abemaciclib in participants with ER+, HER2- advanced breast cancer.

EMBER-3  ·  Phase 3  ·  Enrollment 874  ·  Primary completion June 24, 2024
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the registry-reported EMBER-3 trial data.

1. Trial at a Glance

EMBER-3 was a randomized, parallel, open-label phase 3 trial evaluating three treatment strategies in participants with ER+, HER2- advanced breast cancer. The registry reports 874 enrolled participants, three arms, three registered primary endpoints, and three posted formal primary statistical analyses.

874
Enrolled
ClinicalTrials.gov registry
3
Treatment arms
Parallel randomized design
0.569
Arm C vs Arm A PFS HR
95% CI 0.441–0.733
0.617
ESR1 PFS HR
95% CI 0.464–0.821
FeatureEMBER-3
TrialEMBER-3
NCT identifierNCT04975308
PhasePhase 3
Therapeutic areaOncology
ConditionBreast Neoplasms; Neoplasm Metastasis
PopulationParticipants with ER+, HER2- advanced breast cancer
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment874
Arms3
Trial statusActive, not recruiting
Start dateOctober 4, 2021
Primary completionJune 24, 2024
Lead sponsorEli Lilly and Company
Sponsor typeIndustry

2. Clinical Question

The statistical questions in EMBER-3 are organized around randomized comparisons of investigator-assessed progression-free survival. The registry reports one primary PFS comparison of imlunestrant versus investigator's choice of endocrine therapy, a second primary PFS comparison of imlunestrant plus abemaciclib versus imlunestrant, and a third primary PFS analysis restricted to the ESR1-mutation detected population comparing imlunestrant with investigator's choice of endocrine therapy.

Population

Participants with ER+, HER2- advanced breast cancer, with the ESR1-mutation detected population defined for the third primary analysis as participants in arms A or B with ESR1 mutations detected at baseline.

Intervention

Imlunestrant in Arm A, and imlunestrant plus abemaciclib in Arm C for the combination comparison.

Comparator

Investigator's choice of endocrine therapy in Arm B for the Arm A versus Arm B comparisons, and imlunestrant in Arm A for the Arm C versus Arm A comparison.

Primary questions

How does investigator-assessed PFS compare between the randomized treatment strategies, including the prespecified ESR1-mutation detected population?

3. Trial Design

01
Randomize 874 enrolled
02
Arm A Imlunestrant
03
Arm B Investigator's choice of endocrine therapy
04
Arm C Imlunestrant + abemaciclib
05
Assess PFS Investigator-assessed
ARM A

Imlunestrant

  • Imlunestrant
  • Randomized comparison with investigator's choice of endocrine therapy
  • Randomized comparison with imlunestrant plus abemaciclib
ARM B

Investigator's Choice of Endocrine Therapy

  • Investigator's choice of endocrine therapy
  • Registered intervention includes exemestane
  • Registered intervention includes fulvestrant
ARM C

Imlunestrant + Abemaciclib

  • Imlunestrant
  • Abemaciclib
  • Compared with imlunestrant in the primary Arm C versus Arm A analysis
Allocation
Randomized
Design model
Parallel
Masking
None
Primary purpose
Treatment

The absence of masking is an important design characteristic when interpreting an investigator-assessed endpoint. PFS is a time-to-event outcome, but its event definition includes investigator-assessed disease progression using RECIST version 1.1 criteria. Because treatment assignment was not masked, assessment processes can be a potential source of bias even though the randomized treatment comparison remains the central basis for causal inference.

4. Primary Endpoints

Registered primary endpointTime frameTypeFormal analysis
Investigator-assessed Progression Free Survival (PFS) (Between Arm A and Arm B) Randomization to the date of first documented progression of disease or death from any cause (up to 28 months) Time-to-event Log-rank test; hazard ratio
Investigator-assessed PFS (Between Arm C and Arm A) Randomization to the date of first documented progression of disease or death from any cause (up to 26 months) Time-to-event Log-rank test; hazard ratio
Investigator-assessed PFS in the Estrogen Receptor 1 (ESR1)-Mutation Detected Population (Between Arm A and Arm B) Randomization to the date of first documented progression of disease or death from any cause (up to 28 months) Time-to-event Log-rank test; hazard ratio

Registry endpoint definition

For each of the three registered primary endpoints, PFS was defined as the time from randomization to the date of first documented progression of disease or death from any cause in the absence of disease progression, using Response Evaluation Criteria in Solid Tumors (RECIST) version 1.1 criteria, as assessed by investigator. The registry definition states that progressive disease was defined as at least a 20% increase in the sum of the diameters of target lesions, with reference to the registry definition.

Why the endpoint is time-to-event: PFS does not simply record whether progression occurred. It records when the first qualifying event occurred, while participants who have not experienced progression or death by the relevant follow-up point can contribute censored observations. This is why the primary analyses use survival-analysis methods rather than a simple comparison of proportions.

5. Statistical Methodology

Log-rank test

All three posted primary analyses used the log-rank test. The log-rank test is designed to compare time-to-event distributions between randomized groups while incorporating the timing of events and accommodating right censoring.

Core survival-analysis question
H0: the PFS distributions do not differ between the comparison groups

The registry identifies the hypothesis type for each primary analysis as superiority. The reported effect measure is a hazard ratio, accompanied by a two-sided 95% confidence interval.

Hazard ratio

The effect measure reported for all three primary analyses is the hazard ratio (HR). The hazard ratio compares the estimated instantaneous event rates between the randomized groups over the analyzed follow-up.

Interpretation of a hazard ratio
HR < 1  →  lower estimated instantaneous event rate in the first-listed treatment group

For the Arm C versus Arm A analysis, for example, an HR of 0.569 means the estimated hazard of progression or death was 0.569 times that of Arm A under the reported analysis. It does not mean that 56.9% of patients avoided progression, nor does it directly provide an absolute difference in PFS.

Kaplan-Meier estimation

Although the ClinicalTrials.gov record identifies the formal comparison method as a log-rank test rather than separately reporting Kaplan-Meier estimates, PFS is a time-to-event endpoint for which Kaplan-Meier estimation is the standard descriptive framework. Kaplan-Meier curves estimate the probability of remaining event-free over time while retaining information from censored participants up to their censoring times.

Conceptual Kaplan-Meier estimator
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at event time ti, while ni represents the number at risk immediately before that time.

Censoring

The ClinicalTrials.gov record explicitly identify censored participants. In the Arm A versus Arm B analysis, 94 participants in Arm A and 77 in Arm B were censored. In the Arm C versus Arm A analysis, 99 participants in Arm C and 64 in Arm A were censored. In the ESR1-mutation detected analysis, 29 participants in Arm A and 16 in Arm B were censored.

Censoring is not equivalent to a treatment failure. A censored participant contributes information to the analysis until the time at which the participant is censored. The validity of standard survival analysis depends on assumptions about the relationship between censoring and the event process.

Analysis populations

The first primary analysis used all participants randomly assigned to either Arm A or Arm B, including censored participants. The second used all participants randomly assigned to either Arm A or Arm C concurrently, including censored participants. The ESR1 analysis used participants randomly assigned to Arm A or Arm B who had ESR1 mutations detected at baseline, again including censored participants.

Primary comparisonAnalysis populationCensored observations reported
Arm A vs Arm B All participants randomly assigned to Arm A or Arm B, including censored Arm A = 94; Arm B = 77
Arm C vs Arm A All participants randomly assigned to Arm A or Arm C concurrently, including censored Arm C = 99; Arm A = 64
ESR1-mutation detected: Arm A vs Arm B All participants randomly assigned to Arm A or Arm B with ESR1 mutations detected at baseline, including censored Arm A = 29; Arm B = 16

6. Primary Results

ClinicalTrials.gov reports three formal primary statistical analyses. Each uses a log-rank test and reports a hazard ratio with a two-sided 95% confidence interval. The following results are presented exactly as reported in the ClinicalTrials.gov record.

6.1 Investigator-Assessed PFS: Arm A vs Arm B

The first primary comparison evaluated investigator-assessed PFS between imlunestrant and investigator's choice of endocrine therapy.

Hazard ratio for progression or death

0.867

95% CI: 0.724–1.039   ·   P = 0.1158

Comparison: Arm A, imlunestrant vs Arm B, investigator's choice of endocrine therapy

Analysis featureReported result
EndpointInvestigator-assessed Progression Free Survival (PFS) (Between Arm A and Arm B)
Time frameRandomization to the date of first documented progression of disease or death from any cause (up to 28 months)
MethodLog-rank test
Effect measureHazard ratio
Estimate0.867
95% CI0.724–1.039
P-value0.1158
HypothesisSuperiority
Clinical Biostats interpretation

The estimated hazard ratio of 0.867 is below 1, so the point estimate corresponds to an estimated hazard of progression or death that is approximately 86.7% of the hazard in the investigator's-choice group under the reported analysis. Expressed as a relative difference in the estimated hazard, this corresponds to approximately a 13.3% lower estimated hazard for Arm A.

That interpretation applies to the estimated hazard; it does not mean that 13.3% of patients benefited, that PFS was 13.3% longer, or that the probability of progression was reduced by exactly 13.3% for every participant.

The 95% confidence interval, 0.724–1.039, describes uncertainty around the estimated hazard ratio. Because the interval extends across 1, the data are compatible with both a lower and a higher hazard under the statistical model and sampling framework. The p-value of 0.1158 addresses the statistical evidence against the specified null hypothesis; it is not a measure of the magnitude or clinical importance of the estimated effect.

Because PFS is a time-to-event endpoint, the interpretation also depends on appropriate handling of censoring and on the assumptions underlying the hazard-ratio framework. The ClinicalTrials.gov record does not provide a separate test of proportional hazards, so the HR should not be interpreted as an assumption-free description of the entire PFS experience.

6.2 Investigator-Assessed PFS: Arm C vs Arm A

The second primary comparison evaluated whether adding abemaciclib to imlunestrant changed investigator-assessed PFS relative to imlunestrant alone.

Hazard ratio for progression or death

0.569

95% CI: 0.441–0.733   ·   P < 0.0001

Comparison: Arm C, imlunestrant + abemaciclib vs Arm A, imlunestrant

Analysis featureReported result
EndpointInvestigator-assessed PFS (Between Arm C and Arm A)
Time frameRandomization to the date of first documented progression of disease or death from any cause (up to 26 months)
MethodLog-rank test
Effect measureHazard ratio
Estimate0.569
95% CI0.441–0.733
P-value<0.0001
HypothesisSuperiority
Clinical Biostats interpretation

The hazard ratio of 0.569 indicates that the estimated instantaneous hazard of progression or death in Arm C was 0.569 times that in Arm A under the reported analysis. Put another way, the point estimate corresponds to an approximately 43.1% lower estimated hazard for Arm C relative to Arm A.

This is a relative time-to-event measure. It does not mean that 43.1% of participants avoided progression, that individual patients experienced a 43.1% reduction in their personal risk, or that the median PFS differed by 43.1%.

The 95% confidence interval of 0.441–0.733 gives a range of values compatible with the estimated effect under the analysis framework. The entire interval is below 1, which is consistent with a lower estimated hazard for Arm C across the interval of uncertainty represented by the reported confidence interval.

The p-value of <0.0001 quantifies the statistical evidence against the null hypothesis used for the comparison. It does not tell us that the treatment effect is large, clinically important, or certain. Effect size and precision are communicated by the HR and confidence interval; the p-value addresses evidence against the null.

As with any hazard-ratio analysis, interpretation requires attention to censoring and the proportional-hazards assumption. The ClinicalTrials.gov record does not report a separate assessment of proportional hazards.

6.3 Investigator-Assessed PFS in the ESR1-Mutation Detected Population

The third primary analysis restricted the Arm A versus Arm B comparison to participants who had ESR1 mutations detected at baseline.

Hazard ratio for progression or death

0.617

95% CI: 0.464–0.821   ·   P = 0.0008

Comparison: Arm A, imlunestrant vs Arm B, investigator's choice of endocrine therapy

Analysis featureReported result
EndpointInvestigator-assessed PFS in the Estrogen Receptor 1 (ESR1)-Mutation Detected Population (Between Arm A and Arm B)
Time frameRandomization to the date of first documented progression of disease or death from any cause (up to 28 months)
PopulationParticipants randomly assigned to Arm A or Arm B with ESR1 mutations detected at baseline
MethodLog-rank test
Effect measureHazard ratio
Estimate0.617
95% CI0.464–0.821
P-value0.0008
HypothesisSuperiority
Clinical Biostats interpretation

The hazard ratio of 0.617 means that the estimated instantaneous hazard of progression or death in Arm A was 0.617 times that in Arm B for the ESR1-mutation detected population under the reported analysis. The point estimate therefore corresponds to approximately a 38.3% lower estimated hazard for Arm A relative to Arm B.

The estimate does not mean that 38.3% of participants were protected from progression, nor does it establish an absolute PFS improvement of 38.3%. It is a relative time-to-event measure.

The 95% confidence interval of 0.464–0.821 does not cross 1, indicating that the range of effects represented by this interval is on the lower-hazard side of the null value. The interval is also substantially wider than a point estimate alone can communicate, illustrating why precision should always accompany a hazard ratio.

The p-value of 0.0008 measures the evidence against the relevant null hypothesis; it does not measure the size of the treatment effect. The HR and its confidence interval provide the effect-size information.

This is a population-restricted primary analysis. It should not automatically be interpreted as proof that ESR1 mutation status modifies treatment effect unless a formal interaction analysis or other appropriate comparison of treatment effects supports such a conclusion. The ClinicalTrials.gov record does not report an interaction test.

7. Comparing the Three Primary Analyses

Primary endpointComparisonHR95% CIP-value
Investigator-assessed PFS Arm A vs Arm B 0.867 0.724–1.039 0.1158
Investigator-assessed PFS Arm C vs Arm A 0.569 0.441–0.733 <0.0001
Investigator-assessed PFS, ESR1-mutation detected population Arm A vs Arm B 0.617 0.464–0.821 0.0008

The three analyses answer different questions and should not be collapsed into a single overall treatment effect. The first asks about imlunestrant versus investigator's choice of endocrine therapy in the broader randomized comparison. The second asks about adding abemaciclib to imlunestrant. The third asks the Arm A versus Arm B question within the baseline ESR1-mutation detected population.

There is also an important distinction between the point estimates and the evidence reported by the confidence intervals and p-values. The first estimate is below 1 but has a 95% confidence interval extending above 1 and a p-value of 0.1158. The second and third confidence intervals remain below 1, with p-values of <0.0001 and 0.0008, respectively. These are descriptions of the reported statistical results rather than a ranking of the treatment strategies.

8. Secondary Endpoints and Other Posted Outcomes

The registry profile posted on ClinicalTrials.gov for EMBER-3 reports 21 outcome measures and three posted statistical analyses. However, the ClinicalTrials.gov record identifies only the three formal primary statistical analyses and do not provide numerical results for additional secondary outcome measures.

Registry reporting boundary: this page does not invent secondary efficacy estimates, response rates, survival medians, subgroup hazard ratios, or other outcome statistics that are not contained in the ClinicalTrials.gov record. ClinicalTrials.gov does report 21 outcome measures for EMBER-3, but the numerical statistical analyses provided for this page are the three primary analyses listed above.

For a time-to-event secondary endpoint, an appropriate analysis would ordinarily use a survival-analysis framework such as Kaplan-Meier estimation together with a log-rank comparison and a hazard-ratio model when prespecified by the statistical analysis plan. The exact method, effect estimate, confidence interval, and multiplicity treatment should be taken from the corresponding registered analysis rather than inferred from the primary PFS results.

9. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm as affected participants over participants at risk. These entries should be reproduced as reported rather than converted into additional safety statistics.

ArmSerious adverse events affected / at risk
Arm A: Imlunestrant46/327
Arm B: Investigator's Choice of Endocrine Therapy6/32
Arm B: Investigator's Choice of Endocrine Therapy49/292
Arm C: Imlunestrant + Abemaciclib42/208

The two Arm B entries are preserved separately because the ClinicalTrials.gov record contains them as separate affected/at-risk records. They should not be combined into a new denominator or interpreted as a single safety estimate without the underlying registry context that explains the two entries.

Safety versus efficacy: serious adverse events and PFS answer different questions. The serious-adverse-event figures describe an observed safety category, whereas the PFS analyses compare the timing of progression or death. Neither should be used as a substitute for the other when interpreting the trial.

10. Statistical Methods Explained

Why was a log-rank test used?

PFS is a time-to-event endpoint. A simple comparison of the proportion of participants who progressed would discard information about when progression occurred and how long censored participants were observed. The log-rank test instead compares the event experience over follow-up while accommodating right censoring.

What does an HR of 0.569 mean?

An HR of 0.569 means that, under the reported time-to-event analysis, the estimated instantaneous hazard of progression or death in Arm C was 0.569 times that in Arm A. The corresponding point estimate can be described as approximately a 43.1% lower estimated hazard. It does not mean that 43.1% of patients avoided progression or that PFS increased by exactly 43.1%.

Why does the confidence interval matter?

A hazard ratio is an estimate, not a certainty. The 95% confidence interval describes the uncertainty around that estimate under the statistical model and sampling framework. For the Arm C versus Arm A analysis, the interval is 0.441–0.733; for the Arm A versus Arm B analysis, it is 0.724–1.039. These intervals communicate much more information about precision than the point estimates alone.

Why does a p-value not measure effect size?

A p-value addresses the strength of evidence against a null hypothesis under the specified statistical model. It does not tell us how large the treatment effect is. A hazard ratio and its confidence interval are needed to describe the magnitude and precision of the estimated relative effect.

Why is the ESR1 analysis not automatically an interaction test?

The ESR1-mutation detected analysis is a restricted population comparison. A treatment effect estimated within one biomarker-defined group does not by itself establish that the treatment effect differs from the effect in another group. To demonstrate effect modification, the appropriate question is whether the treatment-by-biomarker interaction is supported by a formal statistical test or equivalent prespecified heterogeneity analysis. No such interaction result is included in the ClinicalTrials.gov record.

What does censoring mean in these PFS analyses?

A censored participant is not automatically classified as having had a favorable or unfavorable outcome. Instead, the participant contributes observed follow-up information until the censoring time. The survival analysis then uses the available event and censoring information to estimate the time-to-event distribution. The validity of that approach depends on assumptions concerning the censoring mechanism.

Why is the absence of masking statistically relevant?

The registry describes EMBER-3 as having no masking. Randomization can balance measured and unmeasured prognostic factors in expectation, but lack of masking can still matter for outcomes involving investigator assessment. Because the primary PFS endpoint includes investigator-assessed progression under RECIST version 1.1, the possibility of assessment-related bias is a relevant limitation when interpreting the reported hazard ratios.

11. Multiplicity and the Three Primary Analyses

EMBER-3 has three registered primary endpoints and three posted primary analyses. That structure is statistically important because multiple confirmatory questions can create a multiplicity problem if each hypothesis is tested independently at the same nominal significance level.

Primary questionReported hypothesis typeReported analysisP-value
Arm A vs Arm B PFS Superiority Log-rank test; HR 0.867 0.1158
Arm C vs Arm A PFS Superiority Log-rank test; HR 0.569 <0.0001
ESR1-mutation detected Arm A vs Arm B PFS Superiority Log-rank test; HR 0.617 0.0008

The ClinicalTrials.gov record identifies these as primary analyses but do not provide an alpha-allocation strategy, gatekeeping procedure, hierarchical testing sequence, or other multiplicity-adjustment details. Therefore, this page does not infer one.

Why this matters: a nominal p-value should not automatically be interpreted as establishing a particular familywise-error-controlled claim when the multiplicity procedure is not provided. The distinction is especially important when a trial has several primary questions rather than one isolated hypothesis.

12. Non-Inferiority, Bayesian Methods, Crossover, and Interim Analysis

The registry-reported EMBER-3 trial data identify all three primary hypotheses as superiority. No non-inferiority margin is reported, so no non-inferiority interpretation is appropriate for these primary analyses.

The registry-reported statistical-method fields identify the log-rank test as the method and the hazard ratio as the effect measure. No Bayesian method is reported. No crossover analysis is reported in the ClinicalTrials.gov record. No interim-analysis or alpha-spending procedure is reported in the ClinicalTrials.gov record.

Design topicWhat the ClinicalTrials.gov record supports
Hypothesis typeSuperiority
Non-inferiority marginNot reported in the ClinicalTrials.gov record
Bayesian analysisNot reported in the ClinicalTrials.gov record
CrossoverNot reported in the ClinicalTrials.gov record
Interim analysisNot reported in the ClinicalTrials.gov record
Alpha spendingNot reported in the ClinicalTrials.gov record
Missing-data imputationNot reported in the ClinicalTrials.gov record

These omissions should not be interpreted as evidence that the underlying protocol or statistical analysis plan contains no such procedures. They mean only that the ClinicalTrials.gov record does not provide them, so they are not reconstructed here.

13. Randomization and Analysis Populations

Randomization is central to the causal interpretation of the primary comparisons. The registry describes EMBER-3 as randomized and parallel, with no masking. The primary analyses were defined around participants randomly assigned to the relevant treatment groups.

Arm A vs Arm B

All participants randomly assigned to either Arm A or Arm B, including censored participants, were included in the reported analysis population.

Arm C vs Arm A

All participants randomly assigned to either Arm A or Arm C concurrently, including censored participants, were included in the reported analysis population.

ESR1 analysis

The analysis was restricted to participants randomly assigned to Arm A or Arm B who had ESR1 mutations detected at baseline.

Why randomization matters

Random assignment provides the basis for comparing outcomes across treatment strategies without relying solely on adjustment for measured baseline characteristics.

The ESR1 analysis illustrates an important distinction: randomization remains relevant within the selected population, but restricting an analysis to a biomarker-defined subgroup generally reduces the available information and can increase statistical uncertainty.

14. Interpretation of Hazard Ratios

Arm A vs Arm B

The reported HR of 0.867 corresponds to a point estimate below 1, but the 95% CI of 0.724–1.039 crosses 1. The appropriate statistical description is therefore the complete estimate, interval, and p-value—not simply the direction of the point estimate.

Arm C vs Arm A

The reported HR of 0.569 indicates a lower estimated hazard in Arm C relative to Arm A. The 95% CI of 0.441–0.733 remains below 1, providing the reported precision around that estimate.

ESR1-mutation detected population

The reported HR of 0.617 indicates a lower estimated hazard in Arm A relative to Arm B within the ESR1-mutation detected population. Its 95% CI of 0.464–0.821 remains below 1.

These three interpretations should remain tied to their respective populations and comparisons. A hazard ratio is not a universal property of a drug independent of the comparison group, endpoint definition, follow-up period, or analysis population.

15. Confidence Intervals and Statistical Precision

Confidence intervals are especially useful for distinguishing a point estimate from the range of uncertainty surrounding it.

ComparisonHR95% CI width and position relative to 1Statistical interpretation
Arm A vs Arm B 0.867 0.724–1.039; crosses 1 The interval includes the null hazard ratio.
Arm C vs Arm A 0.569 0.441–0.733; below 1 The reported interval is entirely below the null hazard ratio.
ESR1 Arm A vs Arm B 0.617 0.464–0.821; below 1 The reported interval is entirely below the null hazard ratio.

The confidence interval does not describe the range of treatment effects experienced by individual patients. It describes uncertainty around the estimated population-level effect under the statistical framework used for the analysis.

16. P-Values in Context

The three reported p-values are 0.1158, <0.0001, and 0.0008. They should be interpreted in the context of the corresponding comparison, endpoint, analysis population, and multiplicity structure.

What a p-value is not
P-value ≠ probability that the treatment is effective

A p-value is calculated under a null hypothesis and reflects the compatibility of the observed data, or more extreme data, with that null under the statistical model. It does not directly provide the probability that the null hypothesis is true.

Similarly, a very small p-value does not imply a large treatment effect. EMBER-3 illustrates why the HR and confidence interval should be read alongside the p-value rather than replaced by it.

17. Missing Data and Imputation

The ClinicalTrials.gov record does not report a missing-data or imputation strategy for the primary PFS analyses. Because PFS is a time-to-event endpoint, missing follow-up and censoring are handled differently from missing continuous or binary measurements.

For a standard survival analysis, the key issue is whether censoring can be treated as sufficiently independent of the future event process conditional on the information used by the analysis. The ClinicalTrials.gov record identifies censored observations but do not provide enough information to evaluate the censoring mechanism or any sensitivity analyses around it.

Interpretation boundary: no imputation method has been added to this page. Doing so would require protocol or statistical-analysis-plan information that is not contained in the ClinicalTrials.gov record.

18. Trial Timeline

October 4, 2021

Trial start

The EMBER-3 trial began on October 4, 2021 according to the registry profile.

June 24, 2024

Primary completion

the ClinicalTrials.gov record lists June 24, 2024 as the primary completion date.

Current registry status

Active, not recruiting

The ClinicalTrials.gov profile classifies EMBER-3 as active, not recruiting.

19. Why This Trial Matters Statistically

EMBER-3 is a useful teaching case because the ClinicalTrials.gov record bring together several central ideas in clinical-trial biostatistics: randomized comparisons, time-to-event endpoints, censoring, log-rank testing, hazard ratios, confidence intervals, p-values, and biomarker-defined analysis populations.

Statistical conceptHow it appears in EMBER-3
RandomizationThe trial is randomized and uses a parallel design.
Three-arm designThe registry describes imlunestrant, investigator's choice of endocrine therapy, and imlunestrant plus abemaciclib.
Time-to-event endpointAll three primary endpoints are investigator-assessed PFS analyses.
Log-rank testThe reported formal method for all three primary statistical analyses.
Hazard ratioThe reported effect measure for all three primary analyses.
Confidence intervalEach primary analysis reports a two-sided 95% CI.
P-valueEach primary analysis reports a p-value.
CensoringCensored participant counts are explicitly reported for each primary comparison.
Biomarker-defined populationThe third primary analysis is restricted to participants with ESR1 mutations detected at baseline.
MultiplicityThree primary analyses create a setting in which the error-control strategy matters.
Open-label assessmentThe registry describes the trial as having no masking, relevant to investigator-assessed PFS.

20. Important Limitations and Interpretation Issues

21. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The three reported primary PFS analyses used log-rank tests and hazard ratios. The Arm A versus Arm B estimate was 0.867 with a 95% CI of 0.724–1.039; the Arm C versus Arm A estimate was 0.569 with a 95% CI of 0.441–0.733; and the ESR1-mutation detected Arm A versus Arm B estimate was 0.617 with a 95% CI of 0.464–0.821.

Clinical interpretation

The statistical results describe differences in the time-to-progression-or-death endpoint between the specified randomized groups. They do not, by themselves, describe every dimension of benefit, toxicity, patient experience, or long-term outcome.

A statistically estimated hazard ratio should therefore be treated as one component of the evidence. Absolute event probabilities, median event times, safety, quality of life, subsequent therapy, and other clinical outcomes can add information when those data are available. The registry-reported EMBER-3 dataset does not provide those additional numerical efficacy measures, so they are not inferred here.

22. What the Three Primary Hazard Ratios Do — and Do Not — Mean

HR 0.867: Arm A vs Arm B

The point estimate indicates a lower estimated hazard in Arm A, corresponding to approximately a 13.3% lower estimated hazard relative to Arm B. The 95% CI of 0.724–1.039 includes 1, so the uncertainty around the estimate is important to the interpretation.

HR 0.569: Arm C vs Arm A

The point estimate corresponds to approximately a 43.1% lower estimated hazard in Arm C relative to Arm A. The 95% CI of 0.441–0.733 remains below 1. This does not mean that every participant experienced a 43.1% reduction in risk.

HR 0.617: ESR1-mutation detected population

The point estimate corresponds to approximately a 38.3% lower estimated hazard in Arm A relative to Arm B among participants with ESR1 mutations detected at baseline. The result is specific to that analysis population and should not automatically be interpreted as evidence of biomarker interaction.

The distinction between relative hazard and is fundamental. A hazard ratio does not directly answer how many additional months an individual patient might experience without progression, what percentage of participants remain progression-free at a particular time point, or how treatment effects vary between individuals.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue learning from the statistical methods

The EMBER-3 results illustrate how randomized treatment comparisons, time-to-event endpoints, hazard ratios, confidence intervals, log-rank testing, censoring, and biomarker-defined populations fit together in clinical-trial analysis.

26. Record Summary

EMBER-3 provides a useful example of a randomized phase 3 trial with a three-arm parallel design and multiple primary time-to-event questions. The registry reports three formal investigator-assessed PFS analyses, all using log-rank testing with hazard ratios and two-sided 95% confidence intervals. The Arm A versus Arm B comparison produced an HR of 0.867 (95% CI 0.724–1.039; P = 0.1158), the Arm C versus Arm A comparison produced an HR of 0.569 (95% CI 0.441–0.733; P < 0.0001), and the ESR1-mutation detected Arm A versus Arm B analysis produced an HR of 0.617 (95% CI 0.464–0.821; P = 0.0008).

The statistical story is therefore not simply a collection of p-values. It includes the randomized comparisons, the precise definition of PFS, censoring, the selected analysis populations, the interpretation of hazard ratios, the uncertainty represented by confidence intervals, and the multiplicity implications of having three primary analyses. The open-label design and investigator assessment are also relevant to interpretation.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. Where the ClinicalTrials.gov record does not report a numerical estimate or methodological detail, this page does not reconstruct one from memory or from an unrelated source.