This page separates reported trial results from statistical interpretation. Numerical results are taken from the publicly posted ClinicalTrials.gov record. The registry provides the official trial record.
1. Trial at a Glance
GALACTIC-HF was a randomized, parallel, triple-masked phase 3 study evaluating omecamtiv mecarbil in chronic heart failure with reduced ejection fraction. The registry reports a primary time-to-event endpoint of time to cardiovascular death or first heart failure event.
| Feature | GALACTIC-HF |
|---|---|
| Trial name | GALACTIC-HF |
| Brief title | Registrational Study With Omecamtiv Mecarbil (AMG 423) to Treat Chronic Heart Failure With Reduced Ejection Fraction |
| Condition | Heart Failure |
| Phase | Phase 3 |
| Design | Randomized, parallel, triple-masked |
| Primary purpose | Treatment |
| Enrollment | 8256 |
| Arms | 2 |
| Interventions | Omecamtiv Mecarbil; Placebo; Standard of Care |
| Lead sponsor | Cytokinetics |
| ClinicalTrials.gov | NCT02929329 |
2. Clinical Question
The statistical question was whether treatment assignment to omecamtiv mecarbil, compared with placebo, was associated with a different time to the composite of cardiovascular death or first heart failure event in the randomized trial population.
Population
The registry describes the study as treating chronic heart failure with reduced ejection fraction.
Intervention
Omecamtiv mecarbil, with standard of care included among the registered interventions.
Comparator
Placebo, with standard of care included among the registered interventions.
Primary question
Does randomized treatment assignment change the time to cardiovascular death or first heart failure event?
3. Trial Design
Placebo comparison group
- Placebo
- Standard of care
- Compared with omecamtiv mecarbil for the registered efficacy analyses
Omecamtiv mecarbil group
- Omecamtiv mecarbil
- Standard of care
- Compared with placebo for the registered efficacy analyses
4. Trial Timeline
Study start
The registry lists 2017-01-06 as the trial start date.
Primary completion
The registry lists 2020-09-14 as the primary completion date.
Primary time-to-event analysis cutoff
The registered primary endpoint time frame extends from randomization to up to the earliest of the last confirmed survival status date or the analysis cut-off date of 07 August 2020.
5. Primary Endpoint
| Endpoint | Registry definition / time frame | Endpoint type |
|---|---|---|
| Time to Cardiovascular Death or First Heart Failure Event | The primary outcome was a composite of a heart-failure event or cardiovascular death, whichever occurred first, in a time-to-event analysis. A heart-failure event was defined as an urgent clinic visit, emergency department visit, or hospitalization for subjectively and objectively worsening heart failure leading to treatment intensification beyond a change in oral diuretic therapy. Time frame: from randomization to up to the earliest of last confirmed survival status date or analysis cut-off date (07 August 2020). | Time-to-event |
The composite endpoint is important statistically because the analysis treats the first qualifying component event as the event of interest. A participant's first event can therefore be either a cardiovascular death or a qualifying heart-failure event, whichever occurs first.
6. Statistical Methodology
Stratified Cox proportional-hazards model
The primary hazard-ratio analysis used a Cox proportional-hazards model with baseline hazards stratified according to randomization setting and geographic region. Treatment group and baseline estimated glomerular filtration rate (eGFR) were included as covariates.
The treatment coefficient is transformed into a hazard ratio. Stratification allows the baseline hazard to differ across the specified randomization setting and geographic-region strata without estimating a separate treatment effect for each stratum.
Stratified log-rank test
The registry also reports a stratified log-rank test for the primary endpoint. The test was stratified by randomization setting and region. Unlike the hazard ratio, the log-rank test produces a hypothesis-test result rather than an effect-size estimate.
Covariate adjustment
The primary Cox model included baseline eGFR and treatment group as covariates, while baseline hazards were stratified by randomization setting and geographic region. This means the reported hazard ratio is not simply an unadjusted ratio of crude event rates.
Competing-risk analysis
The registry also reports a competing-risk subdistribution hazard ratio and associated 95% confidence intervals for treatment. Deaths not included in the endpoint were considered the competing risk. This addresses the fact that a competing event can prevent a participant from subsequently experiencing the endpoint of interest in the usual way.
Analysis population
The primary efficacy analysis was reported in the full analysis set. The registry also identifies intention-to-treat analysis as a concept in the primary analysis text. The treatment comparison therefore remains anchored to randomized treatment assignment rather than being defined only by treatment actually received.
7. Primary Result: Cox Hazard Ratio
Time to cardiovascular death or first heart failure event
95% CI: 0.86–0.99 · P = 0.0252 · Two-sided
Analysis population: Full analysis set · Comparison: Placebo vs Omecamtiv Mecarbil
The registry reports a hazard ratio of 0.92 from the Cox proportional-hazards model. Because the treatment comparison is expressed as omecamtiv mecarbil relative to placebo in the analysis context, an HR below 1 indicates a lower estimated instantaneous event hazard for omecamtiv mecarbil under the fitted model.
What the estimate means: HR 0.92 corresponds to a 8% lower estimated hazard of the composite endpoint under the fitted proportional-hazards model, using the omecamtiv mecarbil versus placebo treatment comparison.
What it does not mean: it does not mean that exactly 8% fewer participants experienced the endpoint, nor does it represent an 8% absolute reduction in event probability. A hazard ratio is a relative model-based time-to-event measure.
Precision: the 95% confidence interval of 0.86–0.99 describes uncertainty around the estimated hazard ratio under the statistical model and sampling framework. It does not describe the range of individual patient effects.
The p-value: P = 0.0252 addresses the statistical evidence against the null hypothesis under the prespecified testing framework. It does not measure the magnitude or clinical importance of the treatment effect.
Important cautions: interpretation of a single Cox hazard ratio depends on the proportional-hazards framework. The analysis also used stratification and covariate adjustment, and censoring and the analysis population affect the information contributing to the estimate.
8. Primary Result: Competing-Risk Cox Analysis
Competing-risk subdistribution analysis
95% CI: 0.86–0.99 · Two-sided
Analysis population: Full analysis set · Competing deaths treated as the competing risk
The registry reports the same hazard-ratio estimate and confidence interval for its competing-risk subdistribution analysis. The purpose of this analysis is different from simply repeating the ordinary Cox calculation: it explicitly accounts for deaths that are not themselves part of the composite endpoint as competing events.
What the estimate means: the reported subdistribution hazard ratio of 0.92 indicates a lower estimated subdistribution hazard for the omecamtiv mecarbil group under the competing-risk framework.
What it does not mean: it is not an absolute probability difference and should not be read as saying that 8% of participants avoided the composite endpoint because of treatment.
Precision: the 95% CI of 0.86–0.99 quantifies statistical uncertainty around this reported estimate. A confidence interval is not a range containing the true effect with 95% probability.
The p-value distinction: no separate p-value is reported for this competing-risk analysis in the ClinicalTrials.gov record. The primary Cox analysis reports P = 0.0252, while this competing-risk result is presented with its estimate and confidence interval.
Interpretive caution: competing-risk methods answer a specific estimand. The presence of a competing event changes how event probabilities and treatment effects should be interpreted compared with an analysis that treats competing events simply as ordinary censoring.
9. Primary Result: Stratified Log-Rank Test
Stratified log-rank comparison
Two-sided superiority analysis
Stratified by randomization setting and region
The registry reports a two-sided P-value of 0.0211 from the stratified log-rank test. This provides a hypothesis-test comparison of the time-to-event distributions while respecting the specified randomization-setting and regional strata.
What the result means: the reported P = 0.0211 provides statistical evidence of a difference between the randomized treatment groups under the stratified log-rank testing framework.
What it does not mean: the p-value does not quantify the size of the treatment effect and does not tell us the probability that the treatment hypothesis is true. The hazard ratio of 0.92 is the separate effect-size estimate.
Precision: the ClinicalTrials.gov record does not report a confidence interval specifically for the log-rank test statistic, so precision should be assessed using the separately reported Cox hazard-ratio confidence interval.
Multiplicity: the primary analysis notes that the overall type I error was 0.05 for two-sided testing across primary and secondary outcomes, with a prespecified testing algorithm. Therefore the p-value should be interpreted in the context of the trial's multiple-outcome testing strategy rather than as an isolated calculation.
10. Secondary Endpoint Results
Time to Cardiovascular Death
Hazard ratio for cardiovascular death
95% CI: 0.92–1.11 · P = 0.8555 · Two-sided
The secondary endpoint was analyzed in the full analysis set using a stratified Cox proportional-hazards model. Baseline hazards were stratified according to randomization setting and geographic region, with treatment group and baseline eGFR as covariates.
An HR of 1.01 is very close to 1.00, meaning the estimated instantaneous hazard of cardiovascular death was similar between the randomized groups under this model. The 95% CI of 0.92–1.11 includes 1.00, and the reported P = 0.8555 does not provide evidence against the null hypothesis under this test. This does not prove that the two treatments are identical; it describes the evidence and uncertainty in this particular analysis.
Time to First Heart Failure Hospitalization
Primary Cox analysis
95% CI: 0.87–1.03 · P = 0.1902 · Two-sided
The registry reports time to first heart failure hospitalization as a secondary time-to-event endpoint. The primary posted analysis used a stratified Cox proportional-hazards model with randomization setting and region as strata and baseline eGFR and treatment group as covariates.
HR 0.95 corresponds to a 5% lower estimated hazard of first heart failure hospitalization under the fitted model, but the 95% CI of 0.87–1.03 includes 1.00. The reported P = 0.1902 therefore does not provide evidence of a statistically detectable difference under this analysis. The confidence interval also shows that the data are compatible with effects on either side of the null within the interval; it does not establish equivalence.
Time to First Heart Failure Hospitalization: Competing-Risk Analysis
| Analysis | Estimate | 95% CI | P-value |
|---|---|---|---|
| Competing-risk subdistribution hazard ratio | 0.96 | 0.88–1.04 | Not reported |
The registry reports a competing-risk subdistribution hazard ratio of 0.96 with a 95% confidence interval of 0.88–1.04. Deaths not included in the endpoint were considered the competing risk. No separate p-value is reported for this competing-risk analysis in the ClinicalTrials.gov record.
Time to All-Cause Death
Hazard ratio for all-cause death
95% CI: 0.92–1.09 · P = 0.9633 · Two-sided
The time to all-cause death analysis used the full analysis set and a Cox proportional-hazards model with baseline hazards stratified according to randomization setting and geographic region and with treatment group and baseline eGFR as covariates.
The estimate of HR 1.00 is centered exactly on the null value. The 95% CI of 0.92–1.09 indicates uncertainty around that estimate, and P = 0.9633 provides no evidence of a detectable difference in this analysis. Again, a nonsignificant result should not be converted into a claim of equivalence unless an equivalence or non-inferiority framework was prespecified and supported by the relevant margin.
11. KCCQ Total Symptom Score at Week 24
The registry reports several analyses of change from baseline in Kansas City Cardiomyopathy Questionnaire Total Symptom Score (KCCQ TSS) at Week 24. These analyses used the full analysis set with available data and report treatment differences as omecamtiv mecarbil minus placebo.
| Population / analysis | Estimate | 95% CI | P-value |
|---|---|---|---|
| Outpatients, LS mean difference | -0.46 | -1.40 to 0.48 | Not reported |
| Inpatients, LS mean difference | 2.50 | 0.54 to 4.46 | Not reported |
| Pooled treatment difference, mixed-effects / random-effects meta-analysis approach | 0.75 | -2.55 to 4.51 | Not reported |
| Outpatients, joint longitudinal and survival sensitivity analysis | -0.71 | -1.62 to 0.20 | Not reported |
| Inpatients, joint longitudinal and survival sensitivity analysis | 2.31 | 0.80 to 3.82 | Not reported |
| Omnibus F-test | Not reported | Not reported | 0.0278 |
The registry also states that, if significance for the primary outcome was determined, change from baseline in KCCQ total symptom score was tested against an alpha of 0.002. The ClinicalTrials.gov record therefore distinguish the omnibus P = 0.0278 from the stricter alpha threshold specified for this endpoint in the testing hierarchy.
Outpatients
The reported LS mean difference was -0.46, with a 95% CI of -1.40 to 0.48. Because the difference is defined as omecamtiv mecarbil minus placebo, a negative value favors the placebo direction for this numerical scale, while a positive value favors the omecamtiv mecarbil direction.
Inpatients
The reported LS mean difference was 2.50, with a 95% CI of 0.54 to 4.46. This is a subgroup-specific estimate and should not be treated as interchangeable with the pooled estimate.
Pooled estimate
The pooled treatment difference was 0.75 with a 95% CI of -2.55 to 4.51 using a random-effects meta-analysis approach.
Omnibus test
The registry reports P = 0.0278 from an omnibus F-test. The registry-reported testing note specifies alpha = 0.002 for KCCQ TSS if the primary outcome met its significance criterion.
12. Missing Data and Sensitivity Analysis
The ClinicalTrials.gov record supports a specific sensitivity analysis addressing missing KCCQ TSS information due to death. Joint longitudinal and survival models were fitted using observed KCCQ TSS values with random subject slopes and intercepts for the longitudinal component.
The longitudinal models included terms for baseline eGFR, region, and treatment by slope. The survival models were fit for all-cause death, with baseline eGFR and treatment in the proportional-hazard component. The registry identifies multiple imputation / missing-data methodology as a statistical concept associated with the KCCQ sensitivity analysis.
A continuous outcome measured at a fixed time can be missing because participants discontinue, are unavailable for assessment, or die before the assessment. When death itself is informative, a standard analysis of observed Week 24 values can answer a narrower question than an analysis that jointly models longitudinal measurements and survival.
The reported sensitivity estimates illustrate why missing-data assumptions matter: the outpatient joint-model estimate was -0.71 with a 95% CI of -1.62 to 0.20, while the inpatient estimate was 2.31 with a 95% CI of 0.80 to 3.82.
13. Multiplicity and Type I Error
The primary analysis notes state that the overall type I error was 0.05 for two-sided testing across primary and secondary outcomes. Control for multiple comparisons was achieved using a testing algorithm. The registry text further states that if the primary outcome met the P-value threshold of 0.05, alpha would be divided unequally between cardiovascular and other outcomes.
Why multiplicity matters
When several hypotheses are tested within the same confirmatory program, treating every p-value as though it were the only test can increase the chance of at least one false-positive finding.
Why the hierarchy matters
The registry-reported KCCQ analysis specifies alpha = 0.002 if significance for the primary outcome was determined. A nominal P-value therefore cannot be interpreted independently of the prespecified testing sequence.
This is especially important for a trial with multiple time-to-event and patient-reported outcomes. The statistical meaning of an individual result depends not only on its numerical p-value but also on where that hypothesis sits in the trial's error-control strategy.
14. Stratification and Covariate Adjustment
The primary and several secondary time-to-event analyses used the same broad modeling structure: baseline hazards stratified by randomization setting and geographic region, with baseline eGFR and treatment group as covariates.
| Component | Role in the reported analysis |
|---|---|
| Randomization setting | Stratification factor for baseline hazards |
| Geographic region | Stratification factor for baseline hazards |
| Baseline eGFR | Covariate in the Cox model |
| Treatment group | Covariate defining the treatment comparison |
| Full analysis set | Primary time-to-event analysis population |
Stratification and covariate adjustment solve different statistical problems. Stratification permits the baseline event process to differ across specified strata, whereas covariate adjustment estimates the treatment effect conditional on the included baseline covariate. Neither technique removes the need to consider whether the Cox model is appropriate for the observed time-to-event process.
15. Statistical Methods Explained
Why use a Cox proportional-hazards model?
The primary endpoint is a time-to-event outcome, so simply comparing the proportion of participants who experienced an event would discard information about when events occurred and how long participants were followed. The Cox model uses event times and censoring information to estimate a relative hazard while allowing the baseline hazard to remain unspecified.
What does HR 0.92 mean?
An HR of 0.92 means the fitted model estimates the instantaneous event hazard for omecamtiv mecarbil at about 92% of the corresponding hazard for placebo, under the model's assumptions. It is equivalent to an 8% lower estimated hazard, but it is not an 8-percentage-point reduction in event probability.
Why was the analysis stratified?
The primary Cox model stratified baseline hazards according to randomization setting and geographic region. This allows those strata to have different underlying hazard patterns without forcing them to share a common baseline hazard function.
Why include baseline eGFR as a covariate?
The registry specifically identifies baseline eGFR as a covariate in the Cox model. Including a prespecified baseline covariate can improve the precision of the treatment comparison and accounts for its relationship with the modeled event process.
What is the difference between a Cox model and a stratified log-rank test?
The stratified log-rank test is primarily a hypothesis test comparing survival distributions while accounting for strata. The Cox model additionally provides an interpretable effect estimate—the hazard ratio—and its confidence interval.
Why perform a competing-risk analysis?
For endpoints in which some deaths are not themselves counted as the endpoint, death can prevent the endpoint from occurring and therefore acts as a competing event. A competing-risk subdistribution analysis provides a different treatment-effect estimand that explicitly accounts for this structure.
Why use a joint longitudinal and survival model for KCCQ sensitivity analysis?
KCCQ TSS is repeatedly related to survival because death can prevent a later questionnaire measurement from being observed. A joint longitudinal and survival model allows the longitudinal score process and survival process to be modeled together rather than treating missing scores caused by death as an ordinary missing-value problem.
16. Interpreting the Primary Result in Context
The primary Cox estimate of HR 0.92 indicates a modest relative reduction in the modeled hazard of the composite endpoint. The confidence interval, 0.86–0.99, quantifies uncertainty around that estimate.
The Cox model reports P = 0.0252 and the stratified log-rank test reports P = 0.0211. These are hypothesis-test results, not measures of effect magnitude. Their interpretation also belongs within the trial's stated type I error and multiple-comparison framework.
The secondary cardiovascular-death analysis reports HR 1.01 (95% CI 0.92–1.11; P = 0.8555), while time to first heart failure hospitalization reports HR 0.95 (95% CI 0.87–1.03; P = 0.1902). These component analyses provide a different view from the composite primary endpoint.
This distinction is central to composite-endpoint interpretation. A statistically detectable result for a composite does not automatically imply that every component shows the same treatment effect. Each component has its own event definition, frequency, censoring structure, and statistical uncertainty.
17. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized arm as affected participants divided by participants at risk.
| Safety measure | Placebo | Omecamtiv Mecarbil |
|---|---|---|
| Serious adverse events, affected / at risk | 2435 / 4101 | 2373 / 4110 |
These are the serious-adverse-event counts and denominators reported in the ClinicalTrials.gov record. They should be kept separate from efficacy analyses because safety and efficacy answer different questions and may use different analysis populations or exposure definitions.
18. What the Confidence Interval Adds
The point estimate is a single summary of the fitted treatment effect. The confidence interval communicates how much statistical uncertainty surrounds that estimate.
The lower and upper confidence limits are not alternative estimates of what individual patients experienced. Instead, they describe the uncertainty in estimating the treatment-effect parameter under the model and sampling assumptions.
The fact that the upper confidence limit is 0.99 means the reported interval remains below the null value of 1.00. That is relevant to the statistical hypothesis test, but it should not be confused with the size of an absolute clinical benefit.
19. Composite Endpoints: Why the Definition Matters
The primary endpoint combines cardiovascular death and a qualifying heart-failure event, with whichever occurs first determining the event in the time-to-event analysis.
One analysis, two event types
A composite endpoint can increase the number of observed events available for analysis, but its interpretation depends on the clinical and statistical relationship between its components.
First event governs
The registry definition specifies cardiovascular death or first heart failure event, whichever occurred first. Later events do not replace the first qualifying event in this primary endpoint.
Component results matter
The separately reported cardiovascular-death and heart-failure-hospitalization analyses help show how the composite relates to its individual components.
Different estimands
The ordinary Cox analysis and competing-risk analysis are not interchangeable calculations. Each targets a particular time-to-event quantity.
20. Planned and Reported Statistical Architecture
| Statistical feature | Reported approach |
|---|---|
| Primary endpoint | Time to cardiovascular death or first heart failure event |
| Primary effect measure | Hazard ratio |
| Primary regression method | Cox proportional-hazards model |
| Primary hypothesis test | Stratified log-rank test |
| Stratification | Randomization setting and geographic region |
| Covariates | Baseline eGFR and treatment group |
| Competing-risk analysis | Subdistribution hazard ratio |
| Continuous outcome analysis | General linear model / mixed-effects model concepts reported for KCCQ |
| Missing-data sensitivity | Joint longitudinal and survival models for KCCQ |
| Multiplicity | Overall two-sided type I error of 0.05 across primary and secondary outcomes with a testing algorithm |
| ITT concept | Identified in primary and secondary analysis text |
21. Important Limitations and Interpretation Issues
- Composite endpoint: the primary outcome combines cardiovascular death and first heart failure event. Interpretation of the composite should therefore be distinguished from interpretation of either component alone.
- Hazard-ratio assumptions: Cox hazard ratios are model-based. A single HR is most straightforward to interpret when the proportional-hazards framework is a reasonable description of the treatment comparison over time.
- Censoring: time-to-event analyses depend on how follow-up and censoring are handled. The registry defines the primary time frame through the earliest of the last confirmed survival status date or the 07 August 2020 analysis cutoff.
- Competing risks: cardiovascular-death and heart-failure-event analyses can involve competing events, which is why the registry includes subdistribution-hazard analyses.
- Multiplicity: the trial used a multiple-comparison testing algorithm, so individual p-values should be interpreted within the stated type I error framework.
- KCCQ missingness: Week 24 KCCQ analyses were based on available data, with additional joint longitudinal and survival sensitivity analyses addressing missing information associated with death.
- Subgroup estimates: the outpatient and inpatient KCCQ estimates are different analysis populations and should not be treated as interchangeable with the pooled treatment estimate.
- Safety inference: the registry-reported serious-adverse-event data provide counts and denominators but do not include a formal comparison, confidence interval, or p-value.
- Registry scope: this analysis is limited to the numerical and methodological information from ClinicalTrials.gov and does not add results from external publications.
22. Why This Trial Matters Statistically
GALACTIC-HF is a useful teaching case because it combines several important clinical-trial methods in one randomized study: a composite time-to-event endpoint, stratified log-rank testing, covariate-adjusted Cox regression, competing-risk analysis, continuous patient-reported outcome analysis, mixed-effects modeling, missing-data sensitivity analysis, and multiplicity control.
| Concept | How it appears in GALACTIC-HF |
|---|---|
| Randomization | Randomized, parallel, 2-arm phase 3 design |
| Triple masking | Registry reports triple masking |
| Time-to-event endpoint | Time to cardiovascular death or first heart failure event |
| Hazard ratio | Primary treatment effect estimated as HR 0.92 |
| Confidence interval | Primary 95% CI 0.86–0.99 |
| Stratified log-rank test | Primary hypothesis test with P = 0.0211 |
| Cox model | Stratified by randomization setting and geographic region |
| Covariate adjustment | Baseline eGFR and treatment group included in the Cox model |
| Competing risks | Subdistribution hazard-ratio analyses reported |
| Mixed-effects modeling | Reported for KCCQ TSS analysis |
| General linear model | Omnibus F-test reported for KCCQ TSS |
| Missing data | Joint longitudinal and survival sensitivity analysis for KCCQ |
| Multiplicity | Two-sided overall type I error of 0.05 across primary and secondary outcomes |
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Calculators
25. Sources
- ClinicalTrials.gov: GALACTIC-HF, NCT02929329.
- PubMed: PMID 41847775.
- PubMed: PMID 41771088.
- PubMed: PMID 41528277.
- PubMed: PMID 41123512.
- PubMed: PMID 38215932.
Continue with the statistical methods
Explore the underlying survival-analysis, regression, longitudinal-data, missing-data, and clinical-trial concepts used in GALACTIC-HF.
26. Record Summary
GALACTIC-HF provides a compact example of how a modern randomized cardiovascular trial can combine several statistical estimands and analytical frameworks. The primary endpoint was a time-to-event composite of cardiovascular death or first heart failure event. Its principal Cox analysis reported HR 0.92 (95% CI 0.86–0.99; P = 0.0252), while the stratified log-rank test reported P = 0.0211. Secondary analyses examined cardiovascular death, first heart failure hospitalization, all-cause death, and change in KCCQ TSS, with competing-risk, mixed-effects, general linear-model, and joint longitudinal-survival approaches represented in the posted analyses.
The most important statistical lesson is that these results should not be reduced to a single p-value. The hazard ratio describes relative treatment effect, the confidence interval describes statistical precision, the log-rank test addresses the treatment-distribution comparison, the competing-risk analysis addresses a different event structure, and the KCCQ analyses address a continuous patient-reported outcome with additional missing-data considerations. Multiplicity further determines how individual hypothesis tests should be interpreted within the overall trial.