This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
NAVIGATE ESUS was a randomized, parallel-group, quadruple-masked phase 3 trial comparing rivaroxaban with acetylsalicylic acid for secondary prevention after embolic stroke of undetermined source and prevention of systemic embolism. The registry reports two primary endpoints, both analyzed as time-to-event outcomes using stratified log-rank testing and hazard ratios from stratified Cox proportional-hazards models.
| Feature | NAVIGATE ESUS |
|---|---|
| Trial name | NAVIGATE ESUS |
| NCT ID | NCT02313909 |
| Therapeutic area | Neurology |
| Condition | Stroke |
| Phase | Phase 3 |
| Status | Terminated |
| Enrollment | 7213 |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Prevention |
| Lead sponsor | Bayer |
| Sponsor type | Industry |
| Start | 2014-12-23 |
| Primary completion | 2018-02-15 |
2. Clinical Question
The primary clinical-statistical question was whether rivaroxaban 15 mg once daily differed from acetylsalicylic acid 100 mg once daily on the registered efficacy and major-bleeding time-to-event outcomes in participants with recent embolic stroke of undetermined source.
Population
Patients represented in the NAVIGATE ESUS trial for the condition of stroke and the registered brief-title population of recent embolic stroke of undetermined source.
Intervention
Rivaroxaban 15 mg once daily, with the registry also identifying rivaroxaban-placebo as part of the masked study regimen.
Comparator
Acetylsalicylic acid 100 mg once daily, with aspirin-placebo identified in the masked study regimen.
Primary question
How did the randomized rivaroxaban-versus-acetylsalicylic-acid comparison perform for the adjudicated composite efficacy outcome and ISTH major bleeding?
3. Trial Design
Rivaroxaban
- Rivaroxaban 15 mg once daily
- Rivaroxaban-placebo identified in the masked regimen
- Compared with acetylsalicylic acid 100 mg once daily
Acetylsalicylic acid
- Acetylsalicylic acid 100 mg once daily
- Aspirin-placebo identified in the masked regimen
- Served as the control group for the randomized comparisons
The design was randomized and parallel, with quadruple masking. That combination is important statistically because randomization establishes the framework for comparing treatment assignments, while masking is intended to reduce the influence of treatment knowledge on trial conduct and outcome assessment.
4. Randomization, Stratification, and Analysis Population
The registry identifies the primary and secondary efficacy analyses as using an intention-to-treat analysis set, defined as including all randomized participants. The registry also identifies stratified analysis among the concepts in the statistical-analysis text.
| Element | Registry-supported description |
|---|---|
| Randomization | Randomized allocation |
| Design | Parallel |
| Masking | Quadruple |
| Primary efficacy population | Intention-to-treat analysis set including all randomized participants |
| Analysis stratification | Stratified analysis; the registry analysis text specifies stratified Cox proportional-hazards modeling and stratified log-rank testing |
| Effect measure | Hazard ratio |
The ITT principle is particularly important for a randomized trial. Once participants are randomized, the comparison remains anchored to that assignment rather than being redefined according to later treatment exposure. In this registry record, that principle applies to both primary endpoints and the reported secondary time-to-event analyses.
5. Primary Endpoints
| Endpoint | Registry definition / time frame | Analysis |
|---|---|---|
| Incidence Rate of the Composite Efficacy Outcome (Adjudicated) | From randomization until the efficacy cut-off date (median 326 days). Components include stroke (ischemic, hemorrhagic, and undefined stroke, TIA with positive neuroimaging) and systemic embolism. Incidence rate was estimated as the number of participants with incident events divided by cumulative at-risk time, with a participant no longer at risk once an incident event occurred. | Stratified log-rank test; stratified Cox proportional-hazards model; HR |
| Incidence Rate of a Major Bleeding Event According to ISTH Criteria (Adjudicated) | From randomization until the efficacy cut-off date (median 326 days). Major bleeding was defined according to ISTH criteria, including fatal bleeding; symptomatic bleeding in a critical area or organ; symptomatic intracranial haemorrhage; or clinically overt bleeding associated with a recent decrease in hemoglobin level. | Stratified log-rank test; stratified Cox proportional-hazards model; HR |
Both registered primary endpoints are time-to-event outcomes. That matters because the analysis is not simply a comparison of the proportion of participants who experienced an event. Follow-up time contributes to the analysis, participants can be censored, and the hazard ratio summarizes the relative event hazard under the fitted Cox model.
6. Statistical Methodology
Incidence rates and time at risk
The registry defines the primary incidence-rate outcomes using the number of participants with incident events divided by cumulative at-risk time. For the composite efficacy outcome, once a participant experienced an incident event, that participant was no longer considered at risk for another incident event for this endpoint.
This expresses event occurrence relative to accumulated participant follow-up rather than simply reporting a percentage of participants. It is therefore a rate-based representation of the time-to-event experience.
Log-rank testing
The registry reports a stratified log-rank test for the comparison of rivaroxaban with acetylsalicylic acid. The log-rank framework compares the experience of the randomized groups across event times while accounting for the ordering of events over follow-up.
The important distinction is that the log-rank test addresses evidence for a difference between survival or event-time distributions, whereas the hazard ratio provides an estimate of the relative event hazard. The two are complementary rather than interchangeable.
Stratified Cox proportional-hazards model
The registry states that risk reduction was estimated with a stratified Cox proportional-hazards model and that hazard ratios with 95% confidence intervals were reported relative to the acetylsalicylic acid arm.
An HR below 1 indicates a lower estimated event hazard for rivaroxaban under the model; an HR above 1 indicates a higher estimated event hazard.
Two-sided inference
The registry reports two-sided 95% confidence intervals and two-sided P-values for the posted analyses. The hypothesis type is recorded as superiority. Thus, the reported inferential framework tests for a difference in either direction rather than restricting the statistical alternative to a single direction.
Intention-to-treat analysis
For each posted efficacy analysis, the registry specifies an intention-to-treat analysis set including all randomized participants. This preserves the randomized comparison as the primary basis for inference and avoids redefining the groups according to post-randomization experience.
7. Primary Result: Composite Efficacy Outcome
The first primary endpoint was the incidence rate of the adjudicated composite efficacy outcome from randomization until the efficacy cut-off date, with a median of 326 days. The comparison was rivaroxaban 15 mg once daily versus acetylsalicylic acid 100 mg once daily.
Hazard ratio for the composite efficacy outcome
95% CI: 0.87–1.33 · P = 0.51884
Stratified log-rank test; hazard ratio estimated with a stratified Cox proportional-hazards model.
The hazard ratio of 1.07 is above 1, corresponding to an estimated hazard that was 7% higher for rivaroxaban relative to acetylsalicylic acid under the fitted model. That simple interpretation should not be turned into a claim of a 7% increase in individual patient risk: a hazard ratio is a relative time-to-event measure, not an individual-level probability.
What the estimate means: HR 1.07 means that the fitted model estimated the instantaneous hazard of the composite efficacy event to be 1.07 times the hazard in the acetylsalicylic acid group. In relative terms, this corresponds to a 7% higher estimated hazard for rivaroxaban.
What it does not mean: it does not mean that 7% more participants necessarily experienced an event, nor does it mean that every individual participant had a 7% higher probability of the outcome.
What the confidence interval says: the 95% CI of 0.87–1.33 describes statistical uncertainty around the estimated hazard ratio. It includes 1, so the interval is compatible with a lower hazard, little difference, or a higher hazard under the model and sampling framework.
Why the P-value is different from effect size: P = 0.51884 measures the strength of evidence against the null hypothesis in the specified testing framework. It does not measure how large or clinically important the treatment effect is.
Important cautions: this was a time-to-event analysis with censoring and a stratified Cox model. Interpretation of a single HR also depends on the proportional-hazards framework. In addition, the study was terminated early after the second interim analysis, which means the available evidence arose in the context of an early stopping decision.
8. Primary Result: ISTH Major Bleeding
The second primary endpoint was the incidence rate of an adjudicated major bleeding event according to International Society on Thrombosis and Haemostasis criteria, measured from randomization until the efficacy cut-off date with a median of 326 days.
Hazard ratio for ISTH major bleeding
95% CI: 1.68–4.39 · P = 0.00002
Stratified log-rank test; hazard ratio estimated with a stratified Cox proportional-hazards model.
The point estimate of 2.72 indicates an estimated instantaneous major-bleeding hazard 2.72 times that of the acetylsalicylic acid group under the fitted model. In relative terms, that corresponds to an estimated hazard approximately 172% higher than the comparator hazard.
What the estimate means: HR 2.72 means that the fitted model estimated the instantaneous hazard of an ISTH major bleeding event to be 2.72 times the hazard in the acetylsalicylic acid group.
What it does not mean: it does not mean that 2.72 times as many participants necessarily experienced major bleeding, and it does not provide an individual patient's absolute probability of major bleeding.
What the confidence interval says: the 95% CI of 1.68–4.39 lies entirely above 1. The estimated relative hazard is therefore separated from the null value within this confidence-interval framework, although the interval also shows substantial uncertainty in the exact magnitude.
Why the P-value is not the effect size: P = 0.00002 indicates strong statistical evidence against the null hypothesis in the reported two-sided testing framework. It does not itself say whether the effect is small, moderate, or large; the hazard ratio and confidence interval provide that information.
Important cautions: the analysis is time-to-event based and depends on censoring and the stratified Cox model. The early termination of the trial is also part of the context in which this estimate was generated.
The contrast between the two primary hazard ratios is statistically instructive. The composite efficacy HR was 1.07, while the ISTH major-bleeding HR was 2.72. These are separate endpoints, with different clinical definitions and different event processes. They should not be collapsed into a single numerical measure of overall benefit or harm.
9. Secondary Efficacy Results
The registry contains ten posted secondary analyses. All use the intention-to-treat analysis set, compare rivaroxaban 15 mg once daily with acetylsalicylic acid 100 mg once daily, and use the same broad time-to-event framework: stratified log-rank testing with hazard ratios estimated from stratified Cox proportional-hazards models.
| Secondary endpoint | HR | 95% CI | P-value |
|---|---|---|---|
| Cardiovascular death, recurrent stroke, systemic embolism and myocardial infarction | 1.06 | 0.87–1.29 | 0.56922 |
| All-cause mortality | 1.26 | 0.87–1.81 | 0.22078 |
| Stroke | 1.08 | 0.87–1.34 | 0.47970 |
| Ischemic stroke | 1.03 | 0.83–1.29 | 0.78738 |
| Disabling stroke | 1.42 | 0.88–2.28 | 0.14822 |
| Cardiovascular death | 1.48 | 0.87–2.52 | 0.14051 |
| Myocardial infarction | 0.74 | 0.39–1.38 | 0.34284 |
The secondary efficacy results illustrate why a collection of hazard ratios should be read as a pattern rather than as a list of isolated P-values. The estimates range from 0.74 for myocardial infarction to 1.48 for cardiovascular death, but each confidence interval is compatible with the null value of 1. The point estimates therefore provide descriptive information about the observed direction and magnitude of the modeled comparisons, while the intervals convey their statistical uncertainty.
Composite cardiovascular outcome
Cardiovascular death, recurrent stroke, systemic embolism and myocardial infarction
95% CI: 0.87–1.29 · P = 0.56922
All-cause mortality
All-cause mortality
95% CI: 0.87–1.81 · P = 0.22078
Individual vascular outcomes
| Outcome | HR | 95% CI | P-value | Statistical reading |
|---|---|---|---|---|
| Stroke | 1.08 | 0.87–1.34 | 0.47970 | Point estimate above 1; CI includes 1 |
| Ischemic stroke | 1.03 | 0.83–1.29 | 0.78738 | Point estimate close to 1; CI includes 1 |
| Disabling stroke | 1.42 | 0.88–2.28 | 0.14822 | Point estimate above 1; relatively wide CI includes 1 |
| Cardiovascular death | 1.48 | 0.87–2.52 | 0.14051 | Point estimate above 1; wide CI includes 1 |
| Myocardial infarction | 0.74 | 0.39–1.38 | 0.34284 | Point estimate below 1; wide CI includes 1 |
In particular, the myocardial-infarction estimate of 0.74 should not be interpreted as evidence of a demonstrated 26% reduction. The confidence interval of 0.39–1.38 is wide and includes both values below and above the null. The estimate is therefore best understood as one component of the broader secondary-outcome pattern rather than as a standalone treatment conclusion.
10. Secondary Bleeding Results
The registry reports three additional bleeding endpoints: life-threatening bleeding, clinically relevant non-major bleeding, and intracranial hemorrhage. All were analyzed as time-to-event outcomes using the same stratified log-rank and stratified Cox framework.
| Secondary bleeding endpoint | HR | 95% CI | P-value |
|---|---|---|---|
| Life-threatening bleeding events | 2.34 | 1.28–4.29 | 0.00443 |
| Clinically relevant non-major bleeding events | 1.51 | 1.13–2.00 | 0.00451 |
| Intracranial hemorrhage | 2.01 | 1.00–4.02 | 0.04409 |
Life-threatening bleeding
Hazard ratio for life-threatening bleeding
95% CI: 1.28–4.29 · P = 0.00443
The HR of 2.34 corresponds to a modeled instantaneous hazard more than twice that of the acetylsalicylic acid group. The 95% CI remains above 1, while its width indicates uncertainty about the exact magnitude.
Clinically relevant non-major bleeding
Hazard ratio for clinically relevant non-major bleeding
95% CI: 1.13–2.00 · P = 0.00451
The HR of 1.51 indicates a 51% higher estimated instantaneous hazard under the fitted model. Again, this is a relative hazard measure, not a 51-percentage-point difference in event probability.
Intracranial hemorrhage
Hazard ratio for intracranial hemorrhage
95% CI: 1.00–4.02 · P = 0.04409
The point estimate is approximately two times the comparator hazard. The lower confidence-limit value is exactly 1.00, making the precision of the estimate particularly important when reading the result rather than relying on the point estimate alone.
11. Safety Results
The registry provides serious adverse-event counts by randomized arm as affected participants divided by participants at risk.
| Safety measure | Rivaroxaban 15 mg OD | Acetylsalicylic acid 100 mg OD |
|---|---|---|
| Serious adverse events | 466 / 3562 | 434 / 3559 |
The serious-adverse-event figures are descriptive affected/at-risk counts and should not be substituted for the adjudicated time-to-event analyses. The primary bleeding endpoint, for example, was analyzed using an incidence rate, stratified log-rank test, and Cox-derived hazard ratio. A simple affected/at-risk ratio answers a different statistical question.
12. Statistical Methods Explained
Why was a log-rank test used?
The primary endpoints are time-to-event outcomes, so participants can experience events at different times and some participants can remain event-free through their available follow-up. A log-rank test is designed to compare event-time distributions between groups while using information across the follow-up period. NAVIGATE ESUS used the stratified form of the test.
What does an HR of 1.07 mean?
An HR of 1.07 means that the fitted model estimated a 1.07-fold instantaneous hazard in the rivaroxaban group relative to the acetylsalicylic acid group. It is equivalent to a 7% higher estimated hazard, not a 7-percentage-point increase in the proportion of participants experiencing an event.
Why report a confidence interval with the hazard ratio?
A point estimate alone hides uncertainty. The 95% CI of 0.87–1.33 around the primary efficacy HR shows that the estimated hazard ratio is not known precisely enough to reduce the result to the single value 1.07. The interval also crosses the null value of 1, which is central to its statistical interpretation.
Why is the P-value not an effect-size measure?
A P-value addresses compatibility with a null hypothesis under the specified statistical framework. It depends on the observed data, variability, and amount of information. The size and direction of the treatment effect are described by the hazard ratio and, where relevant, by absolute event rates or other absolute measures. NAVIGATE ESUS therefore should not be summarized by P-values alone.
What does intention-to-treat contribute?
The ITT analysis set includes all randomized participants. Analyzing according to randomized assignment preserves the treatment comparison created by randomization. It also means that the efficacy analysis remains a comparison of treatment strategies as assigned rather than becoming a comparison of selected participants who happened to remain on treatment.
Why does early termination matter statistically?
The registry states that the study was terminated early at the second interim analysis because there was no efficacy improvement over aspirin and very little chance of showing overall benefit if the study were completed. Early stopping changes the information available for estimation and means the final posted evidence must be understood in the context of the interim monitoring decision rather than as if the planned study had simply run to completion.
13. Understanding the Hazard Ratio Across Endpoints
NAVIGATE ESUS provides a useful demonstration of why hazard ratios should always be interpreted together with their endpoint definitions.
| Endpoint | HR | 95% CI | Direction of point estimate |
|---|---|---|---|
| Composite efficacy outcome | 1.07 | 0.87–1.33 | Above 1 |
| ISTH major bleeding | 2.72 | 1.68–4.39 | Above 1 |
| Cardiovascular death + recurrent stroke + systemic embolism + MI | 1.06 | 0.87–1.29 | Above 1 |
| All-cause mortality | 1.26 | 0.87–1.81 | Above 1 |
| Myocardial infarction | 0.74 | 0.39–1.38 | Below 1 |
| Life-threatening bleeding | 2.34 | 1.28–4.29 | Above 1 |
| Clinically relevant non-major bleeding | 1.51 | 1.13–2.00 | Above 1 |
| Intracranial hemorrhage | 2.01 | 1.00–4.02 | Above 1 |
The table illustrates two important principles. First, an HR above 1 is not inherently good or bad; its meaning depends on whether the endpoint is an efficacy event or an adverse event. Second, the width and position of the confidence interval matter as much as the point estimate. An HR of 0.74 with a CI of 0.39–1.38 carries a very different level of statistical precision from an HR of 1.51 with a CI of 1.13–2.00.
14. What the Primary Efficacy Result Does — and Does Not — Mean
The primary efficacy HR of 1.07 indicates an estimated 7% higher instantaneous hazard for the rivaroxaban group relative to acetylsalicylic acid under the stratified Cox model. The 95% CI of 0.87–1.33 includes 1, and the reported two-sided P-value is 0.51884.
This does not establish that rivaroxaban increases the absolute probability of the composite outcome by 7%, nor does it establish equivalence between the treatments. A nonsignificant superiority test is not itself a formal demonstration of equivalence or non-inferiority.
The interval from 0.87 to 1.33 allows for treatment effects in either direction within the statistical uncertainty represented by the model. It is therefore more informative than the point estimate alone. It tells the reader that the observed estimate should not be treated as a precise measurement of a single underlying hazard ratio.
The primary major-bleeding HR of 2.72 has a 95% CI of 1.68–4.39. This is a different endpoint with a different clinical meaning from the composite efficacy outcome. The two estimates should therefore be reported side by side rather than mathematically combined into an invented net-benefit statistic.
15. Early Termination and Interim Analysis
The registry states that NAVIGATE ESUS was terminated early because the second interim analysis showed no efficacy improvement over aspirin and there was very little chance of showing overall benefit if the study were completed.
Why interim analysis matters
An interim analysis evaluates accumulating trial information before the originally contemplated end of follow-up. Once a study stops early, the resulting estimates reflect the amount of information available at that stopping point.
Why stopping for futility changes context
The registry's reason for termination was lack of efficacy improvement and very little chance of demonstrating overall benefit with continued enrollment or follow-up. The results therefore belong to a trial that did not continue to its originally intended completion.
The ClinicalTrials.gov record does not provide the numerical interim boundary, alpha-spending function, conditional-power threshold, or stopping boundary. Those quantities should not be reconstructed from the posted final estimates. What can be stated from the registry is the occurrence and reason for the second-interim termination.
16. Proportional-Hazards Considerations
The registry reports hazard ratios from stratified Cox proportional-hazards models. A Cox HR is a model-based summary of relative event hazard, and its simplest interpretation assumes that the relative hazards are reasonably represented by a proportional-hazards relationship over the analyzed period.
A hazard ratio compares instantaneous event hazards. It should not automatically be translated into a percentage difference in cumulative incidence or an absolute patient-level risk.
This distinction is especially relevant when interpreting endpoints observed over a median of 326 days. A hazard ratio summarizes the modeled event process over follow-up; it does not directly provide the probability that a particular participant will experience the endpoint by the efficacy cut-off date.
17. Confidence Intervals and Statistical Precision
The registry provides 95% two-sided confidence intervals for all twelve posted statistical analyses. Comparing their widths illustrates differences in statistical precision.
| Endpoint | HR | 95% CI | Interpretive feature |
|---|---|---|---|
| Composite efficacy | 1.07 | 0.87–1.33 | Includes the null and excludes neither direction broadly |
| ISTH major bleeding | 2.72 | 1.68–4.39 | Entire interval above 1 |
| All-cause mortality | 1.26 | 0.87–1.81 | Wide interval including 1 |
| Disabling stroke | 1.42 | 0.88–2.28 | Wide interval including 1 |
| Cardiovascular death | 1.48 | 0.87–2.52 | Wide interval including 1 |
| Myocardial infarction | 0.74 | 0.39–1.38 | Wide interval including 1 |
| Intracranial hemorrhage | 2.01 | 1.00–4.02 | Lower limit reaches 1.00 |
A confidence interval should not be interpreted as the range in which the individual patient's treatment effect lies. It is an interval estimate for the population-level parameter under the statistical model and repeated-sampling framework. Wider intervals indicate less precision; narrower intervals indicate greater precision, all else equal.
18. Secondary Endpoint Interpretation
The secondary analyses should be interpreted in the context of their role as additional endpoints rather than as substitutes for the primary analysis. Several point estimates are above 1, while myocardial infarction has a point estimate below 1. This variation is expected when multiple related outcomes are examined.
Direction is not proof
A point estimate below 1 does not by itself establish a protective effect, just as a point estimate above 1 does not by itself establish increased risk.
Precision matters
The confidence interval determines how much uncertainty surrounds each estimate. Several secondary endpoints have intervals spanning substantial ranges on both sides of 1.
P-values need context
The registry supplies two-sided P-values for the analyses. They should be read alongside the endpoint definition, HR, confidence interval, and analysis population.
Different endpoints answer different questions
Stroke, mortality, myocardial infarction, and bleeding are not interchangeable outcomes. Their statistical estimates should remain tied to their specific definitions.
19. Limitations
- Early termination: the study was terminated at the second interim analysis because the registry states there was no efficacy improvement over aspirin and very little chance of showing overall benefit if the study were completed.
- Early stopping and precision: because the study stopped early, the available information reflects the follow-up and event accumulation at the stopping point rather than a completed study course.
- Hazard-ratio assumptions: the primary and secondary analyses use stratified Cox proportional-hazards models. A single HR is a model-based summary and should not automatically be interpreted as a constant relative risk over all times.
- Time-to-event interpretation: hazard ratios are not equivalent to cumulative risks, risk differences, or individual probabilities.
- Multiple secondary endpoints: the registry reports numerous secondary analyses. Their P-values should not automatically be interpreted as though each were an isolated primary hypothesis without considering the overall analysis structure.
- Secondary endpoint precision: several secondary confidence intervals are wide, particularly for outcomes such as disabling stroke, cardiovascular death, and myocardial infarction.
- Safety versus efficacy populations: the registry provides an ITT analysis set for the posted efficacy analyses, while the serious-adverse-event figures are presented as affected/at-risk counts. These summaries answer different questions.
- Registry-level information: the ClinicalTrials.gov record does not provide every element of the underlying statistical analysis plan, such as numerical interim boundaries, detailed missing-data procedures, or a full description of all stratification variables.
20. What the Registry Does Not Establish From These Data
Several common statistical conclusions should not be inferred simply from the posted NAVIGATE ESUS results.
| Potential overinterpretation | Why it is not justified by the ClinicalTrials.gov record |
|---|---|
| “HR 1.07 proves the treatments are equivalent.” | The registry identifies the hypothesis type as superiority. A nonsignificant superiority comparison is not the same as a formal equivalence or non-inferiority demonstration. |
| “HR 1.07 means 7% more patients had an event.” | The HR is a time-to-event measure, not an absolute event-rate difference. |
| “HR 2.72 means 172% of patients had major bleeding.” | The HR describes relative instantaneous hazard, not a percentage of patients. |
| “P = 0.00002 is a measure of how large the bleeding effect was.” | The P-value measures evidence against the specified null hypothesis; effect magnitude is described by the HR and confidence interval. |
| “The myocardial-infarction HR of 0.74 proves a 26% reduction.” | The 95% CI is 0.39–1.38 and includes 1, so the point estimate alone does not establish a treatment effect. |
| “The early stopping result is identical to a completed-trial result.” | The registry explicitly states that the trial was terminated early at the second interim analysis. |
21. Why This Trial Matters Statistically
NAVIGATE ESUS is a useful teaching case because its registry results bring several core clinical-trial concepts together in one randomized comparison. The primary endpoints are both time-to-event outcomes, the analysis uses stratified log-rank testing and Cox modeling, the efficacy population is ITT, and the trial stopped early after an interim analysis.
| Concept | How it appears in NAVIGATE ESUS |
|---|---|
| Randomization | Randomized parallel-group phase 3 design |
| Blinding | Quadruple masking |
| Intention-to-treat | All randomized participants included in the posted efficacy analysis set |
| Time-to-event endpoints | Both primary endpoints and all posted secondary analyses use time-to-event methodology |
| Hazard ratio | Primary and secondary effect measure |
| Confidence intervals | 95% two-sided intervals reported for all posted statistical analyses |
| Log-rank test | Stratified treatment comparison |
| Cox model | Stratified Cox proportional-hazards model for risk-reduction estimation |
| Interim analysis | Second interim analysis led to early termination |
| Composite endpoint | Primary adjudicated efficacy outcome combines stroke-related events, TIA with positive neuroimaging, and systemic embolism |
| Safety endpoint | ISTH major bleeding was a primary endpoint; additional bleeding outcomes were secondary endpoints |
22. A Statistical Reading of the Full Results
Viewed as a statistical profile rather than as a collection of isolated numbers, the registry record has three prominent features.
First, the primary efficacy comparison produced an HR of 1.07 with a 95% CI of 0.87–1.33 and P = 0.51884. The point estimate is close to the null value of 1, and the confidence interval spans both sides of the null. The result therefore does not provide strong statistical evidence of superiority for rivaroxaban on the registered composite efficacy outcome.
Second, the primary major-bleeding comparison produced an HR of 2.72 with a 95% CI of 1.68–4.39 and P = 0.00002. Unlike the efficacy estimate, the entire confidence interval lies above 1. The result represents a materially different statistical pattern, with the estimated hazard substantially above the comparator and the reported two-sided P-value indicating strong evidence against the null hypothesis.
Third, the secondary results show a mixture of point-estimate directions, but the confidence intervals for the reported efficacy outcomes include 1. The additional bleeding analyses have HRs above 1, with confidence intervals of 1.28–4.29 for life-threatening bleeding, 1.13–2.00 for clinically relevant non-major bleeding, and 1.00–4.02 for intracranial hemorrhage.
These observations should remain connected to the study's design. The trial was randomized and quadruple-masked, used ITT efficacy analyses, applied stratified time-to-event methods, and was terminated early after the second interim analysis. Each of those features contributes to the interpretation of the numerical results.
23. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Sources
- ClinicalTrials.gov: NAVIGATE ESUS, NCT02313909.
- PubMed record: PMID 29766772.
- PubMed record: PMID 29525076.
- PubMed record: PMID 31008276.
- PubMed record: PMID 39358761.
- PubMed record: PMID 34791941.
Continue through the Clinical Biostats statistical pathway
Explore the core survival-analysis, clinical-trial, and statistical-inference methods represented in NAVIGATE ESUS.
26. Record Summary
NAVIGATE ESUS provides a clear example of a randomized clinical trial in which the principal statistical questions are time-to-event questions. The registry reports two primary endpoints: an adjudicated composite efficacy outcome and ISTH major bleeding. Both were analyzed using stratified log-rank testing and hazard ratios from stratified Cox proportional-hazards models in the intention-to-treat population.
The primary efficacy analysis produced an HR of 1.07 (95% CI 0.87–1.33; P = 0.51884), while the primary major-bleeding analysis produced an HR of 2.72 (95% CI 1.68–4.39; P = 0.00002). The secondary efficacy estimates likewise require interpretation through their confidence intervals rather than their point estimates alone, while the additional bleeding analyses produced HRs of 2.34 for life-threatening bleeding, 1.51 for clinically relevant non-major bleeding, and 2.01 for intracranial hemorrhage.
The statistical story is therefore not simply a matter of identifying which P-values are small. It involves understanding the randomized comparison, the ITT population, the adjudicated endpoint definitions, the time-to-event framework, the distinction between hazard ratios and absolute risks, the uncertainty represented by 95% confidence intervals, and the effect of early termination after the second interim analysis.