← Clinical Trials
Stroke Phase 3 Time-to-Event Analysis NCT02313909

NAVIGATE ESUS: Complete Statistical Analysis of Rivaroxaban in Embolic Stroke of Undetermined Source

An independent statistical review of the randomized phase 3 NAVIGATE ESUS trial comparing rivaroxaban 15 mg once daily with acetylsalicylic acid 100 mg once daily in patients with recent embolic stroke of undetermined source, focusing on time-to-event methodology, efficacy, bleeding outcomes, and statistical interpretation.

Trial status: Terminated  ·  Enrollment: 7213  ·  Primary completion: 2018-02-15
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

NAVIGATE ESUS was a randomized, parallel-group, quadruple-masked phase 3 trial comparing rivaroxaban with acetylsalicylic acid for secondary prevention after embolic stroke of undetermined source and prevention of systemic embolism. The registry reports two primary endpoints, both analyzed as time-to-event outcomes using stratified log-rank testing and hazard ratios from stratified Cox proportional-hazards models.

7213
Enrolled
2 randomized arms
1.07
Efficacy HR
95% CI 0.87–1.33
2.72
Major bleeding HR
95% CI 1.68–4.39
326
Median days
To efficacy cut-off
FeatureNAVIGATE ESUS
Trial nameNAVIGATE ESUS
NCT IDNCT02313909
Therapeutic areaNeurology
ConditionStroke
PhasePhase 3
StatusTerminated
Enrollment7213
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposePrevention
Lead sponsorBayer
Sponsor typeIndustry
Start2014-12-23
Primary completion2018-02-15

2. Clinical Question

The primary clinical-statistical question was whether rivaroxaban 15 mg once daily differed from acetylsalicylic acid 100 mg once daily on the registered efficacy and major-bleeding time-to-event outcomes in participants with recent embolic stroke of undetermined source.

Population

Patients represented in the NAVIGATE ESUS trial for the condition of stroke and the registered brief-title population of recent embolic stroke of undetermined source.

Intervention

Rivaroxaban 15 mg once daily, with the registry also identifying rivaroxaban-placebo as part of the masked study regimen.

Comparator

Acetylsalicylic acid 100 mg once daily, with aspirin-placebo identified in the masked study regimen.

Primary question

How did the randomized rivaroxaban-versus-acetylsalicylic-acid comparison perform for the adjudicated composite efficacy outcome and ISTH major bleeding?

3. Trial Design

01
Enroll 7213 participants
02
Randomize 2 parallel treatment groups
03
Mask Quadruple masking
04
Follow Events from randomization
05
Analyze Log-rank + stratified Cox
ARM A

Rivaroxaban

  • Rivaroxaban 15 mg once daily
  • Rivaroxaban-placebo identified in the masked regimen
  • Compared with acetylsalicylic acid 100 mg once daily
ARM B

Acetylsalicylic acid

  • Acetylsalicylic acid 100 mg once daily
  • Aspirin-placebo identified in the masked regimen
  • Served as the control group for the randomized comparisons

The design was randomized and parallel, with quadruple masking. That combination is important statistically because randomization establishes the framework for comparing treatment assignments, while masking is intended to reduce the influence of treatment knowledge on trial conduct and outcome assessment.

Early termination: the registry states that the study was terminated early due to no efficacy improvement over aspirin at the second interim analysis and very little chance of showing overall benefit if the study were completed. This is a design feature that directly affects how the observed results should be interpreted: the available follow-up reflects an early stopping decision rather than completion of the originally intended study course.

4. Randomization, Stratification, and Analysis Population

The registry identifies the primary and secondary efficacy analyses as using an intention-to-treat analysis set, defined as including all randomized participants. The registry also identifies stratified analysis among the concepts in the statistical-analysis text.

ElementRegistry-supported description
RandomizationRandomized allocation
DesignParallel
MaskingQuadruple
Primary efficacy populationIntention-to-treat analysis set including all randomized participants
Analysis stratificationStratified analysis; the registry analysis text specifies stratified Cox proportional-hazards modeling and stratified log-rank testing
Effect measureHazard ratio

The ITT principle is particularly important for a randomized trial. Once participants are randomized, the comparison remains anchored to that assignment rather than being redefined according to later treatment exposure. In this registry record, that principle applies to both primary endpoints and the reported secondary time-to-event analyses.

5. Primary Endpoints

EndpointRegistry definition / time frameAnalysis
Incidence Rate of the Composite Efficacy Outcome (Adjudicated) From randomization until the efficacy cut-off date (median 326 days). Components include stroke (ischemic, hemorrhagic, and undefined stroke, TIA with positive neuroimaging) and systemic embolism. Incidence rate was estimated as the number of participants with incident events divided by cumulative at-risk time, with a participant no longer at risk once an incident event occurred. Stratified log-rank test; stratified Cox proportional-hazards model; HR
Incidence Rate of a Major Bleeding Event According to ISTH Criteria (Adjudicated) From randomization until the efficacy cut-off date (median 326 days). Major bleeding was defined according to ISTH criteria, including fatal bleeding; symptomatic bleeding in a critical area or organ; symptomatic intracranial haemorrhage; or clinically overt bleeding associated with a recent decrease in hemoglobin level. Stratified log-rank test; stratified Cox proportional-hazards model; HR

Both registered primary endpoints are time-to-event outcomes. That matters because the analysis is not simply a comparison of the proportion of participants who experienced an event. Follow-up time contributes to the analysis, participants can be censored, and the hazard ratio summarizes the relative event hazard under the fitted Cox model.

6. Statistical Methodology

Incidence rates and time at risk

The registry defines the primary incidence-rate outcomes using the number of participants with incident events divided by cumulative at-risk time. For the composite efficacy outcome, once a participant experienced an incident event, that participant was no longer considered at risk for another incident event for this endpoint.

Registry outcome unit
event / 100 participant-years

This expresses event occurrence relative to accumulated participant follow-up rather than simply reporting a percentage of participants. It is therefore a rate-based representation of the time-to-event experience.

Log-rank testing

The registry reports a stratified log-rank test for the comparison of rivaroxaban with acetylsalicylic acid. The log-rank framework compares the experience of the randomized groups across event times while accounting for the ordering of events over follow-up.

The important distinction is that the log-rank test addresses evidence for a difference between survival or event-time distributions, whereas the hazard ratio provides an estimate of the relative event hazard. The two are complementary rather than interchangeable.

Stratified Cox proportional-hazards model

The registry states that risk reduction was estimated with a stratified Cox proportional-hazards model and that hazard ratios with 95% confidence intervals were reported relative to the acetylsalicylic acid arm.

Hazard-ratio interpretation
HR = instantaneous event hazard in rivaroxaban ÷ instantaneous event hazard in acetylsalicylic acid

An HR below 1 indicates a lower estimated event hazard for rivaroxaban under the model; an HR above 1 indicates a higher estimated event hazard.

Two-sided inference

The registry reports two-sided 95% confidence intervals and two-sided P-values for the posted analyses. The hypothesis type is recorded as superiority. Thus, the reported inferential framework tests for a difference in either direction rather than restricting the statistical alternative to a single direction.

Intention-to-treat analysis

For each posted efficacy analysis, the registry specifies an intention-to-treat analysis set including all randomized participants. This preserves the randomized comparison as the primary basis for inference and avoids redefining the groups according to post-randomization experience.

7. Primary Result: Composite Efficacy Outcome

The first primary endpoint was the incidence rate of the adjudicated composite efficacy outcome from randomization until the efficacy cut-off date, with a median of 326 days. The comparison was rivaroxaban 15 mg once daily versus acetylsalicylic acid 100 mg once daily.

Hazard ratio for the composite efficacy outcome

1.07

95% CI: 0.87–1.33   ·   P = 0.51884

Stratified log-rank test; hazard ratio estimated with a stratified Cox proportional-hazards model.

The hazard ratio of 1.07 is above 1, corresponding to an estimated hazard that was 7% higher for rivaroxaban relative to acetylsalicylic acid under the fitted model. That simple interpretation should not be turned into a claim of a 7% increase in individual patient risk: a hazard ratio is a relative time-to-event measure, not an individual-level probability.

Clinical Biostats interpretation

What the estimate means: HR 1.07 means that the fitted model estimated the instantaneous hazard of the composite efficacy event to be 1.07 times the hazard in the acetylsalicylic acid group. In relative terms, this corresponds to a 7% higher estimated hazard for rivaroxaban.

What it does not mean: it does not mean that 7% more participants necessarily experienced an event, nor does it mean that every individual participant had a 7% higher probability of the outcome.

What the confidence interval says: the 95% CI of 0.87–1.33 describes statistical uncertainty around the estimated hazard ratio. It includes 1, so the interval is compatible with a lower hazard, little difference, or a higher hazard under the model and sampling framework.

Why the P-value is different from effect size: P = 0.51884 measures the strength of evidence against the null hypothesis in the specified testing framework. It does not measure how large or clinically important the treatment effect is.

Important cautions: this was a time-to-event analysis with censoring and a stratified Cox model. Interpretation of a single HR also depends on the proportional-hazards framework. In addition, the study was terminated early after the second interim analysis, which means the available evidence arose in the context of an early stopping decision.

Reading HR 1.07 together with its interval: the most important statistical feature is not simply that the point estimate is above 1. The 95% CI extends from 0.87 to 1.33, spanning the null value of 1. The reported P-value of 0.51884 is consistent with the absence of strong statistical evidence for a superiority difference in this analysis.

8. Primary Result: ISTH Major Bleeding

The second primary endpoint was the incidence rate of an adjudicated major bleeding event according to International Society on Thrombosis and Haemostasis criteria, measured from randomization until the efficacy cut-off date with a median of 326 days.

Hazard ratio for ISTH major bleeding

2.72

95% CI: 1.68–4.39   ·   P = 0.00002

Stratified log-rank test; hazard ratio estimated with a stratified Cox proportional-hazards model.

The point estimate of 2.72 indicates an estimated instantaneous major-bleeding hazard 2.72 times that of the acetylsalicylic acid group under the fitted model. In relative terms, that corresponds to an estimated hazard approximately 172% higher than the comparator hazard.

Clinical Biostats interpretation

What the estimate means: HR 2.72 means that the fitted model estimated the instantaneous hazard of an ISTH major bleeding event to be 2.72 times the hazard in the acetylsalicylic acid group.

What it does not mean: it does not mean that 2.72 times as many participants necessarily experienced major bleeding, and it does not provide an individual patient's absolute probability of major bleeding.

What the confidence interval says: the 95% CI of 1.68–4.39 lies entirely above 1. The estimated relative hazard is therefore separated from the null value within this confidence-interval framework, although the interval also shows substantial uncertainty in the exact magnitude.

Why the P-value is not the effect size: P = 0.00002 indicates strong statistical evidence against the null hypothesis in the reported two-sided testing framework. It does not itself say whether the effect is small, moderate, or large; the hazard ratio and confidence interval provide that information.

Important cautions: the analysis is time-to-event based and depends on censoring and the stratified Cox model. The early termination of the trial is also part of the context in which this estimate was generated.

The contrast between the two primary hazard ratios is statistically instructive. The composite efficacy HR was 1.07, while the ISTH major-bleeding HR was 2.72. These are separate endpoints, with different clinical definitions and different event processes. They should not be collapsed into a single numerical measure of overall benefit or harm.

9. Secondary Efficacy Results

The registry contains ten posted secondary analyses. All use the intention-to-treat analysis set, compare rivaroxaban 15 mg once daily with acetylsalicylic acid 100 mg once daily, and use the same broad time-to-event framework: stratified log-rank testing with hazard ratios estimated from stratified Cox proportional-hazards models.

Secondary endpoint HR 95% CI P-value
Cardiovascular death, recurrent stroke, systemic embolism and myocardial infarction 1.06 0.87–1.29 0.56922
All-cause mortality 1.26 0.87–1.81 0.22078
Stroke 1.08 0.87–1.34 0.47970
Ischemic stroke 1.03 0.83–1.29 0.78738
Disabling stroke 1.42 0.88–2.28 0.14822
Cardiovascular death 1.48 0.87–2.52 0.14051
Myocardial infarction 0.74 0.39–1.38 0.34284

The secondary efficacy results illustrate why a collection of hazard ratios should be read as a pattern rather than as a list of isolated P-values. The estimates range from 0.74 for myocardial infarction to 1.48 for cardiovascular death, but each confidence interval is compatible with the null value of 1. The point estimates therefore provide descriptive information about the observed direction and magnitude of the modeled comparisons, while the intervals convey their statistical uncertainty.

Composite cardiovascular outcome

Cardiovascular death, recurrent stroke, systemic embolism and myocardial infarction

HR 1.06

95% CI: 0.87–1.29   ·   P = 0.56922

All-cause mortality

All-cause mortality

HR 1.26

95% CI: 0.87–1.81   ·   P = 0.22078

Individual vascular outcomes

OutcomeHR95% CIP-valueStatistical reading
Stroke1.080.87–1.340.47970Point estimate above 1; CI includes 1
Ischemic stroke1.030.83–1.290.78738Point estimate close to 1; CI includes 1
Disabling stroke1.420.88–2.280.14822Point estimate above 1; relatively wide CI includes 1
Cardiovascular death1.480.87–2.520.14051Point estimate above 1; wide CI includes 1
Myocardial infarction0.740.39–1.380.34284Point estimate below 1; wide CI includes 1

In particular, the myocardial-infarction estimate of 0.74 should not be interpreted as evidence of a demonstrated 26% reduction. The confidence interval of 0.39–1.38 is wide and includes both values below and above the null. The estimate is therefore best understood as one component of the broader secondary-outcome pattern rather than as a standalone treatment conclusion.

10. Secondary Bleeding Results

The registry reports three additional bleeding endpoints: life-threatening bleeding, clinically relevant non-major bleeding, and intracranial hemorrhage. All were analyzed as time-to-event outcomes using the same stratified log-rank and stratified Cox framework.

Secondary bleeding endpointHR95% CIP-value
Life-threatening bleeding events2.341.28–4.290.00443
Clinically relevant non-major bleeding events1.511.13–2.000.00451
Intracranial hemorrhage2.011.00–4.020.04409

Life-threatening bleeding

Hazard ratio for life-threatening bleeding

2.34

95% CI: 1.28–4.29   ·   P = 0.00443

The HR of 2.34 corresponds to a modeled instantaneous hazard more than twice that of the acetylsalicylic acid group. The 95% CI remains above 1, while its width indicates uncertainty about the exact magnitude.

Clinically relevant non-major bleeding

Hazard ratio for clinically relevant non-major bleeding

1.51

95% CI: 1.13–2.00   ·   P = 0.00451

The HR of 1.51 indicates a 51% higher estimated instantaneous hazard under the fitted model. Again, this is a relative hazard measure, not a 51-percentage-point difference in event probability.

Intracranial hemorrhage

Hazard ratio for intracranial hemorrhage

2.01

95% CI: 1.00–4.02   ·   P = 0.04409

The point estimate is approximately two times the comparator hazard. The lower confidence-limit value is exactly 1.00, making the precision of the estimate particularly important when reading the result rather than relying on the point estimate alone.

11. Safety Results

The registry provides serious adverse-event counts by randomized arm as affected participants divided by participants at risk.

Safety measureRivaroxaban 15 mg ODAcetylsalicylic acid 100 mg OD
Serious adverse events466 / 3562434 / 3559
Serious adverse events: affected / at risk
Rivaroxaban
466 / 3562
Aspirin
434 / 3559

The serious-adverse-event figures are descriptive affected/at-risk counts and should not be substituted for the adjudicated time-to-event analyses. The primary bleeding endpoint, for example, was analyzed using an incidence rate, stratified log-rank test, and Cox-derived hazard ratio. A simple affected/at-risk ratio answers a different statistical question.

Safety versus efficacy: the trial contains distinct efficacy and safety endpoints. A treatment comparison cannot be summarized adequately by looking at only one efficacy HR or one adverse-event count. The appropriate interpretation keeps endpoint definitions, follow-up, analysis population, and statistical method aligned.

12. Statistical Methods Explained

Why was a log-rank test used?

The primary endpoints are time-to-event outcomes, so participants can experience events at different times and some participants can remain event-free through their available follow-up. A log-rank test is designed to compare event-time distributions between groups while using information across the follow-up period. NAVIGATE ESUS used the stratified form of the test.

What does an HR of 1.07 mean?

An HR of 1.07 means that the fitted model estimated a 1.07-fold instantaneous hazard in the rivaroxaban group relative to the acetylsalicylic acid group. It is equivalent to a 7% higher estimated hazard, not a 7-percentage-point increase in the proportion of participants experiencing an event.

Why report a confidence interval with the hazard ratio?

A point estimate alone hides uncertainty. The 95% CI of 0.87–1.33 around the primary efficacy HR shows that the estimated hazard ratio is not known precisely enough to reduce the result to the single value 1.07. The interval also crosses the null value of 1, which is central to its statistical interpretation.

Why is the P-value not an effect-size measure?

A P-value addresses compatibility with a null hypothesis under the specified statistical framework. It depends on the observed data, variability, and amount of information. The size and direction of the treatment effect are described by the hazard ratio and, where relevant, by absolute event rates or other absolute measures. NAVIGATE ESUS therefore should not be summarized by P-values alone.

What does intention-to-treat contribute?

The ITT analysis set includes all randomized participants. Analyzing according to randomized assignment preserves the treatment comparison created by randomization. It also means that the efficacy analysis remains a comparison of treatment strategies as assigned rather than becoming a comparison of selected participants who happened to remain on treatment.

Why does early termination matter statistically?

The registry states that the study was terminated early at the second interim analysis because there was no efficacy improvement over aspirin and very little chance of showing overall benefit if the study were completed. Early stopping changes the information available for estimation and means the final posted evidence must be understood in the context of the interim monitoring decision rather than as if the planned study had simply run to completion.

13. Understanding the Hazard Ratio Across Endpoints

NAVIGATE ESUS provides a useful demonstration of why hazard ratios should always be interpreted together with their endpoint definitions.

EndpointHR95% CIDirection of point estimate
Composite efficacy outcome1.070.87–1.33Above 1
ISTH major bleeding2.721.68–4.39Above 1
Cardiovascular death + recurrent stroke + systemic embolism + MI1.060.87–1.29Above 1
All-cause mortality1.260.87–1.81Above 1
Myocardial infarction0.740.39–1.38Below 1
Life-threatening bleeding2.341.28–4.29Above 1
Clinically relevant non-major bleeding1.511.13–2.00Above 1
Intracranial hemorrhage2.011.00–4.02Above 1

The table illustrates two important principles. First, an HR above 1 is not inherently good or bad; its meaning depends on whether the endpoint is an efficacy event or an adverse event. Second, the width and position of the confidence interval matter as much as the point estimate. An HR of 0.74 with a CI of 0.39–1.38 carries a very different level of statistical precision from an HR of 1.51 with a CI of 1.13–2.00.

14. What the Primary Efficacy Result Does — and Does Not — Mean

Statistical interpretation

The primary efficacy HR of 1.07 indicates an estimated 7% higher instantaneous hazard for the rivaroxaban group relative to acetylsalicylic acid under the stratified Cox model. The 95% CI of 0.87–1.33 includes 1, and the reported two-sided P-value is 0.51884.

This does not establish that rivaroxaban increases the absolute probability of the composite outcome by 7%, nor does it establish equivalence between the treatments. A nonsignificant superiority test is not itself a formal demonstration of equivalence or non-inferiority.

Why the confidence interval matters

The interval from 0.87 to 1.33 allows for treatment effects in either direction within the statistical uncertainty represented by the model. It is therefore more informative than the point estimate alone. It tells the reader that the observed estimate should not be treated as a precise measurement of a single underlying hazard ratio.

Why the major-bleeding result must remain separate

The primary major-bleeding HR of 2.72 has a 95% CI of 1.68–4.39. This is a different endpoint with a different clinical meaning from the composite efficacy outcome. The two estimates should therefore be reported side by side rather than mathematically combined into an invented net-benefit statistic.

15. Early Termination and Interim Analysis

The registry states that NAVIGATE ESUS was terminated early because the second interim analysis showed no efficacy improvement over aspirin and there was very little chance of showing overall benefit if the study were completed.

Why interim analysis matters

An interim analysis evaluates accumulating trial information before the originally contemplated end of follow-up. Once a study stops early, the resulting estimates reflect the amount of information available at that stopping point.

Why stopping for futility changes context

The registry's reason for termination was lack of efficacy improvement and very little chance of demonstrating overall benefit with continued enrollment or follow-up. The results therefore belong to a trial that did not continue to its originally intended completion.

The ClinicalTrials.gov record does not provide the numerical interim boundary, alpha-spending function, conditional-power threshold, or stopping boundary. Those quantities should not be reconstructed from the posted final estimates. What can be stated from the registry is the occurrence and reason for the second-interim termination.

16. Proportional-Hazards Considerations

The registry reports hazard ratios from stratified Cox proportional-hazards models. A Cox HR is a model-based summary of relative event hazard, and its simplest interpretation assumes that the relative hazards are reasonably represented by a proportional-hazards relationship over the analyzed period.

A useful conceptual distinction
HR ≠ relative risk ≠ risk difference ≠ probability of benefit

A hazard ratio compares instantaneous event hazards. It should not automatically be translated into a percentage difference in cumulative incidence or an absolute patient-level risk.

This distinction is especially relevant when interpreting endpoints observed over a median of 326 days. A hazard ratio summarizes the modeled event process over follow-up; it does not directly provide the probability that a particular participant will experience the endpoint by the efficacy cut-off date.

17. Confidence Intervals and Statistical Precision

The registry provides 95% two-sided confidence intervals for all twelve posted statistical analyses. Comparing their widths illustrates differences in statistical precision.

EndpointHR95% CIInterpretive feature
Composite efficacy1.070.87–1.33Includes the null and excludes neither direction broadly
ISTH major bleeding2.721.68–4.39Entire interval above 1
All-cause mortality1.260.87–1.81Wide interval including 1
Disabling stroke1.420.88–2.28Wide interval including 1
Cardiovascular death1.480.87–2.52Wide interval including 1
Myocardial infarction0.740.39–1.38Wide interval including 1
Intracranial hemorrhage2.011.00–4.02Lower limit reaches 1.00

A confidence interval should not be interpreted as the range in which the individual patient's treatment effect lies. It is an interval estimate for the population-level parameter under the statistical model and repeated-sampling framework. Wider intervals indicate less precision; narrower intervals indicate greater precision, all else equal.

18. Secondary Endpoint Interpretation

The secondary analyses should be interpreted in the context of their role as additional endpoints rather than as substitutes for the primary analysis. Several point estimates are above 1, while myocardial infarction has a point estimate below 1. This variation is expected when multiple related outcomes are examined.

Direction is not proof

A point estimate below 1 does not by itself establish a protective effect, just as a point estimate above 1 does not by itself establish increased risk.

Precision matters

The confidence interval determines how much uncertainty surrounds each estimate. Several secondary endpoints have intervals spanning substantial ranges on both sides of 1.

P-values need context

The registry supplies two-sided P-values for the analyses. They should be read alongside the endpoint definition, HR, confidence interval, and analysis population.

Different endpoints answer different questions

Stroke, mortality, myocardial infarction, and bleeding are not interchangeable outcomes. Their statistical estimates should remain tied to their specific definitions.

19. Limitations

20. What the Registry Does Not Establish From These Data

Several common statistical conclusions should not be inferred simply from the posted NAVIGATE ESUS results.

Potential overinterpretationWhy it is not justified by the ClinicalTrials.gov record
“HR 1.07 proves the treatments are equivalent.”The registry identifies the hypothesis type as superiority. A nonsignificant superiority comparison is not the same as a formal equivalence or non-inferiority demonstration.
“HR 1.07 means 7% more patients had an event.”The HR is a time-to-event measure, not an absolute event-rate difference.
“HR 2.72 means 172% of patients had major bleeding.”The HR describes relative instantaneous hazard, not a percentage of patients.
“P = 0.00002 is a measure of how large the bleeding effect was.”The P-value measures evidence against the specified null hypothesis; effect magnitude is described by the HR and confidence interval.
“The myocardial-infarction HR of 0.74 proves a 26% reduction.”The 95% CI is 0.39–1.38 and includes 1, so the point estimate alone does not establish a treatment effect.
“The early stopping result is identical to a completed-trial result.”The registry explicitly states that the trial was terminated early at the second interim analysis.

21. Why This Trial Matters Statistically

NAVIGATE ESUS is a useful teaching case because its registry results bring several core clinical-trial concepts together in one randomized comparison. The primary endpoints are both time-to-event outcomes, the analysis uses stratified log-rank testing and Cox modeling, the efficacy population is ITT, and the trial stopped early after an interim analysis.

ConceptHow it appears in NAVIGATE ESUS
RandomizationRandomized parallel-group phase 3 design
BlindingQuadruple masking
Intention-to-treatAll randomized participants included in the posted efficacy analysis set
Time-to-event endpointsBoth primary endpoints and all posted secondary analyses use time-to-event methodology
Hazard ratioPrimary and secondary effect measure
Confidence intervals95% two-sided intervals reported for all posted statistical analyses
Log-rank testStratified treatment comparison
Cox modelStratified Cox proportional-hazards model for risk-reduction estimation
Interim analysisSecond interim analysis led to early termination
Composite endpointPrimary adjudicated efficacy outcome combines stroke-related events, TIA with positive neuroimaging, and systemic embolism
Safety endpointISTH major bleeding was a primary endpoint; additional bleeding outcomes were secondary endpoints

22. A Statistical Reading of the Full Results

Viewed as a statistical profile rather than as a collection of isolated numbers, the registry record has three prominent features.

First, the primary efficacy comparison produced an HR of 1.07 with a 95% CI of 0.87–1.33 and P = 0.51884. The point estimate is close to the null value of 1, and the confidence interval spans both sides of the null. The result therefore does not provide strong statistical evidence of superiority for rivaroxaban on the registered composite efficacy outcome.

Second, the primary major-bleeding comparison produced an HR of 2.72 with a 95% CI of 1.68–4.39 and P = 0.00002. Unlike the efficacy estimate, the entire confidence interval lies above 1. The result represents a materially different statistical pattern, with the estimated hazard substantially above the comparator and the reported two-sided P-value indicating strong evidence against the null hypothesis.

Third, the secondary results show a mixture of point-estimate directions, but the confidence intervals for the reported efficacy outcomes include 1. The additional bleeding analyses have HRs above 1, with confidence intervals of 1.28–4.29 for life-threatening bleeding, 1.13–2.00 for clinically relevant non-major bleeding, and 1.00–4.02 for intracranial hemorrhage.

These observations should remain connected to the study's design. The trial was randomized and quadruple-masked, used ITT efficacy analyses, applied stratified time-to-event methods, and was terminated early after the second interim analysis. Each of those features contributes to the interpretation of the numerical results.

23. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue through the Clinical Biostats statistical pathway

Explore the core survival-analysis, clinical-trial, and statistical-inference methods represented in NAVIGATE ESUS.

26. Record Summary

NAVIGATE ESUS provides a clear example of a randomized clinical trial in which the principal statistical questions are time-to-event questions. The registry reports two primary endpoints: an adjudicated composite efficacy outcome and ISTH major bleeding. Both were analyzed using stratified log-rank testing and hazard ratios from stratified Cox proportional-hazards models in the intention-to-treat population.

The primary efficacy analysis produced an HR of 1.07 (95% CI 0.87–1.33; P = 0.51884), while the primary major-bleeding analysis produced an HR of 2.72 (95% CI 1.68–4.39; P = 0.00002). The secondary efficacy estimates likewise require interpretation through their confidence intervals rather than their point estimates alone, while the additional bleeding analyses produced HRs of 2.34 for life-threatening bleeding, 1.51 for clinically relevant non-major bleeding, and 2.01 for intracranial hemorrhage.

The statistical story is therefore not simply a matter of identifying which P-values are small. It involves understanding the randomized comparison, the ITT population, the adjudicated endpoint definitions, the time-to-event framework, the distinction between hazard ratios and absolute risks, the uncertainty represented by 95% confidence intervals, and the effect of early termination after the second interim analysis.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. For NAVIGATE ESUS, that means preserving the registry's endpoint definitions and numerical estimates while explaining what the hazard ratios, confidence intervals, P-values, analysis population, and early stopping decision do and do not establish.