This page separates reported trial results from statistical interpretation. Every numerical trial result presented here is taken from the ClinicalTrials.gov record. The registry provides 12 posted outcome measures and 12 statistical analyses, including two primary-endpoint analyses.
1. Trial at a Glance
APPRAISE-2 was a randomized, parallel, triple-masked phase 3 prevention trial in acute coronary syndrome. It compared apixaban with placebo and used time-to-event analyses for its two primary endpoints and the reported secondary endpoints.
| Feature | APPRAISE-2 |
|---|---|
| Trial | APPRAISE-2 |
| NCT identifier | NCT00831441 |
| Therapeutic area | Cardiovascular |
| Condition | Acute Coronary Syndrome |
| Phase | Phase 3 |
| Status | TERMINATED |
| Enrollment | 7484.0 |
| Allocation | RANDOMIZED |
| Design model | PARALLEL |
| Masking | TRIPLE |
| Primary purpose | PREVENTION |
| Interventions | Apixaban (drug); Placebo (drug) |
| Lead sponsor | Bristol-Myers Squibb |
| Sponsor type | INDUSTRY |
| Start | 2009-03 |
| Primary completion | 2011-03 |
2. Clinical Question
The statistical question was whether assignment to apixaban, compared with placebo, was associated with a different rate of the registered cardiovascular composite outcome during the intended treatment period, while also evaluating confirmed major bleeding during the treatment period.
Population
Participants enrolled in the phase 3 APPRAISE-2 trial with the condition recorded as acute coronary syndrome.
Intervention
Apixaban. The ClinicalTrials.gov record identifies the apixaban arm as Apixaban 5 mg BID.
Comparator
Placebo. The ClinicalTrials.gov record identifies the comparator as Placebo BID.
Primary questions
How did the apixaban and placebo groups compare for the cardiovascular death, myocardial infarction, or ischemic stroke endpoint and for TIMI major bleeding?
3. Trial Design
Apixaban
- Drug intervention: apixaban
- Safety data identify the arm as Apixaban 5 mg BID
- Included in the randomized parallel comparison
- Primary efficacy analysis used all randomized participants
Placebo
- Drug intervention: placebo
- Safety data identify the arm as Placebo BID
- Included in the randomized parallel comparison
- Primary efficacy analysis used all randomized participants
The trial was randomized, parallel, and triple-masked. These design features matter statistically because randomization establishes the treatment comparison before outcomes occur, the parallel structure means participants remain in their assigned treatment groups rather than forming sequential treatment periods, and masking is intended to reduce the influence of treatment knowledge on trial conduct and outcome assessment.
4. Trial Timing and Status
Trial start
The ClinicalTrials.gov record lists 2009-03 as the trial start.
Primary completion
The ClinicalTrials.gov record lists 2011-03 as the primary completion date.
Trial status
The current ClinicalTrials.gov record in the ClinicalTrials.gov record lists the study status as TERMINATED.
The primary efficacy time frame extends from randomization (Day 1) to first cardiovascular event, up to March 2011, approximately 2 years. The primary bleeding endpoint instead begins at first dose and is analyzed during the treatment period.
5. Primary Endpoints
| Endpoint | Registry time frame | Analysis population | Statistical method |
|---|---|---|---|
| Event Rate of Cardiovascular Death, Myocardial Infarction, or Ischemic Stroke During the Intended Treatment Period - Randomized Participants | Randomization (Day 1) to first event (CV death, MI, ischemic stroke), up to March 2011, approximately 2 years | All randomized participants | Cox proportional-hazards model; HR |
| Event Rate of Confirmed Major Bleeding Using TIMI Criteria During the Treatment Period - Treated Participants | From first dose to first occurrence of event (TIMI major bleeding) during Treatment Period (first dose to last dose + 2 days), up to March 2011, approximately 2 years | All participants who received at least one dose of blinded study drug and signed informed consent | Cox proportional-hazards model; HR |
Primary efficacy endpoint
The first primary endpoint was the event rate of cardiovascular death, myocardial infarction, or ischemic stroke during the intended treatment period. The registry definition describes the event rate as the percentage of participants with an event per 100 patient-years, with only events confirmed by the adjudication committee included in the analyses.
Primary safety endpoint
The second primary endpoint was confirmed major bleeding according to Thrombolysis in Myocardial Infarction (TIMI) criteria during the treatment period. The registry-reported definition includes fatal bleeding, intracranial hemorrhage, and clinically overt bleeding with a hemoglobin drop of ≥ 5 grams per deciliter (g/dL), or a ≥15% absolute decrease in hematocrit, with hemoglobin and hematocrit measurements adjusted for transfusions according to the registry definition.
6. Statistical Methodology
Time-to-event analysis
The registry classifies the posted analyses as time-to-event analyses and reports the Cox proportional-hazards model as the statistical method for both primary endpoints and all ten listed secondary analyses.
The Cox model describes the hazard at time t relative to a baseline hazard, with the treatment effect represented through the coefficient β. Exponentiating the treatment coefficient gives a hazard ratio.
Hazard ratio
The hazard ratio is the effect measure reported for every posted statistical analysis in the ClinicalTrials.gov record. An HR of 1 represents equality of the modeled event hazards between the compared groups. Values below 1 indicate a lower modeled hazard for the group represented in the numerator of the HR; values above 1 indicate a higher modeled hazard for that group.
The hazard ratio is a relative time-to-event measure. It is not an absolute risk difference, an event-rate difference, or the probability that a particular participant will experience the event.
Analysis populations
For the primary cardiovascular endpoint, all randomized participants were analyzed. For the primary TIMI major-bleeding endpoint, all participants who received at least one dose of blinded study drug and signed informed consent were analyzed.
One-sided testing
The primary efficacy analysis explicitly states that a test of superiority at the one-sided α = 0.025 significance level was performed. This is distinct from the reported two-sided 95% confidence interval. The confidence interval and the hypothesis test therefore have different stated conventions and should not be treated as though they were interchangeable quantities.
7. Primary Results: Cardiovascular Death, MI, or Ischemic Stroke
The first primary analysis compared the event rate of cardiovascular death, myocardial infarction, or ischemic stroke from randomization through first event, up to March 2011, approximately 2 years. All randomized participants were included.
Reported hazard ratio
95% CI: 0.80–1.11 · P = 0.5094
Cox proportional-hazards model · Superiority hypothesis · One-sided α = 0.025 efficacy test
| Feature | Reported result |
|---|---|
| Groups compared | Placebo vs Apixaban 5 mg BID |
| Analysis population | All randomized participants |
| Endpoint | Cardiovascular death, myocardial infarction, or ischemic stroke |
| Estimate | HR 0.95 |
| 95% confidence interval | 0.80–1.11 |
| P-value | 0.5094 |
| Hypothesis | Superiority |
What the estimate means: the registry reports a hazard ratio of 0.95 for the Placebo vs Apixaban 5 mg BID comparison. Because the ClinicalTrials.gov record does not explicitly state which treatment is coded as the numerator of the reported HR, this page does not translate 0.95 into a directional percentage reduction or increase. The important statistical feature is that the point estimate is close to 1.
What it does not mean: an HR of 0.95 does not mean that 95% of participants avoided the composite endpoint, nor does it mean that an individual patient's risk was reduced by 5%. A hazard ratio is a relative model-based time-to-event measure.
Confidence interval: the 95% CI of 0.80–1.11 spans 1.00. Thus, the interval contains both values below and above the equality value. It expresses uncertainty around the estimated relative hazard; it is not a range containing the individual treatment effects experienced by patients.
P-value: the reported 0.5094 is evidence from the specified hypothesis test, not a measure of effect size. A p-value does not tell us whether an effect is clinically large or small. Here, it should also be read in light of the prespecified one-sided α = 0.025 superiority framework.
Censoring and model assumptions: because this is a time-to-event analysis, participants who do not experience the event during available follow-up can contribute information up to their censoring time. The Cox interpretation also depends on the proportional-hazards framework being reasonably appropriate for the treatment comparison.
8. Primary Results: Confirmed TIMI Major Bleeding
The second primary analysis evaluated confirmed major bleeding during the treatment period, beginning with the first dose and continuing to the first occurrence of TIMI major bleeding within the registry-defined treatment period.
Reported hazard ratio
95% CI: 1.50–4.46 · P = 0.0006
Cox proportional-hazards model · Superiority hypothesis
| Feature | Reported result |
|---|---|
| Groups compared | Placebo vs Apixaban 5 mg BID |
| Analysis population | All participants who received at least one dose of blinded study drug and signed informed consent |
| Endpoint | Confirmed major bleeding using TIMI criteria |
| Estimate | HR 2.59 |
| 95% confidence interval | 1.50–4.46 |
| P-value | 0.0006 |
| Hypothesis | Superiority |
What the estimate means: the reported hazard ratio is 2.59, substantially above 1.00. The direction indicates that the group represented in the numerator of the reported HR had a higher estimated instantaneous hazard of the bleeding endpoint than the reference group.
What it does not mean: an HR of 2.59 does not mean that 259% of participants experienced major bleeding, nor does it mean that every participant had 2.59 times the probability of bleeding. The hazard ratio is a relative time-to-event measure and should not be substituted for an absolute event probability.
Confidence interval: the 95% CI of 1.50–4.46 lies entirely above 1.00. The interval indicates uncertainty around the estimated relative hazard while maintaining the direction of the estimated association throughout the interval.
P-value: the reported 0.0006 is the p-value for the specified equality-of-rates hypothesis. It does not quantify the magnitude of the bleeding effect. The magnitude is conveyed by the hazard ratio and its confidence interval.
Analysis-population caution: this endpoint was not analyzed in exactly the same population as the cardiovascular efficacy endpoint. It was restricted to participants who received at least one dose and signed informed consent. Comparing the two HRs therefore requires attention to their different populations and time origins.
9. Primary Results Side by Side
| Primary endpoint | HR | 95% CI | P-value | Population |
|---|---|---|---|---|
| Cardiovascular death, MI, or ischemic stroke | 0.95 | 0.80–1.11 | 0.5094 | All randomized participants |
| Confirmed TIMI major bleeding | 2.59 | 1.50–4.46 | 0.0006 | Participants receiving at least one dose and signing informed consent |
The two primary analyses illustrate why a clinical trial cannot be summarized responsibly by a single p-value. They address different outcomes, begin at different time origins, use different analysis populations, and have different statistical results. The cardiovascular composite has an HR close to 1 with a confidence interval spanning 1, whereas the reported major-bleeding HR is above 1 with a confidence interval that remains above 1.
10. Secondary Endpoint Results
The registry provides ten secondary statistical analyses. Each uses a Cox proportional-hazards model and reports a hazard ratio, two-sided 95% confidence interval, and p-value.
| Secondary endpoint | HR | 95% CI | P-value |
|---|---|---|---|
| Unstable angina during the intended treatment period | 0.94 | 0.70–1.26 | 0.6702 |
| Stroke during the intended treatment period | 0.90 | 0.57–1.40 | 0.6311 |
| Myocardial infarction during the intended treatment period | 0.93 | 0.76–1.14 | 0.5086 |
| Stent thrombosis during the intended treatment period | 0.73 | 0.47–1.12 | 0.1502 |
| Composite of cardiovascular death, MI, unstable angina, or ischemic stroke | 0.94 | 0.82–1.09 | 0.4317 |
| Composite of cardiovascular death, fatal bleed, MI, or stroke | 0.98 | 0.83–1.15 | 0.8015 |
| Composite of all-cause death, MI, or stroke | 0.99 | 0.85–1.15 | 0.8948 |
| Confirmed major bleeding using ISTH criteria | 2.48 | 1.72–3.58 | <0.0001 |
| Confirmed major bleeding or clinically relevant non-major bleeding using ISTH criteria | 2.64 | 1.87–3.72 | <0.0001 |
| All bleeding reported by the investigator | 2.36 | 2.06–2.70 | <0.0001 |
Several of the efficacy-oriented secondary HRs are below 1, including unstable angina, stroke, myocardial infarction, and stent thrombosis. Their confidence intervals all include 1.00. The three bleeding-related secondary analyses have HRs above 1 and confidence intervals that do not include 1.00.
11. Secondary Efficacy Endpoints
Unstable angina
Reported hazard ratio
95% CI: 0.70–1.26 · P = 0.6702
The analysis used all randomized participants and assessed time from randomization to the first event of unstable angina, up to March 2011, approximately 2 years.
Stroke
Reported hazard ratio
95% CI: 0.57–1.40 · P = 0.6311
The stroke endpoint was analyzed from randomization to first stroke in all randomized participants.
Myocardial infarction
Reported hazard ratio
95% CI: 0.76–1.14 · P = 0.5086
The MI endpoint was analyzed from randomization to first myocardial infarction in all randomized participants.
Stent thrombosis
Reported hazard ratio
95% CI: 0.47–1.12 · P = 0.1502
The point estimate is below 1.00, but the 95% confidence interval extends from below to above 1.00. The ClinicalTrials.gov record therefore illustrates why a point estimate should not be interpreted without its uncertainty interval.
12. Secondary Composite Endpoints
| Composite | HR | 95% CI | P-value |
|---|---|---|---|
| Cardiovascular death, MI, unstable angina, or ischemic stroke | 0.94 | 0.82–1.09 | 0.4317 |
| Cardiovascular death, fatal bleed, MI, or stroke | 0.98 | 0.83–1.15 | 0.8015 |
| All-cause death, MI, or stroke | 0.99 | 0.85–1.15 | 0.8948 |
Composite endpoints combine multiple component events into a single time-to-first-event outcome. Statistically, this can increase the number of observed events, but interpretation depends on what the components represent and how frequently each component occurs. The ClinicalTrials.gov record provides the composite definitions and overall HRs but do not provide component-specific event counts, so this page does not infer which component drove any composite estimate.
13. Secondary Bleeding Endpoints
| Bleeding endpoint | HR | 95% CI | P-value |
|---|---|---|---|
| Confirmed major bleeding using ISTH criteria | 2.48 | 1.72–3.58 | <0.0001 |
| Confirmed major bleeding or clinically relevant non-major bleeding using ISTH criteria | 2.64 | 1.87–3.72 | <0.0001 |
| All bleeding reported by the investigator | 2.36 | 2.06–2.70 | <0.0001 |
These three endpoints demonstrate how changing the endpoint definition can change the estimated treatment effect. The ISTH major-bleeding definition is narrower than a combined major-or-CRNM endpoint, while investigator-reported all bleeding is broader still. The HRs should therefore be compared only with attention to their distinct endpoint definitions.
The three bleeding HRs are all above 1.00: 2.48, 2.64, and 2.36. Their corresponding 95% confidence intervals—1.72–3.58, 1.87–3.72, and 2.06–2.70—remain above 1.00. The registry p-values are reported as <0.0001 for each.
This consistency across related bleeding definitions is descriptively notable, but the endpoints are not statistically independent. They overlap conceptually and are evaluated within a larger family of trial outcomes. The ClinicalTrials.gov record does not specify an adjustment procedure that would justify treating each nominal p-value as a separate confirmatory test.
14. Safety Results
The ClinicalTrials.gov record provides serious adverse events by treatment arm. These are reported as affected participants divided by participants at risk.
| Safety measure | Apixaban 5 mg BID | Placebo BID |
|---|---|---|
| Serious adverse events | 894 / 3672 | 884 / 3643 |
The serious-adverse-event data should be distinguished from the formal bleeding endpoint analyses. Serious adverse events are a broad safety category, whereas the primary bleeding endpoint is a specifically defined time-to-first-event outcome based on TIMI criteria. The affected/at-risk figures should therefore not be substituted for the hazard ratio analyses.
15. Statistical Methods Explained
Why was a Cox proportional-hazards model used?
The registered primary and secondary endpoints are time-to-event outcomes: participants are followed from a defined starting point until the first occurrence of an event or censoring. A Cox proportional-hazards model is designed for this setting because it estimates a relative hazard while accommodating different follow-up times and right censoring.
What does an HR of 0.95 mean?
An HR of 0.95 is a relative estimate close to the equality value of 1.00. It does not mean that the treatment changes every patient's probability by 5%. The interpretation of direction also requires knowing which group is represented in the numerator of the reported hazard ratio. The ClinicalTrials.gov record lists the groups compared as Placebo vs Apixaban 5 mg BID but does not explicitly document the coding convention, so the numerical estimate should not be given an unsupported directional translation.
Why does the 95% CI matter?
The confidence interval communicates the statistical precision of the estimated hazard ratio. For the primary cardiovascular endpoint, the interval is 0.80–1.11, which includes 1.00. For the primary bleeding endpoint, it is 1.50–4.46, which does not include 1.00. The intervals therefore provide information that the point estimates alone cannot provide.
Why doesn't the p-value measure effect size?
A p-value measures how compatible the observed data are with a specified null hypothesis under the chosen testing framework. It is affected by the amount of information in the analysis and does not directly quantify the magnitude or clinical importance of an effect. Effect size is better represented by the hazard ratio together with an absolute measure when one is available.
Why does the cardiovascular endpoint use all randomized participants?
The registry specifies that all randomized participants were analyzed for the primary cardiovascular endpoint. This anchors the efficacy comparison to treatment assignment rather than to treatment received. Such an approach preserves the randomized comparison and avoids defining the primary efficacy population based on post-randomization treatment exposure.
Why is the bleeding endpoint analyzed in treated participants?
The primary bleeding endpoint uses a different analysis population: participants who received at least one dose of blinded study drug and signed informed consent. This reflects the fact that the endpoint is defined during the treatment period beginning at first dose. It also means that the efficacy and bleeding estimates are not based on identical populations.
What is the importance of one-sided testing?
The primary efficacy analysis specifies a superiority test at one-sided α = 0.025. A one-sided test allocates the type I error to a prespecified direction. This matters because the decision threshold must be interpreted according to the stated testing convention rather than by applying an unrelated two-sided rule after seeing the data.
16. Reading the Primary Hazard Ratios Together
| Primary endpoint | Point estimate | Interval relative to 1.00 | Statistical signal |
|---|---|---|---|
| CV death, MI, or ischemic stroke | 0.95 | 0.80–1.11 includes 1.00 | P = 0.5094 |
| TIMI major bleeding | 2.59 | 1.50–4.46 remains above 1.00 | P = 0.0006 |
These results illustrate an important principle in clinical-trial statistics: efficacy and safety endpoints must be analyzed separately. A treatment effect is not summarized by asking whether "the trial was positive" in the abstract. Instead, each prespecified endpoint has its own estimand, analysis population, time origin, event definition, effect estimate, uncertainty interval, and hypothesis test.
The cardiovascular composite has a reported HR close to 1.00 and a confidence interval that crosses 1.00. The bleeding endpoint has a reported HR above 1.00 and a confidence interval entirely above 1.00. These are different statistical findings concerning different outcomes.
It would be inappropriate to combine the two HRs mathematically or describe one as an offsetting percentage of the other. Hazard ratios for different endpoints are not commensurate measures that can be subtracted or averaged into a single benefit-risk score.
17. Endpoint Definitions and Time Origins
| Endpoint family | Starting point | Event | Population |
|---|---|---|---|
| Primary cardiovascular composite | Randomization (Day 1) | First cardiovascular death, MI, or ischemic stroke | All randomized participants |
| Primary TIMI bleeding | First dose | First TIMI major bleeding event | Participants receiving at least one dose and signing informed consent |
| Secondary efficacy endpoints | Randomization (Day 1) | Endpoint-specific first event | All randomized participants |
| Secondary bleeding endpoints | First dose | Endpoint-specific first bleeding event | Participants receiving at least one dose and signing informed consent |
The time origin is not a technical footnote. Changing the time origin changes which portion of follow-up contributes to the analysis. Randomization is appropriate for the randomized efficacy comparison, while the bleeding analyses explicitly begin at first dose because their registered treatment-period definitions are exposure-based.
18. What the Confidence Intervals Tell Us
Primary efficacy
HR 0.95 with 95% CI 0.80–1.11. The interval includes the equality value of 1.00, so the estimate is compatible with relative hazards on either side of equality within the interval.
Primary bleeding
HR 2.59 with 95% CI 1.50–4.46. The interval remains above 1.00, providing a substantially different statistical pattern from the primary cardiovascular endpoint.
Secondary efficacy
Each listed secondary efficacy confidence interval includes 1.00, including the interval for stent thrombosis of 0.47–1.12.
Secondary bleeding
The three reported bleeding confidence intervals remain above 1.00: 1.72–3.58, 1.87–3.72, and 2.06–2.70.
A confidence interval is fundamentally an uncertainty statement. It does not describe how much individual patients vary, and it does not tell us that every value inside the interval is equally probable. It summarizes uncertainty in the estimated parameter under the statistical model and sampling framework.
19. P-values and Statistical Significance
The APPRAISE-2 registry results provide p-values for all twelve posted statistical analyses. The primary efficacy endpoint reports 0.5094, while the primary TIMI major-bleeding endpoint reports 0.0006. The secondary efficacy p-values are 0.6702, 0.6311, 0.5086, 0.1502, 0.4317, 0.8015, and 0.8948. The three secondary bleeding endpoints each report <0.0001.
The hazard ratio describes the estimated relative effect; the confidence interval describes uncertainty around that estimate; the p-value addresses a specified null hypothesis under the chosen testing framework.
This separation is particularly important when several endpoints are tested. A small p-value does not automatically establish clinical importance, and a larger p-value does not prove that two treatments are identical. The confidence interval and the prespecified estimand provide essential context.
20. Multiplicity and Multiple Endpoints
The ClinicalTrials.gov record identifies 2 primary endpoints and 12 outcome measures posted, with 12 statistical analyses posted. The two primary endpoints address distinct dimensions of the trial: a cardiovascular composite and confirmed major bleeding.
| Multiplicity issue | APPRAISE-2 information reported |
|---|---|
| Primary endpoints | 2 |
| Posted outcome measures | 12 |
| Posted statistical analyses | 12 |
| Primary analyses | 2 |
| One-sided efficacy alpha | 0.025 |
| Secondary analyses | 10 |
Multiple endpoints create a multiplicity problem because the more statistical questions are examined, the greater the opportunity for chance findings if every test is interpreted independently at a fixed threshold. The ClinicalTrials.gov record identifies the one-sided α = 0.025 superiority framework for the primary efficacy outcome but do not provide enough information to reconstruct a complete multiplicity-adjustment strategy across all posted outcomes.
21. Randomization and Causal Interpretation
Randomization is the structural feature that makes the treatment comparison fundamentally different from an observational association. Participants were randomized to the two parallel study arms, allowing treatment assignment to be established before the analyzed outcomes occurred.
Before outcomes
Randomization determines treatment assignment and protects the comparison from systematic allocation based on observed or unobserved participant characteristics, subject to the usual properties of randomized trials.
After randomization
Events, treatment exposure, censoring, and follow-up can differ. The analysis strategy determines how those post-randomization observations contribute to the estimated treatment effect.
The primary cardiovascular analysis uses all randomized participants, which keeps the principal efficacy comparison anchored to randomization. The bleeding endpoint uses a treated analysis population, reflecting its treatment-period definition. These choices should be retained when interpreting the corresponding HRs.
22. Cox Proportional-Hazards Assumption
The Cox model is semiparametric with respect to the baseline hazard, but the conventional hazard-ratio interpretation relies on a proportional-hazards framework: the relative hazard between treatment groups is assumed to be reasonably stable over the relevant time scale.
When proportional hazards are a reasonable approximation, one HR can summarize the relative hazard over follow-up. If hazards vary substantially over time, a single HR may conceal important temporal patterns.
The ClinicalTrials.gov record reports Cox proportional-hazards models but do not provide the underlying hazard functions or a proportional-hazards diagnostic. Accordingly, this page does not claim that the assumption was formally demonstrated; it simply explains the assumption attached to the reported model.
23. Kaplan-Meier Estimation and Censoring
The learning pathway for APPRAISE-2 includes Kaplan-Meier estimation, although the ClinicalTrials.gov record identifies Cox proportional-hazards models as the reported method. Kaplan-Meier estimation is the natural descriptive companion to a time-to-event analysis because it estimates the event-free survival function over time.
Here, di is the number of events at time ti and ni is the number at risk immediately before that time.
Censoring allows participants who have not experienced the endpoint during their available follow-up to contribute information until their censoring time. The ClinicalTrials.gov record does not include Kaplan-Meier event curves or median event times, so none are added here.
24. Why the Two Primary Endpoints Should Not Be Collapsed
The cardiovascular endpoint and TIMI major bleeding endpoint represent different clinical event domains. One cannot be converted into the other by comparing their HRs. A hazard ratio of 0.95 for one endpoint and 2.59 for another are estimates on different outcome scales.
Efficacy endpoint
Composite of cardiovascular death, myocardial infarction, or ischemic stroke, measured from randomization.
Bleeding endpoint
Confirmed TIMI major bleeding, measured during the treatment period from first dose.
A statistically literate interpretation therefore keeps the endpoints separate. The trial's evidence consists of multiple endpoint-specific estimates rather than one universal treatment-effect number.
25. Important Limitations and Interpretation Issues
- Early termination: the primary cardiovascular endpoint definition states that the study was terminated early and that the last patient, last visit was in Year 2. Early termination can affect the amount of accumulated information and therefore the precision of estimates.
- Different analysis populations: the cardiovascular primary endpoint uses all randomized participants, while the primary bleeding endpoint uses participants who received at least one dose and signed informed consent.
- Different time origins: cardiovascular outcomes begin at randomization; treatment-period bleeding outcomes begin at first dose.
- Composite endpoint: the primary cardiovascular endpoint combines cardiovascular death, myocardial infarction, and ischemic stroke. The ClinicalTrials.gov record does not provide component-specific event counts.
- Hazard-ratio assumptions: Cox HRs rely on a proportional-hazards framework. The ClinicalTrials.gov record does not provide a formal diagnostic of that assumption.
- Multiplicity: the registry contains two primary endpoints and ten secondary statistical analyses. The ClinicalTrials.gov record does not provide enough detail to reconstruct a complete familywise-error strategy across every posted endpoint.
- Nominal p-values: p-values are hypothesis-test quantities and should not be interpreted as measures of effect size.
- No absolute efficacy event rates reported: the statistical-analysis records provide HRs, confidence intervals, and p-values, but not the arm-specific efficacy event counts needed for a fuller absolute-risk presentation.
- No median event times reported: the ClinicalTrials.gov record does not provide median survival or median time-to-event estimates for the listed endpoints.
- No subgroup analyses reported: the trial data do not provide subgroup-specific treatment estimates, so no subgroup conclusions are added.
26. Why This Trial Matters Statistically
APPRAISE-2 is a useful teaching case because it demonstrates how a randomized cardiovascular trial can generate very different statistical results across efficacy and safety endpoints while using the same fundamental time-to-event modeling framework.
| Concept | How it appears in APPRAISE-2 |
|---|---|
| Randomization | Participants were randomized to apixaban or placebo in a parallel phase 3 design. |
| Blinding | The registry identifies the masking as TRIPLE. |
| Time-to-event endpoints | The primary and secondary analyses use event timing from defined starting points. |
| Cox model | Reported statistical method for all 12 posted analyses. |
| Hazard ratio | Effect measure for the two primary and ten secondary analyses. |
| Confidence interval | Two-sided 95% CIs are reported for every statistical analysis reported. |
| One-sided testing | The primary efficacy analysis specifies a one-sided α = 0.025 superiority test. |
| Analysis populations | Randomized participants for cardiovascular efficacy; treated participants for bleeding. |
| Multiplicity | Two primary endpoints plus ten secondary statistical analyses. |
| Safety analysis | Serious adverse events are reported by treatment arm, while bleeding endpoints receive formal time-to-event analyses. |
The statistical lesson is broader than the individual numerical results. A trial's interpretation depends on matching each estimate to its endpoint definition, time origin, analysis population, model, confidence interval, and hypothesis-testing framework. APPRAISE-2 provides a compact example of that discipline.
27. Statistical Interpretation of the Secondary Results
The secondary efficacy HRs range from 0.73 for stent thrombosis to 0.99 for the all-cause death, MI, or stroke composite. The corresponding 95% confidence intervals all include 1.00. This means the registry-reported point estimates should be interpreted together with substantial uncertainty rather than as standalone evidence of treatment differences.
The secondary bleeding HRs are 2.48, 2.64, and 2.36. Each has a two-sided 95% confidence interval entirely above 1.00, and each has a reported p-value of <0.0001. The consistent direction across these related definitions provides a coherent descriptive pattern within the registry analyses.
The secondary results reinforce the importance of defining the estimand before interpreting a statistical result. "Bleeding" is not one endpoint in this dataset: major bleeding, major or CRNM bleeding, and all investigator-reported bleeding are separately defined outcomes. Their estimates therefore answer related but different questions.
28. What the Primary Efficacy Result Does — and Does Not — Mean
The reported primary cardiovascular HR is 0.95. This is a relative estimate from a Cox proportional-hazards model for the composite of cardiovascular death, myocardial infarction, or ischemic stroke.
The 95% CI is 0.80–1.11. It includes 1.00, so the uncertainty interval encompasses the equality value for the relative hazard.
The p-value is 0.5094. This is a hypothesis-test result under the specified statistical framework. It is not a probability that the null hypothesis is true and is not a measure of the clinical size of the treatment effect.
The HR should not be converted into an individual-level probability or absolute risk without additional information. The ClinicalTrials.gov record does not contain the arm-specific event counts or Kaplan-Meier estimates needed to make that conversion responsibly.
29. What the Primary Bleeding Result Does — and Does Not — Mean
The reported primary bleeding HR is 2.59, with a two-sided 95% CI of 1.50–4.46. The estimate is above the equality value of 1.00.
The entire reported interval remains above 1.00. This means the statistical uncertainty interval does not include equality under the reported model and confidence-interval framework.
The reported p-value is 0.0006. It provides evidence against the equality-of-rates null hypothesis under the specified analysis; it does not say that the effect is exactly 2.59 in every population or for every individual.
The bleeding endpoint was analyzed among participants who received at least one dose and signed informed consent. It therefore should not be described as though it used exactly the same analysis population as the randomized cardiovascular endpoint.
30. Related Tutorials
Learn more about the methods used in this trial:
31. Related Calculators
32. Sources
- ClinicalTrials.gov: APPRAISE-2, NCT00831441.
- Linked publication: PubMed record 31310855.
- Linked publication: PubMed record 29910052.
- Linked publication: PubMed record 29447769.
- Linked publication: PubMed record 28438739.
- Linked publication: PubMed record 26271059.
Continue through the Clinical Biostats statistical library
Use the trial's endpoints and analysis methods as a practical route into survival analysis, hazard ratios, confidence intervals, hypothesis testing, and clinical-trial methodology.
33. Record Summary
APPRAISE-2 provides a clear example of endpoint-specific clinical-trial statistics. The trial was a randomized, parallel, triple-masked phase 3 study in acute coronary syndrome with 7484.0 participants and two treatment arms. Its primary cardiovascular endpoint was analyzed from randomization using all randomized participants, while the primary TIMI major-bleeding endpoint was analyzed from first dose among participants who received at least one dose and signed informed consent.
The statistical analyses posted on ClinicalTrials.gov consistently use Cox proportional-hazards models and hazard ratios. The primary cardiovascular endpoint produced an HR of 0.95 with a 95% CI of 0.80–1.11 and P = 0.5094. The primary TIMI major-bleeding endpoint produced an HR of 2.59 with a 95% CI of 1.50–4.46 and P = 0.0006. The secondary analyses provide additional endpoint-specific estimates, including efficacy HRs of 0.94, 0.90, 0.93, 0.73, 0.94, 0.98, and 0.99, and bleeding HRs of 2.48, 2.64, and 2.36.
The most important statistical lesson is that these results should not be collapsed into a single treatment verdict. A rigorous interpretation preserves the distinction between endpoint definition, time origin, analysis population, hazard ratio, confidence interval, and hypothesis test. That framework allows the reported evidence to be understood without adding unsupported assumptions or treating a p-value as a substitute for the effect estimate.