← Clinical Trials
Stroke / Ischemic Attack Phase 3 Terminated NCT00029146

COSS: Complete Statistical Analysis of Extracranial-Intracranial Bypass Surgery in Carotid Occlusion

An independent statistical review of the randomized phase 3 Carotid Occlusion Surgery Study comparing extracranial-intracranial bypass surgery with best medical therapy in participants with stroke, transient ischemic attack, or cerebral infarction.

Trial start: 2002-07  ·  Primary completion: 2010-06  ·  Lead sponsor: University of North Carolina, Chapel Hill
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

COSS was a randomized, parallel-group, single-masked phase 3 treatment study comparing extracranial-intracranial bypass surgery with best medical therapy. The registered primary endpoint combined ipsilateral ischemic stroke in the first 2 years with stroke and death during the specified 30-day peri-procedural or post-randomization period.

700
Registry enrollment
ClinicalTrials.gov profile
2
Arms
Surgical vs non-surgical
1.7
Primary difference
Estimated 2-year rates
0.78
Primary P-value
2-sided z-statistic
FeatureCOSS
Trial nameCarotid Occlusion Surgery Study
NCT identifierNCT00029146
PhasePhase 3
StatusTerminated
Therapeutic areaNeurology
ConditionsStroke; Ischemic Attack, Transient; Cerebral Infarction
AllocationRandomized
Design modelParallel
MaskingSingle
Primary purposeTreatment
Registered primary endpoints1
Outcome measures posted12
Statistical analyses posted11
Lead sponsorUniversity of North Carolina, Chapel Hill

2. Clinical Question

The central statistical question was whether assignment to extracranial-intracranial bypass surgery produced a different 2-year rate for the registered composite primary outcome compared with assignment to best medical therapy.

Population

Participants in a phase 3 treatment trial involving stroke, transient ischemic attack, or cerebral infarction.

Intervention

Extracranial-intracranial bypass surgery.

Comparator

Best medical therapy.

Primary question

Does randomized assignment to surgery versus non-surgical management produce a different estimated 2-year primary-event rate?

3. Trial Design

COSS used randomized allocation and a parallel design. The statistical analyses were primarily based on the intention-to-treat principle, so participants were analyzed according to the group to which they were originally randomized.

01
Randomize 2 treatment groups
02
Treatment Surgery or best medical therapy
03
Follow 2-year outcome window
04
Estimate Product-limit rates
05
Compare 2-sided statistical test
SURGICAL GROUP

Extracranial-intracranial bypass surgery

  • Extracranial-intracranial bypass surgery was the intervention.
  • The primary endpoint included ipsilateral ischemic stroke in 2 years from randomization and all stroke and death through 30 days post-surgery.
  • The primary analysis retained participants according to their original randomized assignment.
NON-SURGICAL GROUP

Best medical therapy

  • Best medical therapy was the comparator.
  • The primary endpoint included ipsilateral ischemic stroke in 2 years from randomization and all stroke and death through 30 days post-randomization.
  • The primary analysis retained participants according to their original randomized assignment.
Important enrollment distinction: the ClinicalTrials.gov record lists enrollment as 700. The posted primary statistical analysis separately states that the study was terminated early for futility after 195 of the planned 372 participants were enrolled. These are registry fields describing different aspects of the record and should not be silently substituted for one another.

4. Trial Timing and Early Termination

2002-07

Trial start

The registered trial start was July 2002.

Early termination

Stopped for futility

The primary analysis notes state that the study was terminated early for futility after 195 of the planned 372 participants were enrolled. The difference in estimated rates divided by the standard error of that difference was compared to a standard unit normal distribution. Positive indicates lower rate in surgical group

2010-06

Primary completion

The registered primary completion date was June 2010.

Early termination is an important part of the statistical context. A trial stopped for futility generally contains less information than the originally planned experiment, so confidence intervals and the amount of observed information become central to interpretation. The registry's posted primary analysis explicitly identifies futility as the reason for early termination.

5. Primary Endpoint

EndpointRegistry definition / time frameEndpoint type
Primary endpoint Surgical Group: Ipsilateral Ischemic Stroke in 2 Yrs From Randomization and All Stroke & Death Through 30d Post-surgery; Non-surgical Group: Ipsilateral Ischemic Stroke in 2 Yrs From Randomization and All Stroke & Death Through 30d Post-randomization Binary
Time frame within 2 yrs of randomization Percentage of participants

The registry defines the primary endpoint using 2 yr Kaplan-Meier estimates of the proportions, with proportions expressed as percentages for reporting purposes. The registry definition also specifies that an ipsilateral ischemic stroke is the clinical diagnosis of a focal neurological deficit due to cerebral ischemia clinically localizable within the internal carotid artery territory distally to the symptomatic occluded internal carotid artery that lasts for more than 24 hours.

Why the time frame matters: this is not simply a raw proportion of participants with an event. The posted analysis used product-limit estimates of the 2-year rates and their standard errors. That approach is appropriate when follow-up times and censoring make a simple observed proportion inadequate.

6. Statistical Methodology

Product-limit estimation of 2-year rates

The primary analysis states that rates for each group were based on product limit estimates of 2-year rates and their standard errors. In survival analysis, the product-limit estimator is the Kaplan-Meier estimator. It uses the sequence of observed event and censoring times to estimate the probability of remaining event-free through a specified time.

Conceptual Kaplan-Meier form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at time ti, while ni is the number at risk immediately before that time.

Difference in estimated 2-year rates

The primary effect measure was the difference in estimated 2 yr rates. The posted analysis calculated the difference between the two product-limit estimates and divided that difference by the standard error of the difference.

Primary test statistic
z = (estimated ratesurgical − estimated ratenon-surgical) / SE(difference)

The resulting 2-sided z-statistic was compared with a standard unit normal distribution.

Intention-to-treat analysis

The primary efficacy analysis followed the intention-to-treat principle: all participants were analyzed in the group to which they were originally randomized. This preserves the treatment comparison created by randomization and avoids changing the primary comparison simply because participants did not receive or remain on their assigned treatment.

Fisher exact test

Several secondary binary outcomes were analyzed using Fisher exact testing. This is a direct categorical-data method that evaluates the treatment-group comparison without relying on the large-sample approximation used by a conventional chi-square test.

Two-sided t-test

The Summary SS-QOL Score was analyzed using a 2-sided t-test, with the reported effect measure being the difference in final values between groups. This is a fundamentally different estimand from the binary-event analyses: it compares a continuous outcome rather than an estimated event rate.

7. Primary Result

The posted primary analysis compared the surgical and non-surgical groups using a 2-sided z-statistic. The analysis population was intention-to-treat.

Difference in estimated 2-year primary-event rates

1.7 percentage points

95% CI: −10.4 to 13.8   ·   P = 0.78

Effect measure: difference in estimated 2 yr rates

Primary analysis elementReported result
Analysis populationIntention to treat principle; all participants analyzed in the group to which they were originally randomized
Groups comparedSurgical Group vs Non-surgical Group
Method2-sided z-statistic
Effect measureDifference in estimated 2 yr rates
Estimate1.7
95% confidence interval−10.4 to 13.8
P-value0.78
Hypothesis typeSuperiority
Clinical Biostats interpretation

The estimated difference was 1.7 percentage points, with the comparison defined as the surgical-group estimated 2-year rate minus the non-surgical-group estimated 2-year rate. The registry analysis therefore estimated a relatively small point difference, but the estimate is accompanied by substantial uncertainty.

The 95% CI of −10.4 to 13.8 is especially important. It spans zero and covers both a negative difference and a positive difference of appreciable magnitude. In statistical terms, the interval indicates that the data provide limited precision about the underlying difference in 2-year rates.

The P-value of 0.78 is a measure of compatibility between the observed test statistic and the null hypothesis under the specified testing framework. It is not a measure of the size of the treatment effect, the probability that the null hypothesis is true, or the probability that surgery is beneficial or harmful.

The analysis also depends on the product-limit estimates and their standard errors, so censoring and follow-up contribute to the statistical structure of the comparison. Finally, because the study was terminated early for futility, the amount of information available was smaller than the originally planned design.

Educational note: the registry provides the 2-year product-limit rate comparison and its confidence interval, but the ClinicalTrials.gov record does not provide the underlying participant-level event and censoring times. A Kaplan-Meier curve should therefore not be reconstructed from the summary estimate alone.

8. Secondary Endpoint Results

The registry contains multiple secondary analyses. Most use the same general framework as the primary endpoint: comparison of estimated 2-year rates between the randomized groups using a 2-sided z-statistic. The functional outcomes use Fisher exact testing, while the continuous Summary SS-QOL Score uses a 2-sided t-test.

Stroke and mortality outcomes

OutcomeEffect estimate95% CIP-valueMethod
All Stroke 3.5 −9.2 to 16.1 0.59 2-sided z-statistic
Disabling Stroke −3.2 −9.0 to 2.6 0.27 2-sided z-statistic
Fatal Stroke 1.3 −2.5 to 5.2 0.50 2-sided z-statistic
Death 4.0 −1.2 to 9.7 0.13 2-sided z-statistic

For these outcomes, the registry notes that positive differences indicate a lower rate in the surgical group for All Stroke, Fatal Stroke, and Death. For Disabling Stroke, the registry notes that a negative difference indicates a lower rate in the non-surgical group.

Functional outcomes and quality of life

OutcomeEffect estimate95% CIP-valueMethod
Modified Rankin 0-1 −6.6 −20.6 to 7.3 0.41 Fisher Exact
Modified Rankin 0-2 4.4 −8.2 to 16.9 0.70 Fisher Exact
Summary SS-QOL Score −0.24 −0.54 to 0.07 0.13 2-sided t-test
Modified Barthel Index 19-20 Not reported in registry-reported analysis 95% CI reported 0.85 Fisher's Exact Test

The Modified Rankin outcomes and Summary SS-QOL Score used a worst-case imputation for death and missing values. For the Modified Rankin outcomes, the registry specifically states that lower rates are worse because Rankin 0-1 and Rankin 0-2 represent good outcomes.

On-treatment analysis

Primary endpoint, on-treatment analysis

1.5 percentage points

95% CI: −10.7 to 13.7   ·   P = 0.81

2-sided z-statistic; difference in estimated 2 yr rates

This secondary analysis departed from the primary ITT framework. It removed four participants assigned to the surgical group who never underwent surgery and censored on the day of surgery three participants assigned to the nonsurgical group who underwent surgery. The purpose of such an analysis is different from the primary randomized comparison: it attempts to describe outcomes according to treatment actually received rather than preserving the original randomized assignment.

Why the on-treatment analysis is not a replacement for ITT

Removing or censoring participants according to treatment received can change the comparability created by randomization. That is why the primary analysis's ITT principle is important. The on-treatment result is useful as a secondary perspective, but it answers a different question and should not be substituted for the randomized comparison.

Post-hoc analysis

OutcomeEffect estimate95% CIP-valueRole
Any Stroke or Death 6.5 −6.5 to 19.6 .33 Post-hoc

The registry states that the Any Stroke or Death rates were based on product-limit estimates of 2-year rates and their standard errors. Positive indicates a lower rate in the surgical group. The reported P-value is reproduced exactly as .33.

9. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm as affected participants divided by participants at risk.

Safety measureSurgical GroupNon-surgical Group
Serious adverse events 14/97 2/98

These figures should be interpreted as the reported affected/at-risk counts. The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or P-value for this safety measure, so none is inferred here.

Safety and efficacy are different estimands. The primary endpoint is a time-framed clinical outcome analyzed under the randomized treatment comparison. Serious adverse events are exposure-related safety observations. Their numerical values should not be combined into a single benefit-risk statistic without a prespecified framework.

10. Missing Data and Imputation

Missing-data handling is explicitly reported for the Modified Rankin 0-1, Modified Rankin 0-2, Summary SS-QOL Score, and Modified Barthel Index 19-20 outcomes. The time frame for these outcomes is at 2 years after randomization or end of trial, with worst case imputed for death and missing values.

Worst-case imputation

For these outcomes, deaths and missing values were handled by assigning the worst case rather than simply excluding those participants.

Why this matters

Excluding participants with missing outcomes could change the composition of the analyzed population. Worst-case imputation instead incorporates those participants into the binary functional-outcome comparison.

This approach also makes the definition of the analysis explicit: a good functional outcome requires a participant to meet the specified threshold rather than merely have an observed assessment available at the analysis time.

11. Statistical Methods Explained

Why was a product-limit estimate used?

The primary endpoint was defined over a 2-year period, and the registry analysis used product-limit estimates rather than a simple arithmetic proportion. The Kaplan-Meier/product-limit approach allows participants to contribute follow-up information even when their complete 2-year outcome is not observed, provided their observation is handled through the survival-analysis framework.

Why compare the two estimated rates with a z-statistic?

The posted primary analysis calculated the difference between the estimated 2-year rates and divided it by the standard error of that difference. The resulting standardized statistic was compared with a standard unit normal distribution using a 2-sided test. This converts the estimated treatment difference into a scale on which statistical evidence against the null can be evaluated.

What does the primary estimate of 1.7 mean?

The value 1.7 is the reported difference in estimated 2-year rates between the surgical and non-surgical groups, expressed in percentage points. It is not a hazard ratio, relative risk, odds ratio, or percentage change.

Why is the confidence interval more informative than the P-value alone?

The P-value addresses statistical evidence against the specified null hypothesis, whereas the confidence interval describes the precision of the estimated difference under the stated statistical framework. Here, the 95% CI ranges from −10.4 to 13.8, showing that the point estimate of 1.7 is not highly precise.

Why was Fisher exact testing used for some outcomes?

Modified Rankin and Modified Barthel outcomes are binary categorical outcomes. Fisher exact testing provides an exact categorical comparison and is especially useful when sample sizes or cell counts may make large-sample approximations less attractive.

Why was a t-test used for Summary SS-QOL Score?

Summary SS-QOL Score is a continuous measure rather than a binary event. The posted analysis therefore used a 2-sided t-test and reported the difference in final values between the randomized groups.

Why is intention-to-treat analysis important here?

ITT preserves the original randomization. Once participants are selectively removed or reassigned according to treatment received, factors related to treatment adherence or treatment availability can affect group comparability. The primary analysis therefore keeps participants in their originally randomized groups.

12. Primary Result: What the Estimate Does — and Does Not — Mean

Effect estimate

The primary estimate of 1.7 percentage points describes the difference between the estimated 2-year rates in the surgical and non-surgical groups. It is an absolute difference in estimated rates, not a relative treatment effect.

Confidence interval

The 95% CI of −10.4 to 13.8 indicates substantial uncertainty around the estimated difference. It includes zero, so the observed point estimate should not be interpreted as establishing a directional treatment difference by itself.

P-value

The P-value of 0.78 quantifies the statistical evidence against the null under the posted 2-sided z-test. It does not measure the magnitude of the effect, the clinical importance of the effect, or the probability that either treatment is superior.

Censoring and follow-up

Because the endpoint was analyzed with product-limit estimates, the calculation incorporates the timing of events and censoring rather than treating every participant as if identical complete follow-up were available.

Early termination

The primary analysis notes that the study was terminated early for futility. This is an important limitation because the final statistical information available to the analysis was less than that of the planned design.

13. Secondary Results: How to Read the Direction of the Differences

The signs of the reported differences cannot be interpreted without reading the registry's directional convention for each outcome. This is particularly important because some endpoints represent adverse events, while Modified Rankin thresholds represent good functional outcomes.

OutcomeEstimateRegistry directional note
All Stroke 3.5 Positive indicates lower rate in surgical group
Disabling Stroke −3.2 Negative indicates lower rate in non-surgical group
Fatal Stroke 1.3 Positive indicates lower rate in surgical group
Death 4.0 Positive indicates lower rate in surgical group
Modified Rankin 0-1 −6.6 Negative indicates lower rate in non-surgical group; lower rate is worse because Rankin 0-1 indicates a good outcome
Modified Rankin 0-2 4.4 Positive indicates lower rate in surgical group; lower rate is worse because Rankin 0-2 indicates a good outcome
Summary SS-QOL Score −0.24 Negative indicates lower score in non-surgical group; higher score indicates better quality of life

This is a useful reminder that the sign of an effect estimate is not inherently "good" or "bad." Interpretation depends on the endpoint definition and the direction in which the outcome is clinically favorable.

14. Intention-to-Treat Versus On-Treatment Analysis

COSS provides a particularly clear statistical teaching example because both an intention-to-treat primary analysis and an on-treatment secondary analysis were posted.

Primary framework
Intention to treat. Participants remain in the group to which they were originally randomized.
Secondary framework
On-treatment. Four surgical-assigned participants who never underwent surgery were removed, and three nonsurgical-assigned participants who underwent surgery were censored on the day of surgery.

The primary ITT estimate was 1.7 with a 95% CI of −10.4 to 13.8 and P = 0.78. The on-treatment estimate was 1.5 with a 95% CI of −10.7 to 13.7 and P = 0.81.

The similarity of these two numerical estimates does not turn the analyses into interchangeable methods. They answer different statistical questions. ITT estimates the effect of assignment under the randomized design; an on-treatment analysis attempts to reflect treatment actually received but can alter the balance generated by randomization.

15. Early Futility and Statistical Information

The primary analysis explicitly states that the study was terminated early for futility after 195 of the planned 372 participants were enrolled. The difference in estimated rates divided by the standard error of that difference was compared to a standard unit normal distribution. Positive indicates lower rate in surgical group

What futility means statistically

A futility decision is a design-level conclusion that continuing the study was not expected to provide a useful path to demonstrating the prespecified objective under the trial's monitoring framework.

What it does not mean

Futility should not be translated automatically into proof that the treatments are identical. The confidence interval remains important because it describes the range of treatment differences compatible with the observed information.

The ClinicalTrials.gov record does not provide an interim-analysis boundary, alpha-spending scheme, conditional-power calculation, or detailed stopping-rule formula. Those features therefore are not reconstructed here.

16. No Unsupported Design Assumptions

The ClinicalTrials.gov record identifies randomization, a parallel design, single masking, a superiority hypothesis, an intention-to-treat primary analysis, and an early futility termination. They do not report a non-inferiority margin, factorial design, Bayesian method, or a formal multiplicity strategy.

Design topicWhat the ClinicalTrials.gov record supports
SuperiorityYes; the posted primary and secondary analyses specify a superiority hypothesis type.
Non-inferiority marginNot reported in the ClinicalTrials.gov record.
Factorial designNot reported; the design model is parallel.
Bayesian methodsNot reported.
Formal multiplicity strategyNot reported in the ClinicalTrials.gov record.
Interim/futility decisionEarly termination for futility is explicitly reported.
ITT analysisExplicitly reported for the primary analysis and most secondary analyses.

17. Why a P-value Alone Is Insufficient

The primary P-value of 0.78 is easy to quote but incomplete as a statistical description. The more informative presentation includes the effect estimate, its confidence interval, the analysis population, and the method used to obtain the estimate.

ComponentPrimary COSS resultWhat it contributes
Effect estimate1.7Magnitude and direction of the estimated rate difference
95% CI−10.4 to 13.8Precision and range of values compatible with the statistical framework
P-value0.78Evidence against the specified null hypothesis
Analysis populationIntention to treatPreserves original randomized assignment
Method2-sided z-statisticDefines how the primary comparison was statistically tested

A complete interpretation therefore does not ask only whether the P-value crossed a conventional threshold. It asks what was estimated, how precisely it was estimated, which participants contributed to the analysis, and how the endpoint was constructed.

18. Limitations

19. Why This Trial Matters Statistically

COSS is a useful statistical teaching case because the registry results combine time-to-event estimation, binary endpoint comparisons, continuous outcomes, intention-to-treat analysis, explicit missing-data handling, an on-treatment sensitivity analysis, and early termination for futility.

ConceptHow it appears in COSS
RandomizationParticipants were randomized to surgical and non-surgical groups.
Intention-to-treat analysisThe primary analysis retained participants in their originally randomized groups.
Kaplan-Meier / product-limit estimation2-year primary and secondary event rates were based on product-limit estimates.
Difference in ratesThe primary effect measure was the difference in estimated 2-year rates.
Confidence intervalThe primary estimate was accompanied by a 95% CI of −10.4 to 13.8.
Two-sided z-testThe primary analysis standardized the estimated difference by its standard error.
Fisher exact testUsed for Modified Rankin and Modified Barthel categorical outcomes.
t-testUsed for the continuous Summary SS-QOL Score.
Missing-data handlingWorst case was imputed for death and missing values for specified functional and quality-of-life outcomes.
On-treatment analysisA secondary analysis removed or censored participants according to treatment received.
Futility terminationThe primary analysis states that the study stopped early for futility.

20. Related Tutorials

Learn more about the methods used in this trial:

21. Related Calculators

22. Sources

Continue through Clinical Biostats

Use the related statistical tutorials and calculators to explore the methods that appear in randomized clinical-trial analyses.

23. Record Summary

COSS provides a compact example of how a randomized clinical trial can require several distinct statistical frameworks. Its primary endpoint used product-limit estimates of a 2-year rate and a 2-sided z-statistic to compare the randomized groups under the intention-to-treat principle. The reported primary difference was 1.7 percentage points, with a 95% CI of −10.4 to 13.8 and P = 0.78. Secondary analyses extended the statistical assessment to stroke, mortality, functional status, quality of life, and an on-treatment analysis, while the registry also reports a post-hoc Any Stroke or Death analysis. The study was terminated early for futility, making the distinction between the observed estimate, its uncertainty, and the amount of statistical information especially important.

Clinical Biostats methodology: A trial-results page should distinguish the reported statistical result from its interpretation. For COSS, the most informative reading combines the estimated treatment difference, confidence interval, P-value, analysis population, endpoint construction, missing-data rules, and early-termination context rather than relying on any single statistic.