This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
COSS was a randomized, parallel-group, single-masked phase 3 treatment study comparing extracranial-intracranial bypass surgery with best medical therapy. The registered primary endpoint combined ipsilateral ischemic stroke in the first 2 years with stroke and death during the specified 30-day peri-procedural or post-randomization period.
| Feature | COSS |
|---|---|
| Trial name | Carotid Occlusion Surgery Study |
| NCT identifier | NCT00029146 |
| Phase | Phase 3 |
| Status | Terminated |
| Therapeutic area | Neurology |
| Conditions | Stroke; Ischemic Attack, Transient; Cerebral Infarction |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Single |
| Primary purpose | Treatment |
| Registered primary endpoints | 1 |
| Outcome measures posted | 12 |
| Statistical analyses posted | 11 |
| Lead sponsor | University of North Carolina, Chapel Hill |
2. Clinical Question
The central statistical question was whether assignment to extracranial-intracranial bypass surgery produced a different 2-year rate for the registered composite primary outcome compared with assignment to best medical therapy.
Population
Participants in a phase 3 treatment trial involving stroke, transient ischemic attack, or cerebral infarction.
Intervention
Extracranial-intracranial bypass surgery.
Comparator
Best medical therapy.
Primary question
Does randomized assignment to surgery versus non-surgical management produce a different estimated 2-year primary-event rate?
3. Trial Design
COSS used randomized allocation and a parallel design. The statistical analyses were primarily based on the intention-to-treat principle, so participants were analyzed according to the group to which they were originally randomized.
Extracranial-intracranial bypass surgery
- Extracranial-intracranial bypass surgery was the intervention.
- The primary endpoint included ipsilateral ischemic stroke in 2 years from randomization and all stroke and death through 30 days post-surgery.
- The primary analysis retained participants according to their original randomized assignment.
Best medical therapy
- Best medical therapy was the comparator.
- The primary endpoint included ipsilateral ischemic stroke in 2 years from randomization and all stroke and death through 30 days post-randomization.
- The primary analysis retained participants according to their original randomized assignment.
4. Trial Timing and Early Termination
Trial start
The registered trial start was July 2002.
Stopped for futility
The primary analysis notes state that the study was terminated early for futility after 195 of the planned 372 participants were enrolled. The difference in estimated rates divided by the standard error of that difference was compared to a standard unit normal distribution. Positive indicates lower rate in surgical group
Primary completion
The registered primary completion date was June 2010.
Early termination is an important part of the statistical context. A trial stopped for futility generally contains less information than the originally planned experiment, so confidence intervals and the amount of observed information become central to interpretation. The registry's posted primary analysis explicitly identifies futility as the reason for early termination.
5. Primary Endpoint
| Endpoint | Registry definition / time frame | Endpoint type |
|---|---|---|
| Primary endpoint | Surgical Group: Ipsilateral Ischemic Stroke in 2 Yrs From Randomization and All Stroke & Death Through 30d Post-surgery; Non-surgical Group: Ipsilateral Ischemic Stroke in 2 Yrs From Randomization and All Stroke & Death Through 30d Post-randomization | Binary |
| Time frame | within 2 yrs of randomization | Percentage of participants |
The registry defines the primary endpoint using 2 yr Kaplan-Meier estimates of the proportions, with proportions expressed as percentages for reporting purposes. The registry definition also specifies that an ipsilateral ischemic stroke is the clinical diagnosis of a focal neurological deficit due to cerebral ischemia clinically localizable within the internal carotid artery territory distally to the symptomatic occluded internal carotid artery that lasts for more than 24 hours.
6. Statistical Methodology
Product-limit estimation of 2-year rates
The primary analysis states that rates for each group were based on product limit estimates of 2-year rates and their standard errors. In survival analysis, the product-limit estimator is the Kaplan-Meier estimator. It uses the sequence of observed event and censoring times to estimate the probability of remaining event-free through a specified time.
Here, di represents the number of events at time ti, while ni is the number at risk immediately before that time.
Difference in estimated 2-year rates
The primary effect measure was the difference in estimated 2 yr rates. The posted analysis calculated the difference between the two product-limit estimates and divided that difference by the standard error of the difference.
The resulting 2-sided z-statistic was compared with a standard unit normal distribution.
Intention-to-treat analysis
The primary efficacy analysis followed the intention-to-treat principle: all participants were analyzed in the group to which they were originally randomized. This preserves the treatment comparison created by randomization and avoids changing the primary comparison simply because participants did not receive or remain on their assigned treatment.
Fisher exact test
Several secondary binary outcomes were analyzed using Fisher exact testing. This is a direct categorical-data method that evaluates the treatment-group comparison without relying on the large-sample approximation used by a conventional chi-square test.
Two-sided t-test
The Summary SS-QOL Score was analyzed using a 2-sided t-test, with the reported effect measure being the difference in final values between groups. This is a fundamentally different estimand from the binary-event analyses: it compares a continuous outcome rather than an estimated event rate.
7. Primary Result
The posted primary analysis compared the surgical and non-surgical groups using a 2-sided z-statistic. The analysis population was intention-to-treat.
Difference in estimated 2-year primary-event rates
95% CI: −10.4 to 13.8 · P = 0.78
Effect measure: difference in estimated 2 yr rates
| Primary analysis element | Reported result |
|---|---|
| Analysis population | Intention to treat principle; all participants analyzed in the group to which they were originally randomized |
| Groups compared | Surgical Group vs Non-surgical Group |
| Method | 2-sided z-statistic |
| Effect measure | Difference in estimated 2 yr rates |
| Estimate | 1.7 |
| 95% confidence interval | −10.4 to 13.8 |
| P-value | 0.78 |
| Hypothesis type | Superiority |
The estimated difference was 1.7 percentage points, with the comparison defined as the surgical-group estimated 2-year rate minus the non-surgical-group estimated 2-year rate. The registry analysis therefore estimated a relatively small point difference, but the estimate is accompanied by substantial uncertainty.
The 95% CI of −10.4 to 13.8 is especially important. It spans zero and covers both a negative difference and a positive difference of appreciable magnitude. In statistical terms, the interval indicates that the data provide limited precision about the underlying difference in 2-year rates.
The P-value of 0.78 is a measure of compatibility between the observed test statistic and the null hypothesis under the specified testing framework. It is not a measure of the size of the treatment effect, the probability that the null hypothesis is true, or the probability that surgery is beneficial or harmful.
The analysis also depends on the product-limit estimates and their standard errors, so censoring and follow-up contribute to the statistical structure of the comparison. Finally, because the study was terminated early for futility, the amount of information available was smaller than the originally planned design.
8. Secondary Endpoint Results
The registry contains multiple secondary analyses. Most use the same general framework as the primary endpoint: comparison of estimated 2-year rates between the randomized groups using a 2-sided z-statistic. The functional outcomes use Fisher exact testing, while the continuous Summary SS-QOL Score uses a 2-sided t-test.
Stroke and mortality outcomes
| Outcome | Effect estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| All Stroke | 3.5 | −9.2 to 16.1 | 0.59 | 2-sided z-statistic |
| Disabling Stroke | −3.2 | −9.0 to 2.6 | 0.27 | 2-sided z-statistic |
| Fatal Stroke | 1.3 | −2.5 to 5.2 | 0.50 | 2-sided z-statistic |
| Death | 4.0 | −1.2 to 9.7 | 0.13 | 2-sided z-statistic |
For these outcomes, the registry notes that positive differences indicate a lower rate in the surgical group for All Stroke, Fatal Stroke, and Death. For Disabling Stroke, the registry notes that a negative difference indicates a lower rate in the non-surgical group.
Functional outcomes and quality of life
| Outcome | Effect estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| Modified Rankin 0-1 | −6.6 | −20.6 to 7.3 | 0.41 | Fisher Exact |
| Modified Rankin 0-2 | 4.4 | −8.2 to 16.9 | 0.70 | Fisher Exact |
| Summary SS-QOL Score | −0.24 | −0.54 to 0.07 | 0.13 | 2-sided t-test |
| Modified Barthel Index 19-20 | Not reported in registry-reported analysis | 95% CI reported | 0.85 | Fisher's Exact Test |
The Modified Rankin outcomes and Summary SS-QOL Score used a worst-case imputation for death and missing values. For the Modified Rankin outcomes, the registry specifically states that lower rates are worse because Rankin 0-1 and Rankin 0-2 represent good outcomes.
On-treatment analysis
Primary endpoint, on-treatment analysis
95% CI: −10.7 to 13.7 · P = 0.81
2-sided z-statistic; difference in estimated 2 yr rates
This secondary analysis departed from the primary ITT framework. It removed four participants assigned to the surgical group who never underwent surgery and censored on the day of surgery three participants assigned to the nonsurgical group who underwent surgery. The purpose of such an analysis is different from the primary randomized comparison: it attempts to describe outcomes according to treatment actually received rather than preserving the original randomized assignment.
Removing or censoring participants according to treatment received can change the comparability created by randomization. That is why the primary analysis's ITT principle is important. The on-treatment result is useful as a secondary perspective, but it answers a different question and should not be substituted for the randomized comparison.
Post-hoc analysis
| Outcome | Effect estimate | 95% CI | P-value | Role |
|---|---|---|---|---|
| Any Stroke or Death | 6.5 | −6.5 to 19.6 | .33 | Post-hoc |
The registry states that the Any Stroke or Death rates were based on product-limit estimates of 2-year rates and their standard errors. Positive indicates a lower rate in the surgical group. The reported P-value is reproduced exactly as .33.
9. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm as affected participants divided by participants at risk.
| Safety measure | Surgical Group | Non-surgical Group |
|---|---|---|
| Serious adverse events | 14/97 | 2/98 |
These figures should be interpreted as the reported affected/at-risk counts. The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or P-value for this safety measure, so none is inferred here.
10. Missing Data and Imputation
Missing-data handling is explicitly reported for the Modified Rankin 0-1, Modified Rankin 0-2, Summary SS-QOL Score, and Modified Barthel Index 19-20 outcomes. The time frame for these outcomes is at 2 years after randomization or end of trial, with worst case imputed for death and missing values.
Worst-case imputation
For these outcomes, deaths and missing values were handled by assigning the worst case rather than simply excluding those participants.
Why this matters
Excluding participants with missing outcomes could change the composition of the analyzed population. Worst-case imputation instead incorporates those participants into the binary functional-outcome comparison.
This approach also makes the definition of the analysis explicit: a good functional outcome requires a participant to meet the specified threshold rather than merely have an observed assessment available at the analysis time.
11. Statistical Methods Explained
Why was a product-limit estimate used?
The primary endpoint was defined over a 2-year period, and the registry analysis used product-limit estimates rather than a simple arithmetic proportion. The Kaplan-Meier/product-limit approach allows participants to contribute follow-up information even when their complete 2-year outcome is not observed, provided their observation is handled through the survival-analysis framework.
Why compare the two estimated rates with a z-statistic?
The posted primary analysis calculated the difference between the estimated 2-year rates and divided it by the standard error of that difference. The resulting standardized statistic was compared with a standard unit normal distribution using a 2-sided test. This converts the estimated treatment difference into a scale on which statistical evidence against the null can be evaluated.
What does the primary estimate of 1.7 mean?
The value 1.7 is the reported difference in estimated 2-year rates between the surgical and non-surgical groups, expressed in percentage points. It is not a hazard ratio, relative risk, odds ratio, or percentage change.
Why is the confidence interval more informative than the P-value alone?
The P-value addresses statistical evidence against the specified null hypothesis, whereas the confidence interval describes the precision of the estimated difference under the stated statistical framework. Here, the 95% CI ranges from −10.4 to 13.8, showing that the point estimate of 1.7 is not highly precise.
Why was Fisher exact testing used for some outcomes?
Modified Rankin and Modified Barthel outcomes are binary categorical outcomes. Fisher exact testing provides an exact categorical comparison and is especially useful when sample sizes or cell counts may make large-sample approximations less attractive.
Why was a t-test used for Summary SS-QOL Score?
Summary SS-QOL Score is a continuous measure rather than a binary event. The posted analysis therefore used a 2-sided t-test and reported the difference in final values between the randomized groups.
Why is intention-to-treat analysis important here?
ITT preserves the original randomization. Once participants are selectively removed or reassigned according to treatment received, factors related to treatment adherence or treatment availability can affect group comparability. The primary analysis therefore keeps participants in their originally randomized groups.
12. Primary Result: What the Estimate Does — and Does Not — Mean
The primary estimate of 1.7 percentage points describes the difference between the estimated 2-year rates in the surgical and non-surgical groups. It is an absolute difference in estimated rates, not a relative treatment effect.
The 95% CI of −10.4 to 13.8 indicates substantial uncertainty around the estimated difference. It includes zero, so the observed point estimate should not be interpreted as establishing a directional treatment difference by itself.
The P-value of 0.78 quantifies the statistical evidence against the null under the posted 2-sided z-test. It does not measure the magnitude of the effect, the clinical importance of the effect, or the probability that either treatment is superior.
Because the endpoint was analyzed with product-limit estimates, the calculation incorporates the timing of events and censoring rather than treating every participant as if identical complete follow-up were available.
The primary analysis notes that the study was terminated early for futility. This is an important limitation because the final statistical information available to the analysis was less than that of the planned design.
13. Secondary Results: How to Read the Direction of the Differences
The signs of the reported differences cannot be interpreted without reading the registry's directional convention for each outcome. This is particularly important because some endpoints represent adverse events, while Modified Rankin thresholds represent good functional outcomes.
| Outcome | Estimate | Registry directional note |
|---|---|---|
| All Stroke | 3.5 | Positive indicates lower rate in surgical group |
| Disabling Stroke | −3.2 | Negative indicates lower rate in non-surgical group |
| Fatal Stroke | 1.3 | Positive indicates lower rate in surgical group |
| Death | 4.0 | Positive indicates lower rate in surgical group |
| Modified Rankin 0-1 | −6.6 | Negative indicates lower rate in non-surgical group; lower rate is worse because Rankin 0-1 indicates a good outcome |
| Modified Rankin 0-2 | 4.4 | Positive indicates lower rate in surgical group; lower rate is worse because Rankin 0-2 indicates a good outcome |
| Summary SS-QOL Score | −0.24 | Negative indicates lower score in non-surgical group; higher score indicates better quality of life |
This is a useful reminder that the sign of an effect estimate is not inherently "good" or "bad." Interpretation depends on the endpoint definition and the direction in which the outcome is clinically favorable.
14. Intention-to-Treat Versus On-Treatment Analysis
COSS provides a particularly clear statistical teaching example because both an intention-to-treat primary analysis and an on-treatment secondary analysis were posted.
The primary ITT estimate was 1.7 with a 95% CI of −10.4 to 13.8 and P = 0.78. The on-treatment estimate was 1.5 with a 95% CI of −10.7 to 13.7 and P = 0.81.
The similarity of these two numerical estimates does not turn the analyses into interchangeable methods. They answer different statistical questions. ITT estimates the effect of assignment under the randomized design; an on-treatment analysis attempts to reflect treatment actually received but can alter the balance generated by randomization.
15. Early Futility and Statistical Information
The primary analysis explicitly states that the study was terminated early for futility after 195 of the planned 372 participants were enrolled. The difference in estimated rates divided by the standard error of that difference was compared to a standard unit normal distribution. Positive indicates lower rate in surgical group
What futility means statistically
A futility decision is a design-level conclusion that continuing the study was not expected to provide a useful path to demonstrating the prespecified objective under the trial's monitoring framework.
What it does not mean
Futility should not be translated automatically into proof that the treatments are identical. The confidence interval remains important because it describes the range of treatment differences compatible with the observed information.
The ClinicalTrials.gov record does not provide an interim-analysis boundary, alpha-spending scheme, conditional-power calculation, or detailed stopping-rule formula. Those features therefore are not reconstructed here.
16. No Unsupported Design Assumptions
The ClinicalTrials.gov record identifies randomization, a parallel design, single masking, a superiority hypothesis, an intention-to-treat primary analysis, and an early futility termination. They do not report a non-inferiority margin, factorial design, Bayesian method, or a formal multiplicity strategy.
| Design topic | What the ClinicalTrials.gov record supports |
|---|---|
| Superiority | Yes; the posted primary and secondary analyses specify a superiority hypothesis type. |
| Non-inferiority margin | Not reported in the ClinicalTrials.gov record. |
| Factorial design | Not reported; the design model is parallel. |
| Bayesian methods | Not reported. |
| Formal multiplicity strategy | Not reported in the ClinicalTrials.gov record. |
| Interim/futility decision | Early termination for futility is explicitly reported. |
| ITT analysis | Explicitly reported for the primary analysis and most secondary analyses. |
17. Why a P-value Alone Is Insufficient
The primary P-value of 0.78 is easy to quote but incomplete as a statistical description. The more informative presentation includes the effect estimate, its confidence interval, the analysis population, and the method used to obtain the estimate.
| Component | Primary COSS result | What it contributes |
|---|---|---|
| Effect estimate | 1.7 | Magnitude and direction of the estimated rate difference |
| 95% CI | −10.4 to 13.8 | Precision and range of values compatible with the statistical framework |
| P-value | 0.78 | Evidence against the specified null hypothesis |
| Analysis population | Intention to treat | Preserves original randomized assignment |
| Method | 2-sided z-statistic | Defines how the primary comparison was statistically tested |
A complete interpretation therefore does not ask only whether the P-value crossed a conventional threshold. It asks what was estimated, how precisely it was estimated, which participants contributed to the analysis, and how the endpoint was constructed.
18. Limitations
- Early termination: the primary analysis states that the study stopped for futility after 195 of the planned 372 participants were enrolled, reducing the information available relative to the planned design.
- Registry enrollment discrepancy: the ClinicalTrials.gov record lists enrollment as 700, while the primary analysis note describes early termination after 195 of a planned 372. These registry fields should be reported as given rather than reconciled by inference.
- Wide uncertainty: the primary 95% CI extends from −10.4 to 13.8, demonstrating substantial uncertainty around the estimated difference.
- Different estimands: the primary ITT analysis and secondary on-treatment analysis address different questions.
- Missing-data assumptions: functional and quality-of-life outcomes used worst-case imputation for death and missing values, which is an explicit analytical assumption rather than an observation directly measured for every participant.
- Endpoint direction: the meaning of a positive or negative difference varies by outcome, so numerical signs should not be interpreted without the endpoint's clinical direction.
- Safety denominator: the registry-reported serious-adverse-event data are reported only as affected/at-risk counts; no comparative inferential analysis is provided.
- Limited methodological detail: the ClinicalTrials.gov record does not specify a non-inferiority margin, Bayesian method, factorial structure, or formal multiplicity strategy.
19. Why This Trial Matters Statistically
COSS is a useful statistical teaching case because the registry results combine time-to-event estimation, binary endpoint comparisons, continuous outcomes, intention-to-treat analysis, explicit missing-data handling, an on-treatment sensitivity analysis, and early termination for futility.
| Concept | How it appears in COSS |
|---|---|
| Randomization | Participants were randomized to surgical and non-surgical groups. |
| Intention-to-treat analysis | The primary analysis retained participants in their originally randomized groups. |
| Kaplan-Meier / product-limit estimation | 2-year primary and secondary event rates were based on product-limit estimates. |
| Difference in rates | The primary effect measure was the difference in estimated 2-year rates. |
| Confidence interval | The primary estimate was accompanied by a 95% CI of −10.4 to 13.8. |
| Two-sided z-test | The primary analysis standardized the estimated difference by its standard error. |
| Fisher exact test | Used for Modified Rankin and Modified Barthel categorical outcomes. |
| t-test | Used for the continuous Summary SS-QOL Score. |
| Missing-data handling | Worst case was imputed for death and missing values for specified functional and quality-of-life outcomes. |
| On-treatment analysis | A secondary analysis removed or censored participants according to treatment received. |
| Futility termination | The primary analysis states that the study stopped early for futility. |
20. Related Tutorials
Learn more about the methods used in this trial:
21. Related Calculators
22. Sources
- ClinicalTrials.gov: NCT00029146 — Carotid Occlusion Surgery Study.
- Linked publication: PubMed record for PMID 23101451.
- Linked publication: PubMed record for PMID 22068990.
Continue through Clinical Biostats
Use the related statistical tutorials and calculators to explore the methods that appear in randomized clinical-trial analyses.
23. Record Summary
COSS provides a compact example of how a randomized clinical trial can require several distinct statistical frameworks. Its primary endpoint used product-limit estimates of a 2-year rate and a 2-sided z-statistic to compare the randomized groups under the intention-to-treat principle. The reported primary difference was 1.7 percentage points, with a 95% CI of −10.4 to 13.8 and P = 0.78. Secondary analyses extended the statistical assessment to stroke, mortality, functional status, quality of life, and an on-treatment analysis, while the registry also reports a post-hoc Any Stroke or Death analysis. The study was terminated early for futility, making the distinction between the observed estimate, its uncertainty, and the amount of statistical information especially important.