This page separates reported trial results from statistical interpretation. Numerical results, endpoint definitions, analysis populations, and statistical methods on this page are restricted to the trial data posted on ClinicalTrials.gov for CARE-MS I.
1. Trial at a Glance
CARE-MS I was a randomized, parallel, single-masked phase 3 treatment trial comparing alemtuzumab with interferon beta-1a in relapsing-remitting multiple sclerosis. The registry reports 581 participants, two study arms, two primary endpoints, and six posted statistical analyses.
| Feature | CARE-MS I |
|---|---|
| Trial name | CARE-MS I |
| Brief title | Comparison of Alemtuzumab and Rebif® Efficacy in Multiple Sclerosis, Study One |
| Phase | Phase 3 |
| Condition | Multiple Sclerosis, Relapsing-Remitting |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Single |
| Primary purpose | Treatment |
| Enrollment | 581.0 |
| Interventions | Alemtuzumab; Interferon beta-1a |
| Primary endpoints | 2 |
| Outcome measures posted | 6 |
| Statistical analyses posted | 6 |
| Lead sponsor | Genzyme, a Sanofi Company |
| Sponsor type | Industry |
| Trial status | Completed |
| Study start | 2007-08 |
| Primary completion | 2011-04 |
| ClinicalTrials.gov | NCT00530348 |
2. Clinical Question
The trial evaluates two related efficacy questions in relapsing-remitting multiple sclerosis: whether treatment assignment is associated with the occurrence of sustained accumulation of disability over up to 2 years, and whether it is associated with the annualized rate of relapse over up to 2 years.
Population
Participants with Multiple Sclerosis, Relapsing-Remitting, as specified in the registry record.
Intervention
Alemtuzumab, classified in the registry as a biological intervention.
Comparator
Interferon beta-1a, classified in the registry as a biological intervention.
Primary questions
How do alemtuzumab and interferon beta-1a compare with respect to sustained accumulation of disability and annualized relapse rate through up to 2 years?
3. Trial Design
Alemtuzumab
- Biological intervention.
- Compared with interferon beta-1a.
- Primary efficacy analyses used the FAS population.
Interferon beta-1a
- Biological intervention.
- Compared with alemtuzumab.
- Primary efficacy analyses used the FAS population.
4. Endpoints
The registry lists two primary endpoints. Both were evaluated over up to 2 years, but they require different statistical frameworks because one is a time-to-event disability endpoint and the other is a recurrent-event rate endpoint.
| Endpoint | Time frame | Endpoint type | Primary effect measure |
|---|---|---|---|
| Percentage of Participants With Sustained Accumulation of Disability (SAD) | Up to 2 years | Binary / time-to-event analysis | Hazard ratio |
| Annualized Relapse Rate | Up to 2 years | Count / rate | Rate ratio |
Sustained Accumulation of Disability
The registry describes EDSS as an ordinal scale in half-point increments that quantifies disability in participants with MS. It assesses 7 functional systems—visual, brainstem, pyramidal, cerebellar, sensory, bowel/bladder, and cerebral—as well as ambulation. The EDSS total score ranges from 0, representing a normal neurological examination, to 10, representing death due to MS.
Annualized Relapse Rate
The registry defines a relapse as new neurological symptoms or worsening of previous neurological symptoms with an objective change on neurological examination, attributable to multiple sclerosis, lasting for at least 48 hours, present at normal body temperature, and preceded by at least 30 days of clinical stability.
The primary endpoint is the Annualized Relapse Rate over up to 2 years. The statistical analysis posted for the endpoint reports a rate ratio estimated using proportional means regression with robust variance estimation and adjustment for geographic region.
5. Statistical Analysis Populations
| Population | Definition reported in trial data | Role |
|---|---|---|
| FAS | All randomized participants who received at least 1 dose of study drug. | Primary and secondary efficacy analyses |
| EDSS analysis subset | Subset of the FAS population with EDSS assessment at both baseline and end. | Change from baseline in EDSS |
| MSFC analysis subset | Subset of the FAS population with MSFC score assessment at baseline; the registry-reported analysis text describes the Year 2 change analysis. | Change from baseline in MSFC |
| MRI-T2 analysis subset | Subset of the FAS population with T2 volume assessment at both baseline and Year 2. | Percent change in MRI-T2 hyperintense lesion volume |
The use of the FAS population means that the principal efficacy comparisons were not restricted to participants who completed the full planned observation period. However, the ClinicalTrials.gov record also show that some continuous or imaging outcomes were evaluated in subsets of the FAS defined by availability of the relevant assessments.
6. Statistical Methodology
Cox proportional-hazards regression for sustained accumulation of disability
The primary SAD analysis used a Cox proportional-hazards regression model with robust variance estimation. Treatment group and geographic region were included as covariates.
The hazard ratio for treatment is obtained by exponentiating the treatment coefficient. The model compares the instantaneous event rate over follow-up rather than simply comparing the percentages observed at one fixed time point.
Proportional means regression for annualized relapse rate
The annualized relapse-rate analysis used a proportional means regression model with robust variance estimation and covariate adjustment for geographic region. The reported effect measure was a rate ratio.
A rate ratio below 1 indicates a lower estimated relapse rate in the numerator group under the comparison specified in the registry analysis. It is a comparison of rates, not a percentage of participants who relapsed.
Wei-Lachin analysis for repeated measures
The EDSS and MSFC change analyses were reported as using the Wei-Lachin method for non-parametric analysis of repeated measures. The registry reports the Wei-Lachin method.
This distinction matters. The registry method field should be reported as the actual method used in the analysis record rather than silently replacing it with a generic mixed-model label. The underlying statistical problem is longitudinal: measurements from the same participant are related, so treating every repeated observation as independent would generally be inappropriate.
Ranked ANCOVA for MRI-T2 lesion volume
The MRI-T2 endpoint used ranked ANCOVA with covariate adjustment for geographic region and baseline T2 lesion volume. The registry reports this analysis as a ranked ANCOVA.
ANCOVA combines a treatment-group comparison with adjustment for one or more baseline covariates. Here, baseline T2 lesion volume is particularly relevant because the endpoint is a change or percent change in lesion volume and baseline disease burden can influence the follow-up measurement.
Sequential analysis of secondary endpoints
The analysis notes state that secondary endpoints were analyzed sequentially as: proportion of participants relapse free at Year 2, change from baseline in EDSS, percent change from baseline in MRI-T2 hyperintense lesion volume at Year 2, and acquisition of disability measured by MSFC. The ClinicalTrials.gov record does not specify a formal multiplicity-adjustment procedure for this sequence, so no additional claim about familywise error control is made here.
7. Primary Results: Sustained Accumulation of Disability
The primary SAD endpoint was analyzed in the FAS population using Cox proportional-hazards regression. The comparison was between interferon beta-1a and alemtuzumab, with treatment group and geographic region as covariates and robust variance estimation.
Hazard ratio for sustained accumulation of disability
95% CI: 0.40–1.23 · P = 0.2173
Time frame: up to 2 years · FAS population
| Primary endpoint | Analysis | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Percentage of Participants With Sustained Accumulation of Disability (SAD) | Cox proportional-hazards regression | Hazard ratio | 0.70 | 0.40–1.23 | 0.2173 |
The estimated hazard ratio of 0.70 indicates that, under the fitted Cox model, the estimated instantaneous hazard of sustained accumulation of disability for the treatment comparison was 0.70 times the comparator hazard. Expressed as a simple relative interpretation, an HR of 0.70 corresponds to a 30% lower estimated hazard, because 1 − 0.70 = 0.30.
This does not mean that 30% fewer participants developed SAD, and it does not mean that every participant experienced a 30% reduction in personal risk. The hazard ratio is a model-based time-to-event measure.
The 95% confidence interval, 0.40–1.23, is relatively broad and crosses 1.00. It therefore includes values compatible with a lower hazard as well as values compatible with a higher hazard under the statistical model. The P-value of 0.2173 is evidence about compatibility with the tested statistical hypothesis; it is not a measure of the size or clinical importance of the estimated effect.
Because this was a Cox model, interpretation also depends on the proportional-hazards framework. The ClinicalTrials.gov record does not report a formal diagnostic for that assumption, so the HR should be understood as the reported model-based summary rather than as a guarantee that the hazard ratio was constant at every point in time.
8. Primary Results: Annualized Relapse Rate
The second primary endpoint was annualized relapse rate over up to 2 years. The FAS population was analyzed using proportional means regression with robust variance estimation and adjustment for geographic region.
Rate ratio for annualized relapse rate
95% CI: 0.32–0.63 · P < 0.0001
Time frame: up to 2 years · FAS population
| Primary endpoint | Analysis | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Annualized Relapse Rate | Proportional means regression with robust variance estimation | Rate ratio | 0.45 | 0.32–0.63 | <0.0001 |
The rate ratio of 0.45 indicates that the estimated annualized relapse rate under the treatment comparison was 0.45 times the comparator rate. As a simple relative interpretation, this corresponds to a 55% lower estimated relapse rate because 1 − 0.45 = 0.55.
A rate ratio is not the same as a risk ratio. It does not say that 55% of participants avoided relapse, nor does it provide the absolute number of relapses experienced by an individual participant.
The 95% confidence interval of 0.32–0.63 describes uncertainty around the estimated rate ratio. Every value in the interval is below 1.00, indicating that the estimated rate remains below the comparator rate across the reported confidence interval.
The P-value of <0.0001 indicates strong statistical evidence against the null hypothesis specified for the comparison, but the P-value itself does not quantify the size of the treatment effect. The effect size is communicated by the rate ratio and its confidence interval.
The registry reports robust variance estimation and geographic-region adjustment. The ClinicalTrials.gov record does not provide the underlying participant-level relapse counts, exposure times, or model diagnostics, so the page does not attempt to reconstruct an absolute relapse-rate table.
9. Secondary Results: Relapse-Free at Year 2
The first listed secondary endpoint was the Percentage of Participants Who Were Relapse Free at Year 2. It was analyzed in the FAS population using Cox proportional-hazards regression with robust variance estimation and covariate adjustment for geographic region.
Hazard ratio for relapse-free analysis
95% CI: 0.33–0.61 · P < 0.0001
Time frame: Year 2 · FAS population
| Secondary endpoint | Method | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Percentage of Participants Who Were Relapse Free at Year 2 | Cox proportional-hazards regression | HR 0.45 | 0.33–0.61 | <0.0001 |
The analysis note states that the secondary endpoints were analyzed sequentially, beginning with the proportion of participants relapse free at Year 2. The ClinicalTrials.gov record does not provide the corresponding arm-specific percentage values, so the hazard ratio is the principal numerical result reported here.
10. Secondary Results: EDSS Change
The endpoint Change From Baseline in Expanded Disability Status Scale (EDSS) Score at Year 2 was evaluated from baseline to Year 2. The analysis used the Wei-Lachin method for non-parametric analysis of repeated measures.
EDSS change from baseline
Time frame: Baseline, Year 2 · FAS subset with EDSS assessment at both baseline and end-of-study (Year 2).
| Secondary endpoint | Method as reported | Analysis population | P-value |
|---|---|---|---|
| Change From Baseline in EDSS Score at Year 2 | Wei-Lachin | FAS subset with EDSS assessment at both baseline and end-of-study (Year 2). | 0.4188 |
No effect estimate or confidence interval for this endpoint is included in the ClinicalTrials.gov record. The result therefore should not be converted into an assumed difference in EDSS or an assumed direction of effect.
The P-value of 0.4188 is the reported inferential result for the repeated-measures EDSS analysis. It should not be interpreted as the probability that the treatments are equal, and it does not tell us the magnitude of any between-group difference.
Because the ClinicalTrials.gov record does not report an effect estimate or confidence interval, the statistically informative description is limited to the reported P-value and the analysis method. The absence of an effect estimate here is a limitation of the posted statistical result, not a reason to infer one.
11. Secondary Results: MSFC Change
The endpoint Change From Baseline in Multiple Sclerosis Functional Composite (MSFC) Score at Year 2 used the Wei-Lachin method for non-parametric analysis of repeated measures. The outcome was expressed as a Z-score.
MSFC change from baseline
Time frame: Baseline, Year 2 · FAS subset described in the registry analysis
| Secondary endpoint | Method as reported | Outcome unit | P-value |
|---|---|---|---|
| Change From Baseline in MSFC Score at Year 2 | Wei-Lachin | Z-score | 0.0115 |
The reported P-value of 0.0115 indicates statistical evidence against the null hypothesis used for this repeated-measures comparison under the analysis framework reported by the registry.
However, the ClinicalTrials.gov record does not provide the estimated between-group change, its confidence interval, or the arm-specific mean or median changes. Therefore, the P-value should not be used to infer the magnitude of the difference or its clinical importance.
The use of a repeated-measures method also means that the statistical result reflects the longitudinal analysis rather than a simple comparison of two independent Year 2 measurements.
12. Secondary Results: MRI-T2 Hyperintense Lesion Volume
The endpoint Percent Change From Baseline in Magnetic Resonance Imaging Time Constant 2 (MRI-T2) Hyperintense Lesion Volume at Year 2 was analyzed using ranked ANCOVA.
MRI-T2 lesion-volume analysis
Time frame: Baseline, Year 2 · Ranked ANCOVA
| Secondary endpoint | Method | Covariates | P-value |
|---|---|---|---|
| Percent Change From Baseline in MRI-T2 Hyperintense Lesion Volume at Year 2 | Ranked ANCOVA | Geographic region and baseline T2 lesion volume | 0.3080 |
The analysis adjusted for geographic region and baseline T2 lesion volume. Baseline adjustment is useful when a continuous outcome can be influenced by the participant's starting disease burden.
The reported P-value of 0.3080 is an inferential result, not an effect-size measure. The ClinicalTrials.gov record does not include an estimated treatment difference or confidence interval for percent change, so no magnitude or direction of treatment effect should be inferred from the P-value alone.
13. Statistical Methods Explained
Why was a Cox proportional-hazards model used for SAD?
Sustained accumulation of disability is naturally represented as a time-to-event outcome: participants can experience the event at different times, while others may not experience it during the observation period. Cox regression uses the timing of events and accommodates right censoring. The CARE-MS I analysis reported a hazard ratio rather than simply comparing the percentage with SAD at a single time point.
What does the SAD hazard ratio of 0.70 mean?
An HR of 0.70 means that the fitted model estimated an instantaneous hazard approximately 30% lower for the treatment comparison because 1 − 0.70 = 0.30. It does not mean that 30% fewer participants necessarily developed SAD, and it does not describe an individual's personal probability of disability accumulation.
Why is the confidence interval important?
The SAD 95% CI is 0.40–1.23. A confidence interval communicates precision around the estimated effect. Because the interval includes 1.00, the statistical uncertainty is compatible with both a lower and a higher hazard under the model. For the relapse-rate ratio, by contrast, the 95% CI is 0.32–0.63, entirely below 1.00.
Why use a rate ratio for annualized relapse rate?
Relapses are recurrent events and can occur at different rates across participants. A rate ratio compares the estimated rate of events between treatment groups. It should not be confused with a risk ratio, which compares the probability of experiencing an event over a specified period.
Why was geographic region included as a covariate?
The analyses posted on ClinicalTrials.gov explicitly report geographic-region adjustment for the primary SAD and annualized relapse-rate analyses, as well as several secondary analyses. Covariate adjustment can account for prespecified variation associated with a geographic factor while estimating the treatment comparison. The ClinicalTrials.gov record does not specify the number or definition of geographic regions, so no further structure is assumed.
Why use Wei-Lachin for EDSS and MSFC?
The registry reports the Wei-Lachin method for non-parametric analysis of repeated measures. Repeated observations from the same participant are correlated, so longitudinal methods account for the within-participant structure rather than treating repeated observations as unrelated data points.
What does a P-value tell us here?
A P-value quantifies how compatible the observed data are with the null hypothesis under the specified statistical model and testing framework. It does not measure the size of the treatment effect, the probability that a treatment works, or clinical importance. Those questions require the effect estimate, confidence interval, absolute outcomes, and clinical context.
14. Confidence Intervals and Effect Measures
The CARE-MS I primary analyses use two different relative effect measures. Understanding the distinction is essential because the numbers have different statistical meanings.
| Measure | Endpoint | Estimate | 95% CI | Plain-language interpretation |
|---|---|---|---|---|
| Hazard ratio | Sustained accumulation of disability | 0.70 | 0.40–1.23 | Model-based comparison of instantaneous event hazards |
| Rate ratio | Annualized relapse rate | 0.45 | 0.32–0.63 | Comparison of estimated relapse rates |
| Hazard ratio | Relapse free at Year 2 | 0.45 | 0.33–0.61 | Model-based time-to-event comparison |
A relative measure such as an HR or rate ratio can be statistically compelling while still leaving important questions about absolute event frequency unanswered. For example, an HR of 0.70 does not tell us how many participants experienced SAD without knowing the underlying event and censoring information.
The confidence interval provides both an estimate and information about its precision. The P-value provides a test-based measure of evidence against the null hypothesis. Neither one alone describes clinical importance.
15. Covariate Adjustment
Covariate adjustment appears repeatedly in the CARE-MS I statistical analyses. Geographic region was included in the primary Cox model for SAD and in the proportional means regression for annualized relapse rate. The MRI-T2 ranked ANCOVA additionally adjusted for baseline T2 lesion volume.
Geographic region
Used as a covariate in the primary SAD and annualized relapse-rate analyses and in the reported relapse-free analysis.
Baseline T2 lesion volume
Included in the ranked ANCOVA for percent change in MRI-T2 hyperintense lesion volume at Year 2.
Why adjust?
Adjustment can account for prespecified variation related to baseline or geographic characteristics when estimating the treatment comparison.
What adjustment does not do
Covariate adjustment does not turn a model estimate into an absolute treatment effect, and it does not eliminate all statistical uncertainty.
16. Secondary Endpoint Sequence and Multiplicity
The registry-reported analysis notes state that the secondary endpoints were analyzed sequentially. The reported sequence begins with the proportion of participants relapse free at Year 2, followed by change from baseline in EDSS, percent change from baseline in MRI-T2 hyperintense lesion volume at Year 2, and acquisition of disability measured by MSFC.
| Sequence | Secondary endpoint | Reported method | Reported P-value |
|---|---|---|---|
| 1 | Percentage of Participants Who Were Relapse Free at Year 2 | Cox proportional-hazards regression | <0.0001 |
| 2 | Change From Baseline in EDSS Score at Year 2 | Wei-Lachin | 0.4188 |
| 3 | Percent Change From Baseline in MRI-T2 Hyperintense Lesion Volume at Year 2 | Ranked ANCOVA | 0.3080 |
| 4 | Change From Baseline in MSFC Score at Year 2 | Wei-Lachin | 0.0115 |
17. Safety
The ClinicalTrials.gov record reports serious adverse events by arm using affected participants divided by participants at risk. These counts are presented exactly as provided in the registry-derived data.
| Treatment arm | Serious adverse events |
|---|---|
| Interferon Beta-1a | 27/187 |
| Alemtuzumab | 69/376 |
The denominators in this safety summary are the registry-reported at-risk counts and should not be replaced with the overall enrollment of 581.0. The registry-derived ClinicalTrials.gov record do not provide a formal statistical comparison, confidence interval, or P-value for serious adverse events, so none is inferred.
18. Reading the Two Primary Endpoints Together
One of the most useful statistical features of CARE-MS I is that the two primary endpoints represent different dimensions of disease activity.
Disability accumulation
SAD addresses a clinically important time-to-event outcome based on EDSS. Its reported HR was 0.70 with a 95% CI of 0.40–1.23.
Relapse activity
Annualized relapse rate addresses recurrent clinical events. Its reported rate ratio was 0.45 with a 95% CI of 0.32–0.63.
These measures should not be collapsed into a single number. The SAD analysis uses time-to-event methodology and asks when sustained disability accumulation occurs. The annualized relapse-rate analysis uses a rate model and asks about the frequency of relapse events over follow-up.
The difference between their effect measures is therefore not a contradiction. A hazard ratio and a rate ratio summarize different statistical quantities.
19. What the Primary P-values Do — and Do Not — Mean
The reported P-value of 0.2173 does not measure the magnitude of the SAD effect. The relevant magnitude is the HR of 0.70, while the confidence interval of 0.40–1.23 communicates its statistical precision.
The reported P-value of <0.0001 indicates strong statistical evidence against the null hypothesis under the reported model. The size of the estimated effect is communicated by the rate ratio of 0.45 and its 95% CI of 0.32–0.63.
The statistical analyses quantify differences between randomized treatment groups. Whether a treatment's overall clinical profile is appropriate for an individual patient requires considerations beyond these reported statistical estimates.
20. Limitations
- Limited absolute outcome reporting: the statistical analyses posted on ClinicalTrials.gov provide effect measures and P-values but do not provide all arm-specific event percentages, relapse counts, or absolute rates.
- Incomplete confidence intervals for secondary continuous outcomes: EDSS, MSFC, and MRI-T2 analyses are reported with P-values but without corresponding effect estimates and confidence intervals.
- Analysis subsets: some longitudinal outcomes were evaluated in subsets of the FAS population based on availability of the relevant assessments. This differs from simply analyzing every randomized participant for every endpoint.
- Model assumptions: the Cox analysis is based on a proportional-hazards model. The ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption.
- Multiplicity: secondary endpoints were analyzed sequentially, but the ClinicalTrials.gov record does not specify a formal multiplicity-adjustment procedure or alpha allocation for those analyses.
- Safety denominator differences: serious adverse-event counts are reported as affected/at-risk counts of 27/187 and 69/376 rather than as counts based on the overall enrollment.
- No unsupported reconstruction: the available registry-derived data are insufficient to reconstruct Kaplan-Meier curves, arm-specific relapse percentages, or unreported effect estimates without introducing information outside the ClinicalTrials.gov record.
21. Why This Trial Matters Statistically
CARE-MS I is a useful teaching example because the same randomized comparison requires several distinct statistical tools. The primary disability endpoint is analyzed with survival methods, relapse frequency is analyzed as a rate, and secondary continuous or longitudinal outcomes use methods designed for repeated measurements or baseline adjustment.
| Statistical concept | How it appears in CARE-MS I |
|---|---|
| Randomization | The trial used randomized allocation to two parallel treatment groups. |
| FAS analysis | The primary and secondary efficacy analyses used a population consisting of randomized participants who received at least 1 dose of study drug. |
| Time-to-event analysis | SAD and relapse-free-at-Year-2 analyses used Cox proportional-hazards regression. |
| Hazard ratio | Used for SAD and the relapse-free-at-Year-2 endpoint. |
| Rate ratio | Used for annualized relapse rate. |
| Covariate adjustment | Geographic region was included in several analyses; baseline T2 lesion volume was also used in the ranked ANCOVA. |
| Repeated measures | EDSS and MSFC changes were analyzed using the Wei-Lachin method for non-parametric repeated measures. |
| ANCOVA | Ranked ANCOVA was used for percent change in MRI-T2 hyperintense lesion volume. |
| Confidence intervals | Primary hazard-ratio and rate-ratio analyses included two-sided 95% confidence intervals. |
| P-values | Reported for all six posted statistical analyses. |
| Safety analysis | Serious adverse events were reported as affected participants divided by participants at risk for each arm. |
22. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
23. Related Statistical Calculators
Use these calculator pathways to explore the statistical quantities that appear in CARE-MS I:
24. Trial Timeline
Study start
The CARE-MS I trial began in August 2007 according to the ClinicalTrials.gov record.
Randomized treatment comparison
The phase 3 study used randomized allocation, a parallel design, and single masking to compare alemtuzumab with interferon beta-1a.
Primary efficacy assessment window
The two registered primary endpoints were evaluated over up to 2 years.
Primary completion
The ClinicalTrials.gov record identifies April 2011 as the primary completion date.
25. What the Hazard Ratio Does — and Does Not — Mean
The reported SAD HR of 0.70 is a relative measure of instantaneous event hazard under the Cox model. A simple mathematical transformation gives 1 − 0.70 = 0.30, so the model estimate corresponds to a 30% lower estimated hazard.
It does not mean that 30% of participants avoided sustained disability accumulation, that 30% of participants were cured, or that every individual participant experienced exactly a 30% reduction in risk.
The annualized relapse-rate ratio of 0.45 means the estimated relapse rate was 45% of the comparator rate under the reported model. Equivalently, 1 − 0.45 = 0.55, corresponding to a 55% lower estimated rate.
This is a rate comparison rather than a direct comparison of the probability that an individual participant experienced at least one relapse.
The SAD 95% CI of 0.40–1.23 communicates substantial uncertainty around the HR estimate. The annualized relapse-rate 95% CI of 0.32–0.63 communicates a narrower range entirely below 1.00. Confidence intervals should be interpreted together with the underlying model, endpoint definition, and analysis population.
26. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The randomized comparison produced a SAD hazard ratio of 0.70 and an annualized relapse-rate ratio of 0.45. The two primary endpoints used different statistical models because they measure different types of outcomes.
Endpoint-specific interpretation
The relapse-rate analysis provides a rate comparison with a 95% CI of 0.32–0.63 and P < 0.0001, while the SAD analysis provides an HR of 0.70 with a 95% CI of 0.40–1.23 and P = 0.2173.
These findings should not be compressed into a single overall statistical score. The appropriate interpretation depends on the endpoint, effect measure, confidence interval, analysis population, model assumptions, and the distinction between efficacy outcomes and safety outcomes.
27. Limitations of Statistical Interpretation
- Hazard-ratio interpretation: Cox regression summarizes time-to-event information under a proportional-hazards framework. The ClinicalTrials.gov record does not report a formal proportional-hazards diagnostic.
- Different effect measures: HRs and rate ratios cannot be interpreted as interchangeable quantities.
- Incomplete arm-specific data: the statistical analyses posted on ClinicalTrials.gov do not include all underlying event counts, absolute rates, or arm-specific percentages.
- Secondary endpoint estimates: several secondary analyses have P-values but no registry-reported effect estimate or confidence interval.
- Repeated-measures subsets: EDSS and MRI-T2 analyses used subsets of the FAS population based on assessment availability.
- Multiplicity information: the sequence of secondary analyses is reported, but a formal multiplicity procedure is not specified in the ClinicalTrials.gov record.
- Safety: serious adverse-event data are reported as affected/at-risk counts, without a formal between-arm statistical test.
28. Sources
- ClinicalTrials.gov: CARE-MS I, NCT00530348.
- Linked PubMed record: PubMed 39935588.
- Linked PubMed record: PubMed 37745914.
- Linked PubMed record: PubMed 37272540.
- Linked PubMed record: PubMed 36619856.
- Linked PubMed record: PubMed 34882037.
Continue with the statistical methods behind CARE-MS I
Explore the underlying survival-analysis, rate-ratio, covariate-adjustment, repeated-measures, confidence-interval, and P-value concepts used to interpret randomized clinical-trial results.
29. Record Summary
CARE-MS I provides a useful example of how a randomized phase 3 trial can require multiple statistical frameworks within the same efficacy program. The primary endpoint of sustained accumulation of disability was analyzed as a time-to-event outcome using Cox proportional-hazards regression, producing an HR of 0.70 with a 95% CI of 0.40–1.23 and P = 0.2173. Annualized relapse rate was analyzed using proportional means regression and produced a rate ratio of 0.45 with a 95% CI of 0.32–0.63 and P < 0.0001.
The secondary analyses illustrate additional statistical concepts: Cox regression for relapse-free status at Year 2, Wei-Lachin repeated-measures analysis for EDSS and MSFC, and ranked ANCOVA for MRI-T2 lesion-volume change. The reported P-values were <0.0001, 0.4188, 0.0115, and 0.3080, respectively. Importantly, several of these secondary results are reported without corresponding effect estimates or confidence intervals, so their magnitude should not be reconstructed from P-values alone.
The trial is therefore particularly useful for understanding a central principle of clinical-trial statistics: the endpoint determines the estimand, the estimand determines the appropriate effect measure, and the effect measure must be interpreted together with its confidence interval, analysis population, model assumptions, and design context.