This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the registry-reported COVACTA registry data.
1. Trial at a Glance
COVACTA was a randomized, double-blind, parallel phase 3 study evaluating the safety and efficacy of tocilizumab in patients with severe COVID-19 pneumonia. The registry reports an enrollment of 452 participants, two study arms, one registered primary endpoint, 13 posted outcome measures, and 15 posted statistical analyses.
| Feature | COVACTA |
|---|---|
| Trial name | COVACTA |
| ClinicalTrials.gov identifier | NCT04320615 |
| Brief title | A Study to Evaluate the Safety and Efficacy of Tocilizumab in Patients With Severe COVID-19 Pneumonia |
| Therapeutic area | Infectious Disease |
| Condition | COVID-19 Pneumonia |
| Phase | Phase 3 |
| Status | Completed |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Double |
| Primary purpose | Treatment |
| Enrollment | 452 |
| Interventions | Tocilizumab (TCZ); placebo |
| Lead sponsor | Hoffmann-La Roche |
| Sponsor type | Industry |
| Results posted | Yes |
| Outcome measures posted | 13 |
| Statistical analyses posted | 15 |
2. Clinical Question
The clinical question was whether tocilizumab produced a different clinical outcome than placebo in patients with severe COVID-19 pneumonia. The registered primary endpoint assessed clinical status at Day 28 using a 7-category ordinal scale.
Population
Patients with severe COVID-19 pneumonia, as described by the trial's brief title and condition in the ClinicalTrials.gov record.
Intervention
Tocilizumab (TCZ).
Comparator
Placebo.
Primary question
How did clinical status at Day 28 compare between the randomized tocilizumab and placebo groups?
3. Trial Design
Tocilizumab (TCZ)
- Intervention classified as a drug.
- Compared with the placebo arm in the posted statistical analyses.
- mITT analyses grouped participants according to treatment assigned at randomization.
Placebo
- Comparator classified as a drug.
- Compared with the tocilizumab arm in the posted statistical analyses.
- mITT analyses grouped participants according to treatment assigned at randomization.
4. Trial Timeline
Study start
The registry lists 2020-04-03 as the study start date.
Primary completion
The registry lists 2020-06-24 as the primary completion date.
Registry status
The ClinicalTrials.gov record identifies COVACTA as COMPLETED and reports results.
5. Analysis Population
The posted statistical analyses repeatedly identify the modified intent-to-treat (mITT) population. The registry defines this population as all randomized participants who received any amount of study medication, with participants grouped according to the treatment assigned at randomization.
This distinction is important. The analysis population is not simply every person enrolled in the study. The registry-reported definition requires randomization and receipt of some study medication, while preserving treatment assignment as the basis for grouping.
The same mITT definition is reported for the primary analysis and the posted secondary analyses. This creates a consistent analytical population across the efficacy comparisons reported in the ClinicalTrials.gov record.
6. Primary Endpoint
| Endpoint | Registry definition / time frame | Analysis |
|---|---|---|
| Clinical Status Assessed Using a 7-Category Ordinal Scale at Day 28 (Week 4) | Day 28. Clinical status was assessed using a 7-category ordinal scale. The registry definition begins: 1. Discharged (or "ready for discharge"); 2. Non-ICU hospital ward (or "ready for hospital ward") not requiring supplemental oxygen; 3. Non-ICU hospital ward (or "ready for hospital ward") requiring supplemental oxygen; 4. ICU or non-ICU hospital ward, requiring non-invasive ventilation or high-flow oxygen. | Van Elteren test |
The ClinicalTrials.gov record classifies the posted analysis as a binary endpoint for the statistical-analysis record, while the registered endpoint itself is explicitly described as a 7-category ordinal scale. The analysis reported in the registry uses a Van Elteren test and reports a median difference in final values.
7. Primary Result
The registry reports one formal primary-endpoint analysis comparing the tocilizumab mITT arm with the placebo mITT arm at Day 28.
Clinical status at Day 28
95% CI: −2.5 to 0.0 · P = 0.3600
Analysis: Van Elteren test · two-sided 95% confidence interval
| Primary endpoint | Tocilizumab vs placebo | Method | 95% CI | P-value |
|---|---|---|---|---|
| Clinical Status Assessed Using a 7-Category Ordinal Scale at Day 28 (Week 4) | Median difference: −1.0 | Van Elteren test | −2.5 to 0.0 | 0.3600 |
The reported median difference of −1.0 is the registry's reported effect measure for the final clinical-status values, comparing the tocilizumab mITT arm with the placebo mITT arm. Because the underlying endpoint is a 7-category clinical-status scale, the estimate should be interpreted in the context of that ordinal outcome rather than as a hazard ratio or a conventional mean difference.
The 95% confidence interval of −2.5 to 0.0 describes the statistical uncertainty around the reported median-difference estimate under the analysis framework. The interval reaches 0.0, so the registry-reported estimate is compatible with no difference as represented by this effect measure. The confidence interval does not describe the range of individual patient outcomes.
The P-value of 0.3600 is a measure of evidence against the statistical null hypothesis represented by the test; it is not a measure of effect size, clinical importance, or the probability that either treatment is effective. A p-value should therefore be interpreted together with the effect estimate and confidence interval.
The registry identifies the test as two-sided through the confidence-interval specification. The analysis population is the mITT population, not the full enrolled population. Because the registered outcome is ordinal, the choice of a nonparametric method is an important part of the statistical story.
8. Secondary Endpoint Results
The registry contains 14 posted secondary statistical analyses. They cover time-to-event outcomes, categorical outcomes, and final-value comparisons. All of the analyses in the ClinicalTrials.gov record use the mITT population unless a baseline subgroup is explicitly stated in the comparison.
| Secondary endpoint | Time frame | Method | Effect estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Time to Clinical Improvement (TTCI), defined as a National Early Warning Score 2 (NEWS2) of ≤2 maintained for 24 Hours | Up to Day 28 | Log-rank test | HR 1.448 | 1.01 to 2.08 | 0.0443 |
| Time to Improvement of at Least 2 Categories Relative to Baseline on a 7-Category Ordinal Scale of Clinical Status | Up to Day 28 | Log-rank test | HR 1.263 | 0.97 to 1.64 | 0.0820 |
| Time to Hospital Discharge or "Ready for Discharge" | Up to Day 28 | Log-rank test | HR 1.350 | 1.02 to 1.79 | 0.0370 |
| Incidence of Mechanical Ventilation by Day 28 | Up to Day 28 | Cochran-Mantel-Haenszel test | Weighted % difference −6.2 | −13.8 to 1.4 | 0.0996 |
| Incidence of Mechanical Ventilation by Day 28: TCZ - No Mechanical Ventilation at Baseline vs Placebo - No Mechanical Ventilation at Baseline | Up to Day 28 | Cochran-Mantel-Haenszel test | Weighted % difference −8.9 | −20.7 to 3.0 | 0.1355 |
| Ventilator-Free Days to Day 28 | Up to Day 28 | Van Elteren test | Median difference 5.5 | −2.8 to 13.0 | 0.3202 |
| Incidence of Intensive Care Unit (ICU) Stay by Day 28 (Week 4) | Up to Day 28 | Cochran-Mantel-Haenszel test | Weighted % difference −5.7 | −13.7 to 2.2 | 0.1514 |
| Incidence of ICU Stay by Day 28: TCZ - Not in ICU at Baseline vs Placebo - Not in ICU at Baseline | Up to Day 28 | Cochran-Mantel-Haenszel test | Weighted % difference −14.8 | −28.6 to −1.0 | 0.0290 |
| Duration of ICU Stay to Day 28 (Week 4) | Up to Day 28 | Van Elteren test | Median difference −5.8 | −15.0 to 2.9 | 0.0454 |
| Clinical Status Assessed Using a 7-Category Ordinal Scale at Day 14 | Day 14 | Van Elteren test | Median difference −1.0 | −2.0 to 0.5 | 0.0548 |
| Time to Clinical Failure to Day 28 (Week 4) | Up to Day 28 | Log-rank test | HR 0.790 | 0.57 to 1.10 | 0.1627 |
| Mortality Rate at Day 28 (Week 4) | Day 28 | Cochran-Mantel-Haenszel test | Weighted % difference 0.3 | −7.6 to 8.2 | 0.9410 |
| Time to Recovery to Day 28 (Week 4) | Up to Day 28 | Log-rank test | HR 1.307 | 1.00 to 1.72 | 0.0528 |
| Duration of Supplemental Oxygen to Day 28 (Week 4) | Up to Day 28 | Van Elteren test | Median difference −1.5 | −9.0 to 0.5 | 0.0477 |
9. Time-to-Event Results
Several secondary endpoints were analyzed with the log-rank test and reported using hazard ratios. This is a coherent analytical pattern for endpoints defined by the time until an event occurs within the specified follow-up window.
Time to Clinical Improvement
95% CI: 1.01–2.08 · P = 0.0443
Time frame: Up to Day 28
For time to clinical improvement, the reported hazard ratio of 1.448 means that the estimated instantaneous rate of reaching the defined improvement event was 1.448 times as high in the tocilizumab group as in the placebo group under the time-to-event analysis. Expressed as a simple relative interpretation, this is an estimated 44.8% higher instantaneous event rate, not a statement that 44.8% more patients necessarily improved.
The 95% CI of 1.01 to 2.08 indicates uncertainty around the estimated hazard ratio. The lower limit is just above 1.00, while the upper limit is substantially higher, so the interval is not especially narrow. The estimate therefore should not be separated from its uncertainty.
The P-value of 0.0443 describes evidence against the relevant null hypothesis in the reported log-rank analysis. It does not quantify the magnitude or clinical importance of the difference. It also does not establish that the treatment effect is exactly 44.8%.
As with any hazard ratio, interpretation depends on the time-to-event framework and censoring process. A single HR is most straightforward when the relative hazards are reasonably represented by the model over time; the ClinicalTrials.gov record does not provide a separate assessment of the proportional-hazards assumption.
Time to Hospital Discharge or "Ready for Discharge"
95% CI: 1.02–1.79 · P = 0.0370
Time frame: Up to Day 28
The hazard ratio of 1.350 indicates a higher estimated instantaneous rate of reaching the defined discharge or "ready for discharge" event in the tocilizumab group under the reported analysis. The corresponding simple relative interpretation is a 35.0% higher estimated event rate, not a 35.0% absolute increase in the number of discharged participants.
The 95% CI of 1.02 to 1.79 indicates that the estimate has meaningful uncertainty, although the interval lies above 1.00. The p-value of 0.0370 is evidence against the null hypothesis represented by the log-rank test, but it should not be treated as a measure of the size of the treatment effect.
Time to Clinical Failure
95% CI: 0.57–1.10 · P = 0.1627
Time frame: Up to Day 28
The reported hazard ratio of 0.790 corresponds to an estimated instantaneous clinical-failure rate that is 79.0% of the placebo-group rate under the reported time-to-event model. A simple relative interpretation is therefore a 21.0% lower estimated hazard, but this is not an absolute risk reduction and does not mean that 21.0% of participants avoided clinical failure.
The 95% CI of 0.57 to 1.10 spans 1.00, indicating considerable uncertainty about the direction and magnitude of the underlying relative hazard. The P-value of 0.1627 is not a measure of effect size and should be considered alongside this confidence interval.
Time to Recovery
95% CI: 1.00–1.72 · P = 0.0528
Time frame: Up to Day 28
The hazard ratio of 1.307 represents a higher estimated instantaneous rate of the defined recovery event in the tocilizumab group. A simple relative reading is a 30.7% higher estimated event rate, but this should not be converted into a statement about the percentage of patients who recovered.
The 95% CI of 1.00 to 1.72 reaches 1.00 at its lower boundary. The p-value of 0.0528 likewise should not be treated as an effect-size statistic. Together, the estimate and interval indicate that the magnitude of the time-to-recovery difference is uncertain.
10. Categorical Endpoint Results
COVACTA also used the Cochran-Mantel-Haenszel test for several categorical outcomes. The registry reports weighted percentage differences with two-sided 95% confidence intervals.
| Outcome | Weighted % difference | 95% CI | P-value |
|---|---|---|---|
| Incidence of Mechanical Ventilation by Day 28 | −6.2 | −13.8 to 1.4 | 0.0996 |
| Mechanical Ventilation by Day 28 among participants with no mechanical ventilation at baseline | −8.9 | −20.7 to 3.0 | 0.1355 |
| Incidence of ICU Stay by Day 28 | −5.7 | −13.7 to 2.2 | 0.1514 |
| ICU Stay by Day 28 among participants not in ICU at baseline | −14.8 | −28.6 to −1.0 | 0.0290 |
| Mortality Rate at Day 28 | 0.3 | −7.6 to 8.2 | 0.9410 |
A percentage difference is an absolute-scale comparison, unlike a hazard ratio. For example, the reported weighted percentage difference of −6.2 for mechanical ventilation by Day 28 indicates a lower weighted percentage in the tocilizumab group by 6.2 percentage points under the reported analysis framework. It does not mean a 6.2% relative reduction.
The 95% CI of −13.8 to 1.4 crosses 0, illustrating why the point estimate alone is insufficient. The p-value of 0.0996 does not measure the clinical size of the observed difference.
For the subgroup comparison restricted to participants not in the ICU at baseline, the reported weighted percentage difference was −14.8, with a 95% CI of −28.6 to −1.0 and P = 0.0290. This is a baseline-defined subgroup analysis and should not automatically be treated as equivalent to the overall randomized comparison.
11. Nonparametric Comparisons
The registry reports the Van Elteren test for the primary Day 28 clinical-status analysis and for several secondary final-value comparisons. The method is a stratified nonparametric extension of the Wilcoxon rank-sum approach and is useful when comparing ordered or continuous outcomes without requiring a normal-distribution assumption.
| Endpoint | Median difference | 95% CI | P-value |
|---|---|---|---|
| Clinical status at Day 28 | −1.0 | −2.5 to 0.0 | 0.3600 |
| Ventilator-free days to Day 28 | 5.5 | −2.8 to 13.0 | 0.3202 |
| Duration of ICU stay to Day 28 | −5.8 | −15.0 to 2.9 | 0.0454 |
| Clinical status at Day 14 | −1.0 | −2.0 to 0.5 | 0.0548 |
| Duration of supplemental oxygen to Day 28 | −1.5 | −9.0 to 0.5 | 0.0477 |
The word median in these effect measures matters. A median difference is not the same as a difference in means, and it should not be interpreted as though every participant experienced the estimated difference. It summarizes a distributional comparison on the reported final values.
The trial contains outcomes that are naturally ordered, skewed, or otherwise not well summarized by a simple normal-theory mean comparison. A rank-based nonparametric approach can compare treatment groups without requiring the outcome distribution to be normal. The Van Elteren formulation also allows the comparison to account for stratification when the relevant strata are available.
For the reported duration of ICU stay, the median difference was −5.8, with a 95% CI of −15.0 to 2.9 and P = 0.0454. The confidence interval includes 0 even though the reported p-value is below 0.05, illustrating why the exact definition of an effect measure and its confidence interval should be examined rather than reducing the result to a binary "significant/not significant" label.
12. Secondary Result: Ventilator-Free Days
Ventilator-Free Days to Day 28
95% CI: −2.8 to 13.0 · P = 0.3202
Analysis: Van Elteren test
The reported median difference of 5.5 days is a final-value comparison between the tocilizumab and placebo mITT groups. The positive point estimate indicates a higher median ventilator-free-days value in the tocilizumab group under the reported effect-measure convention.
The 95% CI of −2.8 to 13.0 is broad relative to the point estimate and crosses zero. The p-value of 0.3202 does not measure the size of the observed difference. The result should therefore be read as an estimate with substantial uncertainty rather than as a precise statement about the number of ventilator-free days gained by an individual participant.
13. Secondary Result: ICU Outcomes
ICU incidence
Overall weighted percentage difference: −5.7; 95% CI −13.7 to 2.2; P = 0.1514.
Not in ICU at baseline
Weighted percentage difference: −14.8; 95% CI −28.6 to −1.0; P = 0.0290.
Duration of ICU stay
Median difference: −5.8; 95% CI −15.0 to 2.9; P = 0.0454.
Analytical distinction
Incidence is categorical; duration is a final-value comparison. The registry therefore uses different statistical methods and effect measures.
These results illustrate a central principle of clinical-trial statistics: related clinical questions can require different estimands and methods. Whether a participant experiences an ICU stay is a categorical outcome; the duration of ICU stay is a distributional outcome. Treating the two as interchangeable would obscure the statistical structure of the data.
14. Secondary Result: Clinical Status at Day 14
Clinical Status at Day 14
95% CI: −2.0 to 0.5 · P = 0.0548
Analysis: Van Elteren test
The Day 14 endpoint uses the same general clinical-status framework as the registered Day 28 primary endpoint, but the time frame is earlier. The reported estimate is accompanied by a confidence interval that crosses zero and a p-value of 0.0548.
The Day 14 analysis addresses an earlier time point and is a secondary endpoint. It should therefore be interpreted as its own analysis rather than as an alternative definition of the primary endpoint. A p-value near a conventional threshold does not transform a secondary endpoint into a primary endpoint, and it does not establish a treatment effect by itself.
15. Mortality at Day 28
Mortality Rate at Day 28
95% CI: −7.6 to 8.2 · P = 0.9410
Analysis: Cochran-Mantel-Haenszel test
The reported weighted percentage difference of 0.3 is close to zero. The 95% confidence interval extends from −7.6 to 8.2, indicating substantial uncertainty around the magnitude and direction of the treatment-group difference on this scale.
The reported P-value of 0.9410 provides little evidence against the null hypothesis represented by this comparison. It should not, however, be converted into a probability that the treatments are identical, nor does it establish equivalence.
16. Statistical Methodology
Van Elteren test
The Van Elteren test is a stratified, rank-based nonparametric procedure. It extends the logic of the Wilcoxon rank-sum test by allowing comparisons to be performed across strata and then combined. In COVACTA, the registry identifies the Van Elteren test for the primary clinical-status endpoint and several secondary final-value outcomes.
The method is particularly useful when the outcome is ordinal or when a distribution-free comparison is preferred. The exact weighting and stratification details should be taken from the trial's statistical-analysis specification when those details are available.
Log-rank test
The log-rank test is used for comparing time-to-event distributions between treatment groups. COVACTA uses the log-rank test for time to clinical improvement, time to improvement of at least two categories, time to hospital discharge or "ready for discharge," time to clinical failure, and time to recovery.
The approach uses information from the timing of events rather than only whether an event eventually occurred. Participants who have not experienced the event during available follow-up can contribute censored information.
Cochran-Mantel-Haenszel test
The Cochran-Mantel-Haenszel test is a family of methods for comparing categorical outcomes while accounting for stratification. In the registry-reported COVACTA analyses, it is used for mechanical ventilation, ICU stay incidence, and Day 28 mortality.
Hazard ratio
The registry reports hazard ratios for time-to-event endpoints. A hazard ratio compares the estimated instantaneous event rates between groups within the fitted time-to-event framework.
The direction must always be interpreted relative to the event being analyzed. A hazard ratio below 1 is not inherently favorable or unfavorable without knowing whether the event represents improvement, failure, hospitalization, recovery, or another outcome.
Confidence intervals
The posted analyses report two-sided 95% confidence intervals. For hazard ratios, a value of 1 represents the conventional no-relative-hazard-difference point. For difference measures, 0 represents the conventional no-difference point. These reference values are useful for interpreting the interval, but the clinical meaning still depends on the endpoint and estimand.
17. Statistical Methods Explained
Why was the Van Elteren test used?
The registry identifies a Van Elteren test for the primary clinical-status analysis and several secondary final-value outcomes. This is a nonparametric, stratified method that is appropriate for rank-based comparisons of ordered or non-normally distributed outcomes. Its use is consistent with the fact that clinical status is represented by an ordinal scale and several secondary measures are not naturally characterized by a normal-distribution assumption.
What does a hazard ratio of 1.448 mean?
For the time-to-clinical-improvement endpoint, HR 1.448 means the estimated instantaneous rate of reaching the defined improvement event was 1.448 times the corresponding rate in the placebo group under the reported time-to-event analysis. It is not a 44.8 percentage-point increase in improvement, and it is not the probability that a participant improved.
Why is a hazard ratio different from a percentage difference?
A hazard ratio incorporates the timing of events and censoring in a time-to-event analysis. A weighted percentage difference compares categorical outcome percentages. A value of −6.2 therefore has a fundamentally different interpretation from HR 1.448, even though both are treatment-effect measures.
What does a 95% confidence interval tell us?
The interval describes uncertainty around the estimated treatment effect under the statistical model and sampling framework. For HR 1.448, the interval is 1.01 to 2.08. For a weighted percentage difference of −6.2, the interval is −13.8 to 1.4. The relevant no-effect reference value is 1 for a hazard ratio and 0 for a difference measure.
Why does the p-value not measure effect size?
A p-value summarizes how compatible the observed data are with a specified null hypothesis under the test procedure. It depends on the estimate, variability, sample information, and statistical model. It does not tell us how large the treatment effect is. That is why COVACTA's reported p-values should be read together with the effect estimates and confidence intervals.
Why does endpoint direction matter?
For an improvement endpoint, a hazard ratio above 1 may correspond to a higher rate of reaching improvement. For a failure endpoint, a hazard ratio below 1 may correspond to a lower rate of failure. The same numerical hazard ratio can therefore have different clinical interpretations depending on the event definition.
Why should secondary p-values not be treated as a list of independent discoveries?
COVACTA has one registered primary endpoint and multiple secondary analyses. Each secondary analysis asks a different statistical question, and repeated testing creates a broader multiplicity problem when many hypotheses are considered together. A collection of p-values should therefore not be converted into a simple count of "significant" findings without understanding the prespecified testing strategy.
18. Endpoint-by-Endpoint Statistical Map
| Endpoint family | Example endpoint | Method | Effect measure |
|---|---|---|---|
| Ordinal / final value | Clinical status at Day 28 | Van Elteren test | Median difference |
| Time-to-event | Time to clinical improvement | Log-rank test | Hazard ratio |
| Time-to-event | Time to hospital discharge or "ready for discharge" | Log-rank test | Hazard ratio |
| Categorical | Incidence of mechanical ventilation | Cochran-Mantel-Haenszel test | Weighted % difference |
| Categorical | Mortality rate at Day 28 | Cochran-Mantel-Haenszel test | Weighted % difference |
| Nonparametric final value | Ventilator-free days | Van Elteren test | Median difference |
| Nonparametric final value | Duration of supplemental oxygen | Van Elteren test | Median difference |
This method map is one of the most useful statistical features of the trial. COVACTA does not force every clinical outcome into a single analysis framework. Instead, the posted analyses distinguish between time-to-event, categorical, and rank-based outcomes.
19. Safety
The ClinicalTrials.gov record reports serious adverse events by randomized arm using a safety-evaluable population.
| Safety-evaluable arm | Participants with serious adverse events | At risk | Affected / at risk |
|---|---|---|---|
| Placebo | 64 | 143 | 64/143 |
| Tocilizumab (TCZ) | 116 | 295 | 116/295 |
The reported serious-adverse-event figures are presented as affected participants over the safety-evaluable population in each arm. They should not be substituted for the efficacy analysis population because the ClinicalTrials.gov record explicitly identifies these as safety-evaluable populations.
The serious-adverse-event data are exposure-related safety information rather than an efficacy endpoint. The denominators differ between the two reported safety-evaluable groups, so the raw event counts of 64 and 116 should not be compared as though they represented equally sized populations.
The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or p-value for these serious-adverse-event counts. The appropriate interpretation is therefore descriptive rather than a claim of a statistically tested difference.
20. Missing Data, Censoring, and Analysis Assumptions
The ClinicalTrials.gov record identifies mITT populations and time-to-event methods, but they do not provide detailed missing-data or imputation rules for the individual endpoints. It is therefore important not to infer an imputation procedure that is not present in the ClinicalTrials.gov record.
Time-to-event outcomes
Log-rank analyses use event timing and can incorporate censored observations. The ClinicalTrials.gov record does not specify the complete censoring rules.
Final-value outcomes
Van Elteren analyses compare final values using a nonparametric framework. The ClinicalTrials.gov record does not specify a separate imputation strategy.
Categorical outcomes
Cochran-Mantel-Haenszel analyses compare categorical outcomes with a stratified framework. The ClinicalTrials.gov record does not provide additional missing-outcome handling details.
Model assumptions
The ClinicalTrials.gov record does not report a separate proportional-hazards assessment for the hazard-ratio analyses.
21. Multiplicity and Multiple Secondary Analyses
The registry reports one primary endpoint and multiple secondary analyses. This structure matters because the nominal p-value for any individual secondary endpoint does not, by itself, describe the probability of making at least one false-positive conclusion across the complete collection of hypotheses.
| Analysis layer | Registry information | Statistical interpretation |
|---|---|---|
| Primary | 1 registered primary endpoint; 1 formal primary statistical analysis | The primary result should be interpreted as the principal prespecified efficacy comparison represented in the ClinicalTrials.gov record. |
| Secondary | 14 posted secondary statistical analyses | These address multiple additional clinical questions and should be interpreted individually and collectively rather than as isolated tests. |
| Safety | Serious adverse events reported descriptively by arm | The ClinicalTrials.gov record does not provide a formal inferential comparison for these safety counts. |
No non-inferiority margin, factorial design, crossover analysis, Bayesian analysis, or interim-analysis procedure is included in the ClinicalTrials.gov record. Those design topics are therefore not used as assumptions in this analysis.
22. What the Hazard Ratio Does — and Does Not — Mean
The reported hazard ratio of 1.448 for time to clinical improvement means that, under the reported time-to-event analysis, the estimated instantaneous rate of reaching the defined improvement event was 1.448 times that of the placebo group.
It does not mean that 44.8% more participants improved, that 44.8% of participants benefited, or that each individual participant experienced a 44.8% increase in the chance of improvement.
The reported hazard ratio of 0.790 for time to clinical failure corresponds to a 21.0% lower estimated instantaneous failure hazard under a simple relative interpretation. It does not mean that 21.0% of participants avoided clinical failure.
The confidence interval provides information about precision. For HR 0.790, the 95% CI is 0.57 to 1.10, which spans the conventional no-relative-difference value of 1.00. For HR 1.448, the 95% CI is 1.01 to 2.08. The intervals convey information that cannot be recovered from the p-values alone.
23. P-Values in Context
The COVACTA results include p-values ranging from 0.0290 to 0.9410 across the statistical analyses posted on ClinicalTrials.gov. These values should not be interpreted as a ranking of clinical importance.
| P-value | Endpoint | Effect measure | 95% CI |
|---|---|---|---|
| 0.0290 | ICU Stay by Day 28 among participants not in ICU at baseline | Weighted % difference −14.8 | −28.6 to −1.0 |
| 0.0370 | Time to Hospital Discharge or "Ready for Discharge" | HR 1.350 | 1.02 to 1.79 |
| 0.0443 | Time to Clinical Improvement | HR 1.448 | 1.01 to 2.08 |
| 0.0454 | Duration of ICU Stay | Median difference −5.8 | −15.0 to 2.9 |
| 0.0477 | Duration of Supplemental Oxygen | Median difference −1.5 | −9.0 to 0.5 |
| 0.9410 | Mortality Rate at Day 28 | Weighted % difference 0.3 | −7.6 to 8.2 |
The table demonstrates why p-values should not be interpreted without the effect measure and confidence interval. The reported duration of ICU stay analysis has P = 0.0454, for example, while its 95% confidence interval for the median difference extends from −15.0 to 2.9. The inferential details therefore contain more information than the threshold classification alone.
24. Limitations
- Registry-level evidence: this page is restricted to the ClinicalTrials.gov record and does not introduce additional numerical results from publications.
- Primary endpoint structure: the registered endpoint is an ordinal 7-category clinical-status scale, while the ClinicalTrials.gov record labels the outcome type as binary and the outcome unit as percentage of participants. Those registry fields should not be silently harmonized into a different endpoint definition.
- Analysis-population distinction: efficacy analyses use the registry-reported mITT definition, while serious-adverse-event results are reported for safety-evaluable populations.
- Multiple secondary analyses: fourteen secondary statistical analyses create a broader multiplicity context. Nominal p-values should not automatically be interpreted as independently confirmatory findings.
- Hazard-ratio interpretation: hazard ratios describe relative instantaneous event rates within a time-to-event framework and should not be translated directly into absolute risks or individual probabilities.
- Proportional-hazards assumption: the ClinicalTrials.gov record does not report a separate assessment of proportional hazards. A single hazard ratio can be less informative if the relative hazard changes substantially over time.
- Confidence-interval interpretation: a confidence interval describes uncertainty around the estimated effect; it is not a prediction interval for individual patient outcomes.
- Safety inference: the registry-reported serious-adverse-event counts are descriptive and do not include a formal between-arm statistical test in the ClinicalTrials.gov record.
- Incomplete methodological detail: the ClinicalTrials.gov record does not specify detailed imputation, censoring, stratification, or multiplicity procedures beyond the method names and analysis-population definitions provided.
25. Why This Trial Matters Statistically
COVACTA is a useful teaching example because its posted analyses show how one randomized clinical trial can require several distinct statistical approaches. The primary clinical-status outcome uses a nonparametric rank-based method, several secondary outcomes use time-to-event methods, and categorical outcomes use a Cochran-Mantel-Haenszel framework.
| Concept | How it appears in COVACTA |
|---|---|
| Randomization | The trial is identified as randomized, with participants assigned to tocilizumab or placebo. |
| Double blinding | The registry identifies the trial as double-blind. |
| Parallel design | The two interventions are evaluated in parallel study arms. |
| Ordinal clinical status | The primary endpoint uses a 7-category clinical-status scale at Day 28. |
| Van Elteren test | Used for the primary clinical-status analysis and several secondary final-value comparisons. |
| Log-rank test | Used for several time-to-event outcomes through Day 28. |
| Hazard ratio | Used to summarize relative treatment effects for time-to-event outcomes. |
| Cochran-Mantel-Haenszel test | Used for categorical outcomes including mechanical ventilation, ICU stay, and mortality. |
| Confidence intervals | Two-sided 95% intervals accompany the posted effect estimates. |
| mITT analysis | Analyses include randomized participants who received any amount of study medication, grouped according to randomized assignment. |
| Safety population | Serious adverse events are reported using safety-evaluable populations. |
| Multiplicity | One primary endpoint is accompanied by multiple secondary analyses, requiring careful interpretation of nominal p-values. |
26. A Statistical Reading of the Complete Results
The primary analysis reports a median difference of −1.0 for clinical status at Day 28, with a 95% confidence interval of −2.5 to 0.0 and P = 0.3600. The result is based on a Van Elteren test in the mITT population.
The secondary analyses provide a more heterogeneous statistical picture because they address different estimands. Time-to-event analyses report hazard ratios, categorical outcomes report weighted percentage differences, and several final-value outcomes report median differences. These estimates cannot be placed on a single numerical scale or interpreted as though they were interchangeable.
Several time-to-event endpoints have hazard ratios above 1, including time to clinical improvement (1.448), time to hospital discharge or "ready for discharge" (1.350), and time to recovery (1.307). Time to clinical failure has a hazard ratio below 1 (0.790). The meaning of those directions follows from the definitions of the events, not from the numerical values alone.
The categorical analyses similarly require endpoint-specific interpretation. The reported weighted percentage difference for mortality at Day 28 is 0.3, while the comparison among participants not in ICU at baseline for ICU stay has a weighted percentage difference of −14.8. Those are different estimands and should not be combined into a single summary effect.
From a statistical-methodology perspective, the most important lesson is that a clinical trial is not represented by one p-value. COVACTA contains an ordinal primary endpoint, time-to-event secondary endpoints, categorical secondary endpoints, rank-based final-value comparisons, and a separately defined safety population. The interpretation must preserve those distinctions.
27. Related Tutorials
Learn more about the methods used in this trial:
28. Related Calculators
29. Sources
- ClinicalTrials.gov: COVACTA, NCT04320615. Official registry record used for the trial data and statistical-analysis fields presented on this page.
- Linked publication: PubMed record for PMID 35475258.
- Linked publication: PubMed record for PMID 34612846.
- Linked publication: PubMed record for PMID 33631066.
Continue with the statistical methods behind COVACTA
Explore the broader Clinical Biostats tutorial and calculator collections for survival analysis, categorical methods, confidence intervals, randomization, and nonparametric clinical-trial methods.
30. Record Summary
COVACTA provides a useful example of how statistical analysis should follow the structure of the clinical questions being asked. The randomized, double-blind, parallel phase 3 design compares tocilizumab with placebo in patients with severe COVID-19 pneumonia. The registered primary endpoint assesses clinical status at Day 28 using a 7-category ordinal scale, with the posted formal analysis using a Van Elteren test and reporting a median difference of −1.0, a 95% confidence interval of −2.5 to 0.0, and P = 0.3600.
The secondary results demonstrate three principal statistical frameworks: log-rank testing with hazard ratios for time-to-event outcomes, Cochran-Mantel-Haenszel testing with weighted percentage differences for categorical outcomes, and Van Elteren testing with median differences for rank-based final-value comparisons. The safety results use separate safety-evaluable populations and are reported descriptively in the ClinicalTrials.gov record.
The most important statistical lesson is that effect estimates, confidence intervals, p-values, endpoint definitions, and analysis populations must be interpreted together. A hazard ratio is not a percentage difference, a p-value is not an effect size, and a secondary nominal p-value does not automatically represent an independently confirmatory finding. Preserving those distinctions is essential to a precise reading of the COVACTA results.