← Clinical Trials
COVID-19 Pneumonia Phase 3 Double-Blind NCT04320615

COVACTA: Complete Statistical Analysis of Tocilizumab in Severe COVID-19 Pneumonia

An independent statistical review of the randomized phase 3 COVACTA trial evaluating tocilizumab versus placebo in patients with severe COVID-19 pneumonia, with emphasis on the registered clinical-status endpoint, time-to-event analyses, categorical outcomes, nonparametric comparisons, confidence intervals, and interpretation of reported p-values.

COVACTA  ·  Phase 3  ·  Randomized  ·  Double-blind  ·  Completed
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the registry-reported COVACTA registry data.

1. Trial at a Glance

COVACTA was a randomized, double-blind, parallel phase 3 study evaluating the safety and efficacy of tocilizumab in patients with severe COVID-19 pneumonia. The registry reports an enrollment of 452 participants, two study arms, one registered primary endpoint, 13 posted outcome measures, and 15 posted statistical analyses.

452
Enrollment
Randomized phase 3 study
2
Study Arms
Tocilizumab vs placebo
1
Primary Endpoint
Clinical status at Day 28
15
Statistical Analyses
1 primary; 14 secondary
FeatureCOVACTA
Trial nameCOVACTA
ClinicalTrials.gov identifierNCT04320615
Brief titleA Study to Evaluate the Safety and Efficacy of Tocilizumab in Patients With Severe COVID-19 Pneumonia
Therapeutic areaInfectious Disease
ConditionCOVID-19 Pneumonia
PhasePhase 3
StatusCompleted
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
Enrollment452
InterventionsTocilizumab (TCZ); placebo
Lead sponsorHoffmann-La Roche
Sponsor typeIndustry
Results postedYes
Outcome measures posted13
Statistical analyses posted15

2. Clinical Question

The clinical question was whether tocilizumab produced a different clinical outcome than placebo in patients with severe COVID-19 pneumonia. The registered primary endpoint assessed clinical status at Day 28 using a 7-category ordinal scale.

Population

Patients with severe COVID-19 pneumonia, as described by the trial's brief title and condition in the ClinicalTrials.gov record.

Intervention

Tocilizumab (TCZ).

Comparator

Placebo.

Primary question

How did clinical status at Day 28 compare between the randomized tocilizumab and placebo groups?

3. Trial Design

01
Randomize452 participants enrolled
02
Parallel groupsTwo randomized arms
03
Double-blindTocilizumab or placebo
04
AssessClinical and safety outcomes
05
Day 28Primary clinical-status assessment
Allocation
The registry identifies the allocation as randomized.
Design model
The trial used a parallel design with two study arms.
Masking
The registry identifies the trial as double-blind.
Primary purpose
The registry identifies the primary purpose as treatment.
ARM 1

Tocilizumab (TCZ)

  • Intervention classified as a drug.
  • Compared with the placebo arm in the posted statistical analyses.
  • mITT analyses grouped participants according to treatment assigned at randomization.
ARM 2

Placebo

  • Comparator classified as a drug.
  • Compared with the tocilizumab arm in the posted statistical analyses.
  • mITT analyses grouped participants according to treatment assigned at randomization.

4. Trial Timeline

April 3, 2020

Study start

The registry lists 2020-04-03 as the study start date.

June 24, 2020

Primary completion

The registry lists 2020-06-24 as the primary completion date.

Completed

Registry status

The ClinicalTrials.gov record identifies COVACTA as COMPLETED and reports results.

5. Analysis Population

The posted statistical analyses repeatedly identify the modified intent-to-treat (mITT) population. The registry defines this population as all randomized participants who received any amount of study medication, with participants grouped according to the treatment assigned at randomization.

mITT definition reported by the registry
Randomized + received any amount of study medication → analyzed according to randomized treatment assignment

This distinction is important. The analysis population is not simply every person enrolled in the study. The registry-reported definition requires randomization and receipt of some study medication, while preserving treatment assignment as the basis for grouping.

The same mITT definition is reported for the primary analysis and the posted secondary analyses. This creates a consistent analytical population across the efficacy comparisons reported in the ClinicalTrials.gov record.

6. Primary Endpoint

EndpointRegistry definition / time frameAnalysis
Clinical Status Assessed Using a 7-Category Ordinal Scale at Day 28 (Week 4) Day 28. Clinical status was assessed using a 7-category ordinal scale. The registry definition begins: 1. Discharged (or "ready for discharge"); 2. Non-ICU hospital ward (or "ready for hospital ward") not requiring supplemental oxygen; 3. Non-ICU hospital ward (or "ready for hospital ward") requiring supplemental oxygen; 4. ICU or non-ICU hospital ward, requiring non-invasive ventilation or high-flow oxygen. Van Elteren test

The ClinicalTrials.gov record classifies the posted analysis as a binary endpoint for the statistical-analysis record, while the registered endpoint itself is explicitly described as a 7-category ordinal scale. The analysis reported in the registry uses a Van Elteren test and reports a median difference in final values.

Important endpoint distinction: the clinical-status endpoint is ordinal by its registered definition. The ClinicalTrials.gov record labels the posted outcome unit as percentage of participants and the endpoint type as binary, but the reported method is the nonparametric Van Elteren test with a median-difference effect measure. These registry fields should be reported as reported in the registry rather than silently replacing one field with another.

7. Primary Result

The registry reports one formal primary-endpoint analysis comparing the tocilizumab mITT arm with the placebo mITT arm at Day 28.

Clinical status at Day 28

Median difference −1.0

95% CI: −2.5 to 0.0   ·   P = 0.3600

Analysis: Van Elteren test  ·  two-sided 95% confidence interval

Primary endpointTocilizumab vs placeboMethod95% CIP-value
Clinical Status Assessed Using a 7-Category Ordinal Scale at Day 28 (Week 4) Median difference: −1.0 Van Elteren test −2.5 to 0.0 0.3600
Clinical Biostats interpretation

The reported median difference of −1.0 is the registry's reported effect measure for the final clinical-status values, comparing the tocilizumab mITT arm with the placebo mITT arm. Because the underlying endpoint is a 7-category clinical-status scale, the estimate should be interpreted in the context of that ordinal outcome rather than as a hazard ratio or a conventional mean difference.

The 95% confidence interval of −2.5 to 0.0 describes the statistical uncertainty around the reported median-difference estimate under the analysis framework. The interval reaches 0.0, so the registry-reported estimate is compatible with no difference as represented by this effect measure. The confidence interval does not describe the range of individual patient outcomes.

The P-value of 0.3600 is a measure of evidence against the statistical null hypothesis represented by the test; it is not a measure of effect size, clinical importance, or the probability that either treatment is effective. A p-value should therefore be interpreted together with the effect estimate and confidence interval.

The registry identifies the test as two-sided through the confidence-interval specification. The analysis population is the mITT population, not the full enrolled population. Because the registered outcome is ordinal, the choice of a nonparametric method is an important part of the statistical story.

8. Secondary Endpoint Results

The registry contains 14 posted secondary statistical analyses. They cover time-to-event outcomes, categorical outcomes, and final-value comparisons. All of the analyses in the ClinicalTrials.gov record use the mITT population unless a baseline subgroup is explicitly stated in the comparison.

Secondary endpoint Time frame Method Effect estimate 95% CI P-value
Time to Clinical Improvement (TTCI), defined as a National Early Warning Score 2 (NEWS2) of ≤2 maintained for 24 Hours Up to Day 28 Log-rank test HR 1.448 1.01 to 2.08 0.0443
Time to Improvement of at Least 2 Categories Relative to Baseline on a 7-Category Ordinal Scale of Clinical Status Up to Day 28 Log-rank test HR 1.263 0.97 to 1.64 0.0820
Time to Hospital Discharge or "Ready for Discharge" Up to Day 28 Log-rank test HR 1.350 1.02 to 1.79 0.0370
Incidence of Mechanical Ventilation by Day 28 Up to Day 28 Cochran-Mantel-Haenszel test Weighted % difference −6.2 −13.8 to 1.4 0.0996
Incidence of Mechanical Ventilation by Day 28: TCZ - No Mechanical Ventilation at Baseline vs Placebo - No Mechanical Ventilation at Baseline Up to Day 28 Cochran-Mantel-Haenszel test Weighted % difference −8.9 −20.7 to 3.0 0.1355
Ventilator-Free Days to Day 28 Up to Day 28 Van Elteren test Median difference 5.5 −2.8 to 13.0 0.3202
Incidence of Intensive Care Unit (ICU) Stay by Day 28 (Week 4) Up to Day 28 Cochran-Mantel-Haenszel test Weighted % difference −5.7 −13.7 to 2.2 0.1514
Incidence of ICU Stay by Day 28: TCZ - Not in ICU at Baseline vs Placebo - Not in ICU at Baseline Up to Day 28 Cochran-Mantel-Haenszel test Weighted % difference −14.8 −28.6 to −1.0 0.0290
Duration of ICU Stay to Day 28 (Week 4) Up to Day 28 Van Elteren test Median difference −5.8 −15.0 to 2.9 0.0454
Clinical Status Assessed Using a 7-Category Ordinal Scale at Day 14 Day 14 Van Elteren test Median difference −1.0 −2.0 to 0.5 0.0548
Time to Clinical Failure to Day 28 (Week 4) Up to Day 28 Log-rank test HR 0.790 0.57 to 1.10 0.1627
Mortality Rate at Day 28 (Week 4) Day 28 Cochran-Mantel-Haenszel test Weighted % difference 0.3 −7.6 to 8.2 0.9410
Time to Recovery to Day 28 (Week 4) Up to Day 28 Log-rank test HR 1.307 1.00 to 1.72 0.0528
Duration of Supplemental Oxygen to Day 28 (Week 4) Up to Day 28 Van Elteren test Median difference −1.5 −9.0 to 0.5 0.0477
How to read the table: the direction of an estimate depends on the endpoint and the effect measure. A hazard ratio above 1 indicates a higher estimated event rate in the tocilizumab group relative to placebo when the event is the endpoint being modeled; a negative difference indicates a lower tocilizumab-minus-placebo difference for the corresponding difference measure. The clinical meaning of direction therefore has to be established from the endpoint definition rather than from the sign alone.

9. Time-to-Event Results

Several secondary endpoints were analyzed with the log-rank test and reported using hazard ratios. This is a coherent analytical pattern for endpoints defined by the time until an event occurs within the specified follow-up window.

Time to Clinical Improvement

HR 1.448

95% CI: 1.01–2.08   ·   P = 0.0443

Time frame: Up to Day 28

Clinical Biostats interpretation

For time to clinical improvement, the reported hazard ratio of 1.448 means that the estimated instantaneous rate of reaching the defined improvement event was 1.448 times as high in the tocilizumab group as in the placebo group under the time-to-event analysis. Expressed as a simple relative interpretation, this is an estimated 44.8% higher instantaneous event rate, not a statement that 44.8% more patients necessarily improved.

The 95% CI of 1.01 to 2.08 indicates uncertainty around the estimated hazard ratio. The lower limit is just above 1.00, while the upper limit is substantially higher, so the interval is not especially narrow. The estimate therefore should not be separated from its uncertainty.

The P-value of 0.0443 describes evidence against the relevant null hypothesis in the reported log-rank analysis. It does not quantify the magnitude or clinical importance of the difference. It also does not establish that the treatment effect is exactly 44.8%.

As with any hazard ratio, interpretation depends on the time-to-event framework and censoring process. A single HR is most straightforward when the relative hazards are reasonably represented by the model over time; the ClinicalTrials.gov record does not provide a separate assessment of the proportional-hazards assumption.

Time to Hospital Discharge or "Ready for Discharge"

HR 1.350

95% CI: 1.02–1.79   ·   P = 0.0370

Time frame: Up to Day 28

Clinical Biostats interpretation

The hazard ratio of 1.350 indicates a higher estimated instantaneous rate of reaching the defined discharge or "ready for discharge" event in the tocilizumab group under the reported analysis. The corresponding simple relative interpretation is a 35.0% higher estimated event rate, not a 35.0% absolute increase in the number of discharged participants.

The 95% CI of 1.02 to 1.79 indicates that the estimate has meaningful uncertainty, although the interval lies above 1.00. The p-value of 0.0370 is evidence against the null hypothesis represented by the log-rank test, but it should not be treated as a measure of the size of the treatment effect.

Time to Clinical Failure

HR 0.790

95% CI: 0.57–1.10   ·   P = 0.1627

Time frame: Up to Day 28

Clinical Biostats interpretation

The reported hazard ratio of 0.790 corresponds to an estimated instantaneous clinical-failure rate that is 79.0% of the placebo-group rate under the reported time-to-event model. A simple relative interpretation is therefore a 21.0% lower estimated hazard, but this is not an absolute risk reduction and does not mean that 21.0% of participants avoided clinical failure.

The 95% CI of 0.57 to 1.10 spans 1.00, indicating considerable uncertainty about the direction and magnitude of the underlying relative hazard. The P-value of 0.1627 is not a measure of effect size and should be considered alongside this confidence interval.

Time to Recovery

HR 1.307

95% CI: 1.00–1.72   ·   P = 0.0528

Time frame: Up to Day 28

Clinical Biostats interpretation

The hazard ratio of 1.307 represents a higher estimated instantaneous rate of the defined recovery event in the tocilizumab group. A simple relative reading is a 30.7% higher estimated event rate, but this should not be converted into a statement about the percentage of patients who recovered.

The 95% CI of 1.00 to 1.72 reaches 1.00 at its lower boundary. The p-value of 0.0528 likewise should not be treated as an effect-size statistic. Together, the estimate and interval indicate that the magnitude of the time-to-recovery difference is uncertain.

Educational note: a Kaplan-Meier curve cannot be reconstructed responsibly from these summary hazard ratios and confidence intervals alone. The ClinicalTrials.gov record does not contain the individual event and censoring times needed to construct a valid curve.

10. Categorical Endpoint Results

COVACTA also used the Cochran-Mantel-Haenszel test for several categorical outcomes. The registry reports weighted percentage differences with two-sided 95% confidence intervals.

OutcomeWeighted % difference95% CIP-value
Incidence of Mechanical Ventilation by Day 28−6.2−13.8 to 1.40.0996
Mechanical Ventilation by Day 28 among participants with no mechanical ventilation at baseline−8.9−20.7 to 3.00.1355
Incidence of ICU Stay by Day 28−5.7−13.7 to 2.20.1514
ICU Stay by Day 28 among participants not in ICU at baseline−14.8−28.6 to −1.00.0290
Mortality Rate at Day 280.3−7.6 to 8.20.9410
Clinical Biostats interpretation

A percentage difference is an absolute-scale comparison, unlike a hazard ratio. For example, the reported weighted percentage difference of −6.2 for mechanical ventilation by Day 28 indicates a lower weighted percentage in the tocilizumab group by 6.2 percentage points under the reported analysis framework. It does not mean a 6.2% relative reduction.

The 95% CI of −13.8 to 1.4 crosses 0, illustrating why the point estimate alone is insufficient. The p-value of 0.0996 does not measure the clinical size of the observed difference.

For the subgroup comparison restricted to participants not in the ICU at baseline, the reported weighted percentage difference was −14.8, with a 95% CI of −28.6 to −1.0 and P = 0.0290. This is a baseline-defined subgroup analysis and should not automatically be treated as equivalent to the overall randomized comparison.

11. Nonparametric Comparisons

The registry reports the Van Elteren test for the primary Day 28 clinical-status analysis and for several secondary final-value comparisons. The method is a stratified nonparametric extension of the Wilcoxon rank-sum approach and is useful when comparing ordered or continuous outcomes without requiring a normal-distribution assumption.

EndpointMedian difference95% CIP-value
Clinical status at Day 28−1.0−2.5 to 0.00.3600
Ventilator-free days to Day 285.5−2.8 to 13.00.3202
Duration of ICU stay to Day 28−5.8−15.0 to 2.90.0454
Clinical status at Day 14−1.0−2.0 to 0.50.0548
Duration of supplemental oxygen to Day 28−1.5−9.0 to 0.50.0477

The word median in these effect measures matters. A median difference is not the same as a difference in means, and it should not be interpreted as though every participant experienced the estimated difference. It summarizes a distributional comparison on the reported final values.

Why the Van Elteren test is informative here

The trial contains outcomes that are naturally ordered, skewed, or otherwise not well summarized by a simple normal-theory mean comparison. A rank-based nonparametric approach can compare treatment groups without requiring the outcome distribution to be normal. The Van Elteren formulation also allows the comparison to account for stratification when the relevant strata are available.

For the reported duration of ICU stay, the median difference was −5.8, with a 95% CI of −15.0 to 2.9 and P = 0.0454. The confidence interval includes 0 even though the reported p-value is below 0.05, illustrating why the exact definition of an effect measure and its confidence interval should be examined rather than reducing the result to a binary "significant/not significant" label.

12. Secondary Result: Ventilator-Free Days

Ventilator-Free Days to Day 28

Median difference 5.5 days

95% CI: −2.8 to 13.0   ·   P = 0.3202

Analysis: Van Elteren test

Clinical Biostats interpretation

The reported median difference of 5.5 days is a final-value comparison between the tocilizumab and placebo mITT groups. The positive point estimate indicates a higher median ventilator-free-days value in the tocilizumab group under the reported effect-measure convention.

The 95% CI of −2.8 to 13.0 is broad relative to the point estimate and crosses zero. The p-value of 0.3202 does not measure the size of the observed difference. The result should therefore be read as an estimate with substantial uncertainty rather than as a precise statement about the number of ventilator-free days gained by an individual participant.

13. Secondary Result: ICU Outcomes

ICU incidence

Overall weighted percentage difference: −5.7; 95% CI −13.7 to 2.2; P = 0.1514.

Not in ICU at baseline

Weighted percentage difference: −14.8; 95% CI −28.6 to −1.0; P = 0.0290.

Duration of ICU stay

Median difference: −5.8; 95% CI −15.0 to 2.9; P = 0.0454.

Analytical distinction

Incidence is categorical; duration is a final-value comparison. The registry therefore uses different statistical methods and effect measures.

These results illustrate a central principle of clinical-trial statistics: related clinical questions can require different estimands and methods. Whether a participant experiences an ICU stay is a categorical outcome; the duration of ICU stay is a distributional outcome. Treating the two as interchangeable would obscure the statistical structure of the data.

14. Secondary Result: Clinical Status at Day 14

Clinical Status at Day 14

Median difference −1.0

95% CI: −2.0 to 0.5   ·   P = 0.0548

Analysis: Van Elteren test

The Day 14 endpoint uses the same general clinical-status framework as the registered Day 28 primary endpoint, but the time frame is earlier. The reported estimate is accompanied by a confidence interval that crosses zero and a p-value of 0.0548.

Why the Day 14 result should not be substituted for the primary endpoint

The Day 14 analysis addresses an earlier time point and is a secondary endpoint. It should therefore be interpreted as its own analysis rather than as an alternative definition of the primary endpoint. A p-value near a conventional threshold does not transform a secondary endpoint into a primary endpoint, and it does not establish a treatment effect by itself.

15. Mortality at Day 28

Mortality Rate at Day 28

Weighted % difference 0.3

95% CI: −7.6 to 8.2   ·   P = 0.9410

Analysis: Cochran-Mantel-Haenszel test

Clinical Biostats interpretation

The reported weighted percentage difference of 0.3 is close to zero. The 95% confidence interval extends from −7.6 to 8.2, indicating substantial uncertainty around the magnitude and direction of the treatment-group difference on this scale.

The reported P-value of 0.9410 provides little evidence against the null hypothesis represented by this comparison. It should not, however, be converted into a probability that the treatments are identical, nor does it establish equivalence.

16. Statistical Methodology

Van Elteren test

The Van Elteren test is a stratified, rank-based nonparametric procedure. It extends the logic of the Wilcoxon rank-sum test by allowing comparisons to be performed across strata and then combined. In COVACTA, the registry identifies the Van Elteren test for the primary clinical-status endpoint and several secondary final-value outcomes.

Conceptual interpretation
Within-stratum rank comparisons → weighted combination across strata → overall treatment comparison

The method is particularly useful when the outcome is ordinal or when a distribution-free comparison is preferred. The exact weighting and stratification details should be taken from the trial's statistical-analysis specification when those details are available.

Log-rank test

The log-rank test is used for comparing time-to-event distributions between treatment groups. COVACTA uses the log-rank test for time to clinical improvement, time to improvement of at least two categories, time to hospital discharge or "ready for discharge," time to clinical failure, and time to recovery.

Time-to-event framework
Treatment comparison → event times + censoring → survival / event-time comparison

The approach uses information from the timing of events rather than only whether an event eventually occurred. Participants who have not experienced the event during available follow-up can contribute censored information.

Cochran-Mantel-Haenszel test

The Cochran-Mantel-Haenszel test is a family of methods for comparing categorical outcomes while accounting for stratification. In the registry-reported COVACTA analyses, it is used for mechanical ventilation, ICU stay incidence, and Day 28 mortality.

Hazard ratio

The registry reports hazard ratios for time-to-event endpoints. A hazard ratio compares the estimated instantaneous event rates between groups within the fitted time-to-event framework.

Hazard-ratio interpretation
HR > 1 → higher estimated instantaneous event rate in the tocilizumab group for the defined event
HR < 1 → lower estimated instantaneous event rate in the tocilizumab group for the defined event

The direction must always be interpreted relative to the event being analyzed. A hazard ratio below 1 is not inherently favorable or unfavorable without knowing whether the event represents improvement, failure, hospitalization, recovery, or another outcome.

Confidence intervals

The posted analyses report two-sided 95% confidence intervals. For hazard ratios, a value of 1 represents the conventional no-relative-hazard-difference point. For difference measures, 0 represents the conventional no-difference point. These reference values are useful for interpreting the interval, but the clinical meaning still depends on the endpoint and estimand.

17. Statistical Methods Explained

Why was the Van Elteren test used?

The registry identifies a Van Elteren test for the primary clinical-status analysis and several secondary final-value outcomes. This is a nonparametric, stratified method that is appropriate for rank-based comparisons of ordered or non-normally distributed outcomes. Its use is consistent with the fact that clinical status is represented by an ordinal scale and several secondary measures are not naturally characterized by a normal-distribution assumption.

What does a hazard ratio of 1.448 mean?

For the time-to-clinical-improvement endpoint, HR 1.448 means the estimated instantaneous rate of reaching the defined improvement event was 1.448 times the corresponding rate in the placebo group under the reported time-to-event analysis. It is not a 44.8 percentage-point increase in improvement, and it is not the probability that a participant improved.

Why is a hazard ratio different from a percentage difference?

A hazard ratio incorporates the timing of events and censoring in a time-to-event analysis. A weighted percentage difference compares categorical outcome percentages. A value of −6.2 therefore has a fundamentally different interpretation from HR 1.448, even though both are treatment-effect measures.

What does a 95% confidence interval tell us?

The interval describes uncertainty around the estimated treatment effect under the statistical model and sampling framework. For HR 1.448, the interval is 1.01 to 2.08. For a weighted percentage difference of −6.2, the interval is −13.8 to 1.4. The relevant no-effect reference value is 1 for a hazard ratio and 0 for a difference measure.

Why does the p-value not measure effect size?

A p-value summarizes how compatible the observed data are with a specified null hypothesis under the test procedure. It depends on the estimate, variability, sample information, and statistical model. It does not tell us how large the treatment effect is. That is why COVACTA's reported p-values should be read together with the effect estimates and confidence intervals.

Why does endpoint direction matter?

For an improvement endpoint, a hazard ratio above 1 may correspond to a higher rate of reaching improvement. For a failure endpoint, a hazard ratio below 1 may correspond to a lower rate of failure. The same numerical hazard ratio can therefore have different clinical interpretations depending on the event definition.

Why should secondary p-values not be treated as a list of independent discoveries?

COVACTA has one registered primary endpoint and multiple secondary analyses. Each secondary analysis asks a different statistical question, and repeated testing creates a broader multiplicity problem when many hypotheses are considered together. A collection of p-values should therefore not be converted into a simple count of "significant" findings without understanding the prespecified testing strategy.

18. Endpoint-by-Endpoint Statistical Map

Endpoint familyExample endpointMethodEffect measure
Ordinal / final value Clinical status at Day 28 Van Elteren test Median difference
Time-to-event Time to clinical improvement Log-rank test Hazard ratio
Time-to-event Time to hospital discharge or "ready for discharge" Log-rank test Hazard ratio
Categorical Incidence of mechanical ventilation Cochran-Mantel-Haenszel test Weighted % difference
Categorical Mortality rate at Day 28 Cochran-Mantel-Haenszel test Weighted % difference
Nonparametric final value Ventilator-free days Van Elteren test Median difference
Nonparametric final value Duration of supplemental oxygen Van Elteren test Median difference

This method map is one of the most useful statistical features of the trial. COVACTA does not force every clinical outcome into a single analysis framework. Instead, the posted analyses distinguish between time-to-event, categorical, and rank-based outcomes.

19. Safety

The ClinicalTrials.gov record reports serious adverse events by randomized arm using a safety-evaluable population.

Safety-evaluable armParticipants with serious adverse eventsAt riskAffected / at risk
Placebo 64 143 64/143
Tocilizumab (TCZ) 116 295 116/295

The reported serious-adverse-event figures are presented as affected participants over the safety-evaluable population in each arm. They should not be substituted for the efficacy analysis population because the ClinicalTrials.gov record explicitly identifies these as safety-evaluable populations.

Clinical Biostats interpretation

The serious-adverse-event data are exposure-related safety information rather than an efficacy endpoint. The denominators differ between the two reported safety-evaluable groups, so the raw event counts of 64 and 116 should not be compared as though they represented equally sized populations.

The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or p-value for these serious-adverse-event counts. The appropriate interpretation is therefore descriptive rather than a claim of a statistically tested difference.

20. Missing Data, Censoring, and Analysis Assumptions

The ClinicalTrials.gov record identifies mITT populations and time-to-event methods, but they do not provide detailed missing-data or imputation rules for the individual endpoints. It is therefore important not to infer an imputation procedure that is not present in the ClinicalTrials.gov record.

Time-to-event outcomes

Log-rank analyses use event timing and can incorporate censored observations. The ClinicalTrials.gov record does not specify the complete censoring rules.

Final-value outcomes

Van Elteren analyses compare final values using a nonparametric framework. The ClinicalTrials.gov record does not specify a separate imputation strategy.

Categorical outcomes

Cochran-Mantel-Haenszel analyses compare categorical outcomes with a stratified framework. The ClinicalTrials.gov record does not provide additional missing-outcome handling details.

Model assumptions

The ClinicalTrials.gov record does not report a separate proportional-hazards assessment for the hazard-ratio analyses.

21. Multiplicity and Multiple Secondary Analyses

The registry reports one primary endpoint and multiple secondary analyses. This structure matters because the nominal p-value for any individual secondary endpoint does not, by itself, describe the probability of making at least one false-positive conclusion across the complete collection of hypotheses.

Analysis layerRegistry informationStatistical interpretation
Primary 1 registered primary endpoint; 1 formal primary statistical analysis The primary result should be interpreted as the principal prespecified efficacy comparison represented in the ClinicalTrials.gov record.
Secondary 14 posted secondary statistical analyses These address multiple additional clinical questions and should be interpreted individually and collectively rather than as isolated tests.
Safety Serious adverse events reported descriptively by arm The ClinicalTrials.gov record does not provide a formal inferential comparison for these safety counts.

No non-inferiority margin, factorial design, crossover analysis, Bayesian analysis, or interim-analysis procedure is included in the ClinicalTrials.gov record. Those design topics are therefore not used as assumptions in this analysis.

22. What the Hazard Ratio Does — and Does Not — Mean

Example: HR 1.448

The reported hazard ratio of 1.448 for time to clinical improvement means that, under the reported time-to-event analysis, the estimated instantaneous rate of reaching the defined improvement event was 1.448 times that of the placebo group.

It does not mean that 44.8% more participants improved, that 44.8% of participants benefited, or that each individual participant experienced a 44.8% increase in the chance of improvement.

Example: HR 0.790

The reported hazard ratio of 0.790 for time to clinical failure corresponds to a 21.0% lower estimated instantaneous failure hazard under a simple relative interpretation. It does not mean that 21.0% of participants avoided clinical failure.

Why the confidence interval matters

The confidence interval provides information about precision. For HR 0.790, the 95% CI is 0.57 to 1.10, which spans the conventional no-relative-difference value of 1.00. For HR 1.448, the 95% CI is 1.01 to 2.08. The intervals convey information that cannot be recovered from the p-values alone.

23. P-Values in Context

The COVACTA results include p-values ranging from 0.0290 to 0.9410 across the statistical analyses posted on ClinicalTrials.gov. These values should not be interpreted as a ranking of clinical importance.

P-valueEndpointEffect measure95% CI
0.0290ICU Stay by Day 28 among participants not in ICU at baselineWeighted % difference −14.8−28.6 to −1.0
0.0370Time to Hospital Discharge or "Ready for Discharge"HR 1.3501.02 to 1.79
0.0443Time to Clinical ImprovementHR 1.4481.01 to 2.08
0.0454Duration of ICU StayMedian difference −5.8−15.0 to 2.9
0.0477Duration of Supplemental OxygenMedian difference −1.5−9.0 to 0.5
0.9410Mortality Rate at Day 28Weighted % difference 0.3−7.6 to 8.2

The table demonstrates why p-values should not be interpreted without the effect measure and confidence interval. The reported duration of ICU stay analysis has P = 0.0454, for example, while its 95% confidence interval for the median difference extends from −15.0 to 2.9. The inferential details therefore contain more information than the threshold classification alone.

24. Limitations

25. Why This Trial Matters Statistically

COVACTA is a useful teaching example because its posted analyses show how one randomized clinical trial can require several distinct statistical approaches. The primary clinical-status outcome uses a nonparametric rank-based method, several secondary outcomes use time-to-event methods, and categorical outcomes use a Cochran-Mantel-Haenszel framework.

ConceptHow it appears in COVACTA
RandomizationThe trial is identified as randomized, with participants assigned to tocilizumab or placebo.
Double blindingThe registry identifies the trial as double-blind.
Parallel designThe two interventions are evaluated in parallel study arms.
Ordinal clinical statusThe primary endpoint uses a 7-category clinical-status scale at Day 28.
Van Elteren testUsed for the primary clinical-status analysis and several secondary final-value comparisons.
Log-rank testUsed for several time-to-event outcomes through Day 28.
Hazard ratioUsed to summarize relative treatment effects for time-to-event outcomes.
Cochran-Mantel-Haenszel testUsed for categorical outcomes including mechanical ventilation, ICU stay, and mortality.
Confidence intervalsTwo-sided 95% intervals accompany the posted effect estimates.
mITT analysisAnalyses include randomized participants who received any amount of study medication, grouped according to randomized assignment.
Safety populationSerious adverse events are reported using safety-evaluable populations.
MultiplicityOne primary endpoint is accompanied by multiple secondary analyses, requiring careful interpretation of nominal p-values.

26. A Statistical Reading of the Complete Results

The primary analysis reports a median difference of −1.0 for clinical status at Day 28, with a 95% confidence interval of −2.5 to 0.0 and P = 0.3600. The result is based on a Van Elteren test in the mITT population.

The secondary analyses provide a more heterogeneous statistical picture because they address different estimands. Time-to-event analyses report hazard ratios, categorical outcomes report weighted percentage differences, and several final-value outcomes report median differences. These estimates cannot be placed on a single numerical scale or interpreted as though they were interchangeable.

Several time-to-event endpoints have hazard ratios above 1, including time to clinical improvement (1.448), time to hospital discharge or "ready for discharge" (1.350), and time to recovery (1.307). Time to clinical failure has a hazard ratio below 1 (0.790). The meaning of those directions follows from the definitions of the events, not from the numerical values alone.

The categorical analyses similarly require endpoint-specific interpretation. The reported weighted percentage difference for mortality at Day 28 is 0.3, while the comparison among participants not in ICU at baseline for ICU stay has a weighted percentage difference of −14.8. Those are different estimands and should not be combined into a single summary effect.

From a statistical-methodology perspective, the most important lesson is that a clinical trial is not represented by one p-value. COVACTA contains an ordinal primary endpoint, time-to-event secondary endpoints, categorical secondary endpoints, rank-based final-value comparisons, and a separately defined safety population. The interpretation must preserve those distinctions.

27. Related Tutorials

Learn more about the methods used in this trial:

28. Related Calculators

29. Sources

Source discipline: The numerical results on this page are limited to the registry-reported COVACTA trial data. The linked PubMed records are provided as source-navigation links; no additional numerical results from those publications have been introduced into the analysis.

Continue with the statistical methods behind COVACTA

Explore the broader Clinical Biostats tutorial and calculator collections for survival analysis, categorical methods, confidence intervals, randomization, and nonparametric clinical-trial methods.

30. Record Summary

COVACTA provides a useful example of how statistical analysis should follow the structure of the clinical questions being asked. The randomized, double-blind, parallel phase 3 design compares tocilizumab with placebo in patients with severe COVID-19 pneumonia. The registered primary endpoint assesses clinical status at Day 28 using a 7-category ordinal scale, with the posted formal analysis using a Van Elteren test and reporting a median difference of −1.0, a 95% confidence interval of −2.5 to 0.0, and P = 0.3600.

The secondary results demonstrate three principal statistical frameworks: log-rank testing with hazard ratios for time-to-event outcomes, Cochran-Mantel-Haenszel testing with weighted percentage differences for categorical outcomes, and Van Elteren testing with median differences for rank-based final-value comparisons. The safety results use separate safety-evaluable populations and are reported descriptively in the ClinicalTrials.gov record.

The most important statistical lesson is that effect estimates, confidence intervals, p-values, endpoint definitions, and analysis populations must be interpreted together. A hazard ratio is not a percentage difference, a p-value is not an effect size, and a secondary nominal p-value does not automatically represent an independently confirmatory finding. Preserving those distinctions is essential to a precise reading of the COVACTA results.

Clinical Biostats methodology: A trial-results page should reconstruct the statistical structure of the reported evidence rather than reduce a multidimensional clinical trial to a single p-value. The objective is to connect each endpoint to its estimand, analysis method, uncertainty measure, and appropriate interpretation while keeping reported results separate from statistical explanation.