← Clinical Trials
COVID-19 Phase 3 Completed NCT04492475

ACTT-3: Complete Statistical Analysis of Interferon Beta-1a in COVID-19

An independent statistical analysis of the randomized, double-blind phase 3 ACTT-3 trial comparing remdesivir plus interferon beta-1a with remdesivir plus placebo in participants with COVID-19, with emphasis on time to recovery, subgroup hazard ratios, ordinal-scale outcomes, and safety.

Trial period: August 5, 2020 – December 21, 2020  ·  Enrollment: 969  ·  Primary purpose: Treatment
Scope of this record

This page separates reported trial results from statistical interpretation. Every numerical result on this page is taken from the ClinicalTrials.gov record for ACTT-3. The registry provides the official trial record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

ACTT-3 was a completed phase 3 randomized, parallel-group, double-blind clinical trial evaluating remdesivir plus interferon beta-1a versus remdesivir plus placebo in COVID-19. The registry reports 969 enrolled participants and four registered primary endpoints, all centered on time to recovery during Day 1 through Day 29.

969
Enrollment
Phase 3 trial
2
Arms
Parallel randomized design
0.99
Primary HR
95% CI 0.87–1.13
0.880
Primary P-value
Two-sided
FeatureACTT-3
Trial nameAdaptive COVID-19 Treatment Trial 3 (ACTT-3)
NCT identifierNCT04492475
PhasePhase 3
ConditionCOVID-19
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
Enrollment969
Primary endpoints4 registered endpoints
Primary endpoint typeTime-to-event
Primary hypothesis typeSuperiority
Lead sponsorNational Institute of Allergy and Infectious Diseases (NIAID)
Sponsor typeNIH
StatusCompleted

2. Clinical Question

The central statistical question was whether participants assigned to remdesivir plus interferon beta-1a experienced a different time to recovery during Day 1 through Day 29 than participants assigned to remdesivir plus placebo.

Population

Participants enrolled in the phase 3 ACTT-3 trial with the condition COVID-19. The ClinicalTrials.gov record identifies the primary analysis population as the modified intention-to-treat population.

Intervention

Remdesivir plus interferon beta-1a.

Comparator

Remdesivir plus placebo.

Primary question

Does the addition of interferon beta-1a alter time to recovery compared with remdesivir plus placebo?

3. Trial Design

01
Enroll969 participants
02
RandomizeTwo parallel arms
03
MaskDouble-blind design
04
AssessDay 1–Day 29
05
CompareTime-to-event outcomes
ARM A

Remdesivir Plus Interferon Beta-1a

  • Remdesivir
  • Interferon beta-1a
  • Randomized treatment assignment
ARM B

Remdesivir Plus Placebo

  • Remdesivir
  • Placebo
  • Randomized treatment assignment
Allocation
Randomized
Structure
Parallel-group
Masking
Double
Hypothesis
Superiority

The design matters statistically because randomization creates the framework for comparing outcomes between treatment assignments. Double masking is intended to reduce the influence of treatment knowledge on trial conduct and outcome assessment. The parallel structure means the two randomized groups are followed as distinct treatment groups rather than being exposed sequentially to different randomized treatments.

4. Endpoints

The registry lists four primary endpoints. All have the same Day 1 through Day 29 time frame and are based on time to recovery, with additional analyses by race, ethnicity, or sex.

Registered primary endpointTime frameType
Time to Recovery for Participants With Baseline Ordinal Score 4, 5 and 6 Day 1 through Day 29 Time-to-event
Time to Recovery for Participants With Baseline Ordinal Score 4, 5 and 6 by Race Day 1 through Day 29 Time-to-event
Time to Recovery for Participants With Baseline Ordinal Score 4, 5 and 6 by Ethnicity Day 1 through Day 29 Time-to-event
Time to Recovery for Participants With Baseline Ordinal Score 4, 5 and 6 by Sex Day 1 through Day 29 Time-to-event

Definition of recovery

The registry defines the day of recovery as the first day on which the subject satisfies one of three categories from the ordinal scale: 1) not hospitalized, with no limitations on activities; 2) not hospitalized, but with a new or increased limitation on activities and/or a new or increased requirement for home oxygen; or 3) hospitalized, not requiring supplemental oxygen and no longer requiring ongoing medical care.

5. Statistical Analysis Populations

PopulationRegistry descriptionRole
Modified intention-to-treat (mITT) Includes all participants who were randomized. Participants were classified by their randomized treatment assignment and their baseline ordinal score. Primary time-to-event efficacy analyses
mITT with ethnicity reported Includes all randomized participants for whom ethnicity was reported. Ethnicity primary-endpoint analyses
Safety population Includes all participants with available data post baseline, analyzed as treated. Safety analyses

This distinction between randomized efficacy analysis and treated safety analysis is important. In the mITT framework, treatment assignment remains defined by randomization. In the safety population, participants are analyzed according to treatment received, which is more directly connected to exposure-related adverse-event reporting.

6. Statistical Methodology

Log-rank test

The registry reports a log-rank test for the principal time-to-recovery comparison. A log-rank test compares the survival or event-time experience of two groups across the observed follow-up rather than comparing only a single time point.

Core time-to-event idea
H0: treatment groups have the same event-time distribution

For ACTT-3, the event is recovery. A participant who has not experienced the defined recovery event by the relevant end of follow-up contributes information without being treated as if recovery occurred.

Cox proportional-hazards effect measure

The registry reports the treatment effect as a Cox proportional hazard, normalized here as a hazard ratio. The hazard ratio is a relative measure of the instantaneous event rate represented by the fitted Cox model.

Hazard-ratio interpretation
HR < 1  →  lower estimated instantaneous recovery rate in the treatment group

For a recovery endpoint, the direction requires care. Unlike an endpoint where the event is death or disease progression, an event of recovery is generally favorable. Therefore an HR below 1 for time to recovery indicates a lower estimated recovery hazard, not a lower mortality hazard.

Intention-to-treat principle

The registry-reported analysis text explicitly identifies the mITT population and states that participants were classified by their randomized treatment assignment. This preserves the treatment comparison defined by randomization and avoids redefining the primary efficacy comparison simply because treatment exposure changed.

Risk difference for safety

The secondary safety analyses use risk difference as the effect measure. A risk difference is an absolute difference in the proportion experiencing an event between the two compared groups.

Risk-difference interpretation
RD = Riskintervention − Riskcomparator

The registry reports the risk difference in percentage-point units for the adverse-event analyses. Unlike a hazard ratio, an RD describes an absolute difference in observed proportions over the specified safety time frame.

7. Primary Result: Time to Recovery

The principal primary analysis compares remdesivir plus interferon beta-1a with remdesivir plus placebo among participants with baseline ordinal scores 4, 5 and 6. The time frame was Day 1 through Day 29, and the analysis population was the mITT population.

Hazard ratio for time to recovery

0.99

95% CI: 0.87–1.13   ·   P = 0.880

Effect measure: Cox proportional hazard  ·  Analysis: Log Rank

FeatureReported result
EndpointTime to Recovery for Participants With Baseline Ordinal Score 4, 5 and 6
Time frameDay 1 through Day 29
ComparisonRemdesivir Plus Interferon Beta-1a vs Remdesivir Plus Placebo
Analysis populationModified intention-to-treat
Analysis methodLog-rank test
Effect measureCox proportional hazard / hazard ratio
Hazard ratio0.99
95% CI0.87–1.13
P-value0.880
HypothesisSuperiority
Clinical Biostats interpretation

The estimated hazard ratio of 0.99 is very close to 1. Under the reported Cox model, the estimated instantaneous rate of achieving the defined recovery event was approximately the same in the two randomized groups.

The HR does not mean that 99% of participants recovered, nor does it mean that the groups had identical individual recovery times. It is a relative model-based comparison of event rates over the analyzed time-to-event framework.

The 95% CI of 0.87–1.13 describes uncertainty around the estimated hazard ratio. Because the interval includes 1, the data are compatible with a range of relative recovery hazards on either side of the null value.

The P = 0.880 value addresses evidence against the null hypothesis under the reported statistical test; it does not measure the size or clinical importance of an effect. A p-value should therefore be interpreted together with the hazard ratio and confidence interval rather than as a standalone measure of treatment effect.

Because the endpoint is recovery, an HR below 1 has a different substantive direction from an HR below 1 for death or disease progression. The proportional-hazards interpretation also depends on the Cox-model framework; the ClinicalTrials.gov record does not report a separate assessment of the proportional-hazards assumption.

8. Primary Endpoint Analyses by Race

The registry includes race-specific analyses of the same time-to-recovery endpoint. These analyses use the mITT population and the same randomized comparison. Four race-specific estimates are posted.

Race subgroupHR95% CIP-value
Asian participants0.920.59–1.45Not reported
Black and African American participants0.920.66–1.27Not reported
White participants1.020.86–1.21Not reported
Race of Other participants0.840.59–1.19Not reported
Clinical Biostats interpretation

The race-specific hazard ratios range from 0.84 to 1.02, but the confidence intervals are relatively broad compared with the overall estimate. The intervals for all four reported race analyses include 1.

These estimates should be read as subgroup-specific descriptions rather than as proof that the treatment effect differs by race. A difference between subgroup estimates does not itself establish treatment-effect modification; a formal interaction analysis would be needed to make that claim.

The confidence intervals also emphasize precision. For example, the Asian-participant estimate of 0.92 has a 95% CI of 0.59–1.45, which spans a comparatively wide range of possible hazard ratios. The ClinicalTrials.gov record does not provide p-values or an interaction test for these race analyses.

9. Primary Endpoint Analyses by Ethnicity

Two ethnicity-specific analyses are posted for the same Day 1 through Day 29 time-to-recovery endpoint. The ethnicity analysis population includes all randomized participants for whom ethnicity was reported.

Ethnicity subgroupHR95% CIP-value
Not Hispanic or Latino participants1.020.86–1.19Not reported
Hispanic or Latino participants0.830.65–1.05Not reported
Clinical Biostats interpretation

The reported ethnicity-specific estimates are 1.02 for participants who were not Hispanic or Latino and 0.83 for participants who were Hispanic or Latino. Their 95% confidence intervals are 0.86–1.19 and 0.65–1.05, respectively.

These estimates describe the randomized comparison within the reported ethnicity categories. They do not establish that ethnicity changes the treatment effect. The ClinicalTrials.gov record contains no formal interaction p-value, and no subgroup p-values are reported.

The distinction between an estimate and its precision is important: the point estimate of 0.83 alone should not be interpreted as a definitive treatment-effect difference when its confidence interval extends from 0.65 to 1.05.

10. Primary Endpoint Analyses by Sex

The registry also posts two sex-specific analyses of time to recovery from Day 1 through Day 29.

Sex subgroupHR95% CIP-value
Male participants0.930.78–1.10Not reported
Female participants1.050.86–1.29Not reported
Clinical Biostats interpretation

The male-participant estimate is 0.93, while the female-participant estimate is 1.05. Their 95% confidence intervals are 0.78–1.10 and 0.86–1.29.

The two point estimates lie on opposite sides of 1, but that observation alone is not evidence of a sex-by-treatment interaction. The confidence intervals overlap the null value, and the ClinicalTrials.gov record does not report a formal interaction test.

Subgroup analyses are particularly useful for examining consistency and generating questions, but the precision and multiplicity of subgroup comparisons must be considered before treating individual subgroup estimates as independent confirmatory findings.

11. Secondary Time-to-Event Results

The registry also posts several secondary time-to-event analyses. These use the mITT population and report Cox proportional-hazard effect measures with two-sided 95% confidence intervals. No p-values are posted on ClinicalTrials.gov for these analyses in the ClinicalTrials.gov record.

Secondary endpointComparisonHR95% CI
Time to an Improvement of One Category Using an Ordinal Scale Baseline ordinal score 4 and 5 1.04 0.90–1.19
Time to an Improvement of One Category Using an Ordinal Scale Baseline ordinal score 6 0.37 0.21–0.66
Time to an Improvement of Two Categories Using an Ordinal Scale Baseline ordinal score 4 and 5 1.04 0.91–1.19
Time to an Improvement of Two Categories Using an Ordinal Scale Baseline ordinal score 6 0.44 0.24–0.82
Time to Discharge or to a National Early Warning Score (NEWS) of <= 2 and Maintained for 24 Hours, Whichever Occurs First Baseline ordinal score 4 and 5 1.08 0.93–1.26
Time to Discharge or to a National Early Warning Score (NEWS) of <= 2 and Maintained for 24 Hours, Whichever Occurs First Baseline ordinal score 6 0.39 0.21–0.72
Time to Recovery for Patients With a Baseline Ordinal Score of 4 and 5 Baseline ordinal score 4 and 5 1.04 0.90–1.19

The pattern is notable because the analyses restricted to baseline ordinal score 6 have hazard ratios below 1 for improvement by one category, improvement by two categories, and the discharge/NEWS endpoint. By contrast, the corresponding estimates for baseline ordinal scores 4 and 5 are close to 1. These are different endpoints and populations, so they should not be collapsed into a single treatment-effect estimate.

Important direction-of-effect caution: These are recovery and improvement events, not death or progression events. An HR below 1 therefore represents a lower estimated instantaneous rate of achieving the specified recovery/improvement event under the Cox model. The numerical direction should not be interpreted using the usual "HR below 1 means fewer bad events" shorthand.

12. Safety Results

The ClinicalTrials.gov record reports secondary safety analyses for grade 3 and 4 clinical and/or laboratory adverse events and for serious adverse events. These analyses use the safety population, defined as participants with available post-baseline data and analyzed as treated.

Grade 3 and 4 clinical and/or laboratory adverse events

Baseline groupComparisonRisk difference95% CI
Ordinal score 4 and 5 Remdesivir Plus Interferon Beta-1a vs Remdesivir Plus Placebo 7.2 0.9–13.4
Ordinal score 6 Remdesivir Plus Interferon Beta-1a vs Remdesivir Plus Placebo 23.6 0.0–43.9

Serious adverse events

Baseline groupComparisonRisk difference95% CI
Ordinal score 4 and 5 Remdesivir Plus Interferon Beta-1a vs Remdesivir Plus Placebo 1.4 -3.2–6.0
Ordinal score 6 Remdesivir Plus Interferon Beta-1a vs Remdesivir Plus Placebo 35.8 12.3–54.2

The ClinicalTrials.gov record also provide affected/at-risk counts for serious adverse events by arm:

Baseline groupRemdesivir Plus Interferon Beta-1aRemdesivir Plus Placebo
Ordinal score 4 and 565/44258/435
Ordinal score 621/328/31
Clinical Biostats interpretation

The risk-difference framework provides an absolute comparison rather than a time-to-event comparison. For example, the reported serious-adverse-event RD of 35.8 for baseline ordinal score 6 has a 95% CI of 12.3–54.2. The ClinicalTrials.gov record therefore describe a substantial positive difference in the reported event proportions for this comparison, with the entire reported confidence interval above 0.

For baseline ordinal scores 4 and 5, the reported serious-adverse-event RD is 1.4 with a 95% CI of -3.2–6.0. That interval includes 0, illustrating why the confidence interval is essential when interpreting an absolute risk difference.

The safety results should not be combined with efficacy into a single numerical "net benefit" measure. Efficacy and safety endpoints describe different dimensions of the randomized treatment comparison.

13. Statistical Methods Explained

Why use a time-to-event endpoint for recovery?

A binary recovery outcome would record only whether recovery occurred by a selected cutoff. A time-to-event endpoint retains information about when recovery occurred during the Day 1 through Day 29 window. That distinction can make timing clinically and statistically important.

What does an HR of 0.99 mean when the event is recovery?

An HR of 0.99 is close to 1, so the fitted model estimates nearly identical instantaneous recovery hazards between the randomized treatment groups. It should not be translated into a statement that "1% fewer patients recovered." Hazard ratios and absolute recovery proportions are different quantities.

Why is the interpretation of an HR below 1 different here?

For an endpoint such as death, a hazard ratio below 1 generally corresponds to a lower event rate, which is favorable. In ACTT-3, the event is recovery. Therefore an HR below 1 means a lower estimated instantaneous rate of reaching recovery under the model. The meaning of the numerical direction depends on what event is being modeled.

What does the 95% confidence interval tell us?

The confidence interval describes statistical uncertainty around the estimated effect. For the main analysis, the estimate is 0.99 and the 95% CI is 0.87–1.13. The interval includes the null value of 1, so the reported estimate is not precise enough to exclude values on either side of the null under this confidence-interval framework.

Why does the p-value not measure effect size?

The primary p-value is 0.880. A p-value describes how compatible the observed data are with the specified null hypothesis under the statistical test. It does not tell us how large the treatment effect is. The HR and its confidence interval are the quantities that describe the estimated relative effect and its precision.

Why should the race, ethnicity, and sex analyses not be treated as separate trials?

These analyses are subgroup analyses within the randomized trial. The point estimates can differ between subgroups simply because of sampling variation. A formal claim that the treatment effect differs by race, ethnicity, or sex requires an interaction or heterogeneity analysis rather than comparing the subgroup estimates informally.

Why is the analysis population important?

The primary time-to-event analyses use the mITT population, with participants classified according to randomized treatment assignment. The safety analyses instead use participants with available post-baseline data and analyze them as treated. Changing the analysis population changes the estimand and therefore changes what the resulting statistic describes.

14. Confidence Intervals, Null Values, and Precision

ACTT-3 provides a useful example of why effect estimates should be read together with their confidence intervals.

AnalysisEstimate95% CINull value
Primary time to recoveryHR 0.990.87–1.131
Asian participantsHR 0.920.59–1.451
Black and African American participantsHR 0.920.66–1.271
White participantsHR 1.020.86–1.211
Hispanic or Latino participantsHR 0.830.65–1.051
Male participantsHR 0.930.78–1.101
Female participantsHR 1.050.86–1.291

For hazard ratios, the null value is 1. For risk differences, the null value is 0. This distinction is basic but important: an interval that crosses 1 for a hazard ratio is interpreted differently from an interval that crosses 0 for a risk difference.

Relative effect

A hazard ratio compares event rates within a time-to-event model. It is not a percentage of patients who benefited.

Absolute effect

A risk difference compares event proportions directly and expresses their absolute separation in the reported percentage-point scale.

Precision

A narrow confidence interval indicates greater statistical precision than a wide interval, all else being equal.

Statistical significance

A confidence interval crossing its null value and a large p-value are evidence considerations, not measurements of clinical importance.

15. Reading the Primary Analysis Correctly

The main result can be reduced to three related quantities:

ACTT-3 primary statistical result

HR 0.99

95% CI 0.87–1.13   ·   P = 0.880

These numbers answer different questions:

A statistically careful interpretation does not replace these three quantities with a single label. In particular, the p-value should not be treated as a measure of treatment magnitude, and the hazard ratio should not be interpreted as an absolute probability of recovery.

Endpoint-direction reminder: Because the primary event is recovery, an HR of 0.99 should be interpreted as a near-null estimated recovery hazard ratio. It should not be described using the language commonly used for a hazard ratio for death, such as "lower risk of death."

16. Subgroup Analysis: What Can and Cannot Be Concluded

The ACTT-3 registry data provide nine primary-endpoint analyses: one overall analysis, four race-specific analyses, two ethnicity-specific analyses, and two sex-specific analyses.

DimensionNumber of reported analysesRange of HR estimates
Overall10.99
Race40.84–1.02
Ethnicity20.83–1.02
Sex20.93–1.05

There are two separate statistical questions here. The first is whether the treatment comparison within a subgroup differs from the null value. The second, which is often more clinically interesting, is whether the treatment effect itself differs between subgroups. The second question requires an interaction or heterogeneity analysis.

Do not compare subgroup p-values informally. The ClinicalTrials.gov record does not provide subgroup p-values or interaction tests. Therefore the posted subgroup estimates can be described, but a treatment-effect difference by race, ethnicity, or sex should not be inferred from the point estimates alone.

17. Ordinal Outcomes and Baseline Severity

Several secondary endpoints use an ordinal clinical scale and distinguish participants by baseline ordinal score. This is statistically informative because the same numerical treatment assignment can be associated with different event-time estimates in different baseline severity groups.

Baseline groupEndpointHR95% CI
Ordinal score 4 and 5 One-category improvement 1.04 0.90–1.19
Ordinal score 6 One-category improvement 0.37 0.21–0.66
Ordinal score 4 and 5 Two-category improvement 1.04 0.91–1.19
Ordinal score 6 Two-category improvement 0.44 0.24–0.82
Ordinal score 4 and 5 Discharge or NEWS <= 2 maintained for 24 hours 1.08 0.93–1.26
Ordinal score 6 Discharge or NEWS <= 2 maintained for 24 hours 0.39 0.21–0.72

The contrast between baseline ordinal score 4 and 5 versus baseline ordinal score 6 is therefore a major part of the statistical story contained in the posted secondary analyses. However, these are secondary endpoints, and the ClinicalTrials.gov record does not provide formal multiplicity-adjusted p-values for them.

18. Multiplicity and Multiple Analyses

The registry reports four primary endpoints and a total of 20 statistical analyses, including nine primary-endpoint analyses and multiple secondary analyses. That breadth matters when interpreting individual estimates.

Primary endpoints

Four registered primary endpoints are reported, all involving time to recovery and analyses by race, ethnicity, or sex.

Primary analyses

Nine primary-endpoint analyses have posted statistical estimates and confidence intervals.

Secondary analyses

The posted secondary results include ordinal improvement, discharge/NEWS, recovery, and safety outcomes.

Interpretive caution

The ClinicalTrials.gov record does not describe a multiplicity-adjustment procedure for the posted secondary and subgroup estimates.

Multiplicity does not make an estimate meaningless. It changes the evidentiary context in which multiple estimates should be interpreted. When many endpoints and subgroup analyses are examined, the probability of observing an apparently notable result somewhere in the collection increases unless the statistical design explicitly controls that multiplicity.

What the ClinicalTrials.gov record supports: the primary overall result has a reported two-sided p-value of 0.880. The subgroup and secondary analyses reported here have estimates and 95% confidence intervals, but the data do not provide corresponding p-values or a multiplicity-adjustment framework for those analyses.

19. Kaplan-Meier Estimation and Censoring

Kaplan-Meier estimation is a standard way to display and estimate time-to-event distributions. The registry-reported ACTT-3 registry data identify time-to-event endpoints and a log-rank comparison, but they do not explicitly report a Kaplan-Meier method in the normalized statistical-method field.

Educational framework
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents events at an event time and ni represents participants at risk immediately before that time. The method allows participants with incomplete event-time observation to contribute information up to their censoring time.

The important conceptual point is that time-to-event analysis does not require every participant to experience the event during the observation period. This is one reason time-to-recovery analysis differs from a simple comparison of the proportion recovered at one fixed time point.

20. Why the Log-Rank Test and Cox Model Work Together

The ACTT-3 primary analysis illustrates two related but distinct statistical functions.

MethodPrimary roleACTT-3 use
Log-rank testCompare event-time experience between groupsReported primary analysis method
Cox proportional-hazards modelExpress the relative event rate as a hazard ratioReported effect measure

The log-rank test and hazard ratio should not be treated as interchangeable. One is a hypothesis-testing framework for comparing time-to-event experience; the other is an effect measure derived from a proportional-hazards model. Reporting both provides complementary information.

Clinical Biostats interpretation

For ACTT-3, the reported combination of a log-rank analysis with a Cox proportional-hazard effect measure means that the statistical story has both a comparison component and an effect-estimation component. The p-value of 0.880 belongs to the reported hypothesis-testing result, while the HR of 0.99 and its 95% CI of 0.87–1.13 describe the estimated relative effect and its precision.

21. Important Limitations and Interpretation Issues

22. Why This Trial Matters Statistically

ACTT-3 is a useful teaching example because it combines randomized treatment assignment with a clinically meaningful time-to-event endpoint, subgroup analyses, ordinal clinical outcomes, and absolute safety comparisons.

ConceptHow it appears in ACTT-3
RandomizationRandomized parallel-group phase 3 design
BlindingDouble masking
Intention-to-treat analysisPrimary analyses use the modified intention-to-treat population and randomized treatment assignment
Time-to-event endpointTime to recovery during Day 1 through Day 29
Log-rank testReported method for the principal primary comparison
Hazard ratioCox proportional-hazard effect measure
Confidence interval95% two-sided intervals accompany the reported effect estimates
Subgroup analysisPrimary endpoint analyses by race, ethnicity, and sex
Ordinal outcomesOne- and two-category improvement analyses by baseline ordinal score
Risk differenceSecondary safety analyses for grade 3 and 4 AEs and serious AEs
Safety populationParticipants with available post-baseline data, analyzed as treated

23. Overall Statistical Interpretation

The central reported ACTT-3 result is a hazard ratio of 0.99 for time to recovery among participants with baseline ordinal scores 4, 5 and 6, with a two-sided 95% CI of 0.87–1.13 and P = 0.880. Statistically, this is an estimate close to the null value of 1, with a confidence interval that spans the null.

The subgroup analyses provide additional descriptive estimates rather than a separate set of independently randomized treatment comparisons. Race-specific HRs range from 0.84 to 1.02; ethnicity-specific HRs are 1.02 and 0.83; and sex-specific HRs are 0.93 and 1.05. Their confidence intervals should be considered alongside the point estimates.

The secondary analyses introduce a different pattern. For participants with baseline ordinal score 6, the reported hazard ratios are below 1 for one-category improvement, two-category improvement, and the discharge/NEWS endpoint. These estimates are accompanied by 95% confidence intervals of 0.21–0.66, 0.24–0.82, and 0.21–0.72, respectively. The corresponding baseline ordinal score 4 and 5 analyses are much closer to 1. These findings should be kept in their specific endpoint and population contexts rather than summarized as a single global effect.

The safety analyses use risk differences rather than hazard ratios. For baseline ordinal score 4 and 5, the serious-adverse-event RD is 1.4 with a 95% CI of -3.2–6.0. For baseline ordinal score 6, the corresponding RD is 35.8 with a 95% CI of 12.3–54.2. These absolute measures answer a different question from the time-to-recovery analyses.

24. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

25. Related Statistical Calculators

26. Sources

Continue with the statistical methods

Explore the broader methods used to understand randomized clinical trials, time-to-event endpoints, effect measures, and statistical uncertainty.

27. Record Summary

ACTT-3 provides a compact example of how a randomized clinical trial can combine a primary time-to-event endpoint with subgroup analyses and secondary ordinal and safety outcomes. The main analysis used a modified intention-to-treat population, a log-rank comparison, and a Cox proportional-hazard effect measure. The principal reported estimate was HR 0.99 with a two-sided 95% CI of 0.87–1.13 and P = 0.880.

The statistical interpretation should remain closely tied to the endpoint definition. Because recovery is the event, an HR below 1 represents a lower estimated recovery hazard rather than the lower hazard of an adverse event. The race, ethnicity, and sex analyses provide additional subgroup estimates, while the secondary ordinal outcomes show different estimates according to baseline ordinal score. Safety analyses use risk differences and a separate treated safety population.

Clinical Biostats methodology: A trial-results page should not merely repeat registry fields. The goal is to connect the reported estimates to their analysis populations, endpoint definitions, effect measures, confidence intervals, and statistical assumptions while clearly separating reported evidence from educational interpretation.