← Clinical Trials
COVID-19 Phase 2/3 Terminated NCT05011513

EPIC-SR: Complete Statistical Analysis of Nirmatrelvir plus Ritonavir in COVID-19

A statistical review of EPIC-SR (Evaluation of Protease Inhibition for COVID-19 in Standard-Risk Patients), a randomized, quadruple-masked, placebo-controlled phase 2/3 trial of nirmatrelvir 300 mg plus ritonavir 100 mg in standard-risk patients with COVID-19.

Sponsor: Pfizer  ·  Start: 2021-08-25  ·  Primary completion: 2022-07-25  ·  Status: Terminated
About this page

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

EPIC-SR was a randomized, parallel-group, quadruple-masked phase 2/3 trial comparing nirmatrelvir (PF-07321332) 300 mg plus ritonavir 100 mg with placebo in standard-risk patients with COVID-19. Its primary endpoint was the time to sustained alleviation of all targeted COVID-19 signs and symptoms through Day 28, and the posted primary log-rank comparison did not show a statistically significant difference between the groups.

1440
Enrollment
Two randomized arms
0.6027
Primary log-rank P
Time to sustained alleviation
0.802
OR, progression to worsening
95% CI 0.613–1.050
8 vs 13
Participants with SAEs
8/654 vs 13/634
FeatureEPIC-SR
PhasePhase 2/3
ConditionCOVID-19 (standard-risk patients)
DesignRandomized, parallel-group, placebo-controlled
MaskingQuadruple
Primary purposeTreatment
InterventionsPF-07321332 (nirmatrelvir) 300 mg + ritonavir 100 mg; matching placebos
Primary endpointTime to sustained alleviation of overall COVID-19 signs and symptoms through Day 28
Primary analysisLog-rank test, superiority hypothesis, mITT1 population
Enrollment1440
DatesStart 2021-08-25; primary completion 2022-07-25
StatusTerminated; results posted
ClinicalTrials.govNCT05011513
SponsorPfizer (industry)

2. Clinical Question

The central question was whether an oral protease-inhibitor regimen, nirmatrelvir boosted with ritonavir, shortens the time until COVID-19 symptoms are durably alleviated in patients who are not in a high-risk category, compared with placebo. Because standard-risk patients rarely progress to hospitalization, the trial used a symptom-based time-to-event endpoint as the primary outcome rather than a hard clinical event.

Population

Standard-risk patients with COVID-19, as described in the registry title.

Intervention

Nirmatrelvir (PF-07321332) 300 mg plus ritonavir 100 mg.

Comparator

Matching placebo for each component, preserving quadruple masking.

Primary question

Does nirmatrelvir plus ritonavir shorten the time to sustained alleviation of all targeted COVID-19 signs and symptoms through Day 28, compared with placebo?

3. Trial Design

01
Enroll1440 participants
02
RandomizeTwo parallel arms, quadruple-masked
03
Day 1Start of study intervention or placebo
04
Symptom scoringDaily targeted signs and symptoms
05
Day 28End of primary assessment window
ARM 1 · SAFETY AT RISK = 654

Nirmatrelvir + ritonavir

  • Nirmatrelvir (PF-07321332) 300 mg
  • Ritonavir 100 mg
  • Oral protease inhibitor, pharmacokinetically boosted by ritonavir
ARM 2 · SAFETY AT RISK = 634

Placebo

  • Placebo matching nirmatrelvir
  • Placebo matching ritonavir
  • Identical appearance to preserve blinding

The arm headers show the number of participants at risk in the registry's serious adverse event tables. The registry lists two separate placebo interventions, which is typical of a double-dummy style of masking: each active component has its own matching placebo so that participants, care providers, investigators and outcome assessors cannot infer the assignment from the number or appearance of tablets.

Terminated status. The registry records the trial as terminated. Termination affects interpretation even when results are posted: the final analyzed sample may differ from the planned sample, and the achieved statistical power for the primary comparison may differ from the design assumptions. The registry entry summarized here does not give a sample-size calculation or planned power.

4. Analysis Populations

All posted statistical analyses used the modified intent-to-treat population (mITT1): all participants randomly assigned to study intervention who received at least 1 dose, analyzed according to the intervention to which they were randomized. For several secondary endpoints, the registry adds that the "Overall Number of Participants Analyzed" refers to participants evaluable for that outcome measure.

Analysis populationDefinition / role
mITT1Randomized participants who received at least 1 dose; analyzed as randomized. Used for the primary endpoint and the posted secondary analyses.
mITT1, evaluable subsetParticipants in mITT1 evaluable for a given outcome measure (and, for oxygen saturation, at the specified time points).
Safety (adverse events)Serious adverse events reported as affected / at risk: 654 in the nirmatrelvir + ritonavir arm and 634 in the placebo arm.

A modified ITT population differs from a strict intention-to-treat population by excluding randomized participants who never received a dose. When blinding is maintained and dosing failures are rare and unrelated to assignment, this exclusion is unlikely to introduce meaningful bias; its main consequence is that the analysis describes participants who actually started treatment. The "evaluable" qualifier on secondary endpoints is a further restriction and deserves more caution, because evaluability can depend on follow-up and data completeness.

5. Endpoints

Primary endpoint

EndpointRegistry definitionTime frame
Time to Sustained Alleviation of Overall COVID-19 Signs and Symptoms Through Day 28Sustained alleviation is the event occurring on the first 4 consecutive days when all symptoms scored as moderate or severe at enrollment were scored as mild or absent, and those scored mild or absent at enrollment were scored as absent. Missing severity at baseline was considered mild. Time is measured in days from start of study intervention or placebo (Day 1) until sustained alleviation of all targeted signs and symptoms, reported consolidated for overall COVID-19 signs and symptoms.From Day 1 to Day 28

Secondary endpoints with posted statistical analyses

EndpointUnitTime framePosted method
Percentage of participants with severe signs and symptoms of COVID-19 through Day 28Percentage of participantsDay 1 to Day 28Logistic regression; odds ratio
Percentage of participants with progression to worsening status of COVID-19 signs and symptomsPercentage of participantsDay 1 to Day 28Logistic regression; odds ratio
Percentage of participants with resting peripheral oxygen saturation >=95% at Day 1 and Day 5Percentage of participantsDay 1 and Day 5Within-arm odds ratios (Day 5 vs Day 1); Breslow-Day test
Percentage of participants with COVID-19 related hospitalization or death from any cause through Day 28Percentage of participantsDay 1 to Day 28Normal approximation
Number of COVID-19 related medical visits per day through Day 28Medical visits per dayDay 1 to Day 28Negative binomial
Time to sustained resolution of overall COVID-19 signs and symptoms through Day 28DaysDay 1 to Day 28Log-rank

Why a "sustained" definition matters

COVID-19 symptoms fluctuate from day to day. A definition that required only a single day of improvement would register many transient dips as events and would be sensitive to diary noise. Requiring 4 consecutive days of alleviation makes the event more clinically meaningful and more reproducible, at the cost of pushing event times later and increasing the chance that a participant is censored at Day 28 before the event is confirmed. The rule is also relative to each participant's own baseline: symptoms that started moderate or severe need only fall to mild or absent, whereas symptoms that started mild must disappear entirely.

6. Results: Primary Endpoint

Time to sustained alleviation of overall COVID-19 signs and symptoms through Day 28

Log-rank test, nirmatrelvir + ritonavir vs placebo

P = 0.6027

Superiority hypothesis  ·  mITT1 population  ·  Time frame: Day 1 to Day 28

No hazard ratio or confidence interval accompanies the posted log-rank analysis.

Clinical Biostats interpretation

What the result means. The log-rank test compares the whole time-to-alleviation curves of the two arms across the 28-day window. A P-value of 0.6027 means that, if there were truly no difference in the distribution of time to sustained alleviation between arms, differences between the observed curves at least as large as those seen would arise in about 60% of repeated trials. The data are therefore entirely compatible with no difference, and the superiority hypothesis was not supported.

What it does not mean. A non-significant log-rank test is not proof that the two regimens are equivalent. Establishing equivalence or non-inferiority would require a prespecified margin and a confidence interval for an effect measure contained within it. The trial was framed as a superiority comparison, so the only formal conclusion is a failure to demonstrate superiority.

Precision. The registry analysis gives no effect estimate or confidence interval for this endpoint, so the plausible size of any difference in time to alleviation, in either direction, cannot be judged from the posted analysis alone. This is the main practical limitation of a stand-alone log-rank P-value.

Why the P-value is not an effect size. The log-rank statistic combines the size of any difference with the number of events. A large P-value can come from a small true effect, an imprecise estimate, or both. It says nothing on its own about how many days, if any, symptoms were shortened.

Cautions. The log-rank test has the most power when hazards are proportional over time; if curves converge or cross, power falls. Participants who did not reach sustained alleviation by Day 28, or who left follow-up earlier, are censored, and the test assumes censoring is non-informative. The analysis was in the mITT1 population rather than in all randomized participants.

7. Results: Secondary Endpoints

The registry posts formal statistical analyses for six secondary outcome measures. All were framed as superiority comparisons in the mITT1 population. None of the posted between-arm P-values was below 0.05.

Secondary endpointMethodEstimate (95% CI)P-value
Severe signs and symptoms through Day 28Logistic regressionOR 0.819 (0.618–1.084)0.1622
Progression to worsening status of signs and symptomsLogistic regressionOR 0.802 (0.613–1.050)0.1086
COVID-19 related hospitalization or death from any cause through Day 28Normal approximationNot posted0.1796
COVID-19 related medical visits per day through Day 28Negative binomialNot posted0.0971
Time to sustained resolution of overall signs and symptoms through Day 28Log-rankNot posted0.4298
Resting SpO2 >=95%, Day 5 vs Day 1 (within arm)Odds ratio; Breslow-Day testNirmatrelvir + ritonavir: OR 50.333 (13.163–192.472); Placebo: OR 22.224 (8.360–59.080)0.3262 (Breslow-Day)

Severe signs and symptoms through Day 28

Odds ratio, nirmatrelvir + ritonavir vs placebo

0.819

95% CI: 0.618–1.084 (two-sided)  ·  P = 0.1622

Logistic regression with main effects of treatment, geographic region, baseline SARS-CoV-2 serology status and baseline viral load (< 4 vs >= 4 log10 copies/mL)

An odds ratio of 0.819 corresponds to an estimated 18.1% lower odds of severe signs and symptoms in the nirmatrelvir + ritonavir arm, adjusted for the listed covariates. The confidence interval runs from 0.618, a meaningful reduction in odds, to 1.084, a modest increase. Because the interval includes 1, the data are compatible with no effect.

Progression to worsening status of COVID-19 signs and symptoms

Odds ratio, nirmatrelvir + ritonavir vs placebo

0.802

95% CI: 0.613–1.050 (two-sided)  ·  P = 0.1086

Logistic regression with main effects of treatment, geographic region, symptom onset duration (<=3, >3), baseline serology status, vaccination status and baseline viral load

The point estimate implies 19.8% lower adjusted odds of progression to worsening status, but the upper confidence limit of 1.050 leaves no difference, or a small increase in odds, within the range of values compatible with the data. The adjustment set is broader than for the severe-symptoms endpoint and notably includes vaccination status and symptom-onset duration.

Odds ratios are not risk ratios. An odds ratio of 0.802 does not mean that 19.8% fewer participants progressed. Odds ratios and risk ratios are close when the outcome is uncommon, but diverge as the outcome becomes more frequent. Covariate-adjusted odds ratios from logistic regression are also conditional on the covariates in the model, so they are not directly comparable to an unadjusted odds ratio computed from the arm percentages.

Hospitalization or death, medical visits and sustained resolution

For three secondary endpoints, the registry posts a P-value without an accompanying effect estimate:

Resting peripheral oxygen saturation >=95% at Day 1 and Day 5

This endpoint has an unusual analysis structure. Rather than a single between-arm odds ratio, the registry posts a within-arm odds ratio for Day 5 versus Day 1 in each arm: 50.333 (95% CI 13.163–192.472) with nirmatrelvir + ritonavir and 22.224 (95% CI 8.360–59.080) with placebo. A Breslow-Day test, P = 0.3262, is posted with the nirmatrelvir + ritonavir entry.

The Breslow-Day test assesses whether odds ratios are homogeneous across strata. Read in that way, it asks whether the Day 5 versus Day 1 odds ratio differs between the two arms. With P = 0.3262, there is no statistically significant evidence that the improvement in oxygen saturation over time differed by treatment group. Both within-arm odds ratios are large, which indicates that most participants in both arms had oxygen saturation of at least 95% by Day 5, consistent with the natural course of standard-risk illness. The very wide intervals, especially the upper limit of 192.472, reflect sparse cells, which is expected when nearly all participants already meet the threshold.

8. Safety: Serious Adverse Events

ArmParticipants with serious adverse eventsParticipants at risk
Nirmatrelvir 300 mg + Ritonavir 100 mg8654
Placebo13634
Participants with serious adverse events (count)
Nirmatrelvir + ritonavir
8/654
Placebo
13/634

Serious adverse events were infrequent in both arms, with fewer affected participants in the nirmatrelvir + ritonavir arm (8 of 654) than in the placebo arm (13 of 634). Because serious adverse events in COVID-19 trials often include the disease complications themselves, fewer events in the active arm may partly reflect disease course rather than drug tolerability alone. With so few events, any between-arm comparison would be imprecise, and no formal statistical test of serious adverse events is posted.

Participant flow note from the registry: one discontinuation due to an adverse event was captured under "Death" as the reason for discontinuation. The adverse event was COVID-19 pneumonia; the participant died due to that event and discontinued the study. Categorization rules such as this matter when counting discontinuations by reason.

9. Statistical Methodology

Kaplan-Meier estimation and the log-rank test

The primary endpoint and the sustained-resolution endpoint are time-to-event outcomes measured in days from Day 1. Participants who did not achieve the event by Day 28, or whose follow-up ended earlier, are right-censored: they contribute information up to their last observed day. Kaplan-Meier curves are the standard descriptive tool for such data, and the log-rank test compares the curves between arms.

Conceptual form of the log-rank statistic
Z = Σi (O1i − E1i) / √ Σi Vi

At each event day i, the observed number of events in arm 1 (O) is compared with the number expected (E) if both arms shared the same event rate, given the numbers still at risk. Summing across days and standardizing by the variance V gives a statistic that is approximately standard normal under the null hypothesis.

In a symptom-alleviation endpoint the event is desirable, so a treatment benefit would appear as a curve that falls (or, if plotted as cumulative incidence, rises) faster in the active arm. The log-rank test is indifferent to this direction; it detects any systematic difference.

Covariate-adjusted logistic regression

The two binary secondary endpoints with posted odds ratios were analyzed with logistic regression including treatment and prognostic covariates as main effects: geographic region, baseline SARS-CoV-2 serology status and baseline viral load (< 4 vs >= 4 log10 copies/mL), with symptom-onset duration and vaccination status added for the progression endpoint.

Model form
logit P(Y = 1) = β0 + βtrt·Treatment + β2·Region + β3·Serology + β4·ViralLoad + …

The reported odds ratio is exp(βtrt), the multiplicative change in the odds of the outcome for the active arm versus placebo, holding the covariates fixed. The Wald-type confidence interval is exp(βtrt ± 1.96·SE).

In a randomized trial, covariate adjustment is not needed to remove confounding, but adjusting for strongly prognostic baseline factors, such as serology status and viral load in COVID-19, typically increases precision and power for the treatment comparison.

Negative binomial regression

Medical visits per day are counts with many zeros. A Poisson model would assume the variance equals the mean; the negative binomial model adds a dispersion parameter that allows the variance to exceed the mean, avoiding overly narrow confidence intervals and overly small P-values when counts are overdispersed.

Normal approximation for a binary endpoint

For hospitalization or death, a normal-approximation test compares event proportions using a z-statistic built from the difference in proportions and its standard error. This is a large-sample method; with rare events, exact or score-based methods are often preferred for confidence intervals.

Breslow-Day test

The Breslow-Day test is used in stratified 2×2 table analyses to test whether the odds ratio is the same across strata. Here it is applied to the Day 5 versus Day 1 odds ratios in the two arms, so it serves as a test of whether the change over time differs by treatment.

Handling of missing baseline severity

The primary endpoint definition specifies that missing symptom severity at baseline was considered mild. This rule has direct consequences: a symptom treated as mild at baseline must reach "absent" to count toward alleviation, the stricter of the two thresholds. The rule is conservative in the sense that it cannot make alleviation easier to achieve, and it applies equally to both masked arms.

10. Statistical Methods Explained

Why was a log-rank test used for the primary endpoint?

Time to sustained alleviation is measured in days and is right-censored at Day 28 for participants who had not yet reached the event. The log-rank test uses every participant's follow-up time, including censored observations, and compares the entire event-time distribution between arms. A simple comparison of the proportion alleviated at Day 28 would discard information about how quickly symptoms improved.

Does a primary P-value of 0.6027 show that nirmatrelvir plus ritonavir has no effect on symptoms?

No. It shows that the trial did not detect a statistically significant difference in time to sustained alleviation. Absence of evidence of a difference is not evidence of absence. Without a hazard ratio and confidence interval, the range of effects compatible with the data cannot be characterized from the posted analysis, and a small benefit or small harm cannot be excluded.

What does an odds ratio of 0.802 with a 95% CI of 0.613–1.050 mean?

The adjusted odds of progression to worsening status were estimated to be 0.802 times those in the placebo arm, a 19.8% relative reduction in odds. The interval shows that values from a 38.7% reduction in odds to a 5.0% increase are compatible with the data at the 95% level. Because 1 lies inside the interval, the result is not statistically significant, which agrees with P = 0.1086.

Why were the Day 5 versus Day 1 oxygen-saturation odds ratios so large, and what does the Breslow-Day test add?

Within each arm, the odds of having SpO2 of at least 95% were far higher at Day 5 than at Day 1, giving odds ratios of 50.333 and 22.224. That reflects improvement over time in both arms, not a treatment effect. The treatment question is whether that improvement differed between arms, which is what the Breslow-Day homogeneity test addresses; P = 0.3262 gives no evidence that it did. Comparing the two point estimates by eye would be misleading, because their confidence intervals are very wide and overlap heavily.

Why use negative binomial regression for medical visits per day?

Visit counts are overdispersed: most participants have none, while a few have several. The negative binomial model accommodates this extra variability. Using a Poisson model instead would understate the standard error and could produce a spuriously small P-value. The posted P = 0.0971 is above the conventional 0.05 threshold.

With six secondary analyses, how should the nominal P-values be read?

Each additional test carries its own chance of a false-positive result. When the primary endpoint is not statistically significant, secondary endpoints are generally regarded as supportive or exploratory rather than confirmatory, especially if the trial used a hierarchical testing strategy. No multiplicity adjustment is described for these posted analyses, so their P-values are best read as nominal. In EPIC-SR none crossed 0.05, so multiplicity does not change the overall conclusion.

11. Limitations

12. Why This Trial Matters Statistically

EPIC-SR is a useful teaching case because it shows how a well-masked randomized trial can return a clear negative answer on its primary endpoint, and because its posted analyses span several distinct statistical models applied to different endpoint types within one trial.

ConceptHow it appears in EPIC-SR
Time-to-event endpointsTime to sustained alleviation and to sustained resolution, censored at Day 28
Log-rank testPrimary comparison (P = 0.6027) and sustained resolution (P = 0.4298)
Logistic regressionCovariate-adjusted odds ratios for severe symptoms and progression to worsening
Odds ratio and confidence intervalsOR 0.819 (0.618–1.084) and OR 0.802 (0.613–1.050), both including 1
Negative binomial regressionOverdispersed count endpoint: medical visits per day
Normal approximation / z-testHospitalization or death from any cause
Breslow-Day testHomogeneity of Day 5 vs Day 1 oxygen-saturation odds ratios across arms
Modified intention-to-treatmITT1: randomized participants who received at least 1 dose, analyzed as randomized
Superiority vs equivalenceA non-significant superiority test does not establish equivalence
P-values vs effect sizesSeveral analyses posted with P-values but no effect estimates

Statistical interpretation

The primary log-rank comparison did not show superiority of nirmatrelvir plus ritonavir over placebo for time to sustained symptom alleviation, and no posted secondary comparison reached P < 0.05. Odds ratios for severe symptoms and progression were below 1 with intervals including 1.

Clinical interpretation

In a standard-risk population, where most participants improve without treatment, demonstrating faster symptom relief requires a sizeable effect on a variable, patient-reported outcome. The trial did not show one, and uncommon clinical events limit what can be said about hospitalization or death.

13. Related Tutorials

Learn more about the methods used in this trial:

14. Related Calculators

15. Sources

Continue exploring clinical trial statistics

Each trial page connects its endpoints and methods to deeper statistical tutorials and to calculators for working through the same analyses.

16. Record Summary

EPIC-SR randomized standard-risk patients with COVID-19 to nirmatrelvir plus ritonavir or placebo under quadruple masking. The primary log-rank comparison of time to sustained alleviation of symptoms through Day 28 gave P = 0.6027, and no posted secondary comparison reached conventional statistical significance, although odds ratios for severe symptoms (0.819) and progression to worsening (0.802) were below 1. The most useful reading of the trial combines the prespecified superiority framework, the distinction between a non-significant result and demonstrated equivalence, the confidence intervals that are available, and the design features, including termination and the mITT1 population, that affect interpretation.

Clinical Biostats methodology: A trial-results page should not merely repeat the registry entry. The goal is to reconstruct the statistical story of the trial in a standardized format while clearly separating reported evidence from educational interpretation.