← Clinical Trials
HIV-1 PrEP Phase 3 Non-Inferiority NCT02842086

DISCOVER: Complete Statistical Analysis of F/TAF in HIV-1 Pre-Exposure Prophylaxis

An independent statistical review of the randomized phase 3 DISCOVER trial evaluating once-daily F/TAF versus F/TDF for pre-exposure prophylaxis in men and transgender women who have sex with men and are at risk of HIV-1 infection.

Trial status: Terminated  ·  Enrollment: 5399  ·  Primary completion: 2019-01-31
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

DISCOVER was a randomized, double-blind, parallel phase 3 trial evaluating F/TAF versus F/TDF for pre-exposure prophylaxis of HIV-1 infection. The registered primary endpoint was the incidence of HIV-1 infection per 100 person-years, with a prespecified non-inferiority comparison using a rate ratio.

5399
Enrollment
Participants
4
Arms
Two active drugs + two placebo arms
0.468
Primary rate ratio
F/TAF vs F/TDF
1.62
NI margin
Upper CI bound criterion
FeatureDISCOVER
PhasePhase 3
Therapeutic areaInfectious Disease
ConditionPre-Exposure Prophylaxis of HIV-1 Infection
DesignRandomized, double-blind, parallel
AllocationRandomized
Primary purposeTreatment
Enrollment5399
Primary endpointIncidence of HIV-1 Infection Per 100 Person Years (PY)
Primary endpoint typeCount / rate
Lead sponsorGilead Sciences
Sponsor typeIndustry
StatusTerminated
Study datesStart: 2016-09-02; primary completion: 2019-01-31
ClinicalTrials.govNCT02842086

2. Clinical Question

The central statistical question was whether F/TAF was non-inferior to F/TDF for the incidence of HIV-1 infection per 100 person-years. The registry reports this endpoint as the number of participants who became HIV infected after the first dose of study drug divided by the total person-years of follow-up while participants were at risk of HIV infection.

Population

Men and transgender women who have sex with men and are at risk of HIV-1 infection, according to the registered trial title.

Intervention

F/TAF fixed-dose combination once daily for pre-exposure prophylaxis.

Comparator

F/TDF fixed-dose combination once daily for pre-exposure prophylaxis.

Primary question

Is the HIV-1 infection incidence rate with F/TAF sufficiently close to that with F/TDF to satisfy the prespecified non-inferiority criterion?

3. Trial Design

01
Randomize5399 enrolled
02
Double-blindBlinded treatment assignment
03
Parallel groupsF/TAF and F/TDF comparisons
04
Follow-upHIV infection and laboratory outcomes
05
AnalysisRate and continuous outcomes
ACTIVE REGIMEN

F/TAF

  • F/TAF drug intervention
  • Once-daily fixed-dose combination for pre-exposure prophylaxis
  • Registry abbreviation: DVY / Descovy in the serious-adverse-event data
ACTIVE REGIMEN

F/TDF

  • F/TDF drug comparator
  • Once-daily fixed-dose combination for pre-exposure prophylaxis
  • Registry abbreviation: TVD / Truvada in the serious-adverse-event data

The registry lists four interventions: F/TAF, F/TDF, F/TAF placebo, and F/TDF placebo. The overall design is therefore represented in the registry as a four-arm double-blind parallel study, while the primary statistical comparison is between the F/TAF and F/TDF treatment groups.

4. Randomization, Blinding, and Analysis Populations

DISCOVER used randomized allocation and double masking. The primary analysis was performed in the Full Analysis Set, which the registry defines as participants randomized into the study who received at least 1 dose of study drug, were not HIV positive on Day 1, and had at least 1 postbaseline HIV laboratory assessment.

Analysis populationRegistry-supported role
Full Analysis SetPrimary HIV-1 infection incidence analysis; participants were randomized, received at least 1 dose, were not HIV positive on Day 1, and had at least 1 postbaseline HIV laboratory assessment.
Safety Analysis SetUsed for serum creatinine and renal biomarker analyses; generally included randomized participants who received at least 1 dose, with available data for the relevant analysis.
Hip DXA Analysis SetDXA substudy population with randomized treatment, at least 1 dose, and nonmissing hip BMD data as specified by the relevant endpoint.
Spine DXA Analysis SetDXA substudy population with randomized treatment, at least 1 dose, and available spine BMD data as specified by the relevant endpoint.

The distinction between these populations matters. Randomization establishes the principal basis for the efficacy comparison, whereas laboratory and DXA outcomes may be restricted to participants with the required postbaseline measurements. Consequently, estimates for these secondary endpoints describe the corresponding analysis sets rather than automatically the entire enrolled population.

5. Primary Endpoint

EndpointRegistered definition / time framePrimary analysis
Incidence of HIV-1 Infection Per 100 Person Years (PY)When all participants completed minimum follow-up of 48 weeks and at least 50% of the participants completed 96 weeks of follow-up.Poisson regression; rate ratio; non-inferiority

The registry defines the incidence rate as the number of participants who became HIV infected during the study after the first dose of study drug divided by the sum of participants' years of follow-up while at risk of HIV infection. A year is defined as 365.25 days in the registry definition.

6. Statistical Methodology

Poisson regression for infection incidence

The primary endpoint is a rate rather than simply a proportion. Participants can contribute different amounts of follow-up time, so the analysis uses person-years as the exposure scale. Poisson regression is therefore appropriate for modeling the number of observed infections while accounting for the amount of time participants were at risk.

Conceptual rate model
log(E[Y]) = log(PY) + β0 + β1 Treatment

Here, the person-years term functions as an exposure component. The treatment coefficient can be expressed through an exponentiated rate ratio comparing F/TAF with F/TDF.

Non-inferiority framework

The primary hypothesis was non-inferiority. The registry states that non-inferiority of F/TAF to F/TDF would be concluded if the upper bound of the two-sided 95.003% confidence interval for the rate ratio was less than 1.62.

Prespecified non-inferiority criterion

Upper CI bound < 1.62

Rate ratio defined as F/TAF incidence rate divided by F/TDF incidence rate.

The reported primary upper confidence bound was 1.149.

ANOVA and ANCOVA for continuous outcomes

The registry reports ANOVA for hip and spine BMD outcomes and ANCOVA for serum creatinine. The ANCOVA models incorporated baseline serum creatinine as a covariate, while baseline TVD for PrEP and treatment were included as fixed effects. This structure adjusts the comparison for a prespecified baseline measurement rather than comparing raw postbaseline values alone.

Van Elteren testing

The registry used the Van Elteren test for percent changes in urine beta-2-microglobulin-to-creatinine and retinol-binding-protein-to-creatinine ratios. This is a stratified nonparametric approach and is useful when the outcome distribution does not justify relying on a conventional parametric comparison.

Rank analysis of covariance

The number of participants by urine protein and urine protein-to-creatinine-ratio categories was analyzed using rank analysis of covariance. The registry normalizes this method under the ANCOVA family, while explicitly identifying the method as rank analysis of covariance in the analysis description.

7. Primary Result: HIV-1 Infection Incidence

The primary endpoint was analyzed in the Full Analysis Set using Poisson regression. The reported effect measure was the rate ratio for F/TAF versus F/TDF.

HIV-1 infection incidence rate ratio

0.468

95.003% two-sided CI: 0.191–1.149

Hypothesis: non-inferiority  ·  Non-inferiority margin: 1.62

Primary endpointF/TAF vs F/TDF
Effect measureRate ratio
Estimate0.468
Confidence interval95.003% two-sided CI: 0.191–1.149
Non-inferiority margin1.62
Analysis methodPoisson regression with generalized model associated with a Poisson distribution and logarithmic link
Analysis populationFull Analysis Set
Clinical Biostats interpretation

The rate ratio of 0.468 means that the estimated HIV-1 infection incidence rate in the F/TAF group was 0.468 times the corresponding rate in the F/TDF group under the fitted Poisson model. Expressed as a relative rate comparison, this corresponds to an estimated incidence rate about 53.2% lower for F/TAF relative to F/TDF.

The rate ratio is not a probability that an individual participant avoided HIV infection, and it is not a statement that every participant experienced a 53.2% reduction. It is a group-level relative rate estimate that accounts for person-time at risk.

The 95.003% confidence interval of 0.191–1.149 describes uncertainty around the estimated rate ratio under the specified statistical model and sampling framework. It is not a range containing the individual treatment effects experienced by participants.

The key feature for the non-inferiority conclusion is the upper confidence bound. Because 1.149 is below the prespecified margin of 1.62, the registry's stated non-inferiority criterion is satisfied by the reported estimate.

The confidence interval and non-inferiority margin should be interpreted together. The p-value is not the relevant measure of effect size in this framework; the magnitude of the rate ratio and the position of its confidence interval relative to the non-inferiority margin provide the central statistical information.

Why the non-inferiority margin matters

A non-inferiority analysis asks whether the new intervention is not unacceptably worse than the comparator according to a prespecified margin. Here, the margin is 1.62 for the rate ratio. Because the ratio is defined as F/TAF divided by F/TDF, values above 1 indicate a higher estimated infection rate for F/TAF, and the upper confidence bound is therefore the critical boundary for the stated non-inferiority rule.

This is different from simply asking whether a conventional null hypothesis of equal rates has been rejected. Non-inferiority is a margin-based design: the question is whether the data are sufficiently incompatible with an effect worse than the allowed margin.

8. Secondary Result: HIV-1 Infection Incidence at 96 Weeks

The registry also posted a secondary analysis of the same incidence endpoint when all participants had 96 weeks of follow-up after randomization or had permanently discontinued from the study, subject to the registry's stated maximum follow-up framework.

96-week HIV-1 infection incidence rate ratio

0.536

95.003% two-sided CI: 0.227–1.264

Hypothesis: non-inferiority  ·  Non-inferiority margin: 1.62

Clinical Biostats interpretation

The reported rate ratio of 0.536 represents the estimated HIV-1 infection incidence rate in F/TAF relative to F/TDF at the secondary 96-week analysis. As a relative rate measure, it corresponds to an estimated incidence rate about 46.4% lower in the F/TAF group under the model.

The 95.003% confidence interval was 0.227–1.264. Its upper bound, 1.264, remains below the prespecified non-inferiority margin of 1.62. Thus, the registry's stated non-inferiority criterion is also satisfied for this secondary incidence analysis.

The confidence interval is wider than the primary estimate's interval, and the point estimate differs from the primary estimate. Those differences illustrate why a treatment effect should not be reduced to a single number: the estimate depends on the analysis time frame, accumulated follow-up, and statistical uncertainty.

As with the primary result, the rate ratio does not represent an individual participant's probability of infection and does not describe an absolute risk difference.

9. Secondary Results: Bone Mineral Density

Hip BMD at Week 48

Difference in least squares mean percent change

1.142

95% two-sided CI: 0.628–1.655  ·  P < 0.0001

Hip BMD was compared between F/TAF and F/TDF using ANOVA, with baseline TVD for PrEP and treatment as fixed effects. The outcome was percent change from baseline at Week 48 in the blinded phase.

Clinical Biostats interpretation

The reported difference in least squares mean percent change was 1.142 percentage points, with a 95% confidence interval of 0.628–1.655. Because the confidence interval lies above zero, the estimated difference favors a greater percent-change value in the F/TAF group for this endpoint.

The P-value of <0.0001 describes evidence against the null comparison used for this superiority analysis; it does not quantify the size or clinical importance of the difference. The estimate and its confidence interval provide that effect-size information.

Spine BMD at Week 48

Difference in least squares mean percent change

1.567

95% two-sided CI: 0.913–2.220  ·  P < 0.0001

Spine BMD was compared using ANOVA with baseline TVD for PrEP and treatment as fixed effects.

Clinical Biostats interpretation

The estimated difference in least squares mean percent change was 1.567 percentage points, with a 95% confidence interval of 0.913–2.220. The interval excludes zero, while the registry reports P < 0.0001 for the superiority comparison.

These results describe a difference in a continuous biomarker outcome. They do not by themselves establish how an individual participant's BMD would change or what a particular difference means for a clinical outcome.

Hip BMD at Week 96

Difference in least squares mean percent change

1.567

95% two-sided CI: 0.896–2.237  ·  P < 0.0001

Spine BMD at Week 96

Difference in least squares mean percent change

2.253

95% two-sided CI: 1.437–3.069  ·  P < 0.0001

EndpointEffect estimate95% CIP-valueMethod
Hip BMD, Week 481.1420.628–1.655<0.0001ANOVA
Spine BMD, Week 481.5670.913–2.220<0.0001ANOVA
Hip BMD, Week 961.5670.896–2.237<0.0001ANOVA
Spine BMD, Week 962.2531.437–3.069<0.0001ANOVA

The four BMD analyses consistently report positive differences in least squares mean percent change. Statistically, each confidence interval excludes zero and each posted P-value is less than 0.0001. These are separate secondary outcomes, however, and their individual P-values should not automatically be interpreted as though they were isolated from the broader family of analyses reported for the trial.

10. Secondary Results: Serum Creatinine

Week 48

Difference in least squares mean change

-0.02 mg/dL

95% two-sided CI: -0.02 to -0.01  ·  P < 0.0001

The Week 48 analysis used ANCOVA and included baseline TVD for PrEP and treatment as fixed effects, with baseline serum creatinine as a covariate. Participants in the Safety Analysis Set with available data were analyzed.

Clinical Biostats interpretation

The estimated least squares mean difference was -0.02 mg/dL. The negative sign indicates that the estimated change from baseline was lower in the F/TAF group than in the F/TDF group under the fitted ANCOVA model.

The 95% confidence interval of -0.02 to -0.01 mg/dL quantifies uncertainty around that adjusted difference. The P-value of <0.0001 addresses the statistical comparison, not the magnitude or clinical importance of the laboratory difference.

Week 96

Difference in least squares mean change

-0.02 mg/dL

95% two-sided CI: -0.02 to -0.01  ·  P < 0.0001

The Week 96 analysis used the same ANCOVA framework, including baseline TVD for PrEP and treatment as fixed effects and baseline serum creatinine as a covariate.

Clinical Biostats interpretation

The Week 96 estimate and confidence interval are the same as those posted for Week 48: a difference of -0.02 mg/dL, with a 95% CI of -0.02 to -0.01 mg/dL and P < 0.0001.

The repeated appearance of a similar estimate at two time points should not be treated as two independent pieces of evidence without considering the longitudinal relationship between measurements. The registry reports the analyses separately, but the posted information does not provide a covariance structure or longitudinal model that would permit a different interpretation.

11. Secondary Results: Urinary Renal Biomarkers

EndpointTime frameMethodP-valueHypothesis
Percent change in urine beta-2-microglobulin to creatinine ratioBaseline, Week 48Van Elteren test<0.0001Superiority
Percent change in urine retinol binding protein to creatinine ratioBaseline, Week 48Van Elteren test<0.0001Superiority
Number of participants by UP and UPCR categoriesBaseline, Week 48Rank analysis of covariance0.0048Superiority
Percent change in urine beta-2-microglobulin to creatinine ratioBaseline, Week 96Van Elteren test<0.0001Superiority
Percent change in urine RBP to creatinine ratioBaseline, Week 96Van Elteren test<0.0001Superiority
Number of participants by UP and UPCR categoriesBaseline, Week 96Rank analysis of covariance0.2163Superiority

The registry reports statistically significant P-values for both Week 48 biomarker ratio comparisons, both Week 96 biomarker ratio comparisons, and the Week 48 urine protein category analysis. The Week 96 UP/UPCR category analysis has a reported P-value of 0.2163.

Importantly, the registry does not provide an effect estimate or confidence interval for these posted analyses in the ClinicalTrials.gov record. The appropriate interpretation is therefore limited to the reported test results and methods rather than attempting to reconstruct an effect size that was not reported.

Why a Van Elteren test?

The Van Elteren test is a stratified extension of the Wilcoxon rank-sum approach. It compares the distributions of an outcome between treatment groups while incorporating strata. Its use here places less reliance on normality assumptions than a conventional mean-based analysis would.

A P-value from a nonparametric test still does not tell us how large the treatment difference is. For effect-size interpretation, a corresponding difference estimate and confidence interval would be valuable, but those quantities are not contained in the registry analysis data for these endpoints.

12. Secondary Endpoint Results: Consolidated View

EndpointEstimate95% CIP-valueMethod
HIV-1 infection incidence, primary analysisRate ratio 0.4680.191–1.149 (95.003%)Not reported in registry-reported analysisPoisson regression
HIV-1 infection incidence, 96 weeksRate ratio 0.5360.227–1.264 (95.003%)Not reported in registry-reported analysisPoisson regression
Hip BMD, Week 481.1420.628–1.655<0.0001ANOVA
Spine BMD, Week 481.5670.913–2.220<0.0001ANOVA
Serum creatinine, Week 48-0.02 mg/dL-0.02 to -0.01<0.0001ANCOVA
Hip BMD, Week 961.5670.896–2.237<0.0001ANOVA
Spine BMD, Week 962.2531.437–3.069<0.0001ANOVA
Serum creatinine, Week 96-0.02 mg/dL-0.02 to -0.01<0.0001ANCOVA
Urine beta-2-microglobulin/creatinine, Week 48Not reportedNot reported<0.0001Van Elteren test
Urine RBP/creatinine, Week 48Not reportedNot reported<0.0001Van Elteren test
UP/UPCR categories, Week 48Not reportedNot reported0.0048Rank ANCOVA
Urine beta-2-microglobulin/creatinine, Week 96Not reportedNot reported<0.0001Van Elteren test
Urine RBP/creatinine, Week 96Not reportedNot reported<0.0001Van Elteren test
UP/UPCR categories, Week 96Not reportedNot reported0.2163Rank ANCOVA

The registry reports 16 outcome measures and 14 statistical analyses. The ClinicalTrials.gov record identifies one primary endpoint analysis and the secondary analyses summarized above. Where the registry supplies only a P-value, this page does not infer an effect estimate or confidence interval.

13. Safety Results

The ClinicalTrials.gov record provides serious adverse-event counts by the two active treatment groups.

Treatment groupParticipants with serious AEsParticipants at risk
Descovy (DVY)2022694
Truvada (TVD)1862693
Serious adverse events: affected / at risk
Descovy (DVY)
202 / 2694
Truvada (TVD)
186 / 2693

The reported serious-adverse-event counts are 202/2694 for Descovy and 186/2693 for Truvada. The ClinicalTrials.gov record does not provide a formal statistical comparison for these serious-adverse-event counts, so no comparative P-value or confidence interval is assigned here.

Safety interpretation: counts of participants affected by serious adverse events are not interchangeable with a formal risk ratio, rate ratio, or odds ratio. A comparative safety analysis would require the appropriate event definition, follow-up structure, analysis population, and statistical method.

14. Statistical Methods Explained

Why was Poisson regression used for the primary endpoint?

The endpoint is an incidence rate per 100 person-years. Participants can contribute different amounts of time while at risk, so modeling event counts together with person-time exposure is more appropriate than treating every participant as having identical follow-up. The registry specifically identifies Poisson regression and a logarithmic link for the primary analysis.

What does a rate ratio of 0.468 mean?

A rate ratio of 0.468 means that the estimated HIV-1 infection incidence rate under F/TAF was 0.468 times the estimated rate under F/TDF. The calculation is based on incidence rates and person-time, not simply the fraction of participants with an event.

Why is non-inferiority judged against the margin?

The purpose of a non-inferiority analysis is to determine whether the new treatment could be worse than the comparator by more than an unacceptable prespecified amount. Here, the registry defines that amount through a rate-ratio margin of 1.62. The upper confidence bound is therefore compared with 1.62 rather than relying only on whether a conventional equality test produces a small P-value.

Why does the confidence interval matter more than the P-value for the non-inferiority conclusion?

The confidence interval shows both the estimated treatment effect and the uncertainty surrounding it. For this trial, the upper limit of the 95.003% CI is the quantity that must remain below 1.62 under the registry's stated criterion. A P-value, by contrast, is a measure of evidence under a specified null hypothesis and does not directly communicate the magnitude or precision of the treatment effect.

Why was ANCOVA used for serum creatinine?

ANCOVA permits the treatment comparison to be adjusted for baseline serum creatinine. The registry specifies treatment and baseline TVD for PrEP as fixed effects and baseline serum creatinine as a covariate. This can improve the precision and interpretability of the comparison when the baseline measurement is related to the outcome.

Why were ANOVA and least squares means used for BMD?

The BMD endpoints were continuous percent changes from baseline. The registry specifies ANOVA with treatment and baseline TVD for PrEP as fixed effects and reports differences in least squares means. A least squares mean is a model-adjusted group mean rather than simply the arithmetic average of the observed outcome values.

What does a Van Elteren test add?

The Van Elteren test provides a stratified nonparametric comparison. It is useful when rank-based inference is preferred over a normal-theory comparison. In DISCOVER, it was used for the urinary biomarker percent-change outcomes, with the registry reporting P-values but not corresponding effect estimates in the ClinicalTrials.gov record.

15. Confidence Intervals and Statistical Precision

The primary analysis illustrates why confidence intervals should be read as part of the treatment estimate rather than as an afterthought.

Primary rate-ratio interval
Rate ratio = 0.468    95.003% CI = 0.191–1.149

The point estimate is below 1, while the entire reported confidence interval remains below the non-inferiority margin of 1.62.

The interval is compatible with a range of relative rate effects. Its lower and upper limits should not be interpreted as minimum and maximum effects that individual participants could experience. They quantify statistical uncertainty around the estimated group-level rate ratio.

The secondary 96-week analysis provides a useful comparison: the rate ratio is 0.536 with a 95.003% CI of 0.227–1.264. Both upper bounds are below 1.62, but the estimates and interval widths differ. This demonstrates that statistical inference can change with the amount and timing of accumulated follow-up without requiring the underlying trial design to change.

16. Covariate Adjustment in the BMD and Creatinine Analyses

Covariate adjustment is one of the most important methodological distinctions between the trial's continuous-outcome analyses.

BMD models

ANOVA included baseline TVD for PrEP and treatment as fixed effects. The reported effect was the difference in least squares mean percent change.

Creatinine models

ANCOVA included baseline TVD for PrEP and treatment as fixed effects and baseline serum creatinine as a covariate.

Adjustment does not turn an observational comparison into a randomized comparison; rather, it uses information already measured at baseline to estimate the treatment contrast more efficiently within the specified model. Because the trial itself was randomized, the treatment comparison retains the principal advantage of random allocation, while the covariate-adjusted models address the particular continuous endpoints being analyzed.

17. Multiplicity and Multiple Secondary Endpoints

The ClinicalTrials.gov record identifies 1 primary endpoint, 16 outcome measures, and 14 statistical analyses. The analysis set includes the primary non-inferiority comparison plus numerous secondary outcomes covering BMD, serum creatinine, urinary biomarkers, and urine protein categories.

FeatureRegistry-supported informationStatistical implication
Primary endpoint1Primary non-inferiority inference centers on HIV-1 infection incidence.
Outcome measures posted16Multiple secondary outcomes were evaluated.
Statistical analyses posted14Multiple formal analyses are available in the registry.
Hypothesis typesNon-inferiority and superiorityThe interpretation depends on the prespecified hypothesis for each endpoint.

Multiple secondary analyses create a multiplicity issue even when each individual P-value is calculated correctly. A small P-value from one of many secondary analyses does not automatically carry the same confirmatory interpretation as a prespecified primary endpoint. The ClinicalTrials.gov record does not provide a complete multiplicity-adjustment procedure for all secondary endpoints, so this page does not infer one.

18. Intention-to-Treat Principles and Analysis Sets

The primary endpoint used the Full Analysis Set, whose registry definition begins with randomized participants and requires treatment exposure and postbaseline HIV laboratory information. The analysis therefore retains a strong connection to randomized treatment assignment while applying the registry's additional criteria for evaluability.

Secondary laboratory outcomes use more specialized analysis sets. This is particularly important for DXA outcomes, where the registry specifies hip and spine DXA Analysis Sets with required BMD measurements.

Interpretation point: "intention-to-treat" and "available-case" analyses are not interchangeable concepts. Randomization protects the treatment comparison, but restricting an endpoint to participants with available postbaseline measurements can change the population contributing information to that particular analysis.

19. Trial Timeline

2016-09-02

Study start

The DISCOVER trial began on September 2, 2016.

Phase 3

Randomized double-blind comparison

The trial used randomized allocation, double masking, and a parallel design to compare F/TAF with F/TDF for HIV-1 pre-exposure prophylaxis.

2019-01-31

Primary completion

The registered primary completion date was January 31, 2019.

Registry results

Statistical analyses posted

The registry reports 16 outcome measures and 14 statistical analyses, including the primary rate-ratio analysis and multiple secondary laboratory and BMD analyses.

20. Rate Ratios vs Mean Differences

DISCOVER is particularly useful statistically because its reported outcomes use several different effect measures. The primary endpoint uses a rate ratio, while the BMD and serum-creatinine analyses use differences in least squares means.

Effect measureExample in DISCOVERWhat it describes
Rate ratio0.468 for HIV-1 infection incidenceRelative incidence rate in F/TAF compared with F/TDF.
Difference in least squares means1.142 for hip BMD at Week 48Adjusted difference between treatment-group mean percent changes.
Difference in least squares means-0.02 mg/dL for serum creatinineAdjusted difference in change from baseline between groups.
P-value without effect estimate<0.0001 for urinary biomarkersEvidence against the corresponding null comparison; effect magnitude is not reported in the ClinicalTrials.gov record.

These measures cannot be compared numerically as if they were on the same scale. A rate ratio of 0.468 and a mean difference of 1.142 answer fundamentally different questions about different endpoints.

21. Non-Inferiority Logic in Detail

The primary analysis provides a clean example of why non-inferiority should be interpreted through a confidence interval and a clinically prespecified margin.

Point estimate

The rate ratio of 0.468 is below 1, indicating a lower estimated HIV-1 infection incidence rate with F/TAF under the fitted model.

Uncertainty

The 95.003% CI is 0.191–1.149, showing that the plausible range around the estimate is wider than the point estimate alone.

Non-inferiority boundary

The prespecified upper boundary is 1.62. The observed upper confidence bound is 1.149.

Conclusion under the registry rule

Because 1.149 is below 1.62, the stated non-inferiority criterion is met.

This logic differs from a superiority framework. The primary analysis does not require the entire confidence interval to lie below 1 to establish non-inferiority. Instead, it requires the upper confidence limit to remain below the prespecified threshold for unacceptable inferiority.

22. Limitations and Interpretation Issues

23. Why This Trial Matters Statistically

DISCOVER is a useful teaching case because it combines a non-inferiority incidence-rate analysis with several different approaches to continuous and nonparametric secondary outcomes.

ConceptHow it appears in DISCOVER
RandomizationRandomized allocation in a phase 3 parallel-group design.
BlindingDouble masking.
Intention-to-treat principleThe primary analysis is conducted in a Full Analysis Set derived from randomized participants meeting the registry criteria.
Rate ratioPrimary HIV-1 infection incidence comparison: 0.468.
Poisson regressionPrimary incidence-rate analysis using a Poisson model and logarithmic link.
Non-inferiorityUpper 95.003% CI compared with a margin of 1.62.
Confidence intervalPrimary 95.003% CI: 0.191–1.149.
ANOVAHip and spine BMD comparisons.
ANCOVASerum creatinine comparisons with baseline covariate adjustment.
Least squares meanEffect measure for BMD and serum creatinine analyses.
Van Elteren testUrinary beta-2-microglobulin and RBP ratio analyses.
Rank ANCOVAUrine protein and UPCR category analyses.
MultiplicityOne primary endpoint alongside numerous secondary outcomes and analyses.

The statistical lesson is broader than any single result. A clinical trial can contain several legitimate statistical questions, each requiring an effect measure and model appropriate to the outcome. The primary incidence endpoint requires rate-based inference, while BMD and serum creatinine require model-adjusted continuous-outcome comparisons.

24. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The primary rate ratio was 0.468 with a 95.003% CI of 0.191–1.149. Because the upper bound was below the prespecified non-inferiority margin of 1.62, the registry's stated non-inferiority criterion was satisfied.

Endpoint-specific interpretation

Secondary analyses reported differences in BMD and serum creatinine, as well as statistically significant or nonsignificant P-values for urinary biomarkers and protein categories. These endpoints answer different questions and require separate interpretation.

A statistically rigorous summary therefore avoids collapsing the trial into one generalized "positive" or "negative" statement. The primary endpoint has a specific non-inferiority interpretation, while each secondary endpoint has its own estimand, analysis population, effect measure, and uncertainty.

25. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

26. Related Statistical Calculators

27. Sources

Continue through Clinical Biostats

Use the statistical concepts from DISCOVER to explore methods for non-inferiority trials, rate data, covariate adjustment, and clinical-trial inference.

28. Record Summary

DISCOVER provides a compact example of several important clinical-trial statistical principles. The primary endpoint was an HIV-1 infection incidence rate analyzed with Poisson regression, expressed as a rate ratio, and evaluated under a prespecified non-inferiority margin of 1.62. The reported rate ratio was 0.468, with a 95.003% two-sided confidence interval of 0.191–1.149, placing the upper confidence bound below the non-inferiority margin.

The secondary analyses broaden the statistical picture. BMD outcomes used ANOVA and differences in least squares means; serum creatinine used ANCOVA with baseline adjustment; urinary biomarker outcomes used the Van Elteren test; and urine protein category outcomes used rank analysis of covariance. The registry also reports serious adverse-event counts of 202/2694 for Descovy and 186/2693 for Truvada.

Clinical Biostats methodology: A trial-results page should distinguish the estimand, effect measure, confidence interval, hypothesis type, analysis population, and statistical model for each endpoint. DISCOVER illustrates why those distinctions matter: a non-inferiority rate ratio, an adjusted mean difference, and a nonparametric P-value cannot be interpreted as though they were the same statistical object.