This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
DISCOVER was a randomized, double-blind, parallel phase 3 trial evaluating F/TAF versus F/TDF for pre-exposure prophylaxis of HIV-1 infection. The registered primary endpoint was the incidence of HIV-1 infection per 100 person-years, with a prespecified non-inferiority comparison using a rate ratio.
| Feature | DISCOVER |
|---|---|
| Phase | Phase 3 |
| Therapeutic area | Infectious Disease |
| Condition | Pre-Exposure Prophylaxis of HIV-1 Infection |
| Design | Randomized, double-blind, parallel |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 5399 |
| Primary endpoint | Incidence of HIV-1 Infection Per 100 Person Years (PY) |
| Primary endpoint type | Count / rate |
| Lead sponsor | Gilead Sciences |
| Sponsor type | Industry |
| Status | Terminated |
| Study dates | Start: 2016-09-02; primary completion: 2019-01-31 |
| ClinicalTrials.gov | NCT02842086 |
2. Clinical Question
The central statistical question was whether F/TAF was non-inferior to F/TDF for the incidence of HIV-1 infection per 100 person-years. The registry reports this endpoint as the number of participants who became HIV infected after the first dose of study drug divided by the total person-years of follow-up while participants were at risk of HIV infection.
Population
Men and transgender women who have sex with men and are at risk of HIV-1 infection, according to the registered trial title.
Intervention
F/TAF fixed-dose combination once daily for pre-exposure prophylaxis.
Comparator
F/TDF fixed-dose combination once daily for pre-exposure prophylaxis.
Primary question
Is the HIV-1 infection incidence rate with F/TAF sufficiently close to that with F/TDF to satisfy the prespecified non-inferiority criterion?
3. Trial Design
F/TAF
- F/TAF drug intervention
- Once-daily fixed-dose combination for pre-exposure prophylaxis
- Registry abbreviation: DVY / Descovy in the serious-adverse-event data
F/TDF
- F/TDF drug comparator
- Once-daily fixed-dose combination for pre-exposure prophylaxis
- Registry abbreviation: TVD / Truvada in the serious-adverse-event data
The registry lists four interventions: F/TAF, F/TDF, F/TAF placebo, and F/TDF placebo. The overall design is therefore represented in the registry as a four-arm double-blind parallel study, while the primary statistical comparison is between the F/TAF and F/TDF treatment groups.
4. Randomization, Blinding, and Analysis Populations
DISCOVER used randomized allocation and double masking. The primary analysis was performed in the Full Analysis Set, which the registry defines as participants randomized into the study who received at least 1 dose of study drug, were not HIV positive on Day 1, and had at least 1 postbaseline HIV laboratory assessment.
| Analysis population | Registry-supported role |
|---|---|
| Full Analysis Set | Primary HIV-1 infection incidence analysis; participants were randomized, received at least 1 dose, were not HIV positive on Day 1, and had at least 1 postbaseline HIV laboratory assessment. |
| Safety Analysis Set | Used for serum creatinine and renal biomarker analyses; generally included randomized participants who received at least 1 dose, with available data for the relevant analysis. |
| Hip DXA Analysis Set | DXA substudy population with randomized treatment, at least 1 dose, and nonmissing hip BMD data as specified by the relevant endpoint. |
| Spine DXA Analysis Set | DXA substudy population with randomized treatment, at least 1 dose, and available spine BMD data as specified by the relevant endpoint. |
The distinction between these populations matters. Randomization establishes the principal basis for the efficacy comparison, whereas laboratory and DXA outcomes may be restricted to participants with the required postbaseline measurements. Consequently, estimates for these secondary endpoints describe the corresponding analysis sets rather than automatically the entire enrolled population.
5. Primary Endpoint
| Endpoint | Registered definition / time frame | Primary analysis |
|---|---|---|
| Incidence of HIV-1 Infection Per 100 Person Years (PY) | When all participants completed minimum follow-up of 48 weeks and at least 50% of the participants completed 96 weeks of follow-up. | Poisson regression; rate ratio; non-inferiority |
The registry defines the incidence rate as the number of participants who became HIV infected during the study after the first dose of study drug divided by the sum of participants' years of follow-up while at risk of HIV infection. A year is defined as 365.25 days in the registry definition.
6. Statistical Methodology
Poisson regression for infection incidence
The primary endpoint is a rate rather than simply a proportion. Participants can contribute different amounts of follow-up time, so the analysis uses person-years as the exposure scale. Poisson regression is therefore appropriate for modeling the number of observed infections while accounting for the amount of time participants were at risk.
Here, the person-years term functions as an exposure component. The treatment coefficient can be expressed through an exponentiated rate ratio comparing F/TAF with F/TDF.
Non-inferiority framework
The primary hypothesis was non-inferiority. The registry states that non-inferiority of F/TAF to F/TDF would be concluded if the upper bound of the two-sided 95.003% confidence interval for the rate ratio was less than 1.62.
Prespecified non-inferiority criterion
Rate ratio defined as F/TAF incidence rate divided by F/TDF incidence rate.
The reported primary upper confidence bound was 1.149.
ANOVA and ANCOVA for continuous outcomes
The registry reports ANOVA for hip and spine BMD outcomes and ANCOVA for serum creatinine. The ANCOVA models incorporated baseline serum creatinine as a covariate, while baseline TVD for PrEP and treatment were included as fixed effects. This structure adjusts the comparison for a prespecified baseline measurement rather than comparing raw postbaseline values alone.
Van Elteren testing
The registry used the Van Elteren test for percent changes in urine beta-2-microglobulin-to-creatinine and retinol-binding-protein-to-creatinine ratios. This is a stratified nonparametric approach and is useful when the outcome distribution does not justify relying on a conventional parametric comparison.
Rank analysis of covariance
The number of participants by urine protein and urine protein-to-creatinine-ratio categories was analyzed using rank analysis of covariance. The registry normalizes this method under the ANCOVA family, while explicitly identifying the method as rank analysis of covariance in the analysis description.
7. Primary Result: HIV-1 Infection Incidence
The primary endpoint was analyzed in the Full Analysis Set using Poisson regression. The reported effect measure was the rate ratio for F/TAF versus F/TDF.
HIV-1 infection incidence rate ratio
95.003% two-sided CI: 0.191–1.149
Hypothesis: non-inferiority · Non-inferiority margin: 1.62
| Primary endpoint | F/TAF vs F/TDF |
|---|---|
| Effect measure | Rate ratio |
| Estimate | 0.468 |
| Confidence interval | 95.003% two-sided CI: 0.191–1.149 |
| Non-inferiority margin | 1.62 |
| Analysis method | Poisson regression with generalized model associated with a Poisson distribution and logarithmic link |
| Analysis population | Full Analysis Set |
The rate ratio of 0.468 means that the estimated HIV-1 infection incidence rate in the F/TAF group was 0.468 times the corresponding rate in the F/TDF group under the fitted Poisson model. Expressed as a relative rate comparison, this corresponds to an estimated incidence rate about 53.2% lower for F/TAF relative to F/TDF.
The rate ratio is not a probability that an individual participant avoided HIV infection, and it is not a statement that every participant experienced a 53.2% reduction. It is a group-level relative rate estimate that accounts for person-time at risk.
The 95.003% confidence interval of 0.191–1.149 describes uncertainty around the estimated rate ratio under the specified statistical model and sampling framework. It is not a range containing the individual treatment effects experienced by participants.
The key feature for the non-inferiority conclusion is the upper confidence bound. Because 1.149 is below the prespecified margin of 1.62, the registry's stated non-inferiority criterion is satisfied by the reported estimate.
The confidence interval and non-inferiority margin should be interpreted together. The p-value is not the relevant measure of effect size in this framework; the magnitude of the rate ratio and the position of its confidence interval relative to the non-inferiority margin provide the central statistical information.
Why the non-inferiority margin matters
A non-inferiority analysis asks whether the new intervention is not unacceptably worse than the comparator according to a prespecified margin. Here, the margin is 1.62 for the rate ratio. Because the ratio is defined as F/TAF divided by F/TDF, values above 1 indicate a higher estimated infection rate for F/TAF, and the upper confidence bound is therefore the critical boundary for the stated non-inferiority rule.
This is different from simply asking whether a conventional null hypothesis of equal rates has been rejected. Non-inferiority is a margin-based design: the question is whether the data are sufficiently incompatible with an effect worse than the allowed margin.
8. Secondary Result: HIV-1 Infection Incidence at 96 Weeks
The registry also posted a secondary analysis of the same incidence endpoint when all participants had 96 weeks of follow-up after randomization or had permanently discontinued from the study, subject to the registry's stated maximum follow-up framework.
96-week HIV-1 infection incidence rate ratio
95.003% two-sided CI: 0.227–1.264
Hypothesis: non-inferiority · Non-inferiority margin: 1.62
The reported rate ratio of 0.536 represents the estimated HIV-1 infection incidence rate in F/TAF relative to F/TDF at the secondary 96-week analysis. As a relative rate measure, it corresponds to an estimated incidence rate about 46.4% lower in the F/TAF group under the model.
The 95.003% confidence interval was 0.227–1.264. Its upper bound, 1.264, remains below the prespecified non-inferiority margin of 1.62. Thus, the registry's stated non-inferiority criterion is also satisfied for this secondary incidence analysis.
The confidence interval is wider than the primary estimate's interval, and the point estimate differs from the primary estimate. Those differences illustrate why a treatment effect should not be reduced to a single number: the estimate depends on the analysis time frame, accumulated follow-up, and statistical uncertainty.
As with the primary result, the rate ratio does not represent an individual participant's probability of infection and does not describe an absolute risk difference.
9. Secondary Results: Bone Mineral Density
Hip BMD at Week 48
Difference in least squares mean percent change
95% two-sided CI: 0.628–1.655 · P < 0.0001
Hip BMD was compared between F/TAF and F/TDF using ANOVA, with baseline TVD for PrEP and treatment as fixed effects. The outcome was percent change from baseline at Week 48 in the blinded phase.
The reported difference in least squares mean percent change was 1.142 percentage points, with a 95% confidence interval of 0.628–1.655. Because the confidence interval lies above zero, the estimated difference favors a greater percent-change value in the F/TAF group for this endpoint.
The P-value of <0.0001 describes evidence against the null comparison used for this superiority analysis; it does not quantify the size or clinical importance of the difference. The estimate and its confidence interval provide that effect-size information.
Spine BMD at Week 48
Difference in least squares mean percent change
95% two-sided CI: 0.913–2.220 · P < 0.0001
Spine BMD was compared using ANOVA with baseline TVD for PrEP and treatment as fixed effects.
The estimated difference in least squares mean percent change was 1.567 percentage points, with a 95% confidence interval of 0.913–2.220. The interval excludes zero, while the registry reports P < 0.0001 for the superiority comparison.
These results describe a difference in a continuous biomarker outcome. They do not by themselves establish how an individual participant's BMD would change or what a particular difference means for a clinical outcome.
Hip BMD at Week 96
Difference in least squares mean percent change
95% two-sided CI: 0.896–2.237 · P < 0.0001
Spine BMD at Week 96
Difference in least squares mean percent change
95% two-sided CI: 1.437–3.069 · P < 0.0001
| Endpoint | Effect estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| Hip BMD, Week 48 | 1.142 | 0.628–1.655 | <0.0001 | ANOVA |
| Spine BMD, Week 48 | 1.567 | 0.913–2.220 | <0.0001 | ANOVA |
| Hip BMD, Week 96 | 1.567 | 0.896–2.237 | <0.0001 | ANOVA |
| Spine BMD, Week 96 | 2.253 | 1.437–3.069 | <0.0001 | ANOVA |
The four BMD analyses consistently report positive differences in least squares mean percent change. Statistically, each confidence interval excludes zero and each posted P-value is less than 0.0001. These are separate secondary outcomes, however, and their individual P-values should not automatically be interpreted as though they were isolated from the broader family of analyses reported for the trial.
10. Secondary Results: Serum Creatinine
Week 48
Difference in least squares mean change
95% two-sided CI: -0.02 to -0.01 · P < 0.0001
The Week 48 analysis used ANCOVA and included baseline TVD for PrEP and treatment as fixed effects, with baseline serum creatinine as a covariate. Participants in the Safety Analysis Set with available data were analyzed.
The estimated least squares mean difference was -0.02 mg/dL. The negative sign indicates that the estimated change from baseline was lower in the F/TAF group than in the F/TDF group under the fitted ANCOVA model.
The 95% confidence interval of -0.02 to -0.01 mg/dL quantifies uncertainty around that adjusted difference. The P-value of <0.0001 addresses the statistical comparison, not the magnitude or clinical importance of the laboratory difference.
Week 96
Difference in least squares mean change
95% two-sided CI: -0.02 to -0.01 · P < 0.0001
The Week 96 analysis used the same ANCOVA framework, including baseline TVD for PrEP and treatment as fixed effects and baseline serum creatinine as a covariate.
The Week 96 estimate and confidence interval are the same as those posted for Week 48: a difference of -0.02 mg/dL, with a 95% CI of -0.02 to -0.01 mg/dL and P < 0.0001.
The repeated appearance of a similar estimate at two time points should not be treated as two independent pieces of evidence without considering the longitudinal relationship between measurements. The registry reports the analyses separately, but the posted information does not provide a covariance structure or longitudinal model that would permit a different interpretation.
11. Secondary Results: Urinary Renal Biomarkers
| Endpoint | Time frame | Method | P-value | Hypothesis |
|---|---|---|---|---|
| Percent change in urine beta-2-microglobulin to creatinine ratio | Baseline, Week 48 | Van Elteren test | <0.0001 | Superiority |
| Percent change in urine retinol binding protein to creatinine ratio | Baseline, Week 48 | Van Elteren test | <0.0001 | Superiority |
| Number of participants by UP and UPCR categories | Baseline, Week 48 | Rank analysis of covariance | 0.0048 | Superiority |
| Percent change in urine beta-2-microglobulin to creatinine ratio | Baseline, Week 96 | Van Elteren test | <0.0001 | Superiority |
| Percent change in urine RBP to creatinine ratio | Baseline, Week 96 | Van Elteren test | <0.0001 | Superiority |
| Number of participants by UP and UPCR categories | Baseline, Week 96 | Rank analysis of covariance | 0.2163 | Superiority |
The registry reports statistically significant P-values for both Week 48 biomarker ratio comparisons, both Week 96 biomarker ratio comparisons, and the Week 48 urine protein category analysis. The Week 96 UP/UPCR category analysis has a reported P-value of 0.2163.
Importantly, the registry does not provide an effect estimate or confidence interval for these posted analyses in the ClinicalTrials.gov record. The appropriate interpretation is therefore limited to the reported test results and methods rather than attempting to reconstruct an effect size that was not reported.
Why a Van Elteren test?
The Van Elteren test is a stratified extension of the Wilcoxon rank-sum approach. It compares the distributions of an outcome between treatment groups while incorporating strata. Its use here places less reliance on normality assumptions than a conventional mean-based analysis would.
A P-value from a nonparametric test still does not tell us how large the treatment difference is. For effect-size interpretation, a corresponding difference estimate and confidence interval would be valuable, but those quantities are not contained in the registry analysis data for these endpoints.
12. Secondary Endpoint Results: Consolidated View
| Endpoint | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|
| HIV-1 infection incidence, primary analysis | Rate ratio 0.468 | 0.191–1.149 (95.003%) | Not reported in registry-reported analysis | Poisson regression |
| HIV-1 infection incidence, 96 weeks | Rate ratio 0.536 | 0.227–1.264 (95.003%) | Not reported in registry-reported analysis | Poisson regression |
| Hip BMD, Week 48 | 1.142 | 0.628–1.655 | <0.0001 | ANOVA |
| Spine BMD, Week 48 | 1.567 | 0.913–2.220 | <0.0001 | ANOVA |
| Serum creatinine, Week 48 | -0.02 mg/dL | -0.02 to -0.01 | <0.0001 | ANCOVA |
| Hip BMD, Week 96 | 1.567 | 0.896–2.237 | <0.0001 | ANOVA |
| Spine BMD, Week 96 | 2.253 | 1.437–3.069 | <0.0001 | ANOVA |
| Serum creatinine, Week 96 | -0.02 mg/dL | -0.02 to -0.01 | <0.0001 | ANCOVA |
| Urine beta-2-microglobulin/creatinine, Week 48 | Not reported | Not reported | <0.0001 | Van Elteren test |
| Urine RBP/creatinine, Week 48 | Not reported | Not reported | <0.0001 | Van Elteren test |
| UP/UPCR categories, Week 48 | Not reported | Not reported | 0.0048 | Rank ANCOVA |
| Urine beta-2-microglobulin/creatinine, Week 96 | Not reported | Not reported | <0.0001 | Van Elteren test |
| Urine RBP/creatinine, Week 96 | Not reported | Not reported | <0.0001 | Van Elteren test |
| UP/UPCR categories, Week 96 | Not reported | Not reported | 0.2163 | Rank ANCOVA |
The registry reports 16 outcome measures and 14 statistical analyses. The ClinicalTrials.gov record identifies one primary endpoint analysis and the secondary analyses summarized above. Where the registry supplies only a P-value, this page does not infer an effect estimate or confidence interval.
13. Safety Results
The ClinicalTrials.gov record provides serious adverse-event counts by the two active treatment groups.
| Treatment group | Participants with serious AEs | Participants at risk |
|---|---|---|
| Descovy (DVY) | 202 | 2694 |
| Truvada (TVD) | 186 | 2693 |
The reported serious-adverse-event counts are 202/2694 for Descovy and 186/2693 for Truvada. The ClinicalTrials.gov record does not provide a formal statistical comparison for these serious-adverse-event counts, so no comparative P-value or confidence interval is assigned here.
14. Statistical Methods Explained
Why was Poisson regression used for the primary endpoint?
The endpoint is an incidence rate per 100 person-years. Participants can contribute different amounts of time while at risk, so modeling event counts together with person-time exposure is more appropriate than treating every participant as having identical follow-up. The registry specifically identifies Poisson regression and a logarithmic link for the primary analysis.
What does a rate ratio of 0.468 mean?
A rate ratio of 0.468 means that the estimated HIV-1 infection incidence rate under F/TAF was 0.468 times the estimated rate under F/TDF. The calculation is based on incidence rates and person-time, not simply the fraction of participants with an event.
Why is non-inferiority judged against the margin?
The purpose of a non-inferiority analysis is to determine whether the new treatment could be worse than the comparator by more than an unacceptable prespecified amount. Here, the registry defines that amount through a rate-ratio margin of 1.62. The upper confidence bound is therefore compared with 1.62 rather than relying only on whether a conventional equality test produces a small P-value.
Why does the confidence interval matter more than the P-value for the non-inferiority conclusion?
The confidence interval shows both the estimated treatment effect and the uncertainty surrounding it. For this trial, the upper limit of the 95.003% CI is the quantity that must remain below 1.62 under the registry's stated criterion. A P-value, by contrast, is a measure of evidence under a specified null hypothesis and does not directly communicate the magnitude or precision of the treatment effect.
Why was ANCOVA used for serum creatinine?
ANCOVA permits the treatment comparison to be adjusted for baseline serum creatinine. The registry specifies treatment and baseline TVD for PrEP as fixed effects and baseline serum creatinine as a covariate. This can improve the precision and interpretability of the comparison when the baseline measurement is related to the outcome.
Why were ANOVA and least squares means used for BMD?
The BMD endpoints were continuous percent changes from baseline. The registry specifies ANOVA with treatment and baseline TVD for PrEP as fixed effects and reports differences in least squares means. A least squares mean is a model-adjusted group mean rather than simply the arithmetic average of the observed outcome values.
What does a Van Elteren test add?
The Van Elteren test provides a stratified nonparametric comparison. It is useful when rank-based inference is preferred over a normal-theory comparison. In DISCOVER, it was used for the urinary biomarker percent-change outcomes, with the registry reporting P-values but not corresponding effect estimates in the ClinicalTrials.gov record.
15. Confidence Intervals and Statistical Precision
The primary analysis illustrates why confidence intervals should be read as part of the treatment estimate rather than as an afterthought.
The point estimate is below 1, while the entire reported confidence interval remains below the non-inferiority margin of 1.62.
The interval is compatible with a range of relative rate effects. Its lower and upper limits should not be interpreted as minimum and maximum effects that individual participants could experience. They quantify statistical uncertainty around the estimated group-level rate ratio.
The secondary 96-week analysis provides a useful comparison: the rate ratio is 0.536 with a 95.003% CI of 0.227–1.264. Both upper bounds are below 1.62, but the estimates and interval widths differ. This demonstrates that statistical inference can change with the amount and timing of accumulated follow-up without requiring the underlying trial design to change.
16. Covariate Adjustment in the BMD and Creatinine Analyses
Covariate adjustment is one of the most important methodological distinctions between the trial's continuous-outcome analyses.
BMD models
ANOVA included baseline TVD for PrEP and treatment as fixed effects. The reported effect was the difference in least squares mean percent change.
Creatinine models
ANCOVA included baseline TVD for PrEP and treatment as fixed effects and baseline serum creatinine as a covariate.
Adjustment does not turn an observational comparison into a randomized comparison; rather, it uses information already measured at baseline to estimate the treatment contrast more efficiently within the specified model. Because the trial itself was randomized, the treatment comparison retains the principal advantage of random allocation, while the covariate-adjusted models address the particular continuous endpoints being analyzed.
17. Multiplicity and Multiple Secondary Endpoints
The ClinicalTrials.gov record identifies 1 primary endpoint, 16 outcome measures, and 14 statistical analyses. The analysis set includes the primary non-inferiority comparison plus numerous secondary outcomes covering BMD, serum creatinine, urinary biomarkers, and urine protein categories.
| Feature | Registry-supported information | Statistical implication |
|---|---|---|
| Primary endpoint | 1 | Primary non-inferiority inference centers on HIV-1 infection incidence. |
| Outcome measures posted | 16 | Multiple secondary outcomes were evaluated. |
| Statistical analyses posted | 14 | Multiple formal analyses are available in the registry. |
| Hypothesis types | Non-inferiority and superiority | The interpretation depends on the prespecified hypothesis for each endpoint. |
Multiple secondary analyses create a multiplicity issue even when each individual P-value is calculated correctly. A small P-value from one of many secondary analyses does not automatically carry the same confirmatory interpretation as a prespecified primary endpoint. The ClinicalTrials.gov record does not provide a complete multiplicity-adjustment procedure for all secondary endpoints, so this page does not infer one.
18. Intention-to-Treat Principles and Analysis Sets
The primary endpoint used the Full Analysis Set, whose registry definition begins with randomized participants and requires treatment exposure and postbaseline HIV laboratory information. The analysis therefore retains a strong connection to randomized treatment assignment while applying the registry's additional criteria for evaluability.
Secondary laboratory outcomes use more specialized analysis sets. This is particularly important for DXA outcomes, where the registry specifies hip and spine DXA Analysis Sets with required BMD measurements.
19. Trial Timeline
Study start
The DISCOVER trial began on September 2, 2016.
Randomized double-blind comparison
The trial used randomized allocation, double masking, and a parallel design to compare F/TAF with F/TDF for HIV-1 pre-exposure prophylaxis.
Primary completion
The registered primary completion date was January 31, 2019.
Statistical analyses posted
The registry reports 16 outcome measures and 14 statistical analyses, including the primary rate-ratio analysis and multiple secondary laboratory and BMD analyses.
20. Rate Ratios vs Mean Differences
DISCOVER is particularly useful statistically because its reported outcomes use several different effect measures. The primary endpoint uses a rate ratio, while the BMD and serum-creatinine analyses use differences in least squares means.
| Effect measure | Example in DISCOVER | What it describes |
|---|---|---|
| Rate ratio | 0.468 for HIV-1 infection incidence | Relative incidence rate in F/TAF compared with F/TDF. |
| Difference in least squares means | 1.142 for hip BMD at Week 48 | Adjusted difference between treatment-group mean percent changes. |
| Difference in least squares means | -0.02 mg/dL for serum creatinine | Adjusted difference in change from baseline between groups. |
| P-value without effect estimate | <0.0001 for urinary biomarkers | Evidence against the corresponding null comparison; effect magnitude is not reported in the ClinicalTrials.gov record. |
These measures cannot be compared numerically as if they were on the same scale. A rate ratio of 0.468 and a mean difference of 1.142 answer fundamentally different questions about different endpoints.
21. Non-Inferiority Logic in Detail
The primary analysis provides a clean example of why non-inferiority should be interpreted through a confidence interval and a clinically prespecified margin.
Point estimate
The rate ratio of 0.468 is below 1, indicating a lower estimated HIV-1 infection incidence rate with F/TAF under the fitted model.
Uncertainty
The 95.003% CI is 0.191–1.149, showing that the plausible range around the estimate is wider than the point estimate alone.
Non-inferiority boundary
The prespecified upper boundary is 1.62. The observed upper confidence bound is 1.149.
Conclusion under the registry rule
Because 1.149 is below 1.62, the stated non-inferiority criterion is met.
This logic differs from a superiority framework. The primary analysis does not require the entire confidence interval to lie below 1 to establish non-inferiority. Instead, it requires the upper confidence limit to remain below the prespecified threshold for unacceptable inferiority.
22. Limitations and Interpretation Issues
- Non-inferiority is margin-dependent: the conclusion depends directly on the prespecified rate-ratio margin of 1.62. A non-inferiority conclusion does not mean that the two treatments have identical effects.
- Person-time matters: the primary endpoint is an incidence rate per 100 person-years, so it should not be interpreted as a simple percentage of participants with HIV infection.
- Analysis populations differ: the primary analysis uses the Full Analysis Set, while BMD and laboratory analyses use specialized or available-data populations.
- Secondary multiplicity: the registry contains numerous secondary analyses. Individual P-values should not automatically be treated as independent confirmatory findings.
- Effect estimates are not posted for every secondary analysis: several urinary biomarker analyses provide P-values without corresponding estimates or confidence intervals in the ClinicalTrials.gov record.
- Safety comparison: serious adverse-event counts are reported by arm, but the ClinicalTrials.gov record does not provide a formal comparative analysis.
- Continuous endpoints require endpoint-specific interpretation: BMD and serum creatinine differences are not interchangeable with the primary HIV-1 infection rate ratio.
- Registry time frames are endpoint-specific: the primary endpoint and secondary 96-week incidence analysis use different stated follow-up frameworks, so their estimates should not be treated as though they were generated at exactly the same analysis time.
- Clinical meaning versus statistical evidence: a statistically detectable difference in a biomarker does not by itself establish the magnitude of any downstream clinical consequence.
23. Why This Trial Matters Statistically
DISCOVER is a useful teaching case because it combines a non-inferiority incidence-rate analysis with several different approaches to continuous and nonparametric secondary outcomes.
| Concept | How it appears in DISCOVER |
|---|---|
| Randomization | Randomized allocation in a phase 3 parallel-group design. |
| Blinding | Double masking. |
| Intention-to-treat principle | The primary analysis is conducted in a Full Analysis Set derived from randomized participants meeting the registry criteria. |
| Rate ratio | Primary HIV-1 infection incidence comparison: 0.468. |
| Poisson regression | Primary incidence-rate analysis using a Poisson model and logarithmic link. |
| Non-inferiority | Upper 95.003% CI compared with a margin of 1.62. |
| Confidence interval | Primary 95.003% CI: 0.191–1.149. |
| ANOVA | Hip and spine BMD comparisons. |
| ANCOVA | Serum creatinine comparisons with baseline covariate adjustment. |
| Least squares mean | Effect measure for BMD and serum creatinine analyses. |
| Van Elteren test | Urinary beta-2-microglobulin and RBP ratio analyses. |
| Rank ANCOVA | Urine protein and UPCR category analyses. |
| Multiplicity | One primary endpoint alongside numerous secondary outcomes and analyses. |
The statistical lesson is broader than any single result. A clinical trial can contain several legitimate statistical questions, each requiring an effect measure and model appropriate to the outcome. The primary incidence endpoint requires rate-based inference, while BMD and serum creatinine require model-adjusted continuous-outcome comparisons.
24. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The primary rate ratio was 0.468 with a 95.003% CI of 0.191–1.149. Because the upper bound was below the prespecified non-inferiority margin of 1.62, the registry's stated non-inferiority criterion was satisfied.
Endpoint-specific interpretation
Secondary analyses reported differences in BMD and serum creatinine, as well as statistically significant or nonsignificant P-values for urinary biomarkers and protein categories. These endpoints answer different questions and require separate interpretation.
A statistically rigorous summary therefore avoids collapsing the trial into one generalized "positive" or "negative" statement. The primary endpoint has a specific non-inferiority interpretation, while each secondary endpoint has its own estimand, analysis population, effect measure, and uncertainty.
25. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
26. Related Statistical Calculators
27. Sources
- ClinicalTrials.gov: DISCOVER, NCT02842086.
- PubMed record: PMID 32711800.
- PubMed record: PMID 39008999.
- PubMed record: PMID 38312459.
- PubMed record: PMID 37969014.
- PubMed record: PMID 34197772.
Continue through Clinical Biostats
Use the statistical concepts from DISCOVER to explore methods for non-inferiority trials, rate data, covariate adjustment, and clinical-trial inference.
28. Record Summary
DISCOVER provides a compact example of several important clinical-trial statistical principles. The primary endpoint was an HIV-1 infection incidence rate analyzed with Poisson regression, expressed as a rate ratio, and evaluated under a prespecified non-inferiority margin of 1.62. The reported rate ratio was 0.468, with a 95.003% two-sided confidence interval of 0.191–1.149, placing the upper confidence bound below the non-inferiority margin.
The secondary analyses broaden the statistical picture. BMD outcomes used ANOVA and differences in least squares means; serum creatinine used ANCOVA with baseline adjustment; urinary biomarker outcomes used the Van Elteren test; and urine protein category outcomes used rank analysis of covariance. The registry also reports serious adverse-event counts of 202/2694 for Descovy and 186/2693 for Truvada.