← Clinical Trials
HIV Prevention Phase 3 Randomized NCT00557245

Partners PrEP: Complete Statistical Analysis of Pre-Exposure Prophylaxis for HIV-1 Acquisition

An independent statistical review of the randomized phase 3 Partners PrEP trial evaluating daily tenofovir disoproxil fumarate (TDF) and emtricitabine/tenofovir disoproxil fumarate (FTC/TDF) for prevention of HIV-1 acquisition among HIV-1 uninfected participants in HIV-1 discordant partnerships.

University of Washington  ·  Phase 3  ·  Completed  ·  2008-05 to 2013-10
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

Partners PrEP was a randomized, parallel-group, quadruple-masked phase 3 prevention trial evaluating once-daily TDF, once-daily FTC/TDF, and placebo among HIV-1 uninfected participants in HIV-1 discordant partnerships.

4,758
Enrollment
Three-arm trial
3
Arms
TDF · FTC/TDF · placebo
0.33
TDF vs placebo HR
95% CI 0.19–0.56
0.25
FTC/TDF vs placebo HR
95% CI 0.13–0.45
FeaturePartners PrEP
PhasePhase 3
ConditionHIV-1 Infections; HIV Infections
Brief titlePre-Exposure Prophylaxis to Prevent HIV-1 Acquisition Within HIV-1 Discordant Couples
DesignRandomized, parallel
MaskingQuadruple
Primary purposePrevention
Enrollment4,758
InterventionsTenofovir Disoproxil Fumarate (TDF); Emtricitabine/Tenofovir Disoproxil Fumarate (FTC/TDF); Placebo
Primary endpointsIncidence of HIV-1 Seroconversion Among HIV-1 Uninfected Participants; Number of Participants With Serious Adverse Events (SAEs)
Results postedYes
Lead sponsorUniversity of Washington
ClinicalTrials.govNCT00557245

2. Clinical Question

The central statistical question was whether daily TDF or daily FTC/TDF reduced the incidence of HIV-1 acquisition compared with placebo among HIV-1 uninfected participants in HIV-1 discordant partnerships.

Population

HIV-1 uninfected participants within HIV-1 discordant partnerships, as described by the registered primary endpoint.

Intervention 1

Tenofovir Disoproxil Fumarate (TDF).

Intervention 2

Emtricitabine/Tenofovir Disoproxil Fumarate (FTC/TDF).

Comparator

Placebo.

Primary efficacy question

Does once-daily PrEP prevent HIV-1 acquisition, measured by HIV incidence per 100 person-years in each of the three arms?

Primary safety question

How does the number of participants with serious adverse events compare between each active regimen and placebo during follow-up?

3. Trial Design

01
Randomize4,758 participants
02
Three armsTDF · FTC/TDF · placebo
03
MaskedQuadruple masking
04
FollowUp to 36 months
05
AnalyzeHIV-1 seroconversion and safety
ARM 1

TDF

  • Tenofovir Disoproxil Fumarate
  • Daily PrEP intervention
ARM 2

FTC/TDF

  • Emtricitabine/Tenofovir Disoproxil Fumarate
  • Daily PrEP intervention
ARM 3

Placebo

  • Placebo comparator
  • Reference group for the posted Cox analyses
Allocation
Randomized
Design model
Parallel
Masking
Quadruple
Primary purpose
Prevention
Start
2008-05
Primary completion
2013-10

The three-arm structure creates two active-versus-placebo efficacy comparisons. The posted primary analyses use placebo as the reference group, allowing the two active regimens to be interpreted separately rather than as a single pooled intervention.

4. Endpoints

EndpointRegistered definitionTime frameType
Incidence of HIV-1 Seroconversion Among HIV-1 Uninfected Participants The efficacy of once daily PrEP in preventing HIV-1 acquisition among uninfected heterosexuals in HIV-1 discordant partnerships, measured by calculating the HIV incidence per 100 person-years in each of three arms. Up to 36 months Binary
Number of Participants With Serious Adverse Events (SAEs) Safety of daily TDF or FTC/TDF among HIV-1 uninfected individuals randomized to TDF or FTC/TDF compared to those randomized to placebo measured as the number of participants with Serious Adverse Events (SAEs) during follow-up. Up to 36 months Binary

Although the registry classifies the HIV-1 seroconversion endpoint as binary, the posted formal analyses treat it as a time-to-event endpoint because the analysis compares the relative rates of time to first positive HIV-1 serologic test. This distinction is statistically important: a participant's follow-up time contributes information, not merely a final yes/no outcome.

5. Analysis Populations

EndpointAnalysis population
HIV-1 seroconversion All randomized participants, less those who were found to be ineligible (n=11), less those found to be infected at enrollment (n=14), and less those who did not return for any follow-up (n=25).
Serious adverse events Randomized participants, less those found to be ineligible (n=11).
Secondary STI and unprotected-sex analyses Randomized participants, less those found to be ineligible (n=11).
Congenital abnormalities Randomized female participants, less those found to be ineligible (n=7).
Infant growth analyses Infants born to women taking study drug during follow-up.
Why the exclusions matter: The efficacy analysis is not described in the registry as simply "all 4,758 randomized participants." It explicitly excludes participants found to be ineligible, participants infected at enrollment, and participants who did not return for any follow-up. The safety analysis has a different exclusion rule. Keeping these populations distinct prevents an analysis from silently changing its denominator.

6. Primary Efficacy Results

The primary efficacy endpoint was HIV-1 seroconversion among HIV-1 uninfected participants, with follow-up of up to 36 months. The registry reports two formal Cox proportional-hazards analyses, each comparing an active regimen with placebo.

TDF vs Placebo

Hazard ratio for time to first positive HIV-1 serologic test

0.33

95% CI: 0.19–0.56   ·   P < 0.001

Analysis: Cox regression stratified according to site; placebo arm is the reference group.

Clinical Biostats interpretation

A hazard ratio of 0.33 means that the fitted Cox model estimated the relative rate of reaching a first positive HIV-1 serologic test to be 0.33 for TDF compared with placebo. Expressed as a simple relative-hazard interpretation, this corresponds to an estimated 67% lower hazard under the model.

That does not mean that 67% of participants were protected, that 67% fewer participants necessarily experienced seroconversion as a simple proportion, or that an individual participant's probability of infection was reduced by exactly 67%. The hazard ratio is a time-to-event measure that incorporates the timing of events and censoring.

The 95% confidence interval of 0.19–0.56 describes uncertainty around the estimated hazard ratio under the statistical model and sampling framework. It does not describe the range of effects among individual participants.

The P < 0.001 result addresses statistical evidence against the null comparison specified by the analysis. It does not measure the magnitude of the treatment effect; the hazard ratio and its confidence interval provide that information.

The analysis was stratified by site. As with any Cox model, interpretation of a single hazard ratio also depends on the model's proportional-hazards assumption. The ClinicalTrials.gov record does not provide a formal diagnostic of that assumption.

FTC/TDF vs Placebo

Hazard ratio for time to first positive HIV-1 serologic test

0.25

95% CI: 0.13–0.45   ·   P < 0.001

Analysis: Cox regression stratified according to site; placebo arm is the reference group.

Clinical Biostats interpretation

A hazard ratio of 0.25 means that the fitted Cox model estimated the relative rate of reaching a first positive HIV-1 serologic test to be 0.25 for FTC/TDF compared with placebo. As a simple derived interpretation, this corresponds to an estimated 75% lower hazard under the model.

The HR should not be read as a 75-percentage-point reduction in infection probability, nor as a statement that exactly 75% of participants avoided infection because of treatment. It is a relative time-to-event measure.

The 95% confidence interval of 0.13–0.45 indicates the statistical precision of the estimated hazard ratio. The interval concerns the treatment-effect estimate, not individual-level outcomes.

The P < 0.001 value quantifies evidence against the null comparison used in the formal test; it is not an effect-size measure. A small p-value and a large treatment effect are related concepts but are not interchangeable.

The posted analysis is site-stratified Cox regression, and placebo is explicitly the reference group. Interpretation of the HR therefore depends on the time-to-event model and its assumptions, including the proportional-hazards framework.

Primary efficacy comparisonMethodEffect measureEstimate95% CIP-value
TDF vs PlaceboCox proportional-hazards model, stratified by siteHazard ratio0.330.19–0.56<0.001
FTC/TDF vs PlaceboCox proportional-hazards model, stratified by siteHazard ratio0.250.13–0.45<0.001
Educational note: the ClinicalTrials.gov record reports hazard ratios and confidence intervals but do not provide the event-by-event risk-set information needed to reconstruct a valid Kaplan-Meier curve. A Kaplan-Meier graphic should therefore not be fabricated from the summary statistics alone.

7. Primary Safety Results

The second primary endpoint was the number of participants with serious adverse events during follow-up of up to 36 months. The registry reports Fisher exact tests for each active regimen versus placebo.

ComparisonSerious adverse eventsMethodP-value
TDF vs PlaceboTDF: 118/1584
Placebo: 118/1584
Fisher exact test1.00
FTC/TDF vs PlaceboFTC/TDF: 115/1579
Placebo: 118/1584
Fisher exact test0.89
Clinical Biostats interpretation

For TDF versus placebo, the registry reports 118/1584 participants with serious adverse events in each group and a Fisher exact test P-value of 1.00. For FTC/TDF versus placebo, it reports 115/1579 versus 118/1584, with P = 0.89.

These analyses do not provide a posted effect estimate or confidence interval in the ClinicalTrials.gov record. Consequently, the magnitude and precision of a relative or absolute safety effect cannot be quantified from these formal analysis records alone.

A P-value is not a measure of the size of a safety difference. The reported event counts are therefore important alongside the P-values. Likewise, a nonsignificant test should not be translated into a claim that the two treatments have identical safety profiles; it means that the posted test did not provide strong statistical evidence against its null comparison.

Fisher's exact test is appropriate for comparing binary event counts when an exact test of the contingency table is desired. Unlike the Cox analysis for HIV-1 seroconversion, it does not incorporate the timing of an SAE within follow-up.

8. Secondary Endpoint Results

The registry contains additional formal analyses covering infant growth, sexually transmitted infections, unprotected sex, and congenital abnormalities. Several of these analyses report P-values without an effect estimate or confidence interval. Those results are reported here exactly as reported in the registry rather than supplemented with estimates that are not present in the registry data.

Infant Length

ComparisonMethodEffect measureEstimateP-value
FTC/TDF vs PlaceboLinear mixed-effects modelSlope Difference over time0.070.08
TDF vs PlaceboLinear mixed-effects modelSlope difference over timeNot posted0.42

The endpoint was Length Among Infants Born to Female Participants Taking Study Drug, measured as z-score difference per study month over up to 12 months. The analysis population was infants born to women taking study drug during follow-up. The posted FTC/TDF analysis used placebo as the reference group.

Infant Weight

ComparisonMethodEffect measureEstimateP-value
TDF vs PlaceboLinear mixed-effects modelNot postedNot posted0.02
FTC/TDF vs PlaceboLinear mixed-effects modelNot postedNot posted<0.001

The endpoint was Weight Among Infants Born to Female Participants Taking Study Drug, measured as z-score difference per study month over up to 12 months. The ClinicalTrials.gov record provides P-values but no corresponding effect estimates or confidence intervals.

Infant Head Circumference

ComparisonMethodEffect measureEstimateP-value
TDF vs PlaceboLinear mixed-effects modelNot postedNot posted0.35
FTC/TDF vs PlaceboLinear mixed-effects modelNot postedNot posted0.008

The endpoint was Head Circumference Among Infants Born to Female Participants Taking Study Drug, measured as z-score difference per study month over up to 12 months. The ClinicalTrials.gov record identifies the analyses as linear mixed-effects models.

Sexually Transmitted Infection

ComparisonMethodModel detailP-value
TDF vs PlaceboLogistic regressionGeneralized estimating equations with logistic link and robust standard errors to adjust for individual correlation over time0.24
FTC/TDF vs PlaceboLogistic regressionGeneralized estimating equations with logistic link and robust standard errors to adjust for individual correlation over time0.49

Prevalence of Unprotected Sex During Follow-up

ComparisonMethodOutcome unitP-value
TDF vs PlaceboLogistic regression with generalized estimating equations, logistic link, and robust standard errorsPercentage of visits0.32
FTC/TDF vs PlaceboLogistic regression with generalized estimating equations, logistic link, and robust standard errorsPercentage of visits0.66

Congenital Abnormalities

ComparisonMethodModel detailP-value
TDF vs PlaceboLogistic regressionGeneralized estimating equations with logistic link to account for multiple pregnancies and multiple births0.51
FTC/TDF vs PlaceboLogistic regressionGeneralized estimating equations with logistic link to account for multiple pregnancies and multiple births0.86

The congenital-abnormality endpoint was defined as Congenital Abnormalities Among Infants Born to Female Participants Taking Study Drug, measured by number of live-born infants over up to 36 months. The analysis population was randomized female participants, less those found to be ineligible (n=7).

Reading the secondary P-values: Several secondary analyses have P-values without posted effect estimates or confidence intervals. A P-value alone does not establish the direction, magnitude, or practical importance of a difference. The ClinicalTrials.gov record therefore support reporting the test result and method, but not reconstructing an effect size that was not posted.

9. Statistical Methodology

Cox proportional-hazards model

The primary HIV-1 seroconversion comparisons used Cox regression stratified according to site. The model estimated the relative rates of time to first positive HIV-1 serologic test, with placebo as the reference group.

Conceptual form
h(t | X) = h0(t) exp(βX)

For a treatment indicator, exponentiating its coefficient gives a hazard ratio. An HR below 1 indicates a lower estimated instantaneous event rate in the treatment group relative to the reference group.

Site-stratified analysis

The primary Cox analyses were stratified according to site. Stratification allows the baseline hazard to differ across strata while estimating a common treatment effect across those strata. This is useful when the trial has meaningful site-level differences that should not be forced into a single shared baseline hazard.

Fisher exact test

The primary safety endpoint was binary: whether a participant experienced a serious adverse event during follow-up. Fisher's exact test was used for each active regimen versus placebo. The method evaluates the observed two-by-two allocation of participants across event and non-event categories without relying on a large-sample approximation in the same way as a conventional chi-square test.

Logistic regression and generalized estimating equations

Secondary binary outcomes such as sexually transmitted infections and unprotected sex were analyzed with logistic regression. The registry specifically reports generalized estimating equations, a logistic link, and robust standard errors for the STI and unprotected-sex analyses. This accounts for correlation among repeated observations from the same individual over time.

Mixed-effects models

Infant length, weight, and head circumference were analyzed using linear mixed-effects models. The outcome unit was z-score difference per study month. A mixed-effects framework is suited to longitudinal data because repeated measurements from the same infant are correlated rather than statistically independent.

Effect measures

The registry identifies hazard ratio and slope / regression coefficient as effect-measure types used in the posted analyses. The primary HIV-1 analyses provide hazard ratios with confidence intervals. Some secondary analyses provide a slope difference or P-value, while others do not provide a numerical effect estimate in the ClinicalTrials.gov record.

10. Statistical Methods Explained

Why was Cox regression used for HIV-1 seroconversion?

The primary endpoint concerns the time to a first positive HIV-1 serologic test, not merely whether a participant was positive at the end of follow-up. Cox regression uses both event timing and available follow-up time and can appropriately handle participants whose event has not occurred during observed follow-up through censoring.

What does an HR of 0.33 mean?

Relative to placebo, an HR of 0.33 means the fitted model estimated the instantaneous rate of the analyzed HIV-1 seroconversion event to be 33% as large in the TDF group. The complementary interpretation is a 67% lower estimated hazard. It is not a 67-percentage-point change in absolute probability.

Why was the Cox analysis stratified by site?

The registry explicitly states that Cox regression was stratified according to site. In a stratified Cox model, the baseline hazard can vary across sites while the treatment effect is estimated across the strata. This separates site-specific baseline event patterns from the treatment comparison.

Why use Fisher exact testing for serious adverse events?

Serious adverse events were recorded as a binary participant-level outcome. Fisher's exact test directly evaluates the treatment-by-event contingency table. The registry uses it for each active regimen versus placebo and reports P = 1.00 for TDF versus placebo and P = 0.89 for FTC/TDF versus placebo.

Why use generalized estimating equations for STI and unprotected sex?

These outcomes were observed during follow-up, so multiple observations could arise from the same participant. The registry specifies generalized estimating equations with a logistic link and robust standard errors to adjust for individual correlation over time. Treating all repeated observations as independent would ignore that within-person dependence.

Why use mixed-effects models for infant growth?

Length, weight, and head circumference were evaluated over time, with the outcome expressed as a z-score difference per study month. Repeated measurements within an infant tend to be correlated. A linear mixed-effects model provides a framework for representing that longitudinal structure rather than treating every measurement as an unrelated observation.

Why should a P-value not be interpreted as an effect size?

A P-value describes the statistical evidence against a specified null hypothesis under the test's assumptions. It does not tell the reader how large the treatment effect is. The primary HIV-1 analyses illustrate the distinction: the P-values are both <0.001, while the estimated hazard ratios are 0.33 and 0.25 with different confidence intervals.

11. Confidence Intervals and Precision

The two primary HIV-1 analyses provide two-sided 95% confidence intervals. These intervals are particularly important because they show the uncertainty surrounding the point estimates rather than leaving the reader with a single hazard ratio.

ComparisonHR95% CIInterpretive role
TDF vs Placebo0.330.19–0.56Quantifies uncertainty around the estimated relative rate
FTC/TDF vs Placebo0.250.13–0.45Quantifies uncertainty around the estimated relative rate
Precision matters

The point estimate is only one part of the result. A confidence interval communicates how much statistical uncertainty surrounds that estimate. The interval should not be interpreted as saying that 95% of individual participants have hazards within those numerical limits. It is an interval for the underlying treatment-effect parameter under the specified statistical framework.

The ClinicalTrials.gov record does not provide confidence intervals for the primary serious-adverse-event analyses or for the secondary analyses. Those analyses should therefore not be given invented intervals or reverse-engineered estimates on this page.

12. Primary Endpoint Analysis in Context

The HIV-1 endpoint illustrates why the distinction between binary outcomes and time-to-event outcomes matters. The registry defines the endpoint as incidence of HIV-1 seroconversion and describes incidence per 100 person-years. The formal analysis, however, uses Cox regression to estimate the relative rates of time to first positive HIV-1 serologic test.

Binary framing

At a simple level, a participant either experienced HIV-1 seroconversion during the relevant observation period or did not.

Time-to-event framing

The formal analysis also considers when the first positive serologic test occurred and how long participants were observed.

Why censoring matters

Participants can contribute follow-up without experiencing the event during the observed period. Time-to-event methods are designed to use that partial follow-up information.

Why the HR is relative

The hazard ratio compares modeled event rates between treatment groups; it is not an absolute incidence difference or an individual risk prediction.

13. Secondary Analysis Structure

The secondary analyses demonstrate several distinct statistical structures within one trial. The choice of model changes with the outcome and with whether repeated observations need to be accounted for.

OutcomeAnalysis familyKey statistical feature
Infant lengthMixed-effects modelLongitudinal slope difference over time
Infant weightMixed-effects modelRepeated growth measurements
Infant head circumferenceMixed-effects modelRepeated growth measurements
STI during follow-upLogistic regressionGEE, logistic link, robust standard errors
Unprotected sex during follow-upLogistic regressionGEE, logistic link, robust standard errors
Congenital abnormalitiesLogistic regressionGEE for multiple pregnancies and births
Serious adverse eventsFisher exact testBinary participant-level comparison

This is a useful teaching point: there is no single "clinical trial statistical test." A well-designed analysis chooses a model according to the endpoint's measurement scale, timing, correlation structure, and scientific question.

14. Safety by Randomized Arm

ArmParticipants with serious adverse eventsAt risk
Tenofovir Disoproxil Fumarate (TDF)1181584
Emtricitabine/Tenofovir Disoproxil Fumarate (FTC/TDF)1151579
Placebo1181584

The posted safety analyses compare TDF with placebo and FTC/TDF with placebo using Fisher exact tests. The ClinicalTrials.gov record does not report a confidence interval or effect estimate for either comparison, so the counts and P-values are the complete numerical formal-analysis results available for this endpoint in the provided record.

Safety interpretation: Serious adverse events are an important safety endpoint, but this table represents one safety outcome rather than an exhaustive safety profile. The ClinicalTrials.gov record does not provide a complete adverse-event catalogue, severity breakdown, treatment-relatedness assessment, or exposure-adjusted safety analysis.

15. Blinding and Randomization

The trial used randomized allocation and quadruple masking. Randomization is central to the causal interpretation of the active-versus-placebo comparisons because treatment assignment is determined independently of participants' subsequent outcomes.

Randomization
The registry describes allocation as randomized.
Parallel design
The registry identifies the design model as parallel, with three intervention arms.
Quadruple masking
The registry identifies the masking level as quadruple.
Reference group
The primary Cox analyses explicitly use placebo as the reference group.

Masking and randomization address different sources of bias. Randomization establishes the treatment-allocation mechanism, while masking can reduce the influence of knowledge of treatment assignment on behavior, assessment, reporting, or other aspects of follow-up.

16. Multiplicity and Multiple Comparisons

The trial has two active-versus-placebo primary efficacy comparisons for HIV-1 seroconversion: TDF versus placebo and FTC/TDF versus placebo. The ClinicalTrials.gov record identifies both as superiority analyses and report P < 0.001 for each.

Primary comparisonHypothesis typeP-valueWhat is reported
TDF vs PlaceboSuperiority<0.001HR 0.33; two-sided 95% CI 0.19–0.56
FTC/TDF vs PlaceboSuperiority<0.001HR 0.25; two-sided 95% CI 0.13–0.45

Multiple formal comparisons raise an important statistical-design issue: when several hypotheses are tested, the interpretation of the overall false-positive rate depends on the prespecified testing strategy. The ClinicalTrials.gov record identifies the hypothesis type and individual P-values but do not provide an alpha-allocation or multiplicity-adjustment procedure. This page therefore does not infer one.

Important distinction: the presence of two very small P-values does not by itself establish how the trial controlled familywise type I error. That conclusion requires the prespecified statistical-analysis plan or another source that explicitly documents the multiplicity strategy.

17. Missing Data, Censoring, and Follow-up

The ClinicalTrials.gov record explicitly identify participants who did not return for any follow-up as exclusions from the primary HIV-1 seroconversion analysis. They do not specify an imputation procedure for other missing observations, a detailed censoring algorithm, or a missing-data sensitivity analysis.

No follow-up

The HIV-1 analysis excludes participants who did not return for any follow-up.

Censoring

The Cox analysis is a time-to-event analysis and therefore relies on handling incomplete event histories through censoring.

Imputation

The ClinicalTrials.gov record does not report an imputation method for missing measurements.

Interpretive caution

The validity of a time-to-event analysis depends partly on whether censoring is appropriately handled under the model's assumptions.

Because the ClinicalTrials.gov record does not describe the censoring rules in detail, this page does not assume a particular censoring convention beyond the general structure inherent to the reported Cox analysis.

18. Longitudinal Analyses of Infant Outcomes

The infant endpoints provide a different statistical problem from HIV-1 seroconversion. Length, weight, and head circumference were measured longitudinally and expressed as z-score difference per study month. The registry identifies linear mixed-effects models for these analyses.

Conceptual longitudinal model
Yij = β0 + β1Timeij + β2Treatmenti + β3(Time × Treatment)ij + bi + εij

The treatment-by-time component can represent a difference in growth trajectory. The registry's posted infant-length analysis labels its effect measure "Slope Difference over time."

The registry-reported FTC/TDF-versus-placebo length analysis reports an estimated slope difference of 0.07 with P = 0.08. The registry does not provide a confidence interval for that estimate in the ClinicalTrials.gov record. For the remaining infant-growth analyses, P-values are reported but numerical effect estimates are not.

How to interpret a slope difference

A slope difference describes how the modeled outcome trajectory changes over study time between groups. It is not the same quantity as a final-value mean difference. This distinction is important whenever treatment groups are compared repeatedly over time.

19. Why the Different Models Matter

Partners PrEP illustrates how statistical methodology follows the structure of the scientific endpoint.

QuestionData structurePosted methodWhy it fits
When does HIV-1 seroconversion occur?Time-to-eventCox proportional-hazards modelUses event timing and follow-up time
Did a participant experience an SAE?BinaryFisher exact testCompares participant-level event counts
Did STI occur during follow-up?Repeated binary observationsLogistic regression with GEEAccounts for within-person correlation
Did unprotected sex occur at visits?Repeated binary observationsLogistic regression with GEEAccounts for individual correlation over time
How does infant growth change?Repeated continuous measurementsMixed-effects modelModels longitudinal trajectories and within-infant dependence

20. What the Primary Hazard Ratios Do — and Do Not — Mean

Relative effect

The TDF-versus-placebo HR of 0.33 and the FTC/TDF-versus-placebo HR of 0.25 indicate lower estimated rates of the analyzed HIV-1 seroconversion event under the respective Cox models.

Not an absolute risk difference

Neither hazard ratio tells the reader how many additional or fewer infections occurred per 100 participants. The registry endpoint definition references incidence per 100 person-years, but the registry-reported formal analyses report relative hazard ratios rather than an absolute incidence difference.

Not an individual prediction

An HR is a group-level model parameter. It does not predict the exact probability or timing of HIV-1 acquisition for a particular participant.

Not a p-value

The HR communicates effect magnitude on a relative scale. The P-value communicates statistical evidence against the null hypothesis. The confidence interval adds information about precision. These three quantities should be read together.

21. Limitations and Interpretation Issues

22. Why This Trial Matters Statistically

Partners PrEP is a useful teaching case because it places several major clinical-trial methods inside one randomized prevention study. The primary endpoint is analyzed as a time-to-event outcome, while other endpoints require exact categorical testing, repeated binary-data methods, or longitudinal mixed models.

ConceptHow it appears in Partners PrEP
RandomizationRandomized allocation across three parallel arms
BlindingQuadruple masking
Time-to-event analysisTime to first positive HIV-1 serologic test
Hazard ratioPrimary efficacy effect measure
Confidence intervalTwo-sided 95% CIs for the two primary HIV-1 HR estimates
Cox modelSite-stratified analysis of HIV-1 seroconversion
Fisher exact testPrimary serious-adverse-event comparisons
Logistic regressionSTI, unprotected sex, and congenital-abnormality analyses
Generalized estimating equationsAdjustment for repeated observations or multiple pregnancies and births
Mixed-effects modelLongitudinal infant length, weight, and head circumference analyses
MultiplicityTwo active-versus-placebo primary efficacy comparisons

23. Statistical Methods Explained in Practice

Randomization

Creates the treatment-allocation framework underlying the active-versus-placebo comparisons.

Stratified analysis

The primary Cox models were stratified according to site.

Hazard ratio

Summarizes the relative rate of time to first positive HIV-1 serologic test.

Confidence interval

Shows uncertainty around the two posted primary HR estimates.

Repeated binary data

GEE-based logistic models account for individual correlation over time.

Longitudinal modeling

Mixed-effects models evaluate growth trajectories using repeated measurements.

24. Related Tutorials

Learn more about the methods used in this trial:

25. Related Statistical Calculators

26. Sources

Continue through the Clinical Biostats statistical library

Use the trial's statistical methods as a pathway into survival analysis, regression, categorical-data testing, longitudinal modeling, and clinical-trial methodology.

27. Record Summary

Partners PrEP provides a compact example of how clinical-trial endpoints determine statistical methodology. The primary HIV-1 seroconversion endpoint was formally analyzed using site-stratified Cox proportional-hazards models, producing HR 0.33 for TDF versus placebo and HR 0.25 for FTC/TDF versus placebo, each with a two-sided 95% confidence interval and P < 0.001. The primary safety endpoint used Fisher exact tests, while secondary outcomes required logistic regression with generalized estimating equations or linear mixed-effects models.

The most important statistical lesson is that these results should not be reduced to P-values alone. The primary hazard ratios quantify relative treatment effects, the confidence intervals quantify statistical uncertainty, and the time-to-event framework explains why follow-up duration and event timing matter. The safety analyses illustrate a different principle: when the registry supplies event counts and P-values but no effect estimates or confidence intervals, the correct interpretation is to preserve that distinction rather than manufacture additional statistics.

Clinical Biostats methodology: A trial-results page should reconstruct the statistical story of a study without silently filling gaps in the public record. Reported estimates, confidence intervals, P-values, analysis populations, and model choices should remain distinguishable from educational interpretation.