← Clinical Trials
HIV Prevention Phase 3 Double-Blind NCT00458393

iPrEx: Complete Statistical Analysis of Daily TDF/FTC in HIV Prevention

An independent statistical analysis of the randomized phase 3 iPrEx trial evaluating daily TDF/FTC versus placebo for HIV prevention in men, with primary efficacy and safety endpoints, secondary analyses, and the statistical methods used to compare the treatment groups.

2007-06 start  ·  2014-02 primary completion  ·  Completed
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results are restricted to the information posted in the ClinicalTrials.gov record. Where the registry does not provide a requested quantity, that quantity is not inferred from outside publications or reconstructed from other sources.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

iPrEx was a randomized, double-blind, parallel-group phase 3 prevention trial comparing daily TDF/FTC with placebo in a total enrollment of 2499 participants with HIV infections listed as the trial condition. The primary purpose was prevention, and the registry reports five primary endpoints covering HIV seroconversion and laboratory or clinical toxicity.

2499
Enrollment
Phase 3 trial
2
Treatment arms
TDF/FTC vs placebo
0.577
HIV seroconversion HR
95% CI .404–.824
0.002
Primary efficacy P-value
Log-rank analysis
FeatureiPrEx
Trial nameiPrEx
Brief titleEmtricitabine/Tenofovir Disoproxil Fumarate for HIV Prevention in Men
PhasePhase 3
ConditionHIV Infections
DesignRandomized, double-blind, parallel
Primary purposePrevention
Enrollment2499
InterventionsDaily TDF/FTC and placebo
Primary endpoints5 registered endpoints
Outcome measures posted21
Statistical analyses posted22
Lead sponsorNational Institute of Allergy and Infectious Diseases (NIAID)
Sponsor typeNIH
ClinicalTrials.govNCT00458393

2. Clinical Question

The central statistical question was whether participants assigned to daily TDF/FTC differed from those assigned to placebo with respect to confirmed HIV infection and several laboratory and clinical safety outcomes over the registered follow-up periods.

Population

The trial's brief title identifies men being studied for HIV prevention, with HIV infections listed as the condition. The ClinicalTrials.gov record does not provide a more detailed baseline population description.

Intervention

Daily TDF/FTC, listed as a drug intervention.

Comparator

Placebo, listed as a drug intervention.

Primary question

Whether the randomized TDF/FTC and placebo groups differ in confirmed HIV seroconversion and the registered toxicity endpoints.

3. Trial Design

01
Randomize2499 participants
02
AssignTDF/FTC or placebo
03
FollowMonthly or laboratory follow-up
04
AssessHIV and safety endpoints
05
CompareTime-to-event and categorical analyses
Allocation
Randomized allocation in a parallel-group design.
Masking
Double masking.
Primary purpose
Prevention.
Trial status
Completed.
ARM 1

TDF/FTC

  • Daily TDF/FTC.
  • Drug intervention.
  • Compared with placebo across the registered endpoints.
ARM 2

Placebo

  • Placebo.
  • Drug-intervention comparator.
  • Reference group for the reported hazard ratios and risk ratios.

The registry data do not report a factorial design, crossover scheme, non-inferiority margin, Bayesian analysis, or a specific interim-analysis procedure. Those design features are therefore not introduced here.

4. Primary Endpoints

The five registered primary endpoints span one efficacy endpoint and four safety endpoints. The ClinicalTrials.gov record classifies the endpoints as binary or other/unclear, while the posted analyses use either time-to-event or categorical methods depending on the endpoint.

EndpointRegistered definitionTime framePosted analysis
HIV Seroconversion Confirmed HIV infection Monthly follow-up through a median of 1.2 years Log-rank test; HR
Grade 1 or Higher Creatinine Toxicity Creatinine which reach grade 1 (mild, 1.1 to 1.3 local upper limit of normal) or higher by the US Division of AIDS grading table (version 1) or a 50% increase in creatinine from the baseline value. Duration of follow-up, median 1.2 years Fisher exact test; RR
Grade 3 or Higher Phosphorous Toxicity Grade 3 or higher phosphorous toxicity (hypophosphatemia) by the Division of AIDS Grading Table (severe, level at or below 1.9 mg/dL) The entire follow-up period, median 1.2 years Fisher exact test; RR
Grade 2, 3, or 4 Laboratory Adverse Events Number of participants with at least one Grade 2, 3, or 4 laboratory adverse events (moderate, severe or life threatening based on the US Division of AIDS Grading of adverse events, version 1.0). Entire follow-up, median 1.2 years Log-rank test; HR
Grade 2, 3, or 4 Clinical Adverse Events Number of participants with at least 1 Grade 2, 3, or 4 clinical adverse events (moderate, severe or life threatening based on the US Division of AIDS Grading of adverse events, version 1.0). Entire follow-up, median 1.2 years Log-rank test; HR
Endpoint distinction: HIV seroconversion, laboratory adverse events, and clinical adverse events were analyzed with time-to-event methods in the posted analyses, whereas creatinine and phosphorous toxicity were analyzed as categorical outcomes with Fisher exact tests.

5. Statistical Methodology

Time-to-event analysis and the log-rank test

The primary HIV seroconversion analysis used a log-rank test, a standard method for comparing time-to-event distributions between randomized groups. The same method was reported for Grade 2, 3, or 4 laboratory adverse events and Grade 2, 3, or 4 clinical adverse events.

Core survival-analysis idea
Compare the distribution of event times while retaining information from participants who are censored before an event is observed.

For iPrEx, the time-to-event framework is especially relevant to HIV seroconversion because participants were followed repeatedly and the registry specifies monthly follow-up through a median of 1.2 years.

Hazard ratio

The HIV seroconversion analysis reported a hazard ratio of 0.577 for TDF/FTC versus placebo, with placebo as the reference group. A hazard ratio below 1 indicates a lower estimated instantaneous event rate in the TDF/FTC group under the fitted comparison.

Conceptual interpretation
HR < 1  →  lower estimated instantaneous event rate in the TDF/FTC group

The hazard ratio is a relative time-to-event measure. It is not itself an absolute risk difference, a risk ratio, or the percentage of participants who experienced the event.

Fisher exact test

Fisher exact testing was reported for Grade 1 or Higher Creatinine Toxicity and Grade 3 or Higher Phosphorous Toxicity. These endpoints are defined as participant-level toxicity categories, making a two-group categorical comparison appropriate to the endpoint structure.

Mixed-effects models

The registry reports mixed-model analyses for several secondary endpoints, including percentage change in bone mineral density, CD4 count among HIV infected participants, and proportion of missed doses by pill count. Mixed-effects models are useful when observations have a longitudinal or repeated-measure structure because they can model within-participant dependence rather than treating repeated observations as independent.

t-tests

Two-sided t-tests were reported for viral load among HIV infected participants and for the percentage of missed doses by estimate during the computer assisted structured interview. The corresponding effect measure was a mean difference.

Wilcoxon / Mann-Whitney tests

Wilcoxon (Mann-Whitney) testing was reported for the number of condomless sexual partners with HIV positive or unknown status and for total number of sexual partners. This nonparametric approach is useful when a comparison is based on the distribution of a continuous or count-type outcome without relying on the assumptions of a standard two-sample t-test.

Chi-squared testing

A chi-squared test was reported for condomless receptive anal intercourse in the previous 12 weeks with any partners regardless of status. The reported effect measure was risk difference.

Median regression

The body-composition and lipid analyses used median regression for several week-24 percentage-change outcomes. The posted effect measure was a median difference. This is conceptually different from a mean difference: it focuses on the difference between the group-specific medians rather than the arithmetic means.

6. Primary Results: HIV Seroconversion

HIV seroconversion was the primary efficacy endpoint. The registered definition was confirmed HIV infection, assessed with monthly follow-up through a median of 1.2 years.

Hazard ratio for HIV seroconversion

0.577

95% CI: .404–.824   ·   P = 0.002

Analysis: log-rank test; placebo was the reference group.

FeatureReported value
Analysis populationExcludes participants who were HIV+ at enrollment (2 TDF/FTC, 8 Placebo) and those with no follow-up HIV test (25 TDF/FTC and 22 Placebo).
Groups comparedTDF/FTC vs Placebo
EndpointConfirmed HIV infection
Time frameMonthly follow-up through a median of 1.2 years
MethodLog-rank test
Effect measureHazard ratio
Estimate0.577
95% CI.404–.824
P-value0.002
Analysis notesPrimary null hypothesis: Relative hazard of 0.7 or less. Secondary null hypothesis: Relative hazard of 1.0 or less. Stratified by site with Efron correction for ties. Placebo is reference. Results typically quoted as efficacy = 100*(1-HR).
Clinical Biostats interpretation

The reported HR of 0.577 means that the estimated instantaneous hazard of confirmed HIV infection in the TDF/FTC group was 0.577 times the corresponding hazard in the placebo group under the reported time-to-event analysis. Expressed using the registry's stated convention, 100 × (1 − HR) gives an estimated 42.3% lower hazard.

This does not mean that 42.3% of participants avoided HIV infection, that exactly 42.3% of infections were prevented in every individual, or that the absolute probability of infection was reduced by 42.3 percentage points. A hazard ratio is a relative measure of event occurrence over time.

The 95% confidence interval, .404–.824, describes uncertainty around the estimated relative hazard under the analysis framework. It does not describe the range of effects experienced by individual participants. Its width also matters: the point estimate is more informative when considered together with the interval rather than in isolation.

The P-value of 0.002 addresses the statistical evidence under the specified hypothesis framework; it does not measure the magnitude or clinical importance of the effect. A P-value is not a probability that the null hypothesis is true.

The registry also reports that the analysis was stratified by site and used an Efron correction for ties. Because this is a time-to-event analysis, censoring and the assumptions underlying interpretation of a hazard ratio remain relevant. The registry wording concerning the primary and secondary null hypotheses should be preserved rather than silently replaced with a different hypothesis statement.

Why the hazard ratio is the central efficacy measure here

The endpoint is not simply whether an infection was ever observed. Participants had follow-up over time, and the registry specifies monthly follow-up. A participant who remains infection-free through the available observation period contributes information about the timing of the event process even if that participant does not experience seroconversion during follow-up.

This is why a survival-analysis framework can be more informative than a simple comparison of crude proportions when follow-up times or censoring differ. The log-rank test provides the group-comparison component, while the hazard ratio provides a relative effect estimate.

7. Primary Safety Results

The remaining four primary endpoints concern creatinine toxicity, phosphorous toxicity, laboratory adverse events, and clinical adverse events. Each was formally analyzed in the registry.

Grade 1 or Higher Creatinine Toxicity

Risk ratio

1.33

95% CI: 0.79–2.25   ·   P = 0.28

Analysis: Fisher exact test.

FeatureReported value
DefinitionCreatinine reaching grade 1 or higher by the US Division of AIDS grading table (version 1) or a 50% increase in creatinine from baseline.
Time frameDuration of follow-up, median 1.2 years
Analysis populationAll randomized participants with at least 1 follow-up creatinine value
MethodFisher exact test
Effect measureRisk ratio
Estimate1.33
95% CI0.79–2.25
P-value0.28
Clinical Biostats interpretation

A risk ratio of 1.33 indicates that the estimated risk of the registered creatinine-toxicity endpoint was 1.33 times the risk in the placebo group. Because the 95% CI extends from 0.79 to 2.25, the estimate is compatible with both a lower and a higher relative risk under the stated statistical framework.

The P-value of 0.28 does not establish the size of any possible difference. It is a measure of evidence against the specified null hypothesis, not a measure of whether the two groups are clinically identical. The analysis population also requires attention because it includes randomized participants who had at least one follow-up creatinine value rather than simply every enrolled participant.

Grade 3 or Higher Phosphorous Toxicity

Risk ratio

1.3

95% CI: .57–2.96   ·   P = 0.54

Analysis: Fisher exact test.

FeatureReported value
DefinitionGrade 3 or higher phosphorous toxicity (hypophosphatemia) by the Division of AIDS Grading Table; severe, level at or below 1.9 mg/dL.
Time frameThe entire follow-up period, median 1.2 years
Analysis populationAll participants with at least 1 follow-up phosphorus value
MethodFisher exact test
Effect measureRisk ratio
Estimate1.3
95% CI.57–2.96
P-value0.54
Clinical Biostats interpretation

The point estimate of 1.3 is above 1, but the 95% CI of .57–2.96 is wide and includes 1. The interval therefore communicates substantial uncertainty around the estimated risk ratio. The P-value of 0.54 should not be interpreted as evidence that the two risks are exactly equal; it describes the evidence under the stated hypothesis-testing framework.

The endpoint is also specifically defined as grade 3 or higher phosphorous toxicity, so the estimate should not be generalized to every phosphorous measurement or to less severe laboratory abnormalities.

Grade 2, 3, or 4 Laboratory Adverse Events

Hazard ratio

0.89

95% CI: 0.65–1.23   ·   P = 0.50

Analysis: log-rank test.

FeatureReported value
DefinitionParticipants with at least one Grade 2, 3, or 4 laboratory adverse event.
Time frameEntire follow-up, median 1.2 years
Analysis populationParticipants with at least one visit with laboratory values post-baseline
MethodLog-rank test
Effect measureHazard ratio
Estimate0.89
95% CI0.65–1.23
P-value0.50
Clinical Biostats interpretation

The HR of 0.89 is below 1, corresponding to a lower estimated instantaneous event rate in the TDF/FTC group under this analysis. However, the 95% CI of 0.65–1.23 includes 1, so the point estimate should not be treated as evidence of a precisely established relative reduction.

Because the endpoint is defined as time to the specified laboratory adverse-event category, the hazard ratio should be interpreted as a time-to-event measure rather than as a simple ratio of percentages. The P-value of 0.50 is not an effect-size measure.

Grade 2, 3, or 4 Clinical Adverse Events

Hazard ratio

0.99

95% CI: 0.79–1.23   ·   P = 0.92

Analysis: log-rank test.

FeatureReported value
DefinitionParticipants with at least 1 Grade 2, 3, or 4 clinical adverse event.
Time frameEntire follow-up, median 1.2 years
Analysis populationParticipants with at least one follow-up visit
MethodLog-rank test
Effect measureHazard ratio
Estimate0.99
95% CI0.79–1.23
P-value0.92
Clinical Biostats interpretation

An HR of 0.99 is very close to 1, indicating little separation in the estimated instantaneous event rates at the point estimate. The 95% CI of 0.79–1.23 nevertheless represents uncertainty around that estimate, and the interval includes 1.

The P-value of 0.92 should be read as a statistical-test result, not as a probability that the treatment groups are identical. The endpoint also concerns Grade 2, 3, or 4 clinical adverse events specifically; it should not be interpreted as a comparison of every possible adverse event.

8. Primary Results in One Statistical View

Primary endpointMethodEffect measureEstimate95% CIP-value
HIV SeroconversionLog-rankHR0.577.404–.8240.002
Grade 1 or Higher Creatinine ToxicityFisher exactRR1.330.79–2.250.28
Grade 3 or Higher Phosphorous ToxicityFisher exactRR1.3.57–2.960.54
Grade 2, 3, or 4 Laboratory Adverse EventsLog-rankHR0.890.65–1.230.50
Grade 2, 3, or 4 Clinical Adverse EventsLog-rankHR0.990.79–1.230.92

This table illustrates why a trial should not be summarized by a single P-value. iPrEx has multiple prespecified primary endpoints with different outcome structures and different effect measures. The HIV endpoint uses a time-to-event hazard ratio, two toxicity endpoints use risk ratios, and two other toxicity endpoints use hazard ratios.

9. Secondary Endpoint Results

The registry contains a broad set of secondary analyses spanning virologic, metabolic, adherence, sexual-behavior, and sexually transmitted infection outcomes. The reported estimates below are presented using the same endpoint definitions and time frames reported in the ClinicalTrials.gov record.

Hepatitis Flares Among HBV-Infected Persons

This endpoint assessed hepatitis flares among persons with chronic active hepatitis B at enrollment, using quarterly laboratory tests through a median follow-up of 1.2 years.

MethodEffect measureEstimateCIP-value
Fisher exactRisk difference0.00Not reported in the statistical analysis1.00

The registry also lists a second posted Fisher exact analysis of this same endpoint with a P-value of 1.00 and no effect estimate. The available data therefore support reporting the primary posted estimate of 0.00 and the P-value of 1.00 without attempting to derive an additional confidence interval.

Percentage Change in Bone Mineral Density

Time framePopulationMethodEffect measureEstimateP-value
Baseline and week 24 Participants confirmed to be HIV negative at enrollment who consented to participate in the metabolic substudy Mixed-effects model Mean difference (net) -.91 0.001

The registry does not provide a confidence interval for this analysis in the ClinicalTrials.gov record.

Percentage Change in Body Fat

Median difference

-3.8

95% CI: -6.6 to -0.95   ·   P = 0.009

Baseline to week 24; median regression.

The effect measure is a median difference (net). The negative estimate indicates that the reported median percentage change was lower in the TDF/FTC group than in the placebo group by the estimated difference. The confidence interval is entirely below zero, while the P-value is 0.009.

Percentage Change in Fasting Triglycerides

Effect measureEstimate95% CIP-value
Median difference (net)0.0-9.3 to 9.31.00

The estimate is exactly 0.0, while the 95% CI ranges from -9.3 to 9.3. The interval therefore allows for differences in either direction under the statistical model. The P-value of 1.00 is a test result and should not be interpreted as proving that the treatment groups have identical triglyceride responses.

Percent Change in Total Cholesterol

Effect measureEstimate95% CIP-value
Median difference (net)-2.2-5.5 to 1.10.19

The estimated median difference is -2.2, but the confidence interval of -5.5 to 1.1 crosses zero. This illustrates why the point estimate and its uncertainty should be interpreted together.

Viral Load Among HIV Infected Participants

Time frameMethodEffect measureEstimate95% CIP-value
At the time closest to HIV detection Two-sided t-test Mean difference (net) 0.08 -.18 to 0.33 0.56

The analysis included all HIV infections detected during the study, including infections prior to, during, and after study treatment. The mean difference of 0.08 was estimated with a 95% CI from -.18 to 0.33.

Drug Resistance Among HIV Infected Participants

Analysis populationMethodEffect measureEstimateP-value
2 TDF/FTC seroconversions at enrollment and 48 during follow-up; 8 placebo seroconversions at enrollment and 83 during follow-up Fisher exact Risk ratio 1.00 1.00

The registry states that the null hypothesis is that the proportion of mutations is identical. No confidence interval is in the ClinicalTrials.gov record in the provided statistical-analysis record.

CD4 Count Among HIV Infected Participants

Time frameMethodEffect measureEstimate95% CIP-value
At the time infection was detected Mixed-effects model Mean difference (net) -7 -69 to 54 0.32

The analysis included HIV infected participants during the trial, including those who were HIV+ at baseline. The confidence interval spans both negative and positive differences, illustrating the uncertainty around the estimated mean difference of -7 cells per cubic mm.

Adherence Measures

EndpointTime frameMethodEffect measureEstimate95% CIP-value
Proportion of Missed Doses by Pill Count At 24 weeks Mixed-effects model Mean difference (net) -0.005 -0.02 to 0.01 0.53
Percentage of Missed Doses by Estimate During CASI Interview Week 24 Two-sided t-test Mean difference (net) 0.25 -1.1 to 1.6 0.70

These are two different adherence measurements. The pill-count endpoint concerns the proportion of pills not returned among those whose bottles were returned at the week 24 visit. The CASI endpoint concerns the estimated adherence response among participants who answered the adherence question at week 24. They should not be treated as interchangeable measurements.

Sexual-Behavior Measures

EndpointMethodEffect measureEstimate95% CIP-value
Number of Condomless Sexual Partners With HIV Positive or Unknown Status Wilcoxon (Mann-Whitney) Mean difference (final values) 0 Not reported 0.99
Total Number of Sexual Partners Wilcoxon (Mann-Whitney) Median difference (final values) 0 -.42 to .42 0.76
Condomless Receptive Anal Intercourse in the Previous 12 Weeks With Any Partners Regardless of Status. Chi-squared Risk difference -0.008 -0.047 to 0.030 0.68

The registry specifies a null of no difference between the arms for the condomless receptive anal intercourse endpoint. These analyses address reported behavioral outcomes at week 24; they should not be conflated with the time-to-event analysis of confirmed HIV infection.

Sexually Transmitted Infection Endpoints

EndpointMethodEffect measureEstimate95% CIP-value
Incidence of Confirmed Syphilis During Follow-Up Log-rank Hazard ratio 1.13 0.89–1.43 0.30
Incidence of HSV-2 During the Follow-up Period Log-rank Hazard ratio 1.2 0.8–1.7 0.41
Diagnosis of Gonorrhea During the Follow-up Period Log-rank Hazard ratio 0.61 0.34–1.09 0.09

These three outcomes demonstrate the same general time-to-event logic used for HIV seroconversion, but they are separate secondary endpoints. For gonorrhea, for example, an HR of 0.61 is a relative estimate of the event hazard; it is not equivalent to a 39-percentage-point reduction in gonorrhea incidence.

10. Safety: Serious Adverse Events by Arm

The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm as affected participants divided by participants at risk.

ArmAffectedAt riskReported proportion
TDF/FTC 92 1226 92/1226
Placebo 94 1230 94/1230

The requested data provide the affected and at-risk counts but do not provide a formal statistical comparison, confidence interval, or P-value for serious adverse events. Accordingly, no inferential comparison is added here. The counts can be described directly, but a difference in observed counts should not be treated as evidence of a statistically established treatment effect without the corresponding analysis.

11. Secondary Endpoint Statistical Interpretation

Relative versus absolute effects

iPrEx uses several effect measures because its endpoints have different structures. Hazard ratios describe relative event hazards over time; risk ratios compare risks; risk differences describe absolute differences in risk; mean differences compare arithmetic means; and median differences compare medians. These quantities are not interchangeable.

Confidence intervals

A confidence interval provides a range of parameter values compatible with the statistical model and observed data under the stated confidence procedure. For example, the body-fat median difference of -3.8 has a 95% CI of -6.6 to -0.95, whereas the total-cholesterol median difference of -2.2 has a 95% CI of -5.5 to 1.1. The latter interval includes zero, illustrating greater uncertainty about the direction of the underlying difference.

P-values are not effect sizes

The registry contains P-values ranging from 0.001 to 1.00 across the posted analyses. These values describe evidence under particular null hypotheses. They do not directly communicate the magnitude of an effect, its practical importance, or the probability that a treatment works.

Analysis populations matter

The posted analyses use endpoint-specific populations. HIV seroconversion excludes participants HIV+ at enrollment and participants without a follow-up HIV test. Creatinine toxicity requires at least one follow-up creatinine value. Other analyses restrict the population according to the relevant substudy or measurement availability. An effect estimate therefore describes the population actually analyzed, not automatically every person enrolled in the trial.

12. Statistical Methods Explained

Why was a log-rank test used for HIV seroconversion?

HIV seroconversion is a time-to-event endpoint because the analysis concerns when confirmed HIV infection occurs during follow-up. The log-rank test compares the event-time experience of the randomized groups while accommodating censoring. The registry additionally reports stratification by site and an Efron correction for ties.

What does an HR of 0.577 mean?

An HR of 0.577 means that the estimated instantaneous hazard of confirmed HIV infection was 0.577 times that in the placebo group under the reported model. Using the registry's stated efficacy convention, 100 × (1 − 0.577) = 42.3%, so the estimate corresponds to a 42.3% lower estimated hazard. It does not mean that 42.3% of participants were protected or that absolute infection probability fell by 42.3 percentage points.

Why is a Fisher exact test appropriate for the toxicity endpoints?

The creatinine and phosphorous endpoints are defined as participant-level toxicity categories. Fisher exact testing provides an exact categorical comparison between treatment groups and is especially useful when event counts may be limited. The associated effect measure in these analyses is the risk ratio.

What is the difference between a risk ratio and a hazard ratio?

A risk ratio compares the probability of an event over a specified period or under a specified risk definition. A hazard ratio compares instantaneous event rates over time within a time-to-event framework. An HR of 0.61 for gonorrhea, for example, should not be described as a 39-percentage-point difference in gonorrhea risk.

Why were mixed-effects models used for some secondary endpoints?

The registry reports mixed-model analyses for outcomes such as percentage change in bone mineral density, CD4 count, and pill-count adherence. Mixed-effects models are designed for data with repeated or clustered observations and can represent within-participant dependence. The exact model specification is not reported in the normalized registry summary, so additional model assumptions are not inferred here.

Why use a Wilcoxon / Mann-Whitney test for sexual-partner outcomes?

The registry reports Wilcoxon (Mann-Whitney) analyses for counts of sexual partners. This is a rank-based nonparametric comparison and does not require the same distributional assumptions as a conventional two-sample t-test. The interpretation should follow the effect measure actually reported: a final-value mean difference for one endpoint and a final-value median difference for another.

Why is the analysis population important?

A statistical estimate is conditional on who contributed data to that analysis. The HIV analysis excludes participants who were HIV+ at enrollment and those without a follow-up HIV test, while the metabolic and adherence analyses use substudy- or measurement-specific populations. Randomization establishes the treatment comparison, but missing follow-up and endpoint-specific eligibility determine which participants contribute to particular analyses.

13. Multiplicity and Multiple Primary Endpoints

The registry identifies five primary endpoints and reports formal analyses for all five. That structure is statistically important because a trial with multiple primary endpoints does not have the same inferential architecture as a trial with one primary endpoint.

Primary endpointOutcome structureEffect measureTest
HIV SeroconversionTime-to-eventHazard ratioLog-rank
Grade 1 or Higher Creatinine ToxicityBinaryRisk ratioFisher exact
Grade 3 or Higher Phosphorous ToxicityOther / unclearRisk ratioFisher exact
Grade 2, 3, or 4 Laboratory Adverse EventsTime-to-eventHazard ratioLog-rank
Grade 2, 3, or 4 Clinical Adverse EventsTime-to-eventHazard ratioLog-rank

The ClinicalTrials.gov record identifies the hypothesis type as superiority for these analyses, but they do not provide a multiplicity-adjustment procedure, alpha allocation, or hierarchical testing strategy. Consequently, no specific familywise error-control claim is added.

Interpretation caution: the P-value of 0.002 for HIV seroconversion should be read in the context of the five registered primary endpoints and the trial's prespecified statistical plan. The registry summary does not state how multiplicity across the primary endpoints was handled, so the page does not assign an unreported adjusted significance level.

14. Stratification and the HIV Analysis

The HIV seroconversion analysis notes that the analysis was stratified by site and used an Efron correction for ties. Stratification is useful when a trial has a prespecified grouping factor that should be respected during the time-to-event comparison.

Why stratify?

A stratified analysis allows the comparison to account for the site structure rather than treating all event times as if they arose from one completely homogeneous stratum.

Why correct for ties?

Survival data can contain tied event times. The registry specifically reports the Efron correction for ties in the HIV analysis.

Stratification does not transform the hazard ratio into an absolute measure. It changes the statistical framework used to compare event histories while preserving the interpretation of the reported HR as a relative hazard measure.

15. Time-to-Event Interpretation and Censoring

Three of the five primary endpoints were analyzed with log-rank tests and hazard ratios: HIV seroconversion, Grade 2, 3, or 4 laboratory adverse events, and Grade 2, 3, or 4 clinical adverse events.

Why censoring matters
Observed follow-up = event time if the event occurs; otherwise, the available follow-up time contributes until censoring.

A participant who does not experience an endpoint during observed follow-up is not necessarily treated as having an event at the end of follow-up. Instead, the participant contributes information up to the censoring time.

This distinction explains why a hazard ratio cannot be reconstructed simply by comparing the number of observed events without knowing the timing of those events and the censoring structure.

Educational note: the ClinicalTrials.gov record does not contain participant-level event and censoring times or a sufficiently detailed event table to reconstruct a Kaplan-Meier curve. A fabricated survival curve would therefore add information that is not present in the registry data.

16. Secondary Analyses: Longitudinal Data

Several secondary outcomes are measured at baseline and week 24 or otherwise involve repeated measurements. The registry reports mixed-effects models for percentage change in bone mineral density, CD4 count, and pill-count adherence.

EndpointTime frameModelEffect
Percentage Change in Bone Mineral DensityBaseline and week 24Mixed-effects modelMean difference (net) = -.91
CD4 Count Among HIV Infected ParticipantsAt the time infection was detectedMixed-effects modelMean difference (net) = -7
Proportion of Missed Doses by Pill CountAt 24 weeksMixed-effects modelMean difference (net) = -0.005

The key statistical point is that a mixed-effects analysis can account for correlation among measurements from the same participant. This makes the method fundamentally different from simply applying an independent two-sample test to every observed measurement.

For the median-regression analyses, the registry specifically reports median regression, so the interpretation here follows that method rather than assuming an ordinary mean-based regression.

17. Missing Data and Analysis Populations

The ClinicalTrials.gov record demonstrate that endpoint-specific availability matters. The HIV analysis excludes participants without a follow-up HIV test. The creatinine analysis includes randomized participants with at least one follow-up creatinine value. The phosphorus analysis includes participants with at least one follow-up phosphorus value. Substudies impose additional eligibility requirements.

EndpointPopulation restriction reported by the registry
HIV SeroconversionExcludes HIV+ at enrollment and those with no follow-up HIV test.
Creatinine ToxicityRandomized participants with at least 1 follow-up creatinine value.
Phosphorous ToxicityParticipants with at least 1 follow-up phosphorus value.
Bone Mineral DensityHIV-negative participants at enrollment who consented to the metabolic substudy.
Body FatParticipants in the body-composition substudy with a week 24 body scan.
Fasting TriglyceridesMetabolic-substudy participants with a week 24 fasting triglyceride value.
Total CholesterolMetabolic-substudy participants with a week 24 fasting cholesterol measurement.
Adherence by Pill CountParticipants whose bottles were returned at the week 24 visit.

The ClinicalTrials.gov record does not state a general missing-data imputation method. It would therefore be inappropriate to claim that missing observations were imputed using a particular technique such as multiple imputation, last observation carried forward, or pattern-mixture modeling.

18. What the Primary Efficacy Result Does — and Does Not — Mean

What the HR of 0.577 means

The estimated hazard of confirmed HIV infection was lower in the TDF/FTC group than in the placebo group, with an estimated hazard ratio of 0.577. Using the registry's stated efficacy convention, this corresponds to a 42.3% lower estimated hazard.

What it does not mean

The HR is not the percentage of participants who avoided HIV infection, is not an absolute risk difference, and does not imply that every participant experiences the same proportional change in risk.

What the confidence interval adds

The 95% CI of .404–.824 shows uncertainty around the estimated hazard ratio. It is substantially more informative than reporting 0.577 alone because it communicates how precisely the effect was estimated under the analysis framework.

What the P-value adds

The P-value of 0.002 describes statistical evidence against the relevant null hypothesis. It does not measure effect size and should not be converted into a statement such as "99.8% probability that the treatment works."

19. Comparing the Different Effect Measures

Effect measureExample in iPrExInterpretive question
Hazard ratioHIV seroconversion HR 0.577How do estimated event hazards compare over time?
Risk ratioCreatinine toxicity RR 1.33How do the risks of the binary endpoint compare?
Risk differenceCondomless receptive anal intercourse RD -0.008How far apart are the two risks on an absolute scale?
Mean differenceCD4 count difference -7How far apart are the group means?
Median differenceBody-fat difference -3.8How far apart are the group medians?

This is one of the most useful statistical lessons from iPrEx: the effect measure must match the endpoint and analysis method. Treating all estimates as if they were interchangeable "risk reductions" would obscure important differences in what the analyses actually estimate.

20. Limitations

21. Why This Trial Matters Statistically

iPrEx is a particularly useful statistical teaching case because a single randomized trial contains several fundamentally different analysis structures. The primary efficacy endpoint is a time-to-event outcome, while primary safety endpoints include both categorical risk comparisons and additional time-to-event analyses. Secondary outcomes extend the framework to longitudinal mixed models, mean and median differences, nonparametric tests, and categorical risk differences.

Statistical conceptHow it appears in iPrEx
RandomizationParallel randomized comparison of daily TDF/FTC and placebo.
Double maskingDouble-blind trial design.
Time-to-event analysisHIV seroconversion, laboratory adverse events, clinical adverse events, syphilis, HSV-2, and gonorrhea.
Log-rank testUsed for multiple time-to-event comparisons.
Hazard ratioReported for HIV seroconversion and several secondary and safety endpoints.
Fisher exact testUsed for creatinine toxicity, phosphorous toxicity, hepatitis flares, and drug resistance.
Risk ratioUsed for creatinine toxicity, phosphorous toxicity, and drug resistance.
Risk differenceUsed for hepatitis flares and condomless receptive anal intercourse.
Mixed-effects modelUsed for bone mineral density, CD4 count, and pill-count adherence.
t-testUsed for viral load and CASI-based adherence.
Wilcoxon / Mann-WhitneyUsed for sexual-partner outcomes.
Chi-squared testUsed for condomless receptive anal intercourse.
Median regressionUsed for body fat, fasting triglycerides, and total cholesterol percentage changes.
Confidence intervalsReported for many, but not all, posted effect estimates.
Multiple endpointsFive registered primary endpoints plus numerous secondary outcomes.

The statistical value of the trial is therefore not limited to its headline efficacy estimate. It provides a compact example of why endpoint type determines analysis method, why effect measures must be interpreted in their correct units, and why the analysis population and follow-up structure matter.

22. A Statistical Reading of the Primary Endpoint Set

The five primary endpoints can be read as a sequence of related but distinct questions. First, the HIV seroconversion analysis asks about the timing of confirmed infection. The toxicity endpoints then ask whether the randomized groups differ in specified laboratory or clinical safety outcomes. The result is not one combined statistic but a collection of effect estimates tailored to different endpoint definitions.

Efficacy

HIV seroconversion was analyzed as a time-to-event endpoint, producing an HR of 0.577 with a 95% CI of .404–.824 and P = 0.002.

Renal laboratory toxicity

Grade 1 or higher creatinine toxicity was analyzed using Fisher exact testing, with RR 1.33 and 95% CI 0.79–2.25.

Phosphorous toxicity

Grade 3 or higher phosphorous toxicity was analyzed using Fisher exact testing, with RR 1.3 and 95% CI .57–2.96.

Broader safety endpoints

Laboratory and clinical adverse events were analyzed with log-rank methods, with HRs of 0.89 and 0.99 respectively.

This structure is important because the absence of a statistically strong result for one safety endpoint does not logically establish the same conclusion for another endpoint. Each endpoint has its own definition, analysis population, event process, and uncertainty.

23. Clinical Biostats Statistical Takeaways

Takeaway 1 · Start with the endpoint

Before interpreting a P-value or effect estimate, identify exactly what was measured. Confirmed HIV infection, creatinine toxicity, bone mineral density, viral load, and sexual behavior are not statistically interchangeable outcomes.

Takeaway 2 · Match the effect measure to the analysis

The iPrEx registry demonstrates the distinction among hazard ratios, risk ratios, risk differences, mean differences, and median differences. The numerical value of an estimate has meaning only in the context of its effect-measure definition.

Takeaway 3 · Read the confidence interval

The confidence interval communicates uncertainty around an estimate. For example, the HIV HR interval of .404–.824 is narrower than the phosphorous-toxicity RR interval of .57–2.96, reflecting different levels of precision in those estimates.

Takeaway 4 · P-values do not replace effect estimates

A P-value tells you about statistical evidence under a specified null hypothesis. It does not tell you how large an effect is, how precise the estimate is, or whether an effect has a particular practical importance.

Takeaway 5 · Analysis populations are part of the result

The HIV analysis and the metabolic substudy analyses do not use identical populations. Understanding who contributed data is necessary before extending an estimate beyond the analyzed participants.

24. Related Tutorials

Learn more about the methods used in this trial:

25. Related Statistical Calculators

26. Sources

Continue through the Clinical Biostats statistical pathway

Use the related tutorials and calculators to explore the survival, categorical-data, longitudinal, and nonparametric methods represented in this trial.

27. Record Summary

iPrEx provides a broad example of clinical-trial statistical analysis within a single randomized phase 3 prevention study. The primary HIV seroconversion endpoint was analyzed with a site-stratified log-rank framework and reported an HR of 0.577 with a 95% CI of .404–.824 and P = 0.002. The four primary safety endpoints used either Fisher exact or log-rank methods, producing risk-ratio or hazard-ratio estimates with their corresponding uncertainty intervals. Secondary analyses extended the statistical framework to mixed-effects models, t-tests, Wilcoxon / Mann-Whitney testing, chi-squared testing, and median regression.

The most important statistical lesson is that the trial cannot be reduced to one number. Correct interpretation requires keeping the endpoint definition, analysis population, time frame, statistical method, effect measure, confidence interval, and P-value connected. That framework allows the reported results to be understood without treating different statistical quantities as if they represented the same underlying question.

Clinical Biostats methodology: This analysis intentionally distinguishes registry-reported evidence from statistical explanation. Numbers are reproduced from the ClinicalTrials.gov record, and unreported design features, subgroup results, imputation methods, multiplicity procedures, and other quantities are not inferred.