This page separates reported trial results from statistical interpretation. Numerical results are restricted to the information posted in the ClinicalTrials.gov record. Where the registry does not provide a requested quantity, that quantity is not inferred from outside publications or reconstructed from other sources.
1. Trial at a Glance
iPrEx was a randomized, double-blind, parallel-group phase 3 prevention trial comparing daily TDF/FTC with placebo in a total enrollment of 2499 participants with HIV infections listed as the trial condition. The primary purpose was prevention, and the registry reports five primary endpoints covering HIV seroconversion and laboratory or clinical toxicity.
| Feature | iPrEx |
|---|---|
| Trial name | iPrEx |
| Brief title | Emtricitabine/Tenofovir Disoproxil Fumarate for HIV Prevention in Men |
| Phase | Phase 3 |
| Condition | HIV Infections |
| Design | Randomized, double-blind, parallel |
| Primary purpose | Prevention |
| Enrollment | 2499 |
| Interventions | Daily TDF/FTC and placebo |
| Primary endpoints | 5 registered endpoints |
| Outcome measures posted | 21 |
| Statistical analyses posted | 22 |
| Lead sponsor | National Institute of Allergy and Infectious Diseases (NIAID) |
| Sponsor type | NIH |
| ClinicalTrials.gov | NCT00458393 |
2. Clinical Question
The central statistical question was whether participants assigned to daily TDF/FTC differed from those assigned to placebo with respect to confirmed HIV infection and several laboratory and clinical safety outcomes over the registered follow-up periods.
Population
The trial's brief title identifies men being studied for HIV prevention, with HIV infections listed as the condition. The ClinicalTrials.gov record does not provide a more detailed baseline population description.
Intervention
Daily TDF/FTC, listed as a drug intervention.
Comparator
Placebo, listed as a drug intervention.
Primary question
Whether the randomized TDF/FTC and placebo groups differ in confirmed HIV seroconversion and the registered toxicity endpoints.
3. Trial Design
TDF/FTC
- Daily TDF/FTC.
- Drug intervention.
- Compared with placebo across the registered endpoints.
Placebo
- Placebo.
- Drug-intervention comparator.
- Reference group for the reported hazard ratios and risk ratios.
The registry data do not report a factorial design, crossover scheme, non-inferiority margin, Bayesian analysis, or a specific interim-analysis procedure. Those design features are therefore not introduced here.
4. Primary Endpoints
The five registered primary endpoints span one efficacy endpoint and four safety endpoints. The ClinicalTrials.gov record classifies the endpoints as binary or other/unclear, while the posted analyses use either time-to-event or categorical methods depending on the endpoint.
| Endpoint | Registered definition | Time frame | Posted analysis |
|---|---|---|---|
| HIV Seroconversion | Confirmed HIV infection | Monthly follow-up through a median of 1.2 years | Log-rank test; HR |
| Grade 1 or Higher Creatinine Toxicity | Creatinine which reach grade 1 (mild, 1.1 to 1.3 local upper limit of normal) or higher by the US Division of AIDS grading table (version 1) or a 50% increase in creatinine from the baseline value. | Duration of follow-up, median 1.2 years | Fisher exact test; RR |
| Grade 3 or Higher Phosphorous Toxicity | Grade 3 or higher phosphorous toxicity (hypophosphatemia) by the Division of AIDS Grading Table (severe, level at or below 1.9 mg/dL) | The entire follow-up period, median 1.2 years | Fisher exact test; RR |
| Grade 2, 3, or 4 Laboratory Adverse Events | Number of participants with at least one Grade 2, 3, or 4 laboratory adverse events (moderate, severe or life threatening based on the US Division of AIDS Grading of adverse events, version 1.0). | Entire follow-up, median 1.2 years | Log-rank test; HR |
| Grade 2, 3, or 4 Clinical Adverse Events | Number of participants with at least 1 Grade 2, 3, or 4 clinical adverse events (moderate, severe or life threatening based on the US Division of AIDS Grading of adverse events, version 1.0). | Entire follow-up, median 1.2 years | Log-rank test; HR |
5. Statistical Methodology
Time-to-event analysis and the log-rank test
The primary HIV seroconversion analysis used a log-rank test, a standard method for comparing time-to-event distributions between randomized groups. The same method was reported for Grade 2, 3, or 4 laboratory adverse events and Grade 2, 3, or 4 clinical adverse events.
For iPrEx, the time-to-event framework is especially relevant to HIV seroconversion because participants were followed repeatedly and the registry specifies monthly follow-up through a median of 1.2 years.
Hazard ratio
The HIV seroconversion analysis reported a hazard ratio of 0.577 for TDF/FTC versus placebo, with placebo as the reference group. A hazard ratio below 1 indicates a lower estimated instantaneous event rate in the TDF/FTC group under the fitted comparison.
The hazard ratio is a relative time-to-event measure. It is not itself an absolute risk difference, a risk ratio, or the percentage of participants who experienced the event.
Fisher exact test
Fisher exact testing was reported for Grade 1 or Higher Creatinine Toxicity and Grade 3 or Higher Phosphorous Toxicity. These endpoints are defined as participant-level toxicity categories, making a two-group categorical comparison appropriate to the endpoint structure.
Mixed-effects models
The registry reports mixed-model analyses for several secondary endpoints, including percentage change in bone mineral density, CD4 count among HIV infected participants, and proportion of missed doses by pill count. Mixed-effects models are useful when observations have a longitudinal or repeated-measure structure because they can model within-participant dependence rather than treating repeated observations as independent.
t-tests
Two-sided t-tests were reported for viral load among HIV infected participants and for the percentage of missed doses by estimate during the computer assisted structured interview. The corresponding effect measure was a mean difference.
Wilcoxon / Mann-Whitney tests
Wilcoxon (Mann-Whitney) testing was reported for the number of condomless sexual partners with HIV positive or unknown status and for total number of sexual partners. This nonparametric approach is useful when a comparison is based on the distribution of a continuous or count-type outcome without relying on the assumptions of a standard two-sample t-test.
Chi-squared testing
A chi-squared test was reported for condomless receptive anal intercourse in the previous 12 weeks with any partners regardless of status. The reported effect measure was risk difference.
Median regression
The body-composition and lipid analyses used median regression for several week-24 percentage-change outcomes. The posted effect measure was a median difference. This is conceptually different from a mean difference: it focuses on the difference between the group-specific medians rather than the arithmetic means.
6. Primary Results: HIV Seroconversion
HIV seroconversion was the primary efficacy endpoint. The registered definition was confirmed HIV infection, assessed with monthly follow-up through a median of 1.2 years.
Hazard ratio for HIV seroconversion
95% CI: .404–.824 · P = 0.002
Analysis: log-rank test; placebo was the reference group.
| Feature | Reported value |
|---|---|
| Analysis population | Excludes participants who were HIV+ at enrollment (2 TDF/FTC, 8 Placebo) and those with no follow-up HIV test (25 TDF/FTC and 22 Placebo). |
| Groups compared | TDF/FTC vs Placebo |
| Endpoint | Confirmed HIV infection |
| Time frame | Monthly follow-up through a median of 1.2 years |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.577 |
| 95% CI | .404–.824 |
| P-value | 0.002 |
| Analysis notes | Primary null hypothesis: Relative hazard of 0.7 or less. Secondary null hypothesis: Relative hazard of 1.0 or less. Stratified by site with Efron correction for ties. Placebo is reference. Results typically quoted as efficacy = 100*(1-HR). |
The reported HR of 0.577 means that the estimated instantaneous hazard of confirmed HIV infection in the TDF/FTC group was 0.577 times the corresponding hazard in the placebo group under the reported time-to-event analysis. Expressed using the registry's stated convention, 100 × (1 − HR) gives an estimated 42.3% lower hazard.
This does not mean that 42.3% of participants avoided HIV infection, that exactly 42.3% of infections were prevented in every individual, or that the absolute probability of infection was reduced by 42.3 percentage points. A hazard ratio is a relative measure of event occurrence over time.
The 95% confidence interval, .404–.824, describes uncertainty around the estimated relative hazard under the analysis framework. It does not describe the range of effects experienced by individual participants. Its width also matters: the point estimate is more informative when considered together with the interval rather than in isolation.
The P-value of 0.002 addresses the statistical evidence under the specified hypothesis framework; it does not measure the magnitude or clinical importance of the effect. A P-value is not a probability that the null hypothesis is true.
The registry also reports that the analysis was stratified by site and used an Efron correction for ties. Because this is a time-to-event analysis, censoring and the assumptions underlying interpretation of a hazard ratio remain relevant. The registry wording concerning the primary and secondary null hypotheses should be preserved rather than silently replaced with a different hypothesis statement.
Why the hazard ratio is the central efficacy measure here
The endpoint is not simply whether an infection was ever observed. Participants had follow-up over time, and the registry specifies monthly follow-up. A participant who remains infection-free through the available observation period contributes information about the timing of the event process even if that participant does not experience seroconversion during follow-up.
This is why a survival-analysis framework can be more informative than a simple comparison of crude proportions when follow-up times or censoring differ. The log-rank test provides the group-comparison component, while the hazard ratio provides a relative effect estimate.
7. Primary Safety Results
The remaining four primary endpoints concern creatinine toxicity, phosphorous toxicity, laboratory adverse events, and clinical adverse events. Each was formally analyzed in the registry.
Grade 1 or Higher Creatinine Toxicity
Risk ratio
95% CI: 0.79–2.25 · P = 0.28
Analysis: Fisher exact test.
| Feature | Reported value |
|---|---|
| Definition | Creatinine reaching grade 1 or higher by the US Division of AIDS grading table (version 1) or a 50% increase in creatinine from baseline. |
| Time frame | Duration of follow-up, median 1.2 years |
| Analysis population | All randomized participants with at least 1 follow-up creatinine value |
| Method | Fisher exact test |
| Effect measure | Risk ratio |
| Estimate | 1.33 |
| 95% CI | 0.79–2.25 |
| P-value | 0.28 |
A risk ratio of 1.33 indicates that the estimated risk of the registered creatinine-toxicity endpoint was 1.33 times the risk in the placebo group. Because the 95% CI extends from 0.79 to 2.25, the estimate is compatible with both a lower and a higher relative risk under the stated statistical framework.
The P-value of 0.28 does not establish the size of any possible difference. It is a measure of evidence against the specified null hypothesis, not a measure of whether the two groups are clinically identical. The analysis population also requires attention because it includes randomized participants who had at least one follow-up creatinine value rather than simply every enrolled participant.
Grade 3 or Higher Phosphorous Toxicity
Risk ratio
95% CI: .57–2.96 · P = 0.54
Analysis: Fisher exact test.
| Feature | Reported value |
|---|---|
| Definition | Grade 3 or higher phosphorous toxicity (hypophosphatemia) by the Division of AIDS Grading Table; severe, level at or below 1.9 mg/dL. |
| Time frame | The entire follow-up period, median 1.2 years |
| Analysis population | All participants with at least 1 follow-up phosphorus value |
| Method | Fisher exact test |
| Effect measure | Risk ratio |
| Estimate | 1.3 |
| 95% CI | .57–2.96 |
| P-value | 0.54 |
The point estimate of 1.3 is above 1, but the 95% CI of .57–2.96 is wide and includes 1. The interval therefore communicates substantial uncertainty around the estimated risk ratio. The P-value of 0.54 should not be interpreted as evidence that the two risks are exactly equal; it describes the evidence under the stated hypothesis-testing framework.
The endpoint is also specifically defined as grade 3 or higher phosphorous toxicity, so the estimate should not be generalized to every phosphorous measurement or to less severe laboratory abnormalities.
Grade 2, 3, or 4 Laboratory Adverse Events
Hazard ratio
95% CI: 0.65–1.23 · P = 0.50
Analysis: log-rank test.
| Feature | Reported value |
|---|---|
| Definition | Participants with at least one Grade 2, 3, or 4 laboratory adverse event. |
| Time frame | Entire follow-up, median 1.2 years |
| Analysis population | Participants with at least one visit with laboratory values post-baseline |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.89 |
| 95% CI | 0.65–1.23 |
| P-value | 0.50 |
The HR of 0.89 is below 1, corresponding to a lower estimated instantaneous event rate in the TDF/FTC group under this analysis. However, the 95% CI of 0.65–1.23 includes 1, so the point estimate should not be treated as evidence of a precisely established relative reduction.
Because the endpoint is defined as time to the specified laboratory adverse-event category, the hazard ratio should be interpreted as a time-to-event measure rather than as a simple ratio of percentages. The P-value of 0.50 is not an effect-size measure.
Grade 2, 3, or 4 Clinical Adverse Events
Hazard ratio
95% CI: 0.79–1.23 · P = 0.92
Analysis: log-rank test.
| Feature | Reported value |
|---|---|
| Definition | Participants with at least 1 Grade 2, 3, or 4 clinical adverse event. |
| Time frame | Entire follow-up, median 1.2 years |
| Analysis population | Participants with at least one follow-up visit |
| Method | Log-rank test |
| Effect measure | Hazard ratio |
| Estimate | 0.99 |
| 95% CI | 0.79–1.23 |
| P-value | 0.92 |
An HR of 0.99 is very close to 1, indicating little separation in the estimated instantaneous event rates at the point estimate. The 95% CI of 0.79–1.23 nevertheless represents uncertainty around that estimate, and the interval includes 1.
The P-value of 0.92 should be read as a statistical-test result, not as a probability that the treatment groups are identical. The endpoint also concerns Grade 2, 3, or 4 clinical adverse events specifically; it should not be interpreted as a comparison of every possible adverse event.
8. Primary Results in One Statistical View
| Primary endpoint | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| HIV Seroconversion | Log-rank | HR | 0.577 | .404–.824 | 0.002 |
| Grade 1 or Higher Creatinine Toxicity | Fisher exact | RR | 1.33 | 0.79–2.25 | 0.28 |
| Grade 3 or Higher Phosphorous Toxicity | Fisher exact | RR | 1.3 | .57–2.96 | 0.54 |
| Grade 2, 3, or 4 Laboratory Adverse Events | Log-rank | HR | 0.89 | 0.65–1.23 | 0.50 |
| Grade 2, 3, or 4 Clinical Adverse Events | Log-rank | HR | 0.99 | 0.79–1.23 | 0.92 |
This table illustrates why a trial should not be summarized by a single P-value. iPrEx has multiple prespecified primary endpoints with different outcome structures and different effect measures. The HIV endpoint uses a time-to-event hazard ratio, two toxicity endpoints use risk ratios, and two other toxicity endpoints use hazard ratios.
9. Secondary Endpoint Results
The registry contains a broad set of secondary analyses spanning virologic, metabolic, adherence, sexual-behavior, and sexually transmitted infection outcomes. The reported estimates below are presented using the same endpoint definitions and time frames reported in the ClinicalTrials.gov record.
Hepatitis Flares Among HBV-Infected Persons
This endpoint assessed hepatitis flares among persons with chronic active hepatitis B at enrollment, using quarterly laboratory tests through a median follow-up of 1.2 years.
| Method | Effect measure | Estimate | CI | P-value |
|---|---|---|---|---|
| Fisher exact | Risk difference | 0.00 | Not reported in the statistical analysis | 1.00 |
The registry also lists a second posted Fisher exact analysis of this same endpoint with a P-value of 1.00 and no effect estimate. The available data therefore support reporting the primary posted estimate of 0.00 and the P-value of 1.00 without attempting to derive an additional confidence interval.
Percentage Change in Bone Mineral Density
| Time frame | Population | Method | Effect measure | Estimate | P-value |
|---|---|---|---|---|---|
| Baseline and week 24 | Participants confirmed to be HIV negative at enrollment who consented to participate in the metabolic substudy | Mixed-effects model | Mean difference (net) | -.91 | 0.001 |
The registry does not provide a confidence interval for this analysis in the ClinicalTrials.gov record.
Percentage Change in Body Fat
Median difference
95% CI: -6.6 to -0.95 · P = 0.009
Baseline to week 24; median regression.
The effect measure is a median difference (net). The negative estimate indicates that the reported median percentage change was lower in the TDF/FTC group than in the placebo group by the estimated difference. The confidence interval is entirely below zero, while the P-value is 0.009.
Percentage Change in Fasting Triglycerides
| Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|
| Median difference (net) | 0.0 | -9.3 to 9.3 | 1.00 |
The estimate is exactly 0.0, while the 95% CI ranges from -9.3 to 9.3. The interval therefore allows for differences in either direction under the statistical model. The P-value of 1.00 is a test result and should not be interpreted as proving that the treatment groups have identical triglyceride responses.
Percent Change in Total Cholesterol
| Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|
| Median difference (net) | -2.2 | -5.5 to 1.1 | 0.19 |
The estimated median difference is -2.2, but the confidence interval of -5.5 to 1.1 crosses zero. This illustrates why the point estimate and its uncertainty should be interpreted together.
Viral Load Among HIV Infected Participants
| Time frame | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| At the time closest to HIV detection | Two-sided t-test | Mean difference (net) | 0.08 | -.18 to 0.33 | 0.56 |
The analysis included all HIV infections detected during the study, including infections prior to, during, and after study treatment. The mean difference of 0.08 was estimated with a 95% CI from -.18 to 0.33.
Drug Resistance Among HIV Infected Participants
| Analysis population | Method | Effect measure | Estimate | P-value |
|---|---|---|---|---|
| 2 TDF/FTC seroconversions at enrollment and 48 during follow-up; 8 placebo seroconversions at enrollment and 83 during follow-up | Fisher exact | Risk ratio | 1.00 | 1.00 |
The registry states that the null hypothesis is that the proportion of mutations is identical. No confidence interval is in the ClinicalTrials.gov record in the provided statistical-analysis record.
CD4 Count Among HIV Infected Participants
| Time frame | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| At the time infection was detected | Mixed-effects model | Mean difference (net) | -7 | -69 to 54 | 0.32 |
The analysis included HIV infected participants during the trial, including those who were HIV+ at baseline. The confidence interval spans both negative and positive differences, illustrating the uncertainty around the estimated mean difference of -7 cells per cubic mm.
Adherence Measures
| Endpoint | Time frame | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|---|
| Proportion of Missed Doses by Pill Count | At 24 weeks | Mixed-effects model | Mean difference (net) | -0.005 | -0.02 to 0.01 | 0.53 |
| Percentage of Missed Doses by Estimate During CASI Interview | Week 24 | Two-sided t-test | Mean difference (net) | 0.25 | -1.1 to 1.6 | 0.70 |
These are two different adherence measurements. The pill-count endpoint concerns the proportion of pills not returned among those whose bottles were returned at the week 24 visit. The CASI endpoint concerns the estimated adherence response among participants who answered the adherence question at week 24. They should not be treated as interchangeable measurements.
Sexual-Behavior Measures
| Endpoint | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Number of Condomless Sexual Partners With HIV Positive or Unknown Status | Wilcoxon (Mann-Whitney) | Mean difference (final values) | 0 | Not reported | 0.99 |
| Total Number of Sexual Partners | Wilcoxon (Mann-Whitney) | Median difference (final values) | 0 | -.42 to .42 | 0.76 |
| Condomless Receptive Anal Intercourse in the Previous 12 Weeks With Any Partners Regardless of Status. | Chi-squared | Risk difference | -0.008 | -0.047 to 0.030 | 0.68 |
The registry specifies a null of no difference between the arms for the condomless receptive anal intercourse endpoint. These analyses address reported behavioral outcomes at week 24; they should not be conflated with the time-to-event analysis of confirmed HIV infection.
Sexually Transmitted Infection Endpoints
| Endpoint | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Incidence of Confirmed Syphilis During Follow-Up | Log-rank | Hazard ratio | 1.13 | 0.89–1.43 | 0.30 |
| Incidence of HSV-2 During the Follow-up Period | Log-rank | Hazard ratio | 1.2 | 0.8–1.7 | 0.41 |
| Diagnosis of Gonorrhea During the Follow-up Period | Log-rank | Hazard ratio | 0.61 | 0.34–1.09 | 0.09 |
These three outcomes demonstrate the same general time-to-event logic used for HIV seroconversion, but they are separate secondary endpoints. For gonorrhea, for example, an HR of 0.61 is a relative estimate of the event hazard; it is not equivalent to a 39-percentage-point reduction in gonorrhea incidence.
10. Safety: Serious Adverse Events by Arm
The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm as affected participants divided by participants at risk.
| Arm | Affected | At risk | Reported proportion |
|---|---|---|---|
| TDF/FTC | 92 | 1226 | 92/1226 |
| Placebo | 94 | 1230 | 94/1230 |
The requested data provide the affected and at-risk counts but do not provide a formal statistical comparison, confidence interval, or P-value for serious adverse events. Accordingly, no inferential comparison is added here. The counts can be described directly, but a difference in observed counts should not be treated as evidence of a statistically established treatment effect without the corresponding analysis.
11. Secondary Endpoint Statistical Interpretation
iPrEx uses several effect measures because its endpoints have different structures. Hazard ratios describe relative event hazards over time; risk ratios compare risks; risk differences describe absolute differences in risk; mean differences compare arithmetic means; and median differences compare medians. These quantities are not interchangeable.
A confidence interval provides a range of parameter values compatible with the statistical model and observed data under the stated confidence procedure. For example, the body-fat median difference of -3.8 has a 95% CI of -6.6 to -0.95, whereas the total-cholesterol median difference of -2.2 has a 95% CI of -5.5 to 1.1. The latter interval includes zero, illustrating greater uncertainty about the direction of the underlying difference.
The registry contains P-values ranging from 0.001 to 1.00 across the posted analyses. These values describe evidence under particular null hypotheses. They do not directly communicate the magnitude of an effect, its practical importance, or the probability that a treatment works.
The posted analyses use endpoint-specific populations. HIV seroconversion excludes participants HIV+ at enrollment and participants without a follow-up HIV test. Creatinine toxicity requires at least one follow-up creatinine value. Other analyses restrict the population according to the relevant substudy or measurement availability. An effect estimate therefore describes the population actually analyzed, not automatically every person enrolled in the trial.
12. Statistical Methods Explained
Why was a log-rank test used for HIV seroconversion?
HIV seroconversion is a time-to-event endpoint because the analysis concerns when confirmed HIV infection occurs during follow-up. The log-rank test compares the event-time experience of the randomized groups while accommodating censoring. The registry additionally reports stratification by site and an Efron correction for ties.
What does an HR of 0.577 mean?
An HR of 0.577 means that the estimated instantaneous hazard of confirmed HIV infection was 0.577 times that in the placebo group under the reported model. Using the registry's stated efficacy convention, 100 × (1 − 0.577) = 42.3%, so the estimate corresponds to a 42.3% lower estimated hazard. It does not mean that 42.3% of participants were protected or that absolute infection probability fell by 42.3 percentage points.
Why is a Fisher exact test appropriate for the toxicity endpoints?
The creatinine and phosphorous endpoints are defined as participant-level toxicity categories. Fisher exact testing provides an exact categorical comparison between treatment groups and is especially useful when event counts may be limited. The associated effect measure in these analyses is the risk ratio.
What is the difference between a risk ratio and a hazard ratio?
A risk ratio compares the probability of an event over a specified period or under a specified risk definition. A hazard ratio compares instantaneous event rates over time within a time-to-event framework. An HR of 0.61 for gonorrhea, for example, should not be described as a 39-percentage-point difference in gonorrhea risk.
Why were mixed-effects models used for some secondary endpoints?
The registry reports mixed-model analyses for outcomes such as percentage change in bone mineral density, CD4 count, and pill-count adherence. Mixed-effects models are designed for data with repeated or clustered observations and can represent within-participant dependence. The exact model specification is not reported in the normalized registry summary, so additional model assumptions are not inferred here.
Why use a Wilcoxon / Mann-Whitney test for sexual-partner outcomes?
The registry reports Wilcoxon (Mann-Whitney) analyses for counts of sexual partners. This is a rank-based nonparametric comparison and does not require the same distributional assumptions as a conventional two-sample t-test. The interpretation should follow the effect measure actually reported: a final-value mean difference for one endpoint and a final-value median difference for another.
Why is the analysis population important?
A statistical estimate is conditional on who contributed data to that analysis. The HIV analysis excludes participants who were HIV+ at enrollment and those without a follow-up HIV test, while the metabolic and adherence analyses use substudy- or measurement-specific populations. Randomization establishes the treatment comparison, but missing follow-up and endpoint-specific eligibility determine which participants contribute to particular analyses.
13. Multiplicity and Multiple Primary Endpoints
The registry identifies five primary endpoints and reports formal analyses for all five. That structure is statistically important because a trial with multiple primary endpoints does not have the same inferential architecture as a trial with one primary endpoint.
| Primary endpoint | Outcome structure | Effect measure | Test |
|---|---|---|---|
| HIV Seroconversion | Time-to-event | Hazard ratio | Log-rank |
| Grade 1 or Higher Creatinine Toxicity | Binary | Risk ratio | Fisher exact |
| Grade 3 or Higher Phosphorous Toxicity | Other / unclear | Risk ratio | Fisher exact |
| Grade 2, 3, or 4 Laboratory Adverse Events | Time-to-event | Hazard ratio | Log-rank |
| Grade 2, 3, or 4 Clinical Adverse Events | Time-to-event | Hazard ratio | Log-rank |
The ClinicalTrials.gov record identifies the hypothesis type as superiority for these analyses, but they do not provide a multiplicity-adjustment procedure, alpha allocation, or hierarchical testing strategy. Consequently, no specific familywise error-control claim is added.
14. Stratification and the HIV Analysis
The HIV seroconversion analysis notes that the analysis was stratified by site and used an Efron correction for ties. Stratification is useful when a trial has a prespecified grouping factor that should be respected during the time-to-event comparison.
Why stratify?
A stratified analysis allows the comparison to account for the site structure rather than treating all event times as if they arose from one completely homogeneous stratum.
Why correct for ties?
Survival data can contain tied event times. The registry specifically reports the Efron correction for ties in the HIV analysis.
Stratification does not transform the hazard ratio into an absolute measure. It changes the statistical framework used to compare event histories while preserving the interpretation of the reported HR as a relative hazard measure.
15. Time-to-Event Interpretation and Censoring
Three of the five primary endpoints were analyzed with log-rank tests and hazard ratios: HIV seroconversion, Grade 2, 3, or 4 laboratory adverse events, and Grade 2, 3, or 4 clinical adverse events.
A participant who does not experience an endpoint during observed follow-up is not necessarily treated as having an event at the end of follow-up. Instead, the participant contributes information up to the censoring time.
This distinction explains why a hazard ratio cannot be reconstructed simply by comparing the number of observed events without knowing the timing of those events and the censoring structure.
16. Secondary Analyses: Longitudinal Data
Several secondary outcomes are measured at baseline and week 24 or otherwise involve repeated measurements. The registry reports mixed-effects models for percentage change in bone mineral density, CD4 count, and pill-count adherence.
| Endpoint | Time frame | Model | Effect |
|---|---|---|---|
| Percentage Change in Bone Mineral Density | Baseline and week 24 | Mixed-effects model | Mean difference (net) = -.91 |
| CD4 Count Among HIV Infected Participants | At the time infection was detected | Mixed-effects model | Mean difference (net) = -7 |
| Proportion of Missed Doses by Pill Count | At 24 weeks | Mixed-effects model | Mean difference (net) = -0.005 |
The key statistical point is that a mixed-effects analysis can account for correlation among measurements from the same participant. This makes the method fundamentally different from simply applying an independent two-sample test to every observed measurement.
For the median-regression analyses, the registry specifically reports median regression, so the interpretation here follows that method rather than assuming an ordinary mean-based regression.
17. Missing Data and Analysis Populations
The ClinicalTrials.gov record demonstrate that endpoint-specific availability matters. The HIV analysis excludes participants without a follow-up HIV test. The creatinine analysis includes randomized participants with at least one follow-up creatinine value. The phosphorus analysis includes participants with at least one follow-up phosphorus value. Substudies impose additional eligibility requirements.
| Endpoint | Population restriction reported by the registry |
|---|---|
| HIV Seroconversion | Excludes HIV+ at enrollment and those with no follow-up HIV test. |
| Creatinine Toxicity | Randomized participants with at least 1 follow-up creatinine value. |
| Phosphorous Toxicity | Participants with at least 1 follow-up phosphorus value. |
| Bone Mineral Density | HIV-negative participants at enrollment who consented to the metabolic substudy. |
| Body Fat | Participants in the body-composition substudy with a week 24 body scan. |
| Fasting Triglycerides | Metabolic-substudy participants with a week 24 fasting triglyceride value. |
| Total Cholesterol | Metabolic-substudy participants with a week 24 fasting cholesterol measurement. |
| Adherence by Pill Count | Participants whose bottles were returned at the week 24 visit. |
The ClinicalTrials.gov record does not state a general missing-data imputation method. It would therefore be inappropriate to claim that missing observations were imputed using a particular technique such as multiple imputation, last observation carried forward, or pattern-mixture modeling.
18. What the Primary Efficacy Result Does — and Does Not — Mean
The estimated hazard of confirmed HIV infection was lower in the TDF/FTC group than in the placebo group, with an estimated hazard ratio of 0.577. Using the registry's stated efficacy convention, this corresponds to a 42.3% lower estimated hazard.
The HR is not the percentage of participants who avoided HIV infection, is not an absolute risk difference, and does not imply that every participant experiences the same proportional change in risk.
The 95% CI of .404–.824 shows uncertainty around the estimated hazard ratio. It is substantially more informative than reporting 0.577 alone because it communicates how precisely the effect was estimated under the analysis framework.
The P-value of 0.002 describes statistical evidence against the relevant null hypothesis. It does not measure effect size and should not be converted into a statement such as "99.8% probability that the treatment works."
19. Comparing the Different Effect Measures
| Effect measure | Example in iPrEx | Interpretive question |
|---|---|---|
| Hazard ratio | HIV seroconversion HR 0.577 | How do estimated event hazards compare over time? |
| Risk ratio | Creatinine toxicity RR 1.33 | How do the risks of the binary endpoint compare? |
| Risk difference | Condomless receptive anal intercourse RD -0.008 | How far apart are the two risks on an absolute scale? |
| Mean difference | CD4 count difference -7 | How far apart are the group means? |
| Median difference | Body-fat difference -3.8 | How far apart are the group medians? |
This is one of the most useful statistical lessons from iPrEx: the effect measure must match the endpoint and analysis method. Treating all estimates as if they were interchangeable "risk reductions" would obscure important differences in what the analyses actually estimate.
20. Limitations
- Registry-level detail: the ClinicalTrials.gov record provides the posted estimates and methods but do not contain the full statistical analysis plan or participant-level dataset.
- Multiple primary endpoints: five primary endpoints are registered, but the ClinicalTrials.gov record does not specify an overall multiplicity-adjustment strategy.
- Endpoint-specific populations: different analyses use different eligibility and follow-up requirements, so estimates should not automatically be generalized to the entire enrolled population.
- Censoring: time-to-event estimates depend on follow-up and censoring information that is not available at the participant level in the ClinicalTrials.gov record.
- Hazard-ratio interpretation: a single HR summarizes a relative time-to-event comparison and should not be interpreted as an absolute risk difference. The proportional-hazards assumption is also relevant to Cox-type hazard-ratio interpretation, although the registry summary does not provide a diagnostic assessment of that assumption.
- Secondary analyses: the registry contains many secondary endpoints using different methods and populations. Their P-values should be interpreted according to the prespecified statistical plan, which is not fully reproduced in the ClinicalTrials.gov record.
- Missing-data methods: the ClinicalTrials.gov record identifies analysis-population restrictions but does not specify a general imputation strategy.
- Substudy interpretation: metabolic and body-composition analyses use substudy populations rather than the entire trial enrollment.
- Serious adverse events: affected and at-risk counts are available by arm, but no formal statistical comparison is posted on ClinicalTrials.gov for those counts.
- External generalizability: the ClinicalTrials.gov record does not provide a complete baseline-characteristics table, so the extent to which the enrolled population represents other populations cannot be assessed from this dataset alone.
21. Why This Trial Matters Statistically
iPrEx is a particularly useful statistical teaching case because a single randomized trial contains several fundamentally different analysis structures. The primary efficacy endpoint is a time-to-event outcome, while primary safety endpoints include both categorical risk comparisons and additional time-to-event analyses. Secondary outcomes extend the framework to longitudinal mixed models, mean and median differences, nonparametric tests, and categorical risk differences.
| Statistical concept | How it appears in iPrEx |
|---|---|
| Randomization | Parallel randomized comparison of daily TDF/FTC and placebo. |
| Double masking | Double-blind trial design. |
| Time-to-event analysis | HIV seroconversion, laboratory adverse events, clinical adverse events, syphilis, HSV-2, and gonorrhea. |
| Log-rank test | Used for multiple time-to-event comparisons. |
| Hazard ratio | Reported for HIV seroconversion and several secondary and safety endpoints. |
| Fisher exact test | Used for creatinine toxicity, phosphorous toxicity, hepatitis flares, and drug resistance. |
| Risk ratio | Used for creatinine toxicity, phosphorous toxicity, and drug resistance. |
| Risk difference | Used for hepatitis flares and condomless receptive anal intercourse. |
| Mixed-effects model | Used for bone mineral density, CD4 count, and pill-count adherence. |
| t-test | Used for viral load and CASI-based adherence. |
| Wilcoxon / Mann-Whitney | Used for sexual-partner outcomes. |
| Chi-squared test | Used for condomless receptive anal intercourse. |
| Median regression | Used for body fat, fasting triglycerides, and total cholesterol percentage changes. |
| Confidence intervals | Reported for many, but not all, posted effect estimates. |
| Multiple endpoints | Five registered primary endpoints plus numerous secondary outcomes. |
The statistical value of the trial is therefore not limited to its headline efficacy estimate. It provides a compact example of why endpoint type determines analysis method, why effect measures must be interpreted in their correct units, and why the analysis population and follow-up structure matter.
22. A Statistical Reading of the Primary Endpoint Set
The five primary endpoints can be read as a sequence of related but distinct questions. First, the HIV seroconversion analysis asks about the timing of confirmed infection. The toxicity endpoints then ask whether the randomized groups differ in specified laboratory or clinical safety outcomes. The result is not one combined statistic but a collection of effect estimates tailored to different endpoint definitions.
Efficacy
HIV seroconversion was analyzed as a time-to-event endpoint, producing an HR of 0.577 with a 95% CI of .404–.824 and P = 0.002.
Renal laboratory toxicity
Grade 1 or higher creatinine toxicity was analyzed using Fisher exact testing, with RR 1.33 and 95% CI 0.79–2.25.
Phosphorous toxicity
Grade 3 or higher phosphorous toxicity was analyzed using Fisher exact testing, with RR 1.3 and 95% CI .57–2.96.
Broader safety endpoints
Laboratory and clinical adverse events were analyzed with log-rank methods, with HRs of 0.89 and 0.99 respectively.
This structure is important because the absence of a statistically strong result for one safety endpoint does not logically establish the same conclusion for another endpoint. Each endpoint has its own definition, analysis population, event process, and uncertainty.
23. Clinical Biostats Statistical Takeaways
Before interpreting a P-value or effect estimate, identify exactly what was measured. Confirmed HIV infection, creatinine toxicity, bone mineral density, viral load, and sexual behavior are not statistically interchangeable outcomes.
The iPrEx registry demonstrates the distinction among hazard ratios, risk ratios, risk differences, mean differences, and median differences. The numerical value of an estimate has meaning only in the context of its effect-measure definition.
The confidence interval communicates uncertainty around an estimate. For example, the HIV HR interval of .404–.824 is narrower than the phosphorous-toxicity RR interval of .57–2.96, reflecting different levels of precision in those estimates.
A P-value tells you about statistical evidence under a specified null hypothesis. It does not tell you how large an effect is, how precise the estimate is, or whether an effect has a particular practical importance.
The HIV analysis and the metabolic substudy analyses do not use identical populations. Understanding who contributed data is necessary before extending an estimate beyond the analyzed participants.
24. Related Tutorials
Learn more about the methods used in this trial:
25. Related Statistical Calculators
26. Sources
- ClinicalTrials.gov: iPrEx, NCT00458393.
- PubMed record: PMID 21091279.
- PubMed record: PMID 37969014.
- PubMed record: PMID 31192894.
- PubMed record: PMID 29415175.
- PubMed record: PMID 28639995.
Continue through the Clinical Biostats statistical pathway
Use the related tutorials and calculators to explore the survival, categorical-data, longitudinal, and nonparametric methods represented in this trial.
27. Record Summary
iPrEx provides a broad example of clinical-trial statistical analysis within a single randomized phase 3 prevention study. The primary HIV seroconversion endpoint was analyzed with a site-stratified log-rank framework and reported an HR of 0.577 with a 95% CI of .404–.824 and P = 0.002. The four primary safety endpoints used either Fisher exact or log-rank methods, producing risk-ratio or hazard-ratio estimates with their corresponding uncertainty intervals. Secondary analyses extended the statistical framework to mixed-effects models, t-tests, Wilcoxon / Mann-Whitney testing, chi-squared testing, and median regression.
The most important statistical lesson is that the trial cannot be reduced to one number. Correct interpretation requires keeping the endpoint definition, analysis population, time frame, statistical method, effect measure, confidence interval, and P-value connected. That framework allows the reported results to be understood without treating different statistical quantities as if they represented the same underlying question.