← Clinical Trials
HIV Prevention Phase 3 Randomized NCT00625404

FEM-PrEP: Complete Statistical Analysis of Truvada in HIV Prevention

An independent statistical analysis of the randomized phase 3 FEM-PrEP trial evaluating Truvada versus placebo for preventing HIV acquisition in women, with emphasis on time-to-event analysis, safety endpoints, secondary continuous outcomes, and interpretation of reported effect estimates.

FEM-PrEP  ·  Phase 3  ·  Completed  ·  2009-05 to 2012-08
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results and trial characteristics are restricted to the ClinicalTrials.gov record. The page does not add outcome estimates, baseline characteristics, subgroup results, or other findings that are not contained in that record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

FEM-PrEP was a randomized, parallel-group, quadruple-masked phase 3 prevention trial evaluating Truvada versus placebo for HIV acquisition in women. The trial enrolled 2120 participants and had a primary purpose of prevention.

2120
Enrollment
2 treatment arms
3
Phase
Phase 3
0.94
HIV Infection HR
95% CI 0.59–1.52
0.45
Creatinine P-value
Log-rank test
FeatureFEM-PrEP
Trial nameFEM-PrEP
Brief titleFEM-PrEP (Truvada®): Study to Assess the Role of Truvada® in Preventing HIV Acquisition in Women
PhasePhase 3
StatusCompleted
Therapeutic areaInfectious Disease
ConditionHIV Infections
Primary purposePrevention
AllocationRandomized
Design modelParallel
MaskingQuadruple
Enrollment2120
Arms2
InterventionsTruvada (drug); Placebo (other)
Lead sponsorFHI 360
Study period2009-05 through 2012-08
ClinicalTrials.govNCT00625404

2. Clinical Question

The primary clinical question was whether Truvada, compared with placebo, was associated with a different rate of HIV infection during follow-up in women enrolled in this prevention trial. The registry describes HIV infection as a primary time-to-event endpoint measured between enrollment and 52 weeks.

Population

Women participating in a phase 3 prevention study for HIV infection. The ClinicalTrials.gov record specifies 2120 enrolled participants.

Intervention

Truvada.

Comparator

Placebo.

Primary question

How does the hazard of HIV infection compare between the Truvada and placebo arms during the registered follow-up period?

3. Trial Design

01
Randomize 2120 enrolled
02
Parallel arms Truvada vs placebo
03
Quadruple mask Masked study design
04
Follow-up Primary HIV endpoint through 52 weeks
05
Analysis Cox and log-rank methods
TRUVADA ARM

Truvada

  • Intervention classified as a drug
  • Compared directly with placebo
  • Included in the randomized parallel-group comparison
PLACEBO ARM

Placebo

  • Comparator classified as other
  • Compared directly with Truvada
  • Included in the randomized parallel-group comparison

The combination of randomization, parallel allocation, and quadruple masking is statistically important because it separates the treatment assignment from subsequent outcome assessment and participant behavior as far as the trial design permits. Randomization establishes the basis for the between-arm causal comparison, while masking is intended to reduce information and behavior-related biases.

4. Endpoints

The registry lists six primary endpoints. They do not all have the same statistical structure: HIV infection and several laboratory toxicity endpoints are analyzed as time-to-event outcomes, while the adverse-event frequency endpoint is recorded as a count-based safety outcome for which the ClinicalTrials.gov record does not report a formal statistical comparison.

Primary endpointRegistered time frameEndpoint definitionStatistical structure
HIV Infection Cumulative HIV infection between enrollment and 52 weeks HIV Seroconversion, with time to infection refined based on PCR results obtained from stored specimens. Time-to-event
Confirmed Grade 2 or Higher Serum Creatinine Toxicity cumulative toxicity through 52 weeks of product use and 4 weeks post product Repeat specimens were collected to confirm chemistry toxicities. Grade 2 or higher serum creatinine toxicity was defined as ≥1.4 times the upper limit of normal. Time-to-event
Frequency of Adverse Events (AEs) During and Within 4 Weeks After Study Product Administration 10-26 months per site The total number of adverse events in the placebo and Truvada arms during and within 4 weeks after study product administration. Other / unclear
Confirmed Grade 3 or Higher Reduction in Phosphorus Through 52 weeks on product and 4 weeks post-product Repeat specimens were collected to confirm chemistry toxicities. Grade 3 phosphorus reduction was defined as ≤2.4mg/dL. Time-to-event
Confirmed Grade 3 or Higher ALT Elevation Through 52 weeks on product and 4 weeks post-product Grade 3 or higher ALT elevation was defined as ≥ 2.6 times the upper limit of normal. Time-to-event
Confirmed Grade 3 or Higher AST Elevation Through 52 weeks on product and 4 weeks post-product Grade 3 or higher AST elevation was defined as ≥ 2.6 times the upper limit of normal. Time-to-event

5. Statistical Methodology

Cox proportional-hazards model

The primary HIV infection analysis used a proportional-hazards model, with the analysis text specifying stratification on site. The effect measure was a hazard ratio comparing the Truvada arm with the placebo arm.

Primary HIV analysis
HR = hazard of HIV infection in Truvada ÷ hazard of HIV infection in placebo

The reported estimate was HR 0.94 with a two-sided 95% confidence interval of 0.59–1.52.

A hazard ratio is a relative measure of the event rate over time. In a proportional-hazards framework, an HR below 1 corresponds to a lower estimated instantaneous event rate in the treatment group, whereas an HR above 1 corresponds to a higher estimated instantaneous event rate. The interpretation depends on the model and its proportional-hazards assumption.

Stratified analysis

The HIV infection analysis was stratified on site. Stratification allows the baseline hazard to vary across the specified strata while estimating a common treatment effect across those strata. This can be useful when study sites may have materially different underlying event patterns.

Log-rank test

The registry reports log-rank tests for confirmed creatinine toxicity, phosphorus reduction, ALT elevation, and AST elevation. These endpoints are described as rates or time to first event, making a survival-analysis comparison appropriate when participants can have different amounts of follow-up or can be censored before experiencing the event.

Two-sided t-test

The secondary analyses of plasma HIV RNA level, CD4+ T-cell count, and participant-reported change in number of sexual partners used two-sided t-tests. These analyses compare mean values between the randomized groups for the specified continuous outcome.

Mean-difference framework
Mean difference = mean in Truvada group − mean in placebo group

A mean difference describes the average separation between the groups on the analyzed continuous outcome. It is not a hazard ratio and should not be interpreted using time-to-event terminology.

Analysis populations

For the primary HIV infection analysis, the registry specifies an analysis population consisting of all randomized participants who made at least one follow-up visit and were not HIV-positive at enrollment. For the creatinine safety analysis, the Safety Population consisted of randomized women who had at least one follow-up visit and did not return all of their product unused, with the registry noting that only women assessed for a particular safety outcome were included in that analysis.

Why the population definition matters: a hazard ratio calculated in a defined analysis population is not automatically interchangeable with a hazard ratio calculated in every enrolled participant. Exclusions based on enrollment status, follow-up, or safety assessment can change the population contributing information to the estimate.

6. Primary Results: HIV Infection

The primary HIV infection endpoint was cumulative HIV infection between enrollment and 52 weeks. The analysis included all randomized participants who made at least one follow-up visit and were not HIV-positive at enrollment.

Hazard ratio for HIV infection

0.94

95% CI: 0.59–1.52   ·   two-sided

Analysis: Cox proportional-hazards model, stratified on site

The trial was designed to have 90% power to reject the null hypothesis that the HR for infection is > 0.3, according to the registry analysis notes.

Clinical Biostats interpretation

An HR of 0.94 means that, under the reported proportional-hazards model, the estimated instantaneous rate of HIV infection in the Truvada group was approximately 94% of the estimated rate in the placebo group. Expressed as a relative comparison, 0.94 corresponds to an estimated hazard approximately 6% lower in the Truvada arm.

The HR does not mean that 94% of participants remained HIV-free, that 6% of participants were protected, or that an individual participant had exactly a 6% reduction in infection risk. It is a model-based relative measure of the event rate over time.

The two-sided 95% CI of 0.59–1.52 describes uncertainty around the estimated hazard ratio. Its width indicates that the estimate is not highly precise: values below 1 and values above 1 remain within the interval. The interval is not a prediction interval for individual participants and does not describe the range of individual treatment responses.

The p-value is not posted on ClinicalTrials.gov for this primary HIV infection analysis in the ClinicalTrials.gov record. More generally, a p-value would address evidence against a specified null hypothesis; it would not quantify the size or clinical importance of the treatment effect.

Because the analysis uses a Cox proportional-hazards model, interpretation also depends on the proportional-hazards assumption. The ClinicalTrials.gov record does not provide a diagnostic assessment of that assumption.

7. Primary Safety Results

Four of the registered primary safety endpoints have formal statistical analyses in the ClinicalTrials.gov record. Each used a log-rank test comparing the Truvada and placebo groups.

Confirmed Grade 2 or Higher Serum Creatinine Toxicity

Log-rank comparison

P = 0.45

Time to first grade 2 or higher creatinine toxicity

Time frame: cumulative toxicity through 52 weeks of product use and 4 weeks post product

Clinical Biostats interpretation

The reported two-group comparison used a log-rank test based on time to first creatinine toxicity. A p-value of 0.45 does not measure the magnitude of any difference between groups. The ClinicalTrials.gov record does not report a hazard ratio or confidence interval for this endpoint, so the magnitude and precision of the between-group difference cannot be quantified from the provided formal analysis.

The endpoint was defined using a confirmed laboratory toxicity threshold: grade 2 or higher serum creatinine toxicity was defined as ≥1.4 times the upper limit of normal. Because the analysis is time-to-event based, censoring and time at risk are part of the statistical structure.

Confirmed Grade 3 or Higher Reduction in Phosphorus

Log-rank comparison

P = 0.59

Time frame: through 52 weeks on product and 4 weeks post-product

Clinical Biostats interpretation

The log-rank p-value of 0.59 summarizes the reported statistical comparison of the time-to-event distributions for confirmed grade 3 or higher phosphorus reduction. It does not tell us the size of the treatment difference, because no effect estimate or confidence interval is in the ClinicalTrials.gov record.

The registry defines grade 3 phosphorus reduction as ≤2.4mg/dL. Repeat specimens were collected to confirm chemistry toxicities, so the endpoint represents a confirmed toxicity rather than simply a single unconfirmed measurement.

Confirmed Grade 3 or Higher ALT Elevation

Log-rank comparison

P = 0.79

Time frame: through 52 weeks on product and 4 weeks post-product

Clinical Biostats interpretation

The reported log-rank p-value of 0.79 is the formal comparison provided for confirmed grade 3 or higher ALT elevation. It should not be converted into an effect size. Without a hazard ratio, event counts, or confidence interval in the ClinicalTrials.gov record, the statistical magnitude and precision of the difference cannot be characterized further.

The registered definition of grade 3 or higher ALT elevation was ≥ 2.6 times the upper limit of normal.

Confirmed Grade 3 or Higher AST Elevation

Log-rank comparison

P = 0.62

Time frame: through 52 weeks on product and 4 weeks post-product

Clinical Biostats interpretation

The reported log-rank p-value of 0.62 summarizes the comparison of rates between the Truvada and placebo groups for confirmed grade 3 or higher AST elevation. As with the other log-rank results, the p-value alone does not quantify the size of a treatment difference.

The registered grade 3 or higher AST definition was ≥ 2.6 times the upper limit of normal. The ClinicalTrials.gov record does not provide an effect estimate or confidence interval for this endpoint.

Frequency of Adverse Events

The sixth registered primary endpoint was the frequency of adverse events during and within 4 weeks after study product administration, with a time frame of 10-26 months per site. The registry defines the endpoint as the total number of adverse events in the placebo and Truvada arms during and within 4 weeks after study product administration.

No formal comparison is posted on ClinicalTrials.gov for this endpoint. The ClinicalTrials.gov record indicate that results were posted, but no statistical analysis is listed for the adverse-event-frequency endpoint. For a count outcome of this type, the appropriate analysis would depend on the exact estimand and data structure; possible approaches can include comparisons of participant-level event incidence or count models. The ClinicalTrials.gov record does not report a formal statistical comparison, so no treatment-effect inference is added here.

8. Serious Adverse Events by Arm

The registry reports serious adverse events by treatment arm as affected participants divided by participants at risk.

Safety measureTruvada armPlacebo arm
Serious adverse events33/102523/1033
Participants affected by serious adverse events
Truvada
33
Placebo
23

The reported serious-adverse-event figures describe affected participants and the corresponding at-risk denominators. They should not be treated as a formal treatment comparison because the ClinicalTrials.gov record does not provide a p-value, confidence interval, hazard ratio, or other inferential statistic for this safety measure.

Denominator distinction: the serious-adverse-event denominators are 1025 for Truvada and 1033 for placebo. These are the denominators explicitly posted on ClinicalTrials.gov for this safety measure and should not be replaced with the overall enrollment of 2120.

9. Secondary Endpoint Results

Three secondary analyses are reported in the ClinicalTrials.gov record. Unlike the primary HIV infection analysis, these outcomes are continuous measures analyzed using two-sided t-tests.

Plasma HIV RNA Level

Mean difference in final values

0.03

P = 0.89   ·   two-sided t-test

Time frame: up to 16 weeks

The viral-load analysis included 68 women who became infected post-enrollment. Only 48 of these women contributed a specimen sample for analysis at the 16-weeks point: 27 on Truvada and 21 on placebo.

Clinical Biostats interpretation

The reported mean difference of 0.03 is the difference in final plasma HIV RNA levels between the Truvada and placebo groups on the reported log-copies/mL scale. The two-sided t-test produced P = 0.89.

The mean difference is a measure of average separation between groups; it is not a hazard ratio and does not describe the instantaneous rate of an event. The ClinicalTrials.gov record does not provide a confidence interval for this mean difference, so precision cannot be assessed from an interval estimate here.

The analysis population is also much smaller than the overall enrollment because it is restricted to women who became infected and, for the 16-week specimen analysis, to those who contributed a specimen. That distinction is essential when interpreting the estimate.

CD4+ T-cell Count

Mean difference in final values

22.1

P = 0.82   ·   two-sided t-test

Time frame: up to 16 weeks

Clinical Biostats interpretation

The reported mean difference in final CD4+ T-cell count was 22.1 cells/mL, with a two-sided t-test p-value of 0.82. The estimate represents the difference in average final values between the randomized groups for this secondary outcome.

The p-value does not indicate that the mean difference is clinically unimportant, nor does it quantify the size of the difference. The ClinicalTrials.gov record does not include a confidence interval, standard error, or standard deviation, so the precision of the 22.1 estimate cannot be evaluated from the reported information.

Participant Report of Change in Number of Sexual Partners

Mean difference in final values

.01

P = 0.73   ·   two-sided t-test

Time frame: up to 52 weeks

Clinical Biostats interpretation

The reported mean difference was .01 in the mean number of sexual partners, with P = 0.73. The analysis population consisted of women reporting on sexual behavior during follow-up.

The registry analysis text describes this as a t-test for the difference in change in number of sexual partners over time. The reported p-value is evidence about the statistical comparison under that test; it is not a measure of the size of the behavioral difference. No confidence interval is provided in the ClinicalTrials.gov record.

10. Statistical Methods Explained

Why was a Cox proportional-hazards model used for HIV infection?

HIV infection is naturally a time-to-event endpoint because both whether infection occurs and when it occurs are relevant. A Cox model can incorporate different follow-up times and right-censoring while estimating a relative hazard between treatment groups. In FEM-PrEP, the reported model was stratified on site.

What does an HIV infection HR of 0.94 mean?

An HR of 0.94 means that the estimated instantaneous rate of HIV infection in the Truvada group was 0.94 times the corresponding estimated rate in the placebo group under the fitted proportional-hazards model. It does not mean that 94% of participants avoided infection or that every participant received the same relative reduction.

Why does the 0.59–1.52 confidence interval matter?

The confidence interval communicates statistical uncertainty around the HR estimate. Because the registry-reported interval extends from 0.59 to 1.52, it includes values below and above 1. The interval is therefore important context that cannot be replaced by reporting the point estimate of 0.94 alone.

Why were log-rank tests used for the laboratory toxicity endpoints?

The registry analysis describes these toxicities using time to first event or differences in rates between groups. The log-rank test compares the event-time distributions while accounting for the timing of events and censoring, rather than reducing the data to a simple proportion observed at one arbitrary time.

What is the difference between a hazard ratio and a mean difference?

A hazard ratio is a relative measure for a time-to-event outcome. A mean difference compares average values for a continuous outcome. FEM-PrEP contains both structures: HIV infection and several toxicity endpoints use survival-analysis methods, whereas plasma HIV RNA, CD4+ T-cell count, and change in number of sexual partners use t-tests and mean differences.

Why should a p-value not be treated as an effect size?

A p-value is determined by the observed data, the null hypothesis, and the statistical model or test. It does not directly state how large or clinically meaningful the treatment difference is. An effect estimate and, ideally, its confidence interval are needed to describe magnitude and precision. This distinction is especially clear in the FEM-PrEP safety analyses, where several p-values are reported without corresponding effect estimates.

Why does the analysis population matter?

The HIV analysis excludes participants who were HIV-positive at enrollment and requires at least one follow-up visit. The viral-load analysis is narrower still, because it concerns women who became infected and, at the 16-week point, only women who contributed a specimen. Estimates from these populations answer different questions and should not be generalized automatically to all 2120 enrolled participants.

11. Confidence Intervals and the Primary HIV Estimate

Reported primary estimate
HR = 0.94    95% CI = 0.59–1.52

The interval is two-sided and accompanies the Cox proportional-hazards estimate for HIV infection.

The point estimate is only one part of the statistical result. If the HR were reported without its confidence interval, the reader would know the fitted estimate but not the uncertainty associated with it. The interval gives a range of parameter values compatible with the data under the stated statistical framework.

It is also important not to confuse a confidence interval with the range of outcomes that individual participants may experience. A confidence interval concerns uncertainty in the estimated population-level parameter; it does not say that individual hazards or individual treatment effects must fall between 0.59 and 1.52.

Educational distinction: an HR below 1 is not synonymous with proof of efficacy, and an HR above 1 is not synonymous with proof of harm. Statistical interpretation requires the estimate, uncertainty, prespecified hypothesis, analysis population, and study design to be considered together.

12. Time-to-Event Analysis in FEM-PrEP

The trial provides a useful example of why clinical-trial endpoints that occur at different times cannot always be analyzed as ordinary binary outcomes. For HIV infection, a participant can remain infection-free through follow-up, experience infection at a particular time, or contribute information until censoring.

Event

The event for the primary efficacy analysis was HIV seroconversion, with time to infection refined using PCR results from stored specimens.

Time scale

The registered primary time frame was cumulative HIV infection between enrollment and 52 weeks.

Censoring

Time-to-event methods allow participants who do not experience the event during observed follow-up to contribute information up to their censoring time.

Model

The primary HIV analysis used a Cox proportional-hazards model stratified on site.

The same general framework appears in the laboratory toxicity endpoints. For example, the creatinine analysis used a log-rank test based on time to first event. This is conceptually different from simply asking how many participants ever had a toxicity without regard to when it occurred.

13. Primary Endpoint Analysis Map

EndpointFormal methodEffect measure reportedInferential result reported
HIV Infection Cox proportional-hazards model, stratified on site Hazard ratio HR 0.94; 95% CI 0.59–1.52
Confirmed Grade 2 or Higher Serum Creatinine Toxicity Log-rank test Not reported P = 0.45
Frequency of Adverse Events No formal statistical analysis reported Not reported Results posted; no formal comparison reported
Confirmed Grade 3 or Higher Reduction in Phosphorus Log-rank test Not reported P = 0.59
Confirmed Grade 3 or Higher ALT Elevation Log-rank test Not reported P = 0.79
Confirmed Grade 3 or Higher AST Elevation Log-rank test Not reported P = 0.62

This table illustrates an important reporting principle: a trial can contain several formal statistical tests while only some analyses provide a directly interpretable effect estimate with a confidence interval. The absence of an effect estimate in the registry data is a reason to avoid inventing one from a p-value.

14. Secondary Analysis Map

Secondary outcomeTime frameAnalysisEffect measureResult
Plasma HIV RNA Level (HIV-1 Viral Load) up to 16 weeks t-test, 2 sided Mean Difference (Final Values) 0.03; P = 0.89
CD4+ T-cell Count Up to 16 weeks t-test, 2 sided Mean Difference (Final Values) 22.1; P = 0.82
Participant Report of Change in Number of Sexual Partners Up to 52 weeks t-test, 2 sided Mean Difference (Final Values) .01; P = 0.73

The three secondary outcomes demonstrate why endpoint scale matters. Viral load is measured in log copies/mL, CD4+ T-cell count in cells/mL, and participant-reported change in number of sexual partners as mean number of sexual partners. Their numerical mean differences therefore have different substantive meanings and should not be compared simply because all three were analyzed with t-tests.

15. Safety Analysis and Interpretation

The safety endpoints illustrate several distinct statistical questions. Confirmed laboratory toxicities are analyzed as time-to-event outcomes, while overall adverse-event frequency is defined as the total number of adverse events. Serious adverse events are reported as affected participants divided by participants at risk.

Safety conceptFEM-PrEP application
Confirmed laboratory toxicityRepeat specimens were collected to confirm chemistry toxicities.
Creatinine toxicityGrade 2 or higher defined as ≥1.4 times the upper limit of normal.
Phosphorus reductionGrade 3 reduction defined as ≤2.4mg/dL.
ALT elevationGrade 3 or higher defined as ≥ 2.6 times the upper limit of normal.
AST elevationGrade 3 or higher defined as ≥ 2.6 times the upper limit of normal.
Serious adverse eventsTruvada: 33/1025; placebo: 23/1033.

One statistical advantage of analyzing confirmed toxicity as time to first event is that the analysis can incorporate when the event occurred. A participant experiencing toxicity early and a participant experiencing toxicity late are not treated as though they registry-reported identical follow-up histories.

Do not infer a relative risk from the serious-adverse-event counts alone. The ClinicalTrials.gov record provides affected participants and denominators, but they do not provide a formal comparative estimate or confidence interval. A complete inferential analysis would require a prespecified definition of the estimand and the relevant event-time or participant-level data.

16. Multiplicity and Multiple Primary Endpoints

The registry lists 6 primary endpoints, of which the ClinicalTrials.gov record contains formal analyses for five. These endpoints span efficacy and safety domains and do not all use the same statistical method.

Multiplicity is important whenever several hypotheses are tested. If multiple independent hypothesis tests are interpreted simultaneously at an unadjusted significance threshold, the probability of obtaining at least one statistically significant result by chance can exceed the nominal threshold for a single test.

The registry-reported FEM-PrEP data do not provide an alpha-allocation strategy, multiplicity adjustment procedure, hierarchical testing sequence, or other formal familywise-error procedure. Accordingly, this page does not assign a confirmatory multiplicity interpretation beyond the methods explicitly reported in the registry data.

Practical lesson: the presence of several primary endpoints means that each reported p-value should be interpreted in the context of the trial's prespecified statistical plan. The registry extract does not contain enough information to reconstruct that complete multiplicity framework.

17. Blinding and Randomization

FEM-PrEP was randomized, parallel, and quadruple-masked. These design features matter statistically because they address different sources of bias.

Randomization

Random assignment provides the fundamental basis for comparing treatment groups without deliberately assigning treatment according to prognostic characteristics.

Parallel design

Participants remain in their assigned randomized treatment groups rather than serving sequentially as their own treatment and control.

Quadruple masking

The registry classifies the study as quadruple-masked, reducing the opportunity for knowledge of assignment to influence study conduct or assessment.

Prevention setting

The primary purpose was prevention, so the principal efficacy endpoint concerns acquisition of HIV rather than treatment response after established infection.

18. Why Time-to-Event Endpoints Are Different

A conventional binary analysis might classify each participant simply as infected or not infected by a specified date. A time-to-event analysis preserves more information by considering the time at which infection occurs and the amount of observed follow-up for participants who remain event-free.

Conceptual survival function
S(t) = P(T > t)

The survival function represents the probability of remaining event-free beyond time t. In a prevention trial, the event-free state corresponds conceptually to not yet experiencing the event of interest.

The Cox model then expresses the treatment comparison through a hazard ratio rather than directly through a difference in survival probabilities at one selected time point. That makes the reported HR of 0.94 a model-based summary of the relative event rate over time rather than a simple percentage difference.

19. Limitations

20. Why This Trial Matters Statistically

FEM-PrEP is a useful teaching case because it combines randomized prevention-trial design with several different endpoint structures. The primary HIV endpoint is time-to-event and uses a stratified Cox model, while several safety endpoints use log-rank tests and secondary biological or behavioral outcomes use two-sided t-tests.

Statistical conceptHow it appears in FEM-PrEP
RandomizationRandomized allocation to Truvada or placebo.
BlindingQuadruple-masked study design.
Time-to-event analysisHIV infection and several confirmed laboratory toxicity endpoints.
Cox proportional-hazards modelPrimary HIV infection analysis.
Hazard ratioPrimary HIV infection effect measure: HR 0.94.
Confidence interval95% CI 0.59–1.52 for the HIV infection HR.
Stratified analysisPrimary HIV model stratified on site.
Log-rank testFormal comparisons for four laboratory toxicity endpoints.
t-testSecondary analyses of plasma HIV RNA, CD4+ T-cell count, and change in sexual partners.
Analysis populationsDifferent populations are specified for HIV infection, safety, and viral-load analyses.
MultiplicitySix registered primary endpoints create a multiple-endpoint interpretation issue.

21. What the Primary HIV Result Can and Cannot Tell Us

What the estimate says

The reported Cox model produced an HIV infection hazard ratio of 0.94 for Truvada versus placebo, with a two-sided 95% CI of 0.59–1.52.

What the estimate does not say

The HR does not give an absolute probability of infection, an absolute risk difference, a number needed to treat, or the proportion of individual participants who benefited. None of those quantities is reported in the ClinicalTrials.gov record used for this page.

Why the confidence interval matters

The interval is substantially wider than the point estimate alone suggests. It spans values below and above 1, showing that the observed estimate should be interpreted together with substantial statistical uncertainty rather than as a precise single-number description of the treatment comparison.

Why the p-value cannot substitute for the effect estimate

The ClinicalTrials.gov record does not report a p-value for the primary HIV Cox analysis. More generally, even when a p-value is available, it cannot replace the effect estimate and confidence interval because it does not describe the magnitude or precision of the treatment effect.

22. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The primary HIV endpoint was analyzed with a site-stratified Cox proportional-hazards model. The reported HR was 0.94 with a two-sided 95% CI of 0.59–1.52.

Clinical interpretation

Clinical meaning requires consideration of the prevention context, the absolute event experience, treatment exposure, safety, and the prespecified clinical question. The ClinicalTrials.gov record does not provide enough absolute outcome information to construct those additional measures.

Keeping these two levels separate is important. Statistical analysis describes what the observed data support under a specified model. Clinical interpretation asks what those results mean in the context of prevention, patient outcomes, safety, and the intended use of the intervention. The two should not be collapsed into a single numerical conclusion.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Calculators

25. Sources

Continue through the Clinical Biostats statistical library

Explore the underlying survival-analysis, clinical-trial, inference, and continuous-outcome methods used to interpret randomized clinical studies.

26. Record Summary

FEM-PrEP provides a compact but statistically diverse example of randomized clinical-trial analysis. The primary HIV infection endpoint was a time-to-event outcome analyzed with a site-stratified Cox proportional-hazards model, producing an HR of 0.94 with a two-sided 95% CI of 0.59–1.52. Four primary laboratory toxicity endpoints were evaluated with log-rank tests, producing reported p-values of 0.45, 0.59, 0.79, and 0.62. A sixth primary endpoint concerning adverse-event frequency had posted results but no formal statistical comparison in the ClinicalTrials.gov record.

The secondary analyses demonstrate a different statistical structure: two-sided t-tests were used for plasma HIV RNA, CD4+ T-cell count, and participant-reported change in number of sexual partners, with reported mean differences of 0.03, 22.1, and .01, respectively. The corresponding p-values were 0.89, 0.82, and 0.73.

The main statistical lesson is that these numbers cannot be interpreted independently of their endpoint definitions and analysis populations. Hazard ratios describe relative event rates over time; mean differences describe average differences in continuous outcomes; p-values address statistical evidence under a specified test but do not measure effect size; and confidence intervals communicate uncertainty around an estimate. Together, these principles provide a framework for reading the FEM-PrEP results without overstating what the registry data can establish.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. When a registry provides an effect estimate and confidence interval, both should be reported. When it provides only a p-value, the page should not manufacture an effect size or confidence interval. When an endpoint has no formal analysis in the ClinicalTrials.gov record, the absence of that comparison should remain explicit.