This page separates reported trial results from statistical interpretation. Trial-specific numerical results and endpoint definitions on this page are restricted to the ClinicalTrials.gov record for HPTN 083.
1. Trial at a Glance
HPTN 083 was a randomized, parallel, quadruple-masked phase 2/3 prevention trial evaluating injectable cabotegravir compared with TDF/FTC in HIV-uninfected men and transgender women who have sex with men. The trial enrolled 4570 participants and had two primary endpoints.
| Feature | HPTN 083 |
|---|---|
| Trial name | HPTN 083 |
| Phase | 2/3 |
| Condition | HIV Infections |
| Population | HIV-uninfected men and transgender women who have sex with men |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Prevention |
| Enrollment | 4570.0 |
| Primary endpoints | 2 |
| Results posted | Yes |
| Statistical analyses posted | 3 |
| Lead sponsor | National Institute of Allergy and Infectious Diseases (NIAID) |
| Sponsor type | NIH |
| ClinicalTrials.gov | NCT02720094 |
2. Clinical Question
The central question was whether injectable cabotegravir, compared with TDF/FTC, differed in the occurrence of documented incident HIV infections in the trial population. The registered hypothesis type for the formal HIV efficacy analyses was superiority.
Population
HIV-uninfected men and transgender women who have sex with men.
Intervention
Cabotegravir-based prevention, including CAB LA and cabotegravir oral tablet components with the corresponding placebo components.
Comparator
TDF/FTC tablets with placebo for the cabotegravir components.
Primary question
Does cabotegravir produce a different hazard of documented incident HIV infection than TDF/FTC, under the prespecified superiority analysis?
3. Trial Design
Cabotegravir-based intervention
- Cabotegravir oral tablet
- Placebo for TDF/FTC tablets
- CAB LA
- Placebo for CAB LA
TDF/FTC comparator
- TDF/FTC tablets
- Placebo for cabotegravir oral tablet
- Placebo for CAB LA
The ClinicalTrials.gov record identifies the trial as randomized, parallel, and quadruple-masked. Those design features are important statistically because randomization establishes the basis for comparing outcomes between treatment assignments, while masking is intended to reduce the influence of treatment knowledge on trial conduct and assessment.
4. Primary Endpoints
| Endpoint | Registry definition / time frame | Type |
|---|---|---|
| Number of Participants With Documented Incident HIV Infections During Steps 1 and 2 | HIV tests at enrollment, weeks 2, 4, 5, every injection visit (every 8 weeks) and every safety visit (2 weeks after each injection visit). Analyzed through week 153 or the date of DSMB decision to unblind all participants, whichever is earliest. | Binary; analyzed as time-to-event |
| Number of Participants Experiencing Grade 2 or Higher Clinical and Laboratory Adverse Events | Treatment emergent AE* measured with onset date through participant's last study visit, or the date of DSMB decision to unblind all participants, whichever is earliest. Assessed at each visit (injections visits q 8 weeks and safety visits q 8 weeks). | Binary |
5. Analysis Populations
| Population | Definition / role |
|---|---|
| mITT Population (Steps 1 and 2) | Participants who were inappropriately enrolled are excluded. Participants who were found by the EAC to be HIV infected at enrollment are excluded. |
| Injection Step 2 Efficacy | The subgroup of the mITT population who received at least one injection and had at least one HIV result assessed after the first injection visit. |
The primary HIV analysis therefore was not simply a comparison of all 4570 enrolled participants. The registry definition identifies an mITT population and explicitly excludes participants who were inappropriately enrolled or found by the EAC to have HIV infection at enrollment. This distinction matters because the analysis population determines which randomized observations contribute to the estimated treatment effect.
6. Statistical Methodology
Time-to-event analysis
Although the registered HIV endpoint is phrased as a number of participants with documented incident HIV infections, the formal statistical analyses treat the endpoint as a time-to-event outcome. Participants are followed over specified HIV-testing visits, and the analysis incorporates the timing of documented infection rather than reducing the entire follow-up period to a simple event proportion.
Cox proportional-hazards model
The registry identifies a Cox proportional-hazards model as the statistical method for the primary HIV analyses. The principal effect measure is a hazard ratio. In a Cox model, the hazard ratio compares the modeled instantaneous event rates between treatment groups over the analyzed follow-up.
A hazard ratio is a relative time-to-event measure. It is not a risk ratio, absolute risk difference, or percentage of participants who experience an event.
Stratified analysis
One of the posted primary analyses used a Cox proportional-hazards model stratified by region. Stratification allows the baseline hazard to vary across the specified strata while estimating the treatment comparison within the Cox modeling framework.
Bias adjustment for the group-sequential design
The other primary analysis reports a bias-adjusted hazard ratio. The registry states that its hazard ratio, confidence interval, and p-value account for the group-sequential trial design and the early stopping time. This is an important distinction from an ordinary unadjusted final Cox estimate: the reported estimate was specifically adjusted for features of the sequential design.
Kaplan-Meier estimation
For a time-to-event endpoint such as incident HIV infection, Kaplan-Meier estimation provides a way to describe the event-time distribution while accounting for right-censoring. The ClinicalTrials.gov record does not provide a numerical Kaplan-Meier curve or time-specific survival estimates. Accordingly, no such estimates are added here.
7. Primary Efficacy Result: Incident HIV Infection
The primary HIV endpoint was analyzed in the mITT Population (Steps 1 and 2), excluding inappropriately enrolled participants and participants found by the EAC to be HIV infected at enrollment. Two Cox-model analyses are posted for the same primary endpoint.
Bias-adjusted analysis
Bias-adjusted hazard ratio for incident HIV infection
95% CI: 0.18–0.62 · P = 0.0005
Cabotegravir vs TDF/FTC · two-sided 95% CI · superiority hypothesis
The registry identifies this as a bias-adjusted hazard ratio and states that the estimate, confidence interval, and p-value account for the group-sequential design and early stopping time.
A hazard ratio of 0.340 means that, under the fitted time-to-event model, the estimated instantaneous hazard of documented incident HIV infection in the cabotegravir group was about 34% of the hazard in the TDF/FTC group. Equivalently, 1 − 0.340 = 0.660, so the estimate corresponds to an approximately 66% lower estimated hazard for cabotegravir relative to TDF/FTC.
It does not mean that 66% of participants were protected, that 66% fewer participants necessarily became infected, or that every individual had a 66% reduction in personal risk. The hazard ratio is a model-based relative measure that incorporates event timing.
The 95% confidence interval of 0.18–0.62 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of individual treatment effects. Because the entire interval is below 1, the interval is consistent with a lower estimated hazard under the cabotegravir comparison.
The P = 0.0005 value addresses the strength of evidence against the relevant null hypothesis under the prespecified analysis. It does not measure the size of the treatment effect, the probability that the treatment works, or the probability that the null hypothesis is true.
Interpretation also requires attention to the group-sequential design and early stopping. The registry specifically states that the reported estimate, confidence interval, and p-value account for those features. The Cox interpretation additionally depends on the proportional-hazards framework underlying the model.
Unadjusted stratified analysis
Hazard ratio stratified by region
95% CI: 0.18–0.61 · P = 0.0005
Cabotegravir vs TDF/FTC · two-sided 95% CI · superiority hypothesis
This analysis uses a Cox proportional-hazards model stratified by region. Unlike the first result, the registry-reported analysis notes describe this as an unadjusted hazard ratio.
The hazard ratio of 0.328 corresponds to an estimated instantaneous hazard approximately 32.8% as large in the cabotegravir group as in the TDF/FTC group under the stratified Cox model. As a simple derived interpretation, this corresponds to an approximately 67.2% lower estimated hazard.
The estimate does not mean that exactly 67.2% fewer participants experienced HIV infection, nor does it establish that every participant had the same reduction in risk. The distinction between hazard and cumulative risk is essential in time-to-event analysis.
The 95% CI of 0.18–0.61 quantifies uncertainty around the model-based estimate. Its width indicates that the precise magnitude of the relative effect is less certain than the direction of the estimate relative to 1.
The P = 0.0005 value is evidence against the null hypothesis within the specified testing framework; it is not an effect-size metric. The two-sided confidence interval and the p-value should therefore be read alongside the hazard ratio rather than substituted for it.
Because this is a Cox analysis, the interpretation concerns the hazard function and is tied to the model's proportional-hazards framework. Stratification by region allows the baseline hazard to differ by region without requiring a single common baseline hazard across regions.
Comparing the two primary estimates
| Primary analysis | Method | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Bias-adjusted | Cox regression | Bias-adjusted HR | 0.340 | 0.18–0.62 | 0.0005 |
| Stratified | Cox regression stratified by region | HR | 0.328 | 0.18–0.61 | 0.0005 |
The two estimates are close in magnitude, and both have the same reported p-value. They should nevertheless be kept conceptually distinct: the first explicitly accounts for the group-sequential design and early stopping, whereas the second is described as an unadjusted hazard ratio from a Cox model stratified by region.
8. Primary Safety Endpoint
The second registered primary endpoint was the Number of Participants Experiencing Grade 2 or Higher Clinical and Laboratory Adverse Events. The ClinicalTrials.gov record confirms that results were posted for this endpoint and describe the analysis as including Grade 2 or higher clinical and laboratory adverse events during Steps 1 and 2.
| Safety information reported | Cabotegravir | TDF/FTC |
|---|---|---|
| Serious adverse events, affected / at risk | 117 / 2281 | 109 / 2285 |
For a binary safety endpoint, a formal comparison would ordinarily specify the analysis population, event definition, treatment contrast, and an appropriate effect measure such as a risk ratio, risk difference, odds ratio, or model-based estimate. The ClinicalTrials.gov record does not provide such a formal comparison for the Grade 2-or-higher endpoint.
9. Serious Adverse Events by Arm
the ClinicalTrials.gov record reports serious adverse events as affected participants divided by participants at risk:
These counts should be read descriptively. They do not replace the registered Grade 2-or-higher primary safety endpoint and, based on the ClinicalTrials.gov record, do not have a reported confidence interval or p-value attached to a formal between-arm comparison.
10. Secondary Endpoint Result: Incident HIV Infection in Step 2
The registry also posts a secondary analysis of documented incident HIV infections in Step 2. The analysis population was the Injection Step 2 Efficacy population: the subgroup of the mITT population who received at least one injection and had at least one HIV result assessed after the first injection visit.
Hazard ratio for incident HIV infection in Step 2
95% CI: 0.10–0.45 · P < 0.001
Cabotegravir vs TDF/FTC · two-sided 95% CI · superiority hypothesis
The hazard ratio of 0.210 means that the estimated instantaneous hazard of documented incident HIV infection in this Step 2 efficacy analysis was about 21.0% as large in the cabotegravir group as in the TDF/FTC group. As a direct derived interpretation, this corresponds to an approximately 79% lower estimated hazard.
This result applies specifically to the Injection Step 2 Efficacy population defined by the registry. It should not automatically be generalized to every enrolled participant, because participants had to receive at least one injection and have at least one HIV result assessed after the specified injection visit.
The 95% CI of 0.10–0.45 indicates uncertainty around the estimated hazard ratio. The interval remains below 1, but it does not imply that the true treatment effect for every individual lies somewhere between 55% and 90% lower risk. Confidence intervals around hazard ratios describe uncertainty in the population-level model estimate, not individual treatment-response ranges.
The P < 0.001 value provides evidence against the null hypothesis under the reported analysis. It does not quantify the magnitude of the effect and should not be interpreted as the probability that the observed treatment effect is due to chance.
11. How to Read the HIV Hazard Ratios
Relative measure
The hazard ratio compares modeled instantaneous event rates. It is not an absolute percentage of participants infected or protected.
Time matters
HIV infection is observed through repeated testing over follow-up, making time-to-event methods appropriate for incorporating when events occur and when participants are censored.
Confidence interval
The confidence interval describes uncertainty around the estimated treatment effect under the specified statistical framework.
P-value
The p-value addresses evidence against a null hypothesis; it does not measure effect size or clinical importance.
12. Statistical Methods Explained
Why was a Cox proportional-hazards model used?
The primary HIV endpoint is fundamentally a time-to-event outcome. Participants are tested repeatedly over follow-up, and the analysis can incorporate the timing of a documented HIV infection as well as participants who remain without a documented event through their available follow-up. The Cox model provides a way to estimate a relative hazard while allowing the baseline hazard to remain unspecified.
What does an HIV hazard ratio of 0.340 mean?
It means the fitted model estimates the cabotegravir hazard to be 0.340 times the TDF/FTC hazard. A useful derived interpretation is an approximately 66% lower estimated hazard. It does not mean that exactly 66% fewer people became infected or that an individual person's risk was reduced by exactly 66%.
Why is the confidence interval important?
The point estimate alone does not communicate its precision. The 95% CI of 0.18–0.62 shows the uncertainty surrounding the bias-adjusted estimate. A confidence interval is especially useful because it shows both the estimated direction of the effect and the range of parameter values compatible with the statistical model and data under the stated confidence procedure.
Why doesn't the p-value measure the size of the effect?
A p-value is calculated relative to a null hypothesis and the sampling distribution implied by the statistical model. It answers an evidence question under that framework. The hazard ratio and its confidence interval provide the information about estimated effect magnitude and precision.
Why was one analysis stratified by region?
The registry-reported analysis notes state that the unadjusted hazard ratio was based on a Cox proportional-hazards model stratified by region. Stratification permits different baseline hazards across regions while retaining a common treatment-effect parameter across the strata.
Why is the bias-adjusted estimate different from the ordinary stratified estimate?
The registry identifies the first estimate as bias-adjusted and states that its hazard ratio, confidence interval, and p-value account for the group-sequential design and early stopping time. The other estimate is described as an unadjusted hazard ratio from a region-stratified Cox model. The difference therefore reflects different analytical treatment of the trial design, rather than necessarily a different clinical endpoint.
Why does the Step 2 analysis use a different population?
The Step 2 analysis is restricted to participants who received at least one injection and had at least one HIV result assessed after the first injection visit. This creates a more specific analysis population than the overall mITT population. The resulting hazard ratio therefore answers a narrower question and should be labeled accordingly.
13. Interim Analysis and Alpha Spending
The ClinicalTrials.gov record explicitly identify interim analysis / alpha spending as a concept in the primary HIV analysis and state that the bias-adjusted hazard ratio, confidence interval, and p-value account for the group-sequential trial design and early stopping time.
Why interim analysis matters
A group-sequential trial permits accumulating data to be evaluated at planned stages rather than waiting exclusively for the end of the trial.
Why alpha spending matters
When efficacy data are examined repeatedly, the statistical design must account for those looks so that the intended type I error framework is preserved.
The important point for interpreting HPTN 083 is that early stopping is not merely a historical detail. It is part of the statistical context of the primary result. The registry explicitly says that the bias-adjusted analysis accounts for both the group-sequential design and the early stopping time.
14. Randomization and Masking
Randomization is central to the treatment comparison because assignment was randomized rather than selected according to participant characteristics. Under appropriate implementation, randomization provides the principal basis for interpreting differences between treatment groups as differences associated with treatment assignment.
The registry identifies the trial as quadruple-masked. Masking does not replace randomization, but it can reduce the influence of treatment knowledge on participant behavior, investigator decisions, outcome assessment, or other trial processes, depending on which parties are masked.
The validity of this chain depends on the actual conduct of the randomized trial, adherence to the protocol, the prespecified analysis, and appropriate handling of follow-up and censoring.
15. Censoring and Time-to-Event Interpretation
Because the HIV endpoint is analyzed as time-to-event, participants do not all need to experience an HIV infection to contribute information. A participant who remains without a documented infection through available follow-up can contribute information until the point at which their event status is no longer observed.
This is one reason a hazard ratio should not be read as a simple ratio of two event percentages. Time-to-event analysis uses both the occurrence and timing of events and the available follow-up information for participants who do not experience the event during observation.
The ClinicalTrials.gov record specifies repeated HIV testing at enrollment, weeks 2, 4, 5, every injection visit every 8 weeks, and every safety visit 2 weeks after each. These repeated assessments provide the framework within which incident infections were identified.
16. Multiplicity and Multiple Analyses
the ClinicalTrials.gov record reports two primary endpoints and three posted statistical analyses. Two of those statistical analyses are primary analyses of the incident HIV endpoint, while the third is a secondary analysis of incident HIV infection in Step 2.
| Analysis | Role | Population | Reported estimate |
|---|---|---|---|
| Incident HIV infection | Primary | mITT Population (Steps 1 and 2) | Bias-adjusted HR 0.340; 95% CI 0.18–0.62; P = 0.0005 |
| Incident HIV infection | Primary | mITT Population (Steps 1 and 2) | HR 0.328; 95% CI 0.18–0.61; P = 0.0005 |
| Incident HIV infection in Step 2 | Secondary | Injection Step 2 Efficacy population | HR 0.210; 95% CI 0.10–0.45; P < 0.001 |
The ClinicalTrials.gov record does not specify a complete multiplicity hierarchy for every posted analysis or the exact alpha allocation among the two primary endpoints. The page therefore distinguishes primary and secondary analyses without assigning an additional multiplicity interpretation that is not documented in the ClinicalTrials.gov record.
17. What the Results Do Not Tell Us
Not an absolute risk difference
A hazard ratio does not tell us the absolute difference in the probability of HIV infection between groups at a particular time point.
Not individual protection
A population-level hazard ratio does not mean that every participant experienced the same proportional change in risk.
Not a p-value interpretation
A very small p-value does not mean that the treatment effect is large, nor does it give the probability that the treatment is effective.
Not a safety conclusion
The registry-reported serious-AE counts do not constitute a formal analysis of the registered Grade 2-or-higher primary safety endpoint.
18. Limitations
- Registry-level information: this analysis is restricted to the ClinicalTrials.gov record. Important protocol or statistical-analysis-plan details not included in those data are not reconstructed.
- Limited safety statistics: the ClinicalTrials.gov record provides serious adverse events by arm but do not provide a formal statistical comparison for the registered Grade 2-or-higher primary safety endpoint.
- Analysis-population differences: the primary efficacy analysis uses the mITT Population (Steps 1 and 2), while the secondary Step 2 efficacy analysis uses the narrower Injection Step 2 Efficacy population.
- Group-sequential design: early stopping is specifically incorporated into the bias-adjusted primary HIV analysis. The exact interim-monitoring parameters are not reported.
- Proportional-hazards framework: interpretation of a Cox hazard ratio relies on the model's time-to-event framework and its proportional-hazards interpretation. The ClinicalTrials.gov record does not provide a formal assessment of that assumption.
- Subgroup information: the ClinicalTrials.gov record does not provide subgroup-specific estimates, forest plots, or interaction tests, so no subgroup treatment-effect conclusions are presented.
- Missing-data and imputation details: the ClinicalTrials.gov record does not specify an imputation method for the primary HIV analysis. No imputation procedure is therefore attributed to the trial.
19. Why This Trial Matters Statistically
HPTN 083 is a useful statistical teaching case because it connects randomized trial design with a time-to-event endpoint, Cox modeling, stratification, group-sequential monitoring, early stopping, and analysis-population definitions.
| Concept | How it appears in HPTN 083 |
|---|---|
| Randomization | The trial uses randomized allocation in a parallel design. |
| Masking | The registry identifies quadruple masking. |
| Time-to-event endpoint | Incident HIV infection is formally analyzed as a time-to-event outcome. |
| Hazard ratio | The primary and secondary HIV analyses report hazard ratios. |
| Cox model | Cox proportional-hazards regression is the reported analysis method. |
| Stratification | One primary analysis uses a Cox model stratified by region. |
| Confidence intervals | Both primary analyses report two-sided 95% confidence intervals. |
| P-values | The primary HIV analyses report P = 0.0005. |
| Interim analysis | The bias-adjusted primary analysis accounts for the group-sequential design and early stopping. |
| Analysis populations | The primary and secondary HIV analyses use specifically defined populations. |
| Safety endpoints | A Grade 2-or-higher clinical and laboratory AE endpoint is registered as a primary endpoint. |
20. Overall Statistical Interpretation
The primary efficacy analyses report hazard ratios below 1 for documented incident HIV infection when comparing cabotegravir with TDF/FTC. The bias-adjusted estimate is 0.340 with a two-sided 95% CI of 0.18–0.62 and P = 0.0005. The region-stratified Cox analysis reports an HR of 0.328 with a two-sided 95% CI of 0.18–0.61 and P = 0.0005.
The two estimates answer essentially the same primary HIV efficacy question but use different statistical treatments. The first explicitly accounts for the group-sequential design and early stopping; the second is described as an unadjusted Cox hazard ratio stratified by region. Keeping those distinctions visible is important when interpreting the registry's statistical results.
The secondary Step 2 analysis reports an HR of 0.210 with a two-sided 95% CI of 0.10–0.45 and P < 0.001. Because this analysis is restricted to the Injection Step 2 Efficacy population, it answers a narrower question than the primary mITT analysis.
The reported HIV efficacy estimates are relative time-to-event measures. Their interpretation should combine the hazard ratio, confidence interval, p-value, analysis population, and trial-sequential design. No single number fully describes the evidence.
21. Related Tutorials
Learn more about the methods used in this trial:
22. Related Calculators
23. Sources
- ClinicalTrials.gov: HPTN 083 — NCT02720094.
- Linked publication record: PubMed PMID 40590452.
- Linked publication record: PubMed PMID 38783534.
- Linked publication record: PubMed PMID 38709006.
- Linked publication record: PubMed PMID 37952550.
- Linked publication record: PubMed PMID 37783219.
Continue through the Clinical Biostats statistical pathway
Use the related tutorials and calculators to explore the survival-analysis and clinical-trial methods illustrated by HPTN 083.
24. Record Summary
HPTN 083 provides a clear example of how a randomized prevention trial can use time-to-event methodology to analyze incident HIV infection. The ClinicalTrials.gov record reports two primary Cox-model analyses, including a bias-adjusted hazard ratio that accounts for the group-sequential design and early stopping, and a region-stratified Cox analysis. Both primary analyses report hazard ratios below 1 with two-sided 95% confidence intervals and P = 0.0005. A secondary Step 2 efficacy analysis reports a hazard ratio of 0.210 with a two-sided 95% CI of 0.10–0.45 and P < 0.001.
The statistical story is more informative than any individual number. The appropriate interpretation combines the randomized design, quadruple masking, time-to-event endpoint, Cox proportional-hazards model, analysis population, confidence interval, p-value, and group-sequential adjustment. The safety endpoint also illustrates why closely related clinical definitions should not be substituted for one another: the registered primary safety endpoint concerns Grade 2 or higher clinical and laboratory adverse events, whereas the registry-reported arm-level figures describe serious adverse events.