← Clinical Trials
HIV-1 Infection Phase 3 Non-Inferiority NCT01095796

GS-US-236-0102: Complete Statistical Analysis of Stribild in HIV-1 Infection

An independent statistical review of the randomized phase 3 GS-US-236-0102 trial comparing Stribild with Atripla in antiretroviral treatment-naive adults with HIV-1 infection, with emphasis on the Week 48 FDA-defined Snapshot virologic-success endpoint and its non-inferiority framework.

Trial status: Completed  ·  Enrollment: 707  ·  Start: March 2010  ·  Primary completion: August 2011
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results and trial facts are restricted to the ClinicalTrials.gov record for NCT01095796. Where the registry does not provide a value in the ClinicalTrials.gov record, no value is inferred or reconstructed.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

GS-US-236-0102 was a randomized, parallel, quadruple-masked phase 3 trial evaluating the safety and efficacy of Stribild versus Atripla in antiretroviral treatment-naive adults with HIV-1 infection. The registered primary endpoint was binary: the percentage of participants achieving HIV-1 RNA < 50 copies/mL at Week 48 according to the FDA-defined Snapshot analysis.

707
Enrolled
Randomized trial
2
Arms
Parallel design
3.6
Response Difference
Stribild − Atripla
−1.6 to 8.8
95.2% CI
Two-sided
FeatureGS-US-236-0102
Trial nameGS-US-236-0102
ClinicalTrials.gov identifierNCT01095796
Therapeutic areaInfectious Disease
ConditionHIV; HIV Infections
PhasePhase 3
StatusCompleted
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment707
Lead sponsorGilead Sciences
Sponsor typeIndustry

2. Clinical Question

The primary statistical question was whether Stribild could demonstrate non-inferiority to Atripla for Week 48 virologic success, defined by the FDA-defined Snapshot analysis and HIV-1 RNA < 50 copies/mL.

Population

Antiretroviral treatment-naive adults infected with HIV-1, as described by the trial's brief title and registry condition.

Intervention

Stribild, with Stribild placebo also listed among the trial interventions for the masked design.

Comparator

Atripla, with Atripla placebo also listed among the trial interventions for the masked design.

Primary question

Is the Week 48 virologic-success response rate with Stribild less than 12 percentage points worse than the response rate with Atripla?

3. Trial Design

01
Randomize707 participants
02
2 ArmsStribild vs Atripla
03
ParallelRandomized allocation
04
Quadruple MaskMasked trial structure
05
Week 48Snapshot endpoint
Allocation
Randomized allocation to 2 parallel treatment groups.
Masking
Quadruple masking. The registry lists both Stribild placebo and Atripla placebo among the interventions.
Primary purpose
Treatment.
Primary endpoint type
Binary, expressed as the percentage of participants with virologic success.
ARM 1

Stribild

  • Stribild listed as a study intervention.
  • Stribild placebo also listed as an intervention for the masked design.
ARM 2

Atripla

  • Atripla listed as a study intervention.
  • Atripla placebo also listed as an intervention for the masked design.

The trial began in March 2010 and reached primary completion in August 2011. The registry status is Completed.

4. Endpoints

EndpointRegistry definitionTime frameType
Primary endpoint The Percentage of Participants With Virologic Success Using the Food and Drug Administration (FDA)-Defined Snapshot Analysis as Determined by the Achievement of HIV-1 Ribonucleic Acid (RNA) < 50 Copies/mL at Week 48 Week 48 Binary

The ClinicalTrials.gov record identifies 1 primary endpoint and 7 outcome measures posted. Of the posted statistical analyses, 1 is identified as a primary-endpoint analysis and 1 is reported as having an estimate with a confidence interval.

Endpoint interpretation: the primary endpoint is a responder analysis rather than a time-to-event endpoint. Each participant is classified according to the FDA-defined Snapshot framework at Week 48, and the treatment comparison is expressed as a difference in response rates.

5. Statistical Methodology

Primary analysis population

The primary analysis used the Intent-to-treat (ITT) Analysis Set, defined in the registry as participants who were randomized into the study and received at least 1 dose of study drug.

Analyzing participants according to randomized treatment assignment is important because the principal advantage of randomization is preserved in the efficacy comparison. The ITT framework prevents post-randomization treatment behavior from redefining the original treatment groups.

Comparing two response rates

The ClinicalTrials.gov record does not name the statistical test. The effect measure is the difference in response rates, which is typically estimated with a normal-approximation (Wald-type) confidence interval.

Primary effect measure
Difference = Response rateStribild − Response rateAtripla

The reported estimate is 3.6 percentage points. A positive value means the observed response-rate estimate for Stribild was higher than the corresponding estimate for Atripla.

Stratified analysis

The ClinicalTrials.gov record identifies stratified analysis among the other concepts present in the analysis text. The ClinicalTrials.gov record does not provide the specific stratification variables, so none are inferred.

Interim analysis and alpha spending

The ClinicalTrials.gov record also identifies interim analysis / alpha spending as a concept in the analysis text. This indicates that the statistical framework accounted for the possibility of examining accumulating trial information while controlling the relevant type I error. The ClinicalTrials.gov record does not provide the exact alpha-spending function, boundary values, or information fractions, so those details are not reconstructed here.

Non-inferiority framework

The hypothesis type is recorded as non-inferiority or equivalence. The analysis notes specify a non-inferiority framework in which the null hypothesis was that the Stribild response rate at Week 48 was at least 12% worse than the Atripla response rate. The alternative hypothesis was that the Stribild response rate was less than 12% worse than the Atripla response rate.

Prespecified non-inferiority margin

−12%

The null boundary described in the registry is a response-rate difference of −12 percentage points.

The trial was designed so that 700 HIV-1 infected participants randomized in a 1:1 ratio would achieve at least 95% power to establish non-inferiority under the stated Week 48 response-rate framework.

The enrollment listed in the ClinicalTrials.gov record was 707, while the sample-size statement in the posted analysis describes a calculation based on 700 participants. These are different quantities: the former is the registry enrollment value, whereas the latter is the sample-size value explicitly described in the statistical analysis text.

6. Results: Week 48 Virologic Success

The posted primary analysis compares Stribild with Atripla for the percentage of participants achieving HIV-1 RNA < 50 copies/mL at Week 48 under the FDA-defined Snapshot analysis.

Difference in response rates

3.6

95.2% two-sided CI: −1.6 to 8.8

Effect measure: difference in response rates, Stribild vs Atripla

FeaturePosted primary analysis
EndpointWeek 48 virologic success using the FDA-defined Snapshot analysis
Response definitionHIV-1 RNA < 50 copies/mL at Week 48
Analysis populationITT Analysis Set: randomized participants who received at least 1 dose of study drug
Groups comparedStribild vs Atripla
MethodWald / z-test
Effect measureDifference in response rates
Estimate3.6
Confidence interval95.2% two-sided CI, −1.6 to 8.8
Formal p-valueNot reported in the ClinicalTrials.gov record
Hypothesis frameworkNon-inferiority or equivalence
Clinical Biostats interpretation

The estimated difference in response rates was 3.6 percentage points, calculated in the direction Stribild minus Atripla. In descriptive terms, the posted estimate places the Stribild response rate 3.6 percentage points above the Atripla response rate.

That estimate does not mean that every individual participant had a 3.6-percentage-point improvement. It is a group-level difference between two binary response proportions.

The 95.2% two-sided confidence interval of −1.6 to 8.8 describes statistical uncertainty around the estimated response-rate difference under the analysis framework. It does not describe the range of outcomes that an individual participant could experience.

The confidence interval is particularly important in a non-inferiority trial because the relevant question is whether the uncertainty interval is compatible with a treatment difference worse than the prespecified non-inferiority margin. Here, the lower confidence-limit value reported by the registry is −1.6, while the non-inferiority boundary described in the analysis is −12%.

The registry analysis does not provide a formal p-value for this comparison. A p-value would quantify evidence against a specified null hypothesis; it would not itself measure the size of the treatment effect. The estimate and confidence interval are therefore essential to interpreting the magnitude and precision of the comparison.

7. Understanding the Non-Inferiority Result

Non-inferiority trials are fundamentally different from conventional superiority trials. The objective is not necessarily to demonstrate that the experimental treatment produces a larger response. Instead, the objective is to determine whether any loss of efficacy is sufficiently small to remain within a prespecified acceptable margin.

Registry-defined logic
Null: Difference ≤ −12%    |    Alternative: Difference > −12%

The difference is defined in the direction Stribild minus Atripla. A value below −12 percentage points would cross the prespecified non-inferiority boundary described in the registry-reported analysis.

The reported point estimate of 3.6 is above the −12% margin. More importantly for the confidence-interval approach, the lower confidence limit of −1.6 is also above −12%. Thus, based on the registry-reported estimate and confidence interval, the uncertainty interval does not extend to the prespecified non-inferiority boundary.

Non-inferiority is not the same as superiority. An analysis can establish that a treatment is not unacceptably worse without establishing that it is superior. The ClinicalTrials.gov record identifies the hypothesis framework as non-inferiority or equivalence, so the statistical interpretation should remain centered on the prespecified margin rather than treating the positive point estimate of 3.6 as a superiority claim.

Why the margin matters more than a generic p-value

In a non-inferiority trial, a conventional p-value for a null hypothesis of no treatment difference is not sufficient to answer the scientific question. The relevant question is whether the data exclude a treatment disadvantage large enough to cross the non-inferiority margin.

For this trial, the registry-reported analysis defines that disadvantage as 12 percentage points. The confidence interval therefore provides a direct way to examine whether the observed uncertainty is compatible with a difference of −12 percentage points or worse.

8. Sample Size and Power

The statistical analysis states that a total of 700 HIV-1 infected participants, randomized in a 1:1 ratio to 2 groups, would achieve at least 95% power to establish non-inferiority in the Week 48 response-rate difference for HIV-1 RNA < 50 copies/mL under the FDA-defined Snapshot analysis.

Planned sample-size basis

The posted analysis describes a sample-size calculation based on 700 participants randomized 1:1.

Target power

The stated design target was at least 95% power to establish non-inferiority under the specified response-rate framework.

Observed enrollment

the ClinicalTrials.gov record lists enrollment of 707 participants.

What power means

Power is a design property under specified assumptions. It is not the probability that the observed treatment effect is correct or that the null hypothesis is false.

The distinction between planned power and observed results is important. A study's nominal power is calculated before or during trial design under assumptions about response rates, variability, margin, and other design parameters. It should not be interpreted retrospectively as the probability that the trial result is true.

9. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm. These data are presented separately from the primary efficacy endpoint because safety and efficacy answer different questions.

Safety measureStribildAtripla
Serious adverse events69/34850/352
Serious adverse events by arm
Stribild
69/348
Atripla
50/352

The ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse events, confidence intervals, p-values, or a detailed adverse-event classification. Accordingly, no formal comparative safety inference is added here.

Safety interpretation: the denominators posted on ClinicalTrials.gov for serious adverse events are 348 for Stribild and 352 for Atripla. These denominators should not be silently replaced with the overall enrollment of 707 or treated as if they were the primary efficacy ITT denominator.

10. Statistical Methods Explained

Why was a Wald / z-test used?

The primary endpoint is binary: each participant is classified according to whether Week 48 virologic success was achieved under the FDA-defined Snapshot analysis. A comparison of two response rates can therefore be expressed as a difference in proportions. The record does not name the test; a normal-approximation (Wald-type) comparison is the usual approach for this kind of difference.

Conceptually, the procedure evaluates the observed difference relative to its estimated sampling uncertainty. The resulting confidence interval provides a range of plausible values for the underlying response-rate difference under the specified statistical framework.

What does a response-rate difference of 3.6 mean?

The estimate is expressed as Stribild response rate minus Atripla response rate. A value of 3.6 therefore represents an estimated difference of 3.6 percentage points in favor of Stribild on the binary Week 48 response endpoint.

It is not a relative risk, odds ratio, hazard ratio, or individual-level treatment effect. It is an absolute difference between two response proportions.

Why is the confidence interval important in a non-inferiority trial?

The confidence interval allows the reader to examine the range of treatment differences that remain compatible with the observed data. For non-inferiority, the key question is whether the interval reaches the prespecified unacceptable-loss boundary.

Here the registry-reported 95.2% two-sided confidence interval extends from −1.6 to 8.8. The lower limit is above the −12% non-inferiority margin described in the analysis.

Why does a positive estimate not automatically establish superiority?

The point estimate of 3.6 is above zero, but the trial's stated hypothesis framework is non-inferiority or equivalence. A non-inferiority conclusion and a superiority conclusion answer different questions and require different prespecified statistical hypotheses.

Why was an ITT analysis used?

The ITT Analysis Set was defined as participants who were randomized and received at least 1 dose of study drug. Analyzing participants according to randomized assignment helps preserve the benefits of randomization and avoids redefining treatment groups based on later events.

Why does interim analysis and alpha spending matter?

The registry-reported analysis identifies interim analysis and alpha spending as statistical concepts. When trial data can be examined before the final analysis, repeated unadjusted testing can increase the probability of a false-positive finding. An alpha-spending framework is designed to control the relevant overall type I error while permitting prespecified looks at accumulating evidence.

What does stratified analysis add?

Stratification can account for prespecified factors when estimating or testing a treatment effect. The registry-reported analysis identifies stratified analysis but does not provide the actual stratification variables in the ClinicalTrials.gov record. It would therefore be inappropriate to infer those variables.

11. Confidence Intervals, Precision, and the Estimate

Point estimate

The primary analysis estimates a 3.6 percentage-point difference in Week 48 virologic success, in the direction Stribild minus Atripla.

Uncertainty

The reported 95.2% two-sided confidence interval extends from −1.6 to 8.8. The interval contains zero, so the ClinicalTrials.gov record does not support describing the result as a statistically demonstrated positive difference under a conventional zero-difference superiority interpretation.

Non-inferiority interpretation

The more relevant comparison for the registered hypothesis is with the −12% non-inferiority boundary. The registry-reported lower confidence limit of −1.6 is above that boundary.

What the interval does not mean

The confidence interval is not a probability distribution for the treatment effect, and it does not mean that there is a 95.2% probability that the true difference lies between −1.6 and 8.8. It is an interval constructed under the specified frequentist confidence procedure.

12. Non-Inferiority: A Closer Statistical Reading

Non-inferiority designs require unusually careful attention to the direction of the effect measure. Here, the comparison is defined as the Stribild response rate minus the Atripla response rate.

DifferenceInterpretation
PositiveHigher estimated response rate with Stribild.
0No estimated difference in response rates.
Between −12% and 0Stribild has a lower estimated response rate, but the difference remains inside the prespecified non-inferiority margin.
At or below −12%Reaches the non-inferiority boundary described in the posted analysis.

The observed estimate of 3.6 lies above the non-inferiority boundary, and the lower confidence limit of −1.6 also lies above that boundary. This is the central statistical feature of the posted primary analysis.

Do not confuse non-inferiority with equivalence. Equivalence generally requires demonstrating that the treatment difference lies within a prespecified two-sided equivalence interval. The registry classification says "Non-inferiority or equivalence," while the posted analysis notes explicitly describe a one-directional non-inferiority question using the −12% boundary. This page therefore focuses on the documented non-inferiority logic rather than imposing a separate equivalence framework.

13. Analysis Population and Causal Interpretation

The primary analysis population is explicitly defined as participants who were randomized and received at least 1 dose of study drug. This definition combines the randomized allocation principle with a minimum exposure requirement.

Randomization

Randomized allocation creates the structural basis for comparing the two treatment groups without assigning participants according to their later outcomes.

At least 1 dose

The registry definition requires participants to have received at least 1 dose of study drug for inclusion in the stated ITT Analysis Set.

Endpoint classification

The primary outcome is binary, so the analysis compares response proportions rather than event times.

Interpretation boundary

The ClinicalTrials.gov record does not provide enough detail to reconstruct additional analysis populations or sensitivity analyses.

14. Interim Analysis and Alpha Spending

The primary statistical-analysis record identifies interim analysis / alpha spending as an analysis concept. This matters because a clinical trial can accumulate information over time, creating the possibility of examining results before the final analysis.

If interim efficacy analyses are conducted without statistical adjustment, repeated testing can increase the overall probability of incorrectly rejecting a true null hypothesis. Alpha spending provides a framework for distributing the allowable type I error across planned analyses.

General principle
Total type I error = allocated error across prespecified information looks

The ClinicalTrials.gov record identifies alpha spending as part of the analysis framework but do not state the exact spending function, boundary values, timing of each look, or numerical alpha allocation. Those details are therefore not reconstructed.

This distinction is important when interpreting a reported confidence interval or p-value from a group-sequential trial. The inferential procedure must correspond to the prespecified monitoring design rather than treating every interim and final analysis as an independent hypothesis test.

15. Limitations

The registry limitation field states: “There were no limitations affecting the analysis or results.” That statement is retained as a registry-specific caveat. In addition, there are important boundaries on what can be concluded from the registry-reported statistical data.

16. Why This Trial Matters Statistically

GS-US-236-0102 is a useful teaching case because it combines several fundamental ideas in clinical-trial statistics: randomized allocation, masking, a binary efficacy endpoint, ITT analysis, response-rate differences, confidence intervals, a Wald/z-test, stratified analysis, interim monitoring, alpha spending, and non-inferiority methodology.

ConceptHow it appears in GS-US-236-0102
RandomizationParticipants were randomized in a parallel phase 3 trial.
BlindingThe trial used quadruple masking.
Binary endpointWeek 48 virologic success was defined as HIV-1 RNA < 50 copies/mL under the FDA-defined Snapshot analysis.
ITT analysisThe primary analysis used participants randomized and receiving at least 1 dose of study drug.
Response-rate differenceThe effect measure was the difference between Stribild and Atripla response rates.
Wald / z-testThe record does not name the test used for the primary analysis.
Confidence intervalA 95.2% two-sided confidence interval was reported for the response-rate difference.
Non-inferiorityThe analysis used a prespecified 12% non-inferiority margin.
PowerThe analysis states that 700 participants randomized 1:1 would provide at least 95% power under the stated assumptions.
Stratified analysisStratified analysis is identified among the statistical concepts in the posted analysis text.
Interim analysisInterim analysis / alpha spending is identified in the posted analysis text.
SafetySerious adverse events are reported as affected/at-risk counts by treatment arm.

17. What the Primary Estimate Does — and Does Not — Mean

The estimate

The reported difference of 3.6 means that the estimated Week 48 virologic-success response rate was 3.6 percentage points higher for Stribild than for Atripla, using the direction specified by the effect measure.

It is not a relative effect

The value 3.6 is a difference in response rates. It should not be described as a 3.6% relative increase, a 3.6-fold effect, a hazard ratio, or an odds ratio.

The confidence interval

The interval from −1.6 to 8.8 indicates that the data are compatible with response-rate differences ranging from a modest disadvantage for Stribild to a larger advantage, under the stated confidence procedure.

The non-inferiority question

The key design boundary is −12%. Because the registry-reported lower confidence limit is −1.6, the interval remains above the prespecified non-inferiority margin.

The p-value

A formal p-value is not included in the ClinicalTrials.gov record. It therefore should not be reconstructed from the confidence interval or reported as though it were part of the registry data.

18. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The posted primary analysis estimated a 3.6 percentage-point difference in Week 48 virologic-success response rates, with a 95.2% two-sided confidence interval from −1.6 to 8.8. The non-inferiority framework used a −12% boundary.

Clinical interpretation

The endpoint measures whether participants achieved HIV-1 RNA < 50 copies/mL at Week 48 under the FDA-defined Snapshot analysis. The ClinicalTrials.gov record supports interpreting the result through the prespecified non-inferiority framework rather than through a generic superiority comparison.

Keeping these interpretations separate is important. The statistical result describes an estimated difference and its uncertainty. Clinical interpretation considers what that endpoint and margin mean in the context of the trial question.

19. A Worked Reading of the Primary Result

StepQuestionReading of the ClinicalTrials.gov record
1What was measured?Percentage of participants with Week 48 virologic success under the FDA-defined Snapshot analysis.
2How was it summarized?Difference in response rates.
3What was estimated?3.6 percentage points, Stribild minus Atripla.
4How precise was the estimate?95.2% two-sided CI: −1.6 to 8.8.
5What was the non-inferiority boundary?−12 percentage points.
6Does the confidence interval reach the boundary?No. Its lower limit is −1.6, which is above −12.
7Was a p-value reported?No formal p-value is reported in the ClinicalTrials.gov record.

This sequence illustrates why non-inferiority interpretation cannot be reduced to asking whether the estimated difference is positive. The margin and the confidence interval are central to the inferential question.

20. Sources

The linked publications are provided as references associated with the trial record. The numerical trial content on this page is restricted to the ClinicalTrials.gov record rather than reconstructed from those publications.

21. Related Tutorials

Learn more about the methods used in this trial:

22. Related Calculators

Continue through the Clinical Biostats statistical library

Explore clinical-trial methods, statistical tutorials, and calculators that build from the concepts illustrated by GS-US-236-0102.

23. Record Summary

GS-US-236-0102 provides a focused example of non-inferiority methodology for a binary clinical-trial endpoint. The phase 3 trial was randomized, parallel, and quadruple-masked, with 707 participants enrolled and 2 treatment arms. The primary endpoint was Week 48 virologic success under the FDA-defined Snapshot analysis, defined by achievement of HIV-1 RNA < 50 copies/mL.

The posted primary analysis used an ITT Analysis Set consisting of randomized participants who received at least 1 dose of study drug. The effect was expressed as a difference in response rates; the record does not name the test. The estimated difference was 3.6, with a 95.2% two-sided confidence interval of −1.6 to 8.8.

The most important statistical feature is the non-inferiority framework. The registry-reported analysis defines the null hypothesis as the Stribild response rate being at least 12% worse than the Atripla response rate, with the alternative defined as a difference less than 12% worse. The lower confidence limit of −1.6 is above the −12% boundary, which is the central feature for interpreting the posted non-inferiority analysis.

The registry also identifies interim analysis / alpha spending and stratified analysis as concepts in the statistical analysis. The ClinicalTrials.gov record does not provide enough detail to reconstruct the exact monitoring boundaries, alpha-spending function, or stratification variables, so those elements are described without adding undocumented numerical assumptions.

For safety, serious adverse events are reported as 69/348 for Stribild and 50/352 for Atripla. The ClinicalTrials.gov record does not provide a formal statistical comparison of these safety counts, so they are presented descriptively rather than converted into an unreported inferential result.

Clinical Biostats methodology: The purpose of a trial-analysis page is not simply to repeat the registry. It is to explain how the endpoint, estimand, analysis population, effect measure, confidence interval, hypothesis framework, and design features fit together while maintaining a strict boundary between reported results and statistical interpretation.