This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
GS-US-236-0103 was a randomized, parallel-group, quadruple-masked phase 3 trial evaluating the safety and efficacy of Stribild versus ritonavir-boosted atazanavir plus Truvada in antiretroviral treatment-naive adults infected with HIV-1.
| Feature | GS-US-236-0103 |
|---|---|
| Trial identifier | NCT01106586 |
| Phase | Phase 3 |
| Therapeutic area | Infectious Disease |
| Condition | HIV; HIV Infections |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 708.0 |
| Primary endpoint | Binary |
| Results posted | Yes |
| Statistical analyses posted | 1 |
| Lead sponsor | Gilead Sciences |
| Sponsor type | Industry |
2. Clinical Question
The central statistical question was whether Stribild was non-inferior to ritonavir-boosted atazanavir plus Truvada with respect to the percentage of participants achieving HIV-1 RNA < 50 copies/mL at Week 48 under the FDA-defined Snapshot Analysis.
Population
Antiretroviral treatment-naive adults infected with HIV-1, as described in the trial's brief title and registry record.
Intervention
Stribild.
Comparator
Ritonavir-boosted atazanavir plus Truvada, referred to in the statistical analysis as the ATV/r + FTC/TDF group.
Primary question
Is the Week 48 virologic-success response rate with Stribild no more than 12 percentage points worse than that with ATV/r + FTC/TDF?
3. Trial Design
The registry describes a randomized, parallel-group, quadruple-masked phase 3 treatment trial. The registry-reported analysis specifies an intention-to-treat analysis set consisting of participants who were randomized and received at least one dose of study drug.
Stribild
- Stribild was the randomized study treatment.
- Participants were evaluated for virologic success using the FDA-defined Snapshot Analysis.
- The primary efficacy analysis used the ITT analysis set.
ATV/r + FTC/TDF
- Ritonavir-boosted atazanavir plus Truvada was the comparator regimen.
- The primary comparison used the Week 48 virologic-success response rate.
- The reported treatment effect was expressed as a difference in response rates.
4. Endpoints
The trial has one registered primary endpoint. It is a binary endpoint evaluated at a fixed Week 48 time point rather than a time-to-event endpoint.
| Endpoint | Registry definition | Time frame | Type |
|---|---|---|---|
| Primary | The Percentage of Participants With Virologic Success Using the Food and Drug Administration (FDA)-Defined Snapshot Analysis as Determined by the Achievement of HIV-1 Ribonucleic Acid (RNA) < 50 Copies/mL at Week 48 | Week 48 | Binary |
The endpoint's binary structure is important. Each participant contributes to a treatment-group response rate under the prespecified Snapshot Analysis definition, and the primary treatment comparison is the difference between those response rates.
5. Statistical Methodology
Intention-to-treat analysis
The posted primary analysis used an ITT analysis set defined as participants who were randomized into the study and received at least one dose of study drug. Participants were therefore retained in the analysis according to their randomized treatment group rather than being reassigned according to later treatment behavior.
Difference in response rates
The reported effect measure was the difference in response rates. For this analysis, the estimated difference was 3.0 percentage points for Stribild versus ATV/r + FTC/TDF.
A positive estimate indicates a higher observed response rate in the Stribild group under this definition. The estimate of 3.0 therefore corresponds to a 3.0-percentage-point difference in the direction of Stribild.
Wald / z-test
The ClinicalTrials.gov record does not name the statistical test. For a difference in binary response rates, a normal-approximation (Wald-type) approach uses an estimated difference and its standard error to construct a confidence interval and conduct the corresponding hypothesis test.
Stratified analysis
The registry-reported analysis text identifies stratified analysis as a concept used in the statistical analysis. The ClinicalTrials.gov record does not identify the individual stratification variables, so no additional stratification factors are specified here.
Interim analysis and alpha spending
The registry-reported analysis text identifies interim analysis / alpha spending as a statistical concept associated with the analysis. Alpha spending is used in group-sequential settings to control the overall type I error when accumulating data are examined at more than one planned information time. The ClinicalTrials.gov record does not provide the specific spending function or interim boundary, so those details are not reconstructed here.
Non-inferiority framework
The primary hypothesis was explicitly non-inferiority. The null hypothesis was that the Stribild group was at least 12% worse than the ATV/r + Truvada group with respect to the Week 48 response rate. The alternative hypothesis was that the Stribild response rate was less than 12% worse.
For a response-rate difference defined as Stribild minus comparator, the non-inferiority margin is therefore −12 percentage points. The reported two-sided 95.2% confidence interval has a lower bound of −1.9, which is above −12.
6. Sample Size, Power, and the Non-Inferiority Margin
The registry-reported statistical analysis states that a total of 700 HIV-1 infected participants, randomized in a 1:1 ratio to two groups, would achieve at least 95% power to establish non-inferiority in the Week 48 response-rate difference under the FDA-defined Snapshot Analysis.
The enrolled sample size was 708.0, compared with the 700 participants described in the registry-reported power statement. The difference between those figures is a direct descriptive comparison of the reported enrollment and the sample size used in the stated power calculation; it does not by itself establish anything about achieved power.
7. Primary Results
The registry contains one posted formal statistical analysis for the primary endpoint. It compares Stribild with ATV/r + FTC/TDF using the Week 48 FDA-defined Snapshot Analysis endpoint.
Week 48 Virologic Success
Difference in response rates
95.2% two-sided CI: −1.9 to 7.8
Effect measure: difference in response rates · Analysis population: ITT
| Primary endpoint | Estimate | Confidence interval | Hypothesis framework |
|---|---|---|---|
| Week 48 virologic success | 3.0 | 95.2% CI −1.9 to 7.8 | Non-inferiority; margin 12% |
The ClinicalTrials.gov record does not provide a formal P-value for this posted analysis. Accordingly, no P-value is reported here. The statistical evidence that can be evaluated directly from the ClinicalTrials.gov record is the estimated response-rate difference and its confidence interval relative to the prespecified non-inferiority margin.
The estimated difference of 3.0 means that the observed response-rate difference favored Stribild by 3.0 percentage points under the registered endpoint definition. It does not mean that every patient had a 3.0-percentage-point individual improvement, and it does not describe the magnitude of response in either group separately.
The 95.2% two-sided confidence interval of −1.9 to 7.8 describes uncertainty around the estimated difference. Importantly for a non-inferiority question, the lower confidence bound of −1.9 is above the prespecified −12 percentage-point boundary. On the stated margin-based framework, the reported interval is therefore consistent with non-inferiority.
The confidence interval should not be confused with a range containing the effects of individual patients. It quantifies uncertainty around the treatment-group difference under the analysis framework.
A P-value measures evidence against a specified null hypothesis; it does not measure the size or clinical importance of the treatment effect. Because the registry analysis does not report the P-value, it would be inappropriate to manufacture one from the reported numbers.
The interpretation also depends on the analysis population, the Snapshot Analysis definition, the stated non-inferiority margin, and the statistical procedure used. The binary endpoint does not require the proportional-hazards assumption that would be relevant to a Cox hazard-ratio analysis.
8. Understanding the Non-Inferiority Result
The key statistical feature of this trial is that the primary question was not simply whether Stribild was superior. The trial was designed to determine whether its response rate was sufficiently close to that of the comparator that a clinically important loss of efficacy, defined by the 12% margin, could be excluded.
The graphical display is conceptual rather than a scale-accurate reconstruction of the statistical interval. The essential comparison is numerical: the lower confidence bound of −1.9 remains above the non-inferiority boundary of −12.
What the result supports
The reported confidence interval does not extend to the prespecified threshold representing a 12-percentage-point disadvantage for Stribild.
What the result does not show
It does not establish that Stribild is superior on the primary endpoint, because the ClinicalTrials.gov record is presented within a non-inferiority framework rather than a superiority framework.
9. Statistical Methods Explained
Why was a difference in response rates used?
The primary endpoint is binary: participants either meet the FDA-defined Week 48 virologic-success definition or they do not. A difference in response rates directly compares the proportion meeting that endpoint between the two randomized groups. It is therefore an interpretable absolute measure for a binary endpoint.
What does an estimate of 3.0 mean?
The reported estimate of 3.0 is the response-rate difference for Stribild versus ATV/r + FTC/TDF. With the treatment difference defined in that direction, a positive value indicates a higher response rate in the Stribild group under the registered endpoint definition.
Why is the non-inferiority margin more important than simple similarity?
Two treatments can have similar observed response rates by chance even when the uncertainty around their difference is substantial. Non-inferiority therefore requires a prespecified tolerance for the amount of possible disadvantage that remains acceptable. Here that boundary was 12 percentage points. The confidence interval is interpreted relative to that boundary rather than merely by asking whether the point estimate is close to zero.
Why does the confidence interval extend below zero but still support non-inferiority?
The interval from −1.9 to 7.8 allows for the possibility that Stribild's response rate was somewhat lower than the comparator's. However, the lower limit does not approach the prespecified non-inferiority boundary of −12. The distinction is fundamental: a result can be compatible with a small disadvantage while still excluding the larger disadvantage that the non-inferiority margin was designed to rule out.
What is a Wald / z-test doing here?
A Wald-type analysis uses an estimated treatment difference and its estimated uncertainty to form a standardized test statistic. For a binary endpoint, the resulting framework can be used to construct a confidence interval for the response-rate difference and to evaluate a prespecified hypothesis. The registry identifies the method as a Wald / z-test but does not provide the complete computational formula or all variance details.
Why does ITT analysis matter in a randomized trial?
The registry-reported primary analysis used participants who were randomized and received at least one dose of study drug. An ITT-oriented framework preserves the randomized treatment comparison rather than redefining treatment groups based on subsequent behavior. This is particularly important when the purpose of the analysis is to estimate the effect associated with assignment to the randomized strategy.
What is the role of stratified analysis?
The statistical analysis text identifies stratified analysis as one of the concepts used in the analysis. Stratification can account for prespecified design factors when comparing treatment groups, potentially improving the precision or validity of the treatment comparison. Because the ClinicalTrials.gov record does not specify the individual strata, this page does not assign particular variables to them.
10. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm. These figures are presented as affected participants divided by participants at risk.
| Safety measure | Stribild | ATV/r + FTC/TDF |
|---|---|---|
| Serious adverse events | 58/353 | 62/355 |
The serious-adverse-event counts should be interpreted separately from the primary virologic efficacy endpoint. They describe a safety outcome and do not constitute a statistical composite of efficacy and safety.
11. Interim Analysis and Alpha Spending
The registry-reported statistical analysis identifies interim analysis / alpha spending among the concepts in the analysis text. This is particularly relevant when a trial permits formal examination of accumulating data before the final analysis.
Why alpha spending exists
If the same hypothesis is repeatedly tested as data accumulate, the probability of a false-positive conclusion can exceed the nominal type I error rate. Alpha-spending methods allocate the allowable error across planned looks at the data.
What cannot be reconstructed here
The ClinicalTrials.gov record does not identify the exact alpha-spending function, interim boundary, information fraction, or decision threshold. Those details are therefore not inferred.
This distinction matters because an interim analysis is not simply an earlier version of the final analysis. The statistical threshold used at each look can depend on the prespecified group-sequential design.
12. Stratification and Randomization
Randomization is a central source of protection against systematic differences between treatment groups. In this trial, the registry describes the allocation as randomized and the design model as parallel.
The analysis text also identifies stratified analysis. Stratification is most useful when its variables are defined before the analysis and incorporated consistently into the statistical comparison. Because the ClinicalTrials.gov record does not list those variables, no unsupported reconstruction of the randomization strata is made here.
13. Non-Inferiority Versus Superiority
The distinction between non-inferiority and superiority is essential for interpreting this result.
| Question | Non-inferiority interpretation |
|---|---|
| Primary concern | Is the new treatment no more than the prespecified acceptable amount worse? |
| Margin in this trial | 12% worse in the Week 48 response rate |
| Reported point estimate | 3.0 |
| Reported two-sided confidence interval | −1.9 to 7.8 |
| Key comparison | Lower confidence bound versus the −12 percentage-point boundary |
Because the lower confidence bound is −1.9 rather than −12 or below, the interval excludes the degree of disadvantage specified by the non-inferiority margin. That is the central statistical logic of the reported result.
14. Clinical Biostats Interpretation of the Primary Analysis
The reported response-rate difference was 3.0 in favor of Stribild when expressed as Stribild minus ATV/r + FTC/TDF. This is an absolute difference in the probability of meeting the Week 48 virologic-success definition, not a relative risk and not an odds ratio.
The 95.2% two-sided confidence interval of −1.9 to 7.8 provides the principal measure of precision posted on ClinicalTrials.gov for the formal primary analysis. Its width indicates that the point estimate should not be interpreted as an exact treatment effect.
The lower bound of −1.9 is substantially above the prespecified −12 percentage-point non-inferiority boundary. Therefore, the reported interval excludes a disadvantage as large as the trial's stated margin.
The ClinicalTrials.gov record does not contain a P-value for the primary analysis. A P-value should not be reverse-engineered from the confidence interval because the requested data rules require reported values to be used as given.
The primary analysis does not, from the ClinicalTrials.gov record alone, establish the size of either group's absolute response rate, the duration of virologic success, long-term outcomes, or treatment effects in subgroups. Those quantities are not reported in the ClinicalTrials.gov record.
15. Limitations
- Limited posted statistical detail: the ClinicalTrials.gov record contains one formal statistical analysis for the primary endpoint. The detailed computational specification of the Wald / z-test is not provided.
- Missing P-value: the registry-reported analysis reports an estimate and confidence interval but does not provide a formal P-value.
- Stratification details: stratified analysis is identified, but the individual stratification variables are not included in the ClinicalTrials.gov record.
- Interim-analysis details: interim analysis / alpha spending is identified as a concept, but the exact alpha-spending function and interim boundaries are not provided.
- Endpoint scope: the formal result concerns one Week 48 binary virologic-success endpoint. It should not be generalized to outcomes that are not included in the ClinicalTrials.gov record.
- Analysis population: the posted primary analysis uses an ITT analysis set defined as randomized participants who received at least one dose of study drug. That definition should not be silently replaced with a different efficacy population.
- Safety information: serious adverse events are reported by arm, but additional safety endpoints and exposure details are not provided in the ClinicalTrials.gov record.
- Non-inferiority interpretation: the conclusion depends on the prespecified 12% margin and the validity of the underlying endpoint and analysis assumptions. Non-inferiority does not establish equality or superiority.
16. Why This Trial Matters Statistically
GS-US-236-0103 is a useful teaching example because it illustrates how a randomized phase 3 trial can use a binary endpoint and a difference in response rates to answer a non-inferiority question rather than a conventional superiority question.
| Concept | How it appears in GS-US-236-0103 |
|---|---|
| Randomization | The trial uses randomized parallel-group allocation. |
| Blinding | The registry describes the study as quadruple-masked. |
| Binary endpoint | Week 48 virologic success is defined as achievement of HIV-1 RNA < 50 copies/mL using the FDA-defined Snapshot Analysis. |
| ITT analysis | The primary analysis uses participants who were randomized and received at least one dose of study drug. |
| Difference in response rates | The treatment effect is reported as a response-rate difference. |
| Wald / z-test | The record does not name the test used for the formal primary analysis. |
| Confidence interval | A 95.2% two-sided confidence interval accompanies the treatment-effect estimate. |
| Non-inferiority | The analysis uses a 12% margin representing the maximum allowed disadvantage in response rate. |
| Interim analysis | Interim analysis / alpha spending is identified in the analysis text. |
| Stratified analysis | Stratified analysis is identified among the statistical concepts in the posted analysis. |
| Safety comparison | Serious adverse events are reported as affected participants divided by participants at risk for each arm. |
17. Statistical Methods Explained: A Worked Non-Inferiority Framework
Step 1: Define the endpoint
The endpoint is the percentage of participants achieving HIV-1 RNA < 50 copies/mL at Week 48 under the FDA-defined Snapshot Analysis. Because the endpoint classifies participants into response versus non-response, it is binary.
Step 2: Define the treatment effect
The analysis uses a difference in response rates, with the registry-reported estimate of 3.0. In directional form, the effect is represented as:
The reported estimate is Δ = 3.0.
Step 3: Define the non-inferiority margin
The trial's analysis defines a 12% margin. With the treatment difference expressed as Stribild minus comparator, the corresponding lower boundary is −12 percentage points.
Step 4: Estimate uncertainty
The posted analysis supplies a 95.2% two-sided confidence interval from −1.9 to 7.8. This interval describes uncertainty around the estimated treatment difference.
Step 5: Compare the interval with the margin
The lower confidence bound is −1.9, which is above −12. Thus, the reported interval does not include a treatment disadvantage as large as the prespecified non-inferiority margin.
Step 6: Avoid overinterpreting the result
The analysis supports a margin-based non-inferiority interpretation. It does not automatically establish superiority, exact equality of response rates, long-term clinical equivalence, or equivalent safety across all possible outcomes.
18. Related Tutorials
Learn more about the methods used in this trial:
19. Related Calculators
20. Sources
- ClinicalTrials.gov: NCT01106586.
- PubMed: PMID 22748590.
- PubMed: PMID 23337366.
- PubMed: PMID 24346640.
Continue through the Clinical Biostats statistical library
Use the related tutorials and calculators to explore non-inferiority design, confidence intervals, randomized-trial analysis, and Wald / z-test methodology.
21. Record Summary
GS-US-236-0103 provides a clear example of a randomized phase 3 non-inferiority analysis for a binary clinical-trial endpoint. The primary endpoint was Week 48 virologic success under the FDA-defined Snapshot Analysis, and the formal analysis used an ITT analysis set, a difference in response rates, a Wald / z-test framework, and a prespecified 12% non-inferiority margin. The reported estimate was 3.0 with a 95.2% two-sided confidence interval of −1.9 to 7.8.
The central statistical interpretation is driven by the relationship between the confidence interval and the non-inferiority margin. Because the lower bound of −1.9 is above the −12 percentage-point boundary, the reported interval is consistent with the stated non-inferiority criterion. The result should nevertheless be interpreted within the specific endpoint definition, ITT population, masking, stratified-analysis framework, and interim-analysis considerations described in the ClinicalTrials.gov record.