This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
SPRING-2 was a randomized, parallel, phase 3 trial comparing dolutegravir 50 mg once daily with raltegravir 400 mg twice daily in participants with Human Immunodeficiency Virus I infection. The registered primary endpoint was the percentage of participants with HIV-1 RNA below 50 copies/mL through Week 48.
| Feature | SPRING-2 |
|---|---|
| Phase | Phase 3 |
| Condition | Infection, Human Immunodeficiency Virus I |
| Design | Randomized, parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 828 |
| Arms | 2 |
| Lead sponsor | ViiV Healthcare |
| Sponsor type | Industry |
| Trial status | Completed |
| Start | October 19, 2010 |
| Primary completion | February 6, 2012 |
| ClinicalTrials.gov | NCT01227824 |
2. Clinical Question
The central question was whether dolutegravir 50 mg once daily could achieve the prespecified Week 48 virologic endpoint at least as well as raltegravir 400 mg twice daily, using a non-inferiority framework.
Population
Participants enrolled in a phase 3 trial for Infection, Human Immunodeficiency Virus I.
Intervention
GSK1349572 (dolutegravir) 50 mg once daily.
Comparator
Raltegravir 400 mg twice daily.
Primary question
Is the difference in the Week 48 virologic response compatible with non-inferiority of dolutegravir relative to raltegravir?
3. Trial Design
GSK1349572 / Dolutegravir
- GSK1349572 (dolutegravir) 50 mg once daily
- Background treatment included ABC/3TC or TDF/FTC
- GSK1349572 placebo was among the registered study interventions
Raltegravir
- Raltegravir 400 mg twice daily
- Background treatment included ABC/3TC or TDF/FTC
- Raltegravir placebo was among the registered study interventions
The registered design identifies the allocation as randomized and the model as parallel. Masking was quadruple, and the primary purpose was treatment. The trial had two arms and enrolled 828 participants.
4. Primary Endpoint
| Endpoint | Registry definition / time frame | Analysis |
|---|---|---|
| Percentage of Participants With HIV-1 RNA <50 Copies/mL Through Week 48 | Percentage of participants with plasma HIV-1 RNA <50 c/mL assessed using the Missing, Switch or Discontinuation = Failure (MSDF), as codified by the FDA snapshot algorithm. The algorithm treats participants without HIV-1 RNA data as non-responders. | Baseline up to Week 48; binary endpoint; ITT-E population |
5. Statistical Methodology
Binary endpoint and risk difference
The primary endpoint is binary: each participant is classified according to whether HIV-1 RNA is below 50 copies/mL through Week 48 under the registered MSDF snapshot algorithm. The reported treatment effect is a difference in percentage, normalized here as a risk difference.
A positive value means that the percentage meeting the virologic endpoint was higher in the dolutegravir group than in the raltegravir group. It is an absolute difference in percentage points, not a relative percentage increase.
Cochran-Mantel-Haenszel analysis
The primary analysis used the Cochran-Mantel-Haenszel (CMH) test. The registry analysis states that the CMH stratified analysis was adjusted for two baseline stratification factors: baseline HIV-1 RNA and background dual NRTI.
For a binary endpoint, stratification allows the treatment comparison to account for the prespecified strata rather than simply pooling all observations into one unadjusted 2 × 2 table. The resulting treatment comparison is therefore conditional on the stratification structure specified for the analysis.
Intention-to-treat efficacy population
The posted primary analysis was based on the ITT-E Population. An intention-to-treat approach maintains randomized treatment assignment as the basis for the efficacy comparison. This is particularly important in a randomized trial because treatment assignment, rather than subsequent treatment behavior, is the variable created by randomization.
Missing, Switch or Discontinuation = Failure
The endpoint used an MSDF approach through the FDA snapshot algorithm. The registry definition explicitly states that participants without HIV-1 RNA data were treated as non-responders. This means the Week 48 percentage is not simply the percentage among participants with an observed Week 48 laboratory result.
Why this matters
Participants who are missing the required endpoint measurement do not disappear from the analysis. Under the registry-reported definition, absence of HIV-1 RNA data leads to classification as a non-responder.
What it does not solve
A snapshot rule does not make missing data irrelevant. The resulting endpoint still depends on the prespecified rules governing missingness, switching, and discontinuation.
6. Primary Result
The posted formal statistical analysis compared dolutegravir 50 mg once daily with raltegravir 400 mg twice daily for the percentage of participants with HIV-1 RNA below 50 copies/mL through Week 48.
Difference in percentage: DTG − RAL
95% CI: −2.2 to 7.1 · Two-sided confidence interval
Analysis population: ITT-E · Cochran-Mantel-Haenszel stratified analysis
| Primary analysis feature | Reported value |
|---|---|
| Endpoint | HIV-1 RNA <50 copies/mL through Week 48 |
| Analysis population | ITT-E Population |
| Groups compared | DTG 50 mg once a day vs RTG 400 mg BID |
| Method | Cochran-Mantel-Haenszel test |
| Effect measure | Difference in percentage / risk difference |
| Estimate | 2.5 |
| 95% CI | −2.2 to 7.1 |
| Confidence interval | Two-sided |
| Non-inferiority margin | −10 percentage points |
| Formal p-value | Not reported in the ClinicalTrials.gov record |
The estimated difference of 2.5 percentage points means that the estimated percentage achieving HIV-1 RNA below 50 copies/mL through Week 48 was 2.5 percentage points higher with dolutegravir than with raltegravir in the reported stratified analysis.
This does not mean that an individual participant had a 2.5% higher probability of response, nor does it establish that every subgroup had the same difference. It is an aggregate treatment-group comparison in the specified ITT-E analysis.
The 95% confidence interval of −2.2 to 7.1 expresses uncertainty around the estimated treatment difference. It includes zero, so the interval is compatible with a small advantage for either treatment as well as a larger positive difference for dolutegravir. The interval is also the key quantity for the prespecified non-inferiority assessment.
No formal p-value is provided in the registry statistical-analysis record. A p-value would address evidence against a specified null hypothesis; it would not measure the size or clinical importance of the treatment difference.
7. Non-Inferiority Logic
The registry analysis explicitly defines the non-inferiority criterion: non-inferiority could be concluded if the lower bound of a two-sided 95% confidence interval for the difference in percentages between DTG and RAL was greater than −10%.
Reported lower bound = −2.2. The reported lower bound is above the specified −10% non-inferiority margin.
Under the registry's stated decision rule, the relevant comparison is not whether the confidence interval excludes zero. A non-inferiority analysis asks whether the data are sufficiently incompatible with a loss of efficacy as large as the prespecified margin.
Here, the lower confidence limit is −2.2 percentage points, whereas the non-inferiority margin is −10 percentage points. Because −2.2 is greater than −10, the confidence interval does not extend to the prespecified non-inferiority boundary.
This is an important distinction from a conventional superiority test. A confidence interval that crosses zero can still support non-inferiority when the entire interval remains above the non-inferiority margin.
The result should not be interpreted as proving that the two treatments are identical. Non-inferiority means that the observed uncertainty is sufficiently bounded to exclude a loss larger than the prespecified clinically relevant margin under the trial's analysis framework.
8. Why the Confidence Interval Is More Informative Than a Single Number
The point estimate of 2.5 gives one estimate of the treatment difference, but the confidence interval supplies the uncertainty needed to interpret that estimate.
Point estimate
The estimate of 2.5 percentage points favors dolutegravir numerically in the reported treatment comparison.
Lower boundary
The lower limit of −2.2 is above the −10 percentage-point non-inferiority margin.
Upper boundary
The upper limit of 7.1 indicates that the data are also compatible with a larger positive difference in favor of dolutegravir.
Zero is different from the NI margin
Zero represents no estimated percentage difference. The −10 boundary represents the prespecified maximum acceptable loss for non-inferiority.
This distinction is fundamental in non-inferiority trials. Testing only whether the confidence interval excludes zero would answer a superiority question, whereas the registered analysis is explicitly framed around the −10% non-inferiority margin.
9. Stratification and Covariate Adjustment
The primary CMH analysis was adjusted for two baseline stratification factors: baseline HIV-1 RNA and background dual NRTI.
| Stratification factor | Role in the primary analysis |
|---|---|
| Baseline HIV-1 RNA | Baseline stratification factor incorporated into the CMH analysis |
| Background dual NRTI | Baseline stratification factor incorporated into the CMH analysis |
Stratification is especially useful when an important baseline variable is expected to be associated with the outcome. Rather than treating all participants as if they came from one homogeneous population, the CMH framework combines information across strata while accounting for the stratification variables.
The exact weighting and test statistic depend on the observed stratum-specific 2 × 2 tables. The ClinicalTrials.gov record identifies the method and adjustment factors but do not provide those underlying tables.
10. Analysis Population and the Meaning of ITT-E
The primary analysis was conducted in the ITT-E Population. The important statistical principle is that participants remain associated with their randomized treatment assignment for the efficacy comparison.
Why ITT matters
Randomization creates the basis for a fair comparison. An ITT analysis preserves that assignment rather than redefining treatment groups according to what happened after randomization.
Why it matters in non-inferiority trials
Non-inferiority trials require particular care because inappropriate handling of discontinuations or protocol deviations can sometimes make groups appear more similar. The prespecified endpoint algorithm and analysis population therefore matter greatly.
The ClinicalTrials.gov record identifies the analysis population and the MSDF snapshot framework, but they do not provide a separate per-protocol analysis or a detailed protocol-deviation analysis. No such analysis is added here.
11. Statistical Methods Explained
Why was a Cochran-Mantel-Haenszel test used?
The primary endpoint is binary, so the treatment groups can be represented through response and non-response counts. Because the analysis was prespecified to account for baseline HIV-1 RNA and background dual NRTI, a CMH analysis provides a way to compare treatment groups while respecting those strata rather than relying solely on an unadjusted comparison.
What does a risk difference of 2.5 mean?
A risk difference of 2.5 means that the estimated percentage meeting the primary virologic endpoint was 2.5 percentage points higher in the DTG group than in the RAL group. It is an absolute treatment difference. It should not be confused with a 2.5-fold increase, a 2.5% relative increase, or a hazard ratio of 2.5.
Why is the non-inferiority margin more important than whether the CI crosses zero?
The trial was framed around whether dolutegravir could be ruled out as being worse than raltegravir by more than 10 percentage points. Therefore, the critical boundary is −10, not zero. The confidence interval can include zero and still remain entirely above the non-inferiority margin.
What does the −2.2 lower confidence limit tell us?
It is the most conservative end of the reported two-sided 95% confidence interval for the treatment difference. Because −2.2 is still above −10, the interval does not include a treatment difference as unfavorable as the prespecified non-inferiority margin.
Why does the MSDF rule matter?
The primary endpoint is not simply a laboratory measurement among participants who happened to have an available Week 48 value. The registry definition states that participants without HIV-1 RNA data were treated as non-responders. The treatment comparison therefore incorporates the prespecified handling of missing endpoint data.
Why adjust for baseline HIV-1 RNA and background dual NRTI?
These were the baseline stratification factors identified in the posted analysis. Incorporating them into the CMH analysis accounts for the randomized trial's stratification structure and produces the reported adjusted treatment comparison.
Why does a non-inferiority result not prove that the treatments are identical?
Non-inferiority is a bounded-loss conclusion. It asks whether the evidence excludes a treatment disadvantage beyond the prespecified margin. The confidence interval here extends from −2.2 to 7.1, so the analysis permits a range of treatment differences. It does not establish exact equality of efficacy.
12. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm. These figures are reported separately from the primary efficacy analysis because safety and efficacy answer different statistical questions.
| Group | Serious adverse events affected | At risk |
|---|---|---|
| DTG 50 mg Once a Day | 41 | 411 |
| RTG 400 mg BID | 45 | 411 |
| DTG 50 mg Once a Day (Open-label) | 21 | 338 |
The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or p-value for these serious adverse-event counts. They therefore should be read descriptively rather than treated as a formal hypothesis test.
13. Other Registered Interventions
ClinicalTrials.gov lists the following interventions in the ClinicalTrials.gov record: GSK1349572 (dolutegravir), raltegravir, GSK1349572 placebo, ABC/3TC, TDF/FTC, and raltegravir placebo.
| Intervention listed in registry data | Type |
|---|---|
| GSK1349572 (dolutegravir) | Drug |
| Raltegravir | Drug |
| GSK1349572 Placebo | Other |
| ABC/3TC | Other |
| TDF/FTC | Other |
| Raltegravir Placebo | Other |
The presence of these interventions in the registry is consistent with the quadruple-masked design and the use of background dual NRTI treatment. The ClinicalTrials.gov record does not provide enough detail to reconstruct a more granular treatment-administration schedule, so no additional regimen details are inferred.
14. Design Features That Shape Interpretation
What is not supported by the ClinicalTrials.gov record?
The ClinicalTrials.gov record does not report a factorial design, crossover analysis, Bayesian method, interim analysis, formal multiplicity procedure, or imputation method beyond the registered MSDF snapshot handling of missing endpoint data. Those topics are therefore not presented as trial methods here.
15. What the Primary Estimate Does — and Does Not — Mean
The reported 2.5 percentage-point risk difference is the estimated absolute difference in the percentage achieving HIV-1 RNA below 50 copies/mL through Week 48, with DTG compared with RAL, using the posted stratified CMH analysis.
It does not mean that exactly 2.5 additional participants out of every 100 would respond in every population. It is a trial-level estimate subject to statistical uncertainty and to the endpoint's prespecified MSDF classification rules.
The 95% CI of −2.2 to 7.1 quantifies uncertainty around the treatment difference. Its lower end is slightly below zero, while its upper end is positive. The interval therefore includes the possibility of a small negative difference as well as a positive difference.
The important comparison is between the lower confidence limit, −2.2, and the prespecified margin, −10. The lower limit remains above the non-inferiority boundary.
The ClinicalTrials.gov record does not report a p-value for this primary analysis. A p-value, if reported elsewhere, would quantify evidence relative to a specified null hypothesis; it would not replace the confidence-interval comparison required by the stated non-inferiority framework.
16. Limitations and Interpretation Issues
- Registry-level result detail: the registry-reported statistical analysis contains one formal primary analysis, with the treatment difference and confidence interval, but does not provide the underlying response counts by treatment arm.
- Missing endpoint data: the registered MSDF snapshot treats participants without HIV-1 RNA data as non-responders, so the reported percentage is influenced by the prespecified missing-data classification.
- Non-inferiority interpretation: the conclusion depends on the prespecified −10 percentage-point margin and the two-sided 95% confidence interval, not on whether the interval excludes zero.
- Stratification: the primary estimate is adjusted for baseline HIV-1 RNA and background dual NRTI. It should not be interpreted as an unadjusted pooled percentage difference.
- Safety denominator differences: serious adverse events are reported with the registry-reported arm-specific affected/at-risk counts and should not be substituted for the ITT-E efficacy population.
- Limited posted analysis set: the ClinicalTrials.gov record contains one formal statistical analysis. No additional primary or secondary formal comparisons are introduced from outside the ClinicalTrials.gov record.
17. Why This Trial Matters Statistically
SPRING-2 is a useful teaching example because its primary analysis illustrates an important distinction in clinical-trial statistics: non-inferiority is a confidence-interval problem built around a clinically specified margin, not simply a search for a statistically significant difference between treatments.
| Concept | How it appears in SPRING-2 |
|---|---|
| Randomization | Participants were randomized to two parallel treatment arms. |
| Quadruple masking | The registry identifies the trial as quadruple-masked. |
| Binary endpoint | The primary endpoint is the percentage with HIV-1 RNA below 50 copies/mL through Week 48. |
| ITT analysis | The posted primary analysis used the ITT-E population. |
| Cochran-Mantel-Haenszel test | The primary comparison used a stratified CMH analysis. |
| Covariate adjustment | The CMH analysis adjusted for baseline HIV-1 RNA and background dual NRTI. |
| Risk difference | The treatment effect was reported as a difference in percentages. |
| Confidence interval | The treatment difference was accompanied by a two-sided 95% CI. |
| Non-inferiority margin | The lower confidence bound was compared with a −10 percentage-point margin. |
| Missing-data handling | The MSDF snapshot algorithm treated participants without HIV-1 RNA data as non-responders. |
18. A Step-by-Step Reading of the Primary Analysis
Define the binary outcome
The endpoint is whether HIV-1 RNA is below 50 copies/mL through Week 48, using the registered MSDF snapshot framework.
Use the ITT-E population
The primary analysis follows randomized treatment assignment in the ITT-E population.
Respect the stratification factors
The CMH analysis adjusts for baseline HIV-1 RNA and background dual NRTI.
Calculate the treatment difference
The reported difference in percentage is 2.5 percentage points for DTG minus RAL.
Construct the two-sided 95% CI
The reported interval extends from −2.2 to 7.1 percentage points.
Compare the lower bound with −10
The lower confidence limit of −2.2 remains above the prespecified −10 percentage-point non-inferiority margin.
19. Primary Analysis in Statistical Notation
The reported estimate is Δ = 2.5 percentage points.
The reported lower bound is −2.2, which lies above the prespecified −10 percentage-point margin.
This notation makes the logic transparent. The point estimate describes the observed treatment contrast, the confidence interval describes uncertainty, and the non-inferiority margin defines the clinically relevant boundary against which that uncertainty is evaluated.
20. Results and What They Do Not Establish
What the result establishes statistically
The posted primary analysis estimates a 2.5 percentage-point difference, with a two-sided 95% CI from −2.2 to 7.1, using the specified ITT-E and stratified CMH framework.
What it does not establish
The result does not show that the treatments have identical efficacy, nor does it establish that every patient or every subgroup experiences the same treatment difference.
What the CI contributes
The interval provides the uncertainty needed to assess the non-inferiority margin and shows that both small negative and positive treatment differences remain compatible with the estimate.
What the ClinicalTrials.gov record does not provide
The ClinicalTrials.gov record does not provide a formal p-value, underlying response counts, or additional formal primary-endpoint analyses.
21. Related Tutorials
Learn more about the methods used in this trial:
22. Related Statistical Calculators
23. Sources
- ClinicalTrials.gov: SPRING-2, NCT01227824.
- Linked publication: PubMed record for PMID 24074642.
Continue with Clinical Biostats statistical methods
Explore the trial-design and analysis concepts that connect directly to the SPRING-2 primary endpoint, including non-inferiority, confidence intervals, risk differences, stratified analysis, and categorical-data methods.
24. Record Summary
SPRING-2 provides a compact but important example of how a non-inferiority clinical trial can be analyzed for a binary virologic endpoint. The primary comparison used the ITT-E population and a Cochran-Mantel-Haenszel stratified analysis adjusted for baseline HIV-1 RNA and background dual NRTI. The reported treatment difference was 2.5 percentage points, with a two-sided 95% confidence interval of −2.2 to 7.1. The key statistical question is therefore not whether the interval excludes zero, but whether its lower bound remains above the prespecified −10 percentage-point non-inferiority margin.