← Clinical Trials
HIV-1 Infection Phase 3 Non-Inferiority NCT01227824

SPRING-2: Complete Statistical Analysis of Dolutegravir in HIV-1 Infection

An independent statistical analysis of the randomized phase 3 SPRING-2 trial comparing GSK1349572 (dolutegravir) 50 mg once daily with raltegravir 400 mg twice daily, with the primary virologic endpoint analyzed using a stratified Cochran-Mantel-Haenszel method.

Trial status: Completed  ·  Enrollment: 828  ·  Primary completion: February 6, 2012
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

SPRING-2 was a randomized, parallel, phase 3 trial comparing dolutegravir 50 mg once daily with raltegravir 400 mg twice daily in participants with Human Immunodeficiency Virus I infection. The registered primary endpoint was the percentage of participants with HIV-1 RNA below 50 copies/mL through Week 48.

828
Enrollment
2 treatment arms
2
Arms
Parallel design
2.5
Risk difference
DTG − RAL percentage points
95%
Confidence interval
−2.2 to 7.1
FeatureSPRING-2
PhasePhase 3
ConditionInfection, Human Immunodeficiency Virus I
DesignRandomized, parallel
MaskingQuadruple
Primary purposeTreatment
Enrollment828
Arms2
Lead sponsorViiV Healthcare
Sponsor typeIndustry
Trial statusCompleted
StartOctober 19, 2010
Primary completionFebruary 6, 2012
ClinicalTrials.govNCT01227824

2. Clinical Question

The central question was whether dolutegravir 50 mg once daily could achieve the prespecified Week 48 virologic endpoint at least as well as raltegravir 400 mg twice daily, using a non-inferiority framework.

Population

Participants enrolled in a phase 3 trial for Infection, Human Immunodeficiency Virus I.

Intervention

GSK1349572 (dolutegravir) 50 mg once daily.

Comparator

Raltegravir 400 mg twice daily.

Primary question

Is the difference in the Week 48 virologic response compatible with non-inferiority of dolutegravir relative to raltegravir?

3. Trial Design

01
Randomize828 participants
02
2 armsParallel allocation
03
Quadruple maskStudy treatment and placebo components
04
Week 48Primary virologic endpoint
05
CompareStratified CMH analysis
ARM A

GSK1349572 / Dolutegravir

  • GSK1349572 (dolutegravir) 50 mg once daily
  • Background treatment included ABC/3TC or TDF/FTC
  • GSK1349572 placebo was among the registered study interventions
ARM B

Raltegravir

  • Raltegravir 400 mg twice daily
  • Background treatment included ABC/3TC or TDF/FTC
  • Raltegravir placebo was among the registered study interventions

The registered design identifies the allocation as randomized and the model as parallel. Masking was quadruple, and the primary purpose was treatment. The trial had two arms and enrolled 828 participants.

Statistical design feature: the primary analysis was not an unadjusted comparison of two percentages. The registry analysis used a Cochran-Mantel-Haenszel test adjusted for baseline HIV-1 RNA and background dual NRTI, making the analysis explicitly stratified and covariate-adjusted.

4. Primary Endpoint

EndpointRegistry definition / time frameAnalysis
Percentage of Participants With HIV-1 RNA <50 Copies/mL Through Week 48 Percentage of participants with plasma HIV-1 RNA <50 c/mL assessed using the Missing, Switch or Discontinuation = Failure (MSDF), as codified by the FDA snapshot algorithm. The algorithm treats participants without HIV-1 RNA data as non-responders. Baseline up to Week 48; binary endpoint; ITT-E population

5. Statistical Methodology

Binary endpoint and risk difference

The primary endpoint is binary: each participant is classified according to whether HIV-1 RNA is below 50 copies/mL through Week 48 under the registered MSDF snapshot algorithm. The reported treatment effect is a difference in percentage, normalized here as a risk difference.

Reported effect measure
Risk difference = Percentage responding with DTG − Percentage responding with RAL

A positive value means that the percentage meeting the virologic endpoint was higher in the dolutegravir group than in the raltegravir group. It is an absolute difference in percentage points, not a relative percentage increase.

Cochran-Mantel-Haenszel analysis

The primary analysis used the Cochran-Mantel-Haenszel (CMH) test. The registry analysis states that the CMH stratified analysis was adjusted for two baseline stratification factors: baseline HIV-1 RNA and background dual NRTI.

For a binary endpoint, stratification allows the treatment comparison to account for the prespecified strata rather than simply pooling all observations into one unadjusted 2 × 2 table. The resulting treatment comparison is therefore conditional on the stratification structure specified for the analysis.

Intention-to-treat efficacy population

The posted primary analysis was based on the ITT-E Population. An intention-to-treat approach maintains randomized treatment assignment as the basis for the efficacy comparison. This is particularly important in a randomized trial because treatment assignment, rather than subsequent treatment behavior, is the variable created by randomization.

Missing, Switch or Discontinuation = Failure

The endpoint used an MSDF approach through the FDA snapshot algorithm. The registry definition explicitly states that participants without HIV-1 RNA data were treated as non-responders. This means the Week 48 percentage is not simply the percentage among participants with an observed Week 48 laboratory result.

Why this matters

Participants who are missing the required endpoint measurement do not disappear from the analysis. Under the registry-reported definition, absence of HIV-1 RNA data leads to classification as a non-responder.

What it does not solve

A snapshot rule does not make missing data irrelevant. The resulting endpoint still depends on the prespecified rules governing missingness, switching, and discontinuation.

6. Primary Result

The posted formal statistical analysis compared dolutegravir 50 mg once daily with raltegravir 400 mg twice daily for the percentage of participants with HIV-1 RNA below 50 copies/mL through Week 48.

Difference in percentage: DTG − RAL

2.5 percentage points

95% CI: −2.2 to 7.1   ·   Two-sided confidence interval

Analysis population: ITT-E  ·  Cochran-Mantel-Haenszel stratified analysis

Primary analysis featureReported value
EndpointHIV-1 RNA <50 copies/mL through Week 48
Analysis populationITT-E Population
Groups comparedDTG 50 mg once a day vs RTG 400 mg BID
MethodCochran-Mantel-Haenszel test
Effect measureDifference in percentage / risk difference
Estimate2.5
95% CI−2.2 to 7.1
Confidence intervalTwo-sided
Non-inferiority margin−10 percentage points
Formal p-valueNot reported in the ClinicalTrials.gov record
Clinical Biostats interpretation

The estimated difference of 2.5 percentage points means that the estimated percentage achieving HIV-1 RNA below 50 copies/mL through Week 48 was 2.5 percentage points higher with dolutegravir than with raltegravir in the reported stratified analysis.

This does not mean that an individual participant had a 2.5% higher probability of response, nor does it establish that every subgroup had the same difference. It is an aggregate treatment-group comparison in the specified ITT-E analysis.

The 95% confidence interval of −2.2 to 7.1 expresses uncertainty around the estimated treatment difference. It includes zero, so the interval is compatible with a small advantage for either treatment as well as a larger positive difference for dolutegravir. The interval is also the key quantity for the prespecified non-inferiority assessment.

No formal p-value is provided in the registry statistical-analysis record. A p-value would address evidence against a specified null hypothesis; it would not measure the size or clinical importance of the treatment difference.

7. Non-Inferiority Logic

The registry analysis explicitly defines the non-inferiority criterion: non-inferiority could be concluded if the lower bound of a two-sided 95% confidence interval for the difference in percentages between DTG and RAL was greater than −10%.

Prespecified decision rule
Lower 95% CI bound > −10%

Reported lower bound = −2.2. The reported lower bound is above the specified −10% non-inferiority margin.

Clinical Biostats interpretation

Under the registry's stated decision rule, the relevant comparison is not whether the confidence interval excludes zero. A non-inferiority analysis asks whether the data are sufficiently incompatible with a loss of efficacy as large as the prespecified margin.

Here, the lower confidence limit is −2.2 percentage points, whereas the non-inferiority margin is −10 percentage points. Because −2.2 is greater than −10, the confidence interval does not extend to the prespecified non-inferiority boundary.

This is an important distinction from a conventional superiority test. A confidence interval that crosses zero can still support non-inferiority when the entire interval remains above the non-inferiority margin.

The result should not be interpreted as proving that the two treatments are identical. Non-inferiority means that the observed uncertainty is sufficiently bounded to exclude a loss larger than the prespecified clinically relevant margin under the trial's analysis framework.

8. Why the Confidence Interval Is More Informative Than a Single Number

The point estimate of 2.5 gives one estimate of the treatment difference, but the confidence interval supplies the uncertainty needed to interpret that estimate.

Point estimate

The estimate of 2.5 percentage points favors dolutegravir numerically in the reported treatment comparison.

Lower boundary

The lower limit of −2.2 is above the −10 percentage-point non-inferiority margin.

Upper boundary

The upper limit of 7.1 indicates that the data are also compatible with a larger positive difference in favor of dolutegravir.

Zero is different from the NI margin

Zero represents no estimated percentage difference. The −10 boundary represents the prespecified maximum acceptable loss for non-inferiority.

This distinction is fundamental in non-inferiority trials. Testing only whether the confidence interval excludes zero would answer a superiority question, whereas the registered analysis is explicitly framed around the −10% non-inferiority margin.

9. Stratification and Covariate Adjustment

The primary CMH analysis was adjusted for two baseline stratification factors: baseline HIV-1 RNA and background dual NRTI.

Stratification factorRole in the primary analysis
Baseline HIV-1 RNABaseline stratification factor incorporated into the CMH analysis
Background dual NRTIBaseline stratification factor incorporated into the CMH analysis

Stratification is especially useful when an important baseline variable is expected to be associated with the outcome. Rather than treating all participants as if they came from one homogeneous population, the CMH framework combines information across strata while accounting for the stratification variables.

Conceptual CMH structure
Adjusted treatment comparison = weighted combination of treatment contrasts across prespecified strata

The exact weighting and test statistic depend on the observed stratum-specific 2 × 2 tables. The ClinicalTrials.gov record identifies the method and adjustment factors but do not provide those underlying tables.

10. Analysis Population and the Meaning of ITT-E

The primary analysis was conducted in the ITT-E Population. The important statistical principle is that participants remain associated with their randomized treatment assignment for the efficacy comparison.

Why ITT matters

Randomization creates the basis for a fair comparison. An ITT analysis preserves that assignment rather than redefining treatment groups according to what happened after randomization.

Why it matters in non-inferiority trials

Non-inferiority trials require particular care because inappropriate handling of discontinuations or protocol deviations can sometimes make groups appear more similar. The prespecified endpoint algorithm and analysis population therefore matter greatly.

The ClinicalTrials.gov record identifies the analysis population and the MSDF snapshot framework, but they do not provide a separate per-protocol analysis or a detailed protocol-deviation analysis. No such analysis is added here.

11. Statistical Methods Explained

Why was a Cochran-Mantel-Haenszel test used?

The primary endpoint is binary, so the treatment groups can be represented through response and non-response counts. Because the analysis was prespecified to account for baseline HIV-1 RNA and background dual NRTI, a CMH analysis provides a way to compare treatment groups while respecting those strata rather than relying solely on an unadjusted comparison.

What does a risk difference of 2.5 mean?

A risk difference of 2.5 means that the estimated percentage meeting the primary virologic endpoint was 2.5 percentage points higher in the DTG group than in the RAL group. It is an absolute treatment difference. It should not be confused with a 2.5-fold increase, a 2.5% relative increase, or a hazard ratio of 2.5.

Why is the non-inferiority margin more important than whether the CI crosses zero?

The trial was framed around whether dolutegravir could be ruled out as being worse than raltegravir by more than 10 percentage points. Therefore, the critical boundary is −10, not zero. The confidence interval can include zero and still remain entirely above the non-inferiority margin.

What does the −2.2 lower confidence limit tell us?

It is the most conservative end of the reported two-sided 95% confidence interval for the treatment difference. Because −2.2 is still above −10, the interval does not include a treatment difference as unfavorable as the prespecified non-inferiority margin.

Why does the MSDF rule matter?

The primary endpoint is not simply a laboratory measurement among participants who happened to have an available Week 48 value. The registry definition states that participants without HIV-1 RNA data were treated as non-responders. The treatment comparison therefore incorporates the prespecified handling of missing endpoint data.

Why adjust for baseline HIV-1 RNA and background dual NRTI?

These were the baseline stratification factors identified in the posted analysis. Incorporating them into the CMH analysis accounts for the randomized trial's stratification structure and produces the reported adjusted treatment comparison.

Why does a non-inferiority result not prove that the treatments are identical?

Non-inferiority is a bounded-loss conclusion. It asks whether the evidence excludes a treatment disadvantage beyond the prespecified margin. The confidence interval here extends from −2.2 to 7.1, so the analysis permits a range of treatment differences. It does not establish exact equality of efficacy.

12. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm. These figures are reported separately from the primary efficacy analysis because safety and efficacy answer different statistical questions.

GroupSerious adverse events affectedAt risk
DTG 50 mg Once a Day41411
RTG 400 mg BID45411
DTG 50 mg Once a Day (Open-label)21338

The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or p-value for these serious adverse-event counts. They therefore should be read descriptively rather than treated as a formal hypothesis test.

Important denominator distinction: the primary efficacy analysis is based on the ITT-E population, while the serious-adverse-event information is reported using the affected/at-risk denominators posted on ClinicalTrials.gov for each safety group. The safety denominators should not be substituted for the efficacy analysis population.

13. Other Registered Interventions

ClinicalTrials.gov lists the following interventions in the ClinicalTrials.gov record: GSK1349572 (dolutegravir), raltegravir, GSK1349572 placebo, ABC/3TC, TDF/FTC, and raltegravir placebo.

Intervention listed in registry dataType
GSK1349572 (dolutegravir)Drug
RaltegravirDrug
GSK1349572 PlaceboOther
ABC/3TCOther
TDF/FTCOther
Raltegravir PlaceboOther

The presence of these interventions in the registry is consistent with the quadruple-masked design and the use of background dual NRTI treatment. The ClinicalTrials.gov record does not provide enough detail to reconstruct a more granular treatment-administration schedule, so no additional regimen details are inferred.

14. Design Features That Shape Interpretation

Randomization
The trial was randomized, creating the principal basis for a comparative efficacy analysis.
Parallel design
The registry identifies a parallel design with two treatment arms rather than a crossover or factorial design.
Quadruple masking
Masking was recorded as quadruple, reducing opportunities for knowledge of treatment assignment to influence trial conduct or assessment.
Non-inferiority
The primary hypothesis was framed around a −10 percentage-point non-inferiority margin.

What is not supported by the ClinicalTrials.gov record?

The ClinicalTrials.gov record does not report a factorial design, crossover analysis, Bayesian method, interim analysis, formal multiplicity procedure, or imputation method beyond the registered MSDF snapshot handling of missing endpoint data. Those topics are therefore not presented as trial methods here.

15. What the Primary Estimate Does — and Does Not — Mean

Effect estimate

The reported 2.5 percentage-point risk difference is the estimated absolute difference in the percentage achieving HIV-1 RNA below 50 copies/mL through Week 48, with DTG compared with RAL, using the posted stratified CMH analysis.

It does not mean that exactly 2.5 additional participants out of every 100 would respond in every population. It is a trial-level estimate subject to statistical uncertainty and to the endpoint's prespecified MSDF classification rules.

Confidence interval

The 95% CI of −2.2 to 7.1 quantifies uncertainty around the treatment difference. Its lower end is slightly below zero, while its upper end is positive. The interval therefore includes the possibility of a small negative difference as well as a positive difference.

Non-inferiority interpretation

The important comparison is between the lower confidence limit, −2.2, and the prespecified margin, −10. The lower limit remains above the non-inferiority boundary.

P-value interpretation

The ClinicalTrials.gov record does not report a p-value for this primary analysis. A p-value, if reported elsewhere, would quantify evidence relative to a specified null hypothesis; it would not replace the confidence-interval comparison required by the stated non-inferiority framework.

16. Limitations and Interpretation Issues

17. Why This Trial Matters Statistically

SPRING-2 is a useful teaching example because its primary analysis illustrates an important distinction in clinical-trial statistics: non-inferiority is a confidence-interval problem built around a clinically specified margin, not simply a search for a statistically significant difference between treatments.

ConceptHow it appears in SPRING-2
RandomizationParticipants were randomized to two parallel treatment arms.
Quadruple maskingThe registry identifies the trial as quadruple-masked.
Binary endpointThe primary endpoint is the percentage with HIV-1 RNA below 50 copies/mL through Week 48.
ITT analysisThe posted primary analysis used the ITT-E population.
Cochran-Mantel-Haenszel testThe primary comparison used a stratified CMH analysis.
Covariate adjustmentThe CMH analysis adjusted for baseline HIV-1 RNA and background dual NRTI.
Risk differenceThe treatment effect was reported as a difference in percentages.
Confidence intervalThe treatment difference was accompanied by a two-sided 95% CI.
Non-inferiority marginThe lower confidence bound was compared with a −10 percentage-point margin.
Missing-data handlingThe MSDF snapshot algorithm treated participants without HIV-1 RNA data as non-responders.

18. A Step-by-Step Reading of the Primary Analysis

Step 1 · Endpoint

Define the binary outcome

The endpoint is whether HIV-1 RNA is below 50 copies/mL through Week 48, using the registered MSDF snapshot framework.

Step 2 · Population

Use the ITT-E population

The primary analysis follows randomized treatment assignment in the ITT-E population.

Step 3 · Adjustment

Respect the stratification factors

The CMH analysis adjusts for baseline HIV-1 RNA and background dual NRTI.

Step 4 · Estimate

Calculate the treatment difference

The reported difference in percentage is 2.5 percentage points for DTG minus RAL.

Step 5 · Uncertainty

Construct the two-sided 95% CI

The reported interval extends from −2.2 to 7.1 percentage points.

Step 6 · Non-inferiority

Compare the lower bound with −10

The lower confidence limit of −2.2 remains above the prespecified −10 percentage-point non-inferiority margin.

19. Primary Analysis in Statistical Notation

Treatment contrast
Δ = pDTG − pRAL

The reported estimate is Δ = 2.5 percentage points.

Non-inferiority criterion
Lower bound of 95% CI(Δ) > −10

The reported lower bound is −2.2, which lies above the prespecified −10 percentage-point margin.

This notation makes the logic transparent. The point estimate describes the observed treatment contrast, the confidence interval describes uncertainty, and the non-inferiority margin defines the clinically relevant boundary against which that uncertainty is evaluated.

20. Results and What They Do Not Establish

What the result establishes statistically

The posted primary analysis estimates a 2.5 percentage-point difference, with a two-sided 95% CI from −2.2 to 7.1, using the specified ITT-E and stratified CMH framework.

What it does not establish

The result does not show that the treatments have identical efficacy, nor does it establish that every patient or every subgroup experiences the same treatment difference.

What the CI contributes

The interval provides the uncertainty needed to assess the non-inferiority margin and shows that both small negative and positive treatment differences remain compatible with the estimate.

What the ClinicalTrials.gov record does not provide

The ClinicalTrials.gov record does not provide a formal p-value, underlying response counts, or additional formal primary-endpoint analyses.

21. Related Tutorials

Learn more about the methods used in this trial:

22. Related Statistical Calculators

23. Sources

Continue with Clinical Biostats statistical methods

Explore the trial-design and analysis concepts that connect directly to the SPRING-2 primary endpoint, including non-inferiority, confidence intervals, risk differences, stratified analysis, and categorical-data methods.

24. Record Summary

SPRING-2 provides a compact but important example of how a non-inferiority clinical trial can be analyzed for a binary virologic endpoint. The primary comparison used the ITT-E population and a Cochran-Mantel-Haenszel stratified analysis adjusted for baseline HIV-1 RNA and background dual NRTI. The reported treatment difference was 2.5 percentage points, with a two-sided 95% confidence interval of −2.2 to 7.1. The key statistical question is therefore not whether the interval excludes zero, but whether its lower bound remains above the prespecified −10 percentage-point non-inferiority margin.

Clinical Biostats methodology: The most informative reading of a non-inferiority result combines the treatment-effect estimate, its confidence interval, the prespecified non-inferiority margin, the analysis population, the missing-data rule, and the stratification method. A single point estimate cannot convey all of that statistical information.