← Clinical Trial Results
HIV Infection Phase 3 Non-Inferiority NCT01231516

SAILING: Complete Statistical Analysis of GSK1349572 in HIV Infection

An educational statistical analysis of the randomized phase 3 SAILING trial comparing GSK1349572 50 mg once daily with raltegravir 400 mg twice daily, each with an investigator-selected background regimen, in antiretroviral-experienced, integrase inhibitor-naive adults with HIV infection.

Completed  ·  Randomized parallel-group design  ·  Enrollment 724  ·  Results posted
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical trial results on this page are restricted to the ClinicalTrials.gov record and the linked publication identifier in the trial data.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

SAILING was a randomized, parallel phase 3 trial comparing GSK1349572 with raltegravir, each administered with an investigator-selected background regimen, in antiretroviral-experienced, integrase inhibitor-naive adults with HIV infection. The registry reports a binary primary endpoint at Week 48 and a prespecified non-inferiority framework based on the difference in the percentage of participants achieving HIV-1 RNA <50 c/mL.

724
Enrolled
2 treatment arms
48
Primary time point
Week 48
7.4
Risk difference
DTG − RAL, percentage points
0.030
P-value
Cochran-Mantel-Haenszel
FeatureSAILING
Trial nameSAILING
PhasePhase 3
Therapeutic areaInfectious Disease
ConditionInfection, Human Immunodeficiency Virus; HIV Infections
AllocationRandomized
Design modelParallel
MaskingQuadruple
Primary purposeTreatment
Enrollment724
Primary endpointPercentage of participants with HIV-1 RNA <50 copies/mL at Week 48
Primary endpoint typeBinary
Primary statistical methodCochran-Mantel-Haenszel test
Effect measureRisk difference / difference in percentage
Hypothesis typeNon-inferiority or equivalence
Trial statusCompleted
StartOctober 26, 2010
Primary completionFebruary 4, 2013
Lead sponsorViiV Healthcare
Sponsor typeIndustry
ClinicalTrials.govNCT01231516

2. Clinical Question

The primary question was whether GSK1349572 50 mg once daily, given with an investigator-selected background regimen, was non-inferior to raltegravir 400 mg twice daily with an investigator-selected background regimen with respect to the percentage of participants achieving HIV-1 RNA <50 copies/mL at Week 48.

Population

Antiretroviral-experienced, integrase inhibitor-naive adults with HIV infection.

Intervention

GSK1349572 50 mg once daily, with an investigator-selected background regimen.

Comparator

Raltegravir 400 mg twice daily, with an investigator-selected background regimen.

Primary question

Is the Week 48 virologic response with GSK1349572 non-inferior to that with raltegravir under the prespecified statistical criterion?

3. Trial Design

01
Randomize724 participants
02
Assign2 parallel treatment groups
03
BlindQuadruple masking
04
AssessHIV-1 RNA at Week 48
05
CompareAdjusted percentage difference
ARM A · GSK1349572

GSK1349572 50 mg once daily

  • GSK1349572 50 mg once daily
  • Investigator-selected background regimen
  • GSK1349572 placebo was among the registered interventions
ARM B · RAL

Raltegravir 400 mg twice daily

  • Raltegravir 400 mg twice daily
  • Investigator-selected background regimen
  • Raltegravir placebo was among the registered interventions

The randomized parallel-group structure is important statistically because treatment assignment creates the basis for comparing the Week 48 binary outcome. The quadruple-masked design is also relevant: masking can reduce the potential for knowledge of treatment assignment to influence trial conduct and outcome assessment.

Trial chronology

October 26, 2010

Trial start

The registered study start date was October 26, 2010.

February 4, 2013

Primary completion

The registered primary completion date was February 4, 2013.

Completed

Results posted

The ClinicalTrials.gov record reports that results were posted, including one formal statistical analysis for the primary endpoint.

4. Endpoints

The registry identifies one primary endpoint. It is a binary virologic outcome assessed at Week 48.

EndpointRegistry definition / time frameEndpoint type
Percentage of Participants With HIV-1 Ribonucleic Acid (RNA) <50 Copies/Milliliter (c/mL) at Week 48 At Week 48 Binary

The registered definition specifies that the percentage was assessed using the Missing, Switch or Discontinuation = Failure (MSDF) approach, codified by the FDA “snapshot” algorithm. Under the registry definition, participants without HIV-1 RNA at Week 48 were treated as nonresponders, as were participants who switched treatment or discontinued under the circumstances specified by the algorithm.

Endpoint interpretation: this is a responder/nonresponder endpoint at a fixed Week 48 time point. It is therefore fundamentally different from a time-to-event endpoint such as overall survival: the analysis compares proportions achieving a defined virologic threshold at one prespecified assessment point.

5. Analysis Population and Treatment Comparison

Analysis featureRegistry information
Analysis populationModified Intent-To-Treat Exposed (mITT-E) Population
Population principleRandomized participants who received at least one dose of investigational product, with the stated site exclusion
Groups comparedDTG 50 mg OD vs RAL 400 mg BID
Primary endpointHIV-1 RNA <50 c/mL at Week 48
Statistical methodCochran-Mantel-Haenszel
Effect measureDifference in percentage / risk difference
AdjustmentBaseline stratification factors

The analysis population deserves attention because the phrase intent-to-treat does not automatically mean that every randomized participant is present in every analysis population. Here, the registry specifically reports an mITT-E population defined by exposure to investigational product and a stated site exclusion. The interpretation of the numerical result therefore belongs to that analysis population rather than being silently generalized to every enrolled participant.

6. Statistical Methodology

Binary responder analysis

The primary outcome is binary: a participant either meets the HIV-1 RNA <50 c/mL criterion at Week 48 under the registered MSDF/snapshot framework or is classified as a nonresponder under that framework. The natural statistical object is therefore a difference between two response proportions.

Risk difference
Risk difference = P(response | DTG) − P(response | RAL)

A positive risk difference means that the observed response percentage is higher in the GSK1349572 group than in the raltegravir group. The estimate is expressed in percentage points when the component probabilities are expressed as percentages.

Cochran-Mantel-Haenszel test

The posted primary analysis used the Cochran-Mantel-Haenszel (CMH) method. The analysis text states that the adjusted difference in proportion was based on the difference in percentage and adjusted for baseline stratification factors.

For a stratified binary outcome, the CMH framework allows the comparison between treatment groups to account for prespecified strata rather than treating all observations as though they belonged to one homogeneous unstratified table. This is particularly relevant when baseline characteristics used for stratification may be associated with the response outcome.

Baseline stratification factors

The registry identifies three baseline stratification factors used for the adjusted analysis:

FactorStrata
HIV-1 RNA≤50000 vs >50000 c/mL
Darunavir-ritonavir use without primary protease inhibitor mutationsYes vs no
Phenotypic susceptibility score to background regimen2 vs <2

These factors were incorporated into the adjusted difference in proportions. The statistical purpose is not to “fix” an imbalance after the fact; rather, stratified adjustment incorporates information from the prespecified baseline structure into the comparison.

Non-inferiority framework

The trial's posted hypothesis type is non-inferiority or equivalence. The registry provides a specific non-inferiority rule: non-inferiority of DTG 50 mg and raltegravir at Week 48 can be concluded if the lower bound of a two-sided 95% confidence interval for the difference in percentages (DTG − RAL) is greater than −12%.

Prespecified non-inferiority criterion

Lower 95% CI bound > −12%

The relevant comparison is the two-sided 95% confidence interval for the difference in Week 48 response percentages.

If non-inferiority were established, the registry states that superiority would then be tested at the nominal 5% level based on a pre-specified strategy.

This is a central feature of the SAILING analysis. In a non-inferiority trial, the question is not initially whether the treatment has a statistically significant positive difference. The question is whether the data are sufficiently compatible with a loss no larger than the prespecified non-inferiority margin.

7. Primary Result: Week 48 Virologic Response

The posted formal analysis compares DTG 50 mg once daily with RAL 400 mg twice daily for the percentage of participants with HIV-1 RNA <50 c/mL at Week 48.

Adjusted difference in Week 48 response

7.4 percentage points

95% CI: 0.7 to 14.2   ·   P = 0.030

Effect measure: difference in percentage (DTG − RAL)

Primary endpointComparisonEstimate95% CIP-valueMethod
Percentage of participants with HIV-1 RNA <50 c/mL at Week 48 DTG 50 mg OD vs RAL 400 mg BID 7.4 percentage points 0.7 to 14.2 0.030 Cochran-Mantel-Haenszel
Clinical Biostats interpretation

The estimated adjusted difference was 7.4 percentage points, defined as the Week 48 response percentage for DTG minus the corresponding response percentage for RAL. Because the estimate is positive, the estimated response percentage was higher in the DTG group in this analysis.

The estimate does not mean that 7.4% of all participants benefited, nor does it mean that each individual participant had a 7.4% change in probability. It is a group-level difference in the proportion meeting the prespecified binary response criterion.

The two-sided 95% confidence interval extends from 0.7 to 14.2 percentage points. This interval describes uncertainty around the estimated adjusted difference under the analysis framework; it is not a range containing the treatment effect for individual participants.

The P-value of 0.030 measures the compatibility of the observed data with the statistical null hypothesis used by the CMH analysis. It is not a measure of the size or clinical importance of the effect. The size of the estimated difference is conveyed by the 7.4 percentage-point estimate and its confidence interval.

For non-inferiority, the critical issue is the lower confidence-limit comparison with the prespecified −12% margin. The reported lower limit of 0.7% is above that margin. The registry therefore provides the statistical information needed to evaluate the stated non-inferiority criterion.

Visualizing the reported difference

Reported adjusted difference
Point estimate
7.4
Lower CI bound
0.7
Upper CI bound
14.2

The chart is a visual representation of the reported numerical estimate and confidence limits; it does not reconstruct participant-level response percentages or the underlying stratified contingency tables.

8. How the Non-Inferiority Logic Works

Non-inferiority requires a different reading strategy from a conventional superiority test. Let the treatment difference be defined as:

Direction used in the registry
Difference = percentage responding with DTG − percentage responding with RAL

Under this direction, increasingly negative values favor the possibility that DTG is meaningfully worse than RAL. The prespecified non-inferiority margin is −12%.

The registry states that non-inferiority can be concluded when the lower bound of the two-sided 95% confidence interval is greater than −12%. The reported interval is 0.7 to 14.2%. Its lower bound therefore lies above the non-inferiority margin.

What the margin means

The −12% margin defines the largest loss of response that the trial's non-inferiority criterion was designed to rule out, using the registry's stated difference scale.

Why the lower bound matters

The lower confidence limit represents the unfavorable end of the uncertainty interval for the DTG-minus-RAL difference. Non-inferiority depends on this limit remaining above −12%.

Why the point estimate is not enough

A positive point estimate alone does not establish non-inferiority. The confidence interval determines how much uncertainty remains about the treatment difference.

What superiority adds

The registry states that if non-inferiority were established, superiority would be tested at the nominal 5% level according to a prespecified strategy.

9. Statistical Methods Explained

Why was the Cochran-Mantel-Haenszel test used?

The primary outcome is binary and the analysis incorporated baseline stratification factors. The CMH framework is designed for comparing treatment groups across strata, producing an adjusted comparison rather than ignoring the stratified structure. In SAILING, the registry explicitly identifies CMH as the posted primary method.

What does a risk difference of 7.4 mean?

A risk difference of 7.4 percentage points means that the adjusted percentage meeting the Week 48 HIV-1 RNA <50 c/mL criterion was estimated to be 7.4 percentage points higher with DTG than with RAL. It is an absolute difference between two proportions, not a relative percentage increase and not an individual-level probability.

Why is the confidence interval central to non-inferiority?

Non-inferiority asks whether the data rule out an unacceptably unfavorable treatment difference. That requires looking at the entire uncertainty interval, especially the lower bound when the treatment difference is defined as DTG minus RAL. Here the lower bound is 0.7%, which is above the −12% margin specified by the registry.

Why does the P-value not measure effect size?

The P-value summarizes how compatible the observed data are with a specified null hypothesis under the statistical model. It does not say that the treatment effect is “3.0%” or that a P-value of 0.030 represents a 3% improvement. The effect size is the 7.4 percentage-point difference, while the confidence interval communicates uncertainty around that estimate.

Why adjust for baseline stratification factors?

The registry states that the analysis was adjusted for HIV-1 RNA category, darunavir-ritonavir use without primary protease inhibitor mutations, and phenotypic susceptibility score to the background regimen. Stratified adjustment incorporates these prespecified baseline factors into the treatment comparison and can account for differences in outcome rates across strata.

Why does the analysis population matter?

The posted analysis used the mITT-E population rather than simply labeling the result as an analysis of all 724 enrolled participants. Because eligibility for the analysis included receipt of at least one dose of investigational product and a specified site exclusion, the numerical result should be interpreted in the population actually analyzed.

10. Understanding the ClinicalTrials.gov Endpoint Algorithm

The Week 48 endpoint uses the Missing, Switch or Discontinuation = Failure approach, described in the registry as the FDA “snapshot” algorithm. This is more than a technical detail: it determines how participants with missing Week 48 measurements or treatment changes contribute to the binary outcome.

Observed response

A participant with the required HIV-1 RNA result below the registered threshold at Week 48 can contribute as a responder under the endpoint definition.

Missing measurement

The registered algorithm treats participants without HIV-1 RNA at Week 48 as nonresponders.

Switching treatment

The endpoint framework treats treatment switching according to the MSDF/snapshot failure rules described in the registry definition.

Discontinuation

Discontinuation is incorporated into the endpoint through the registry's failure classification rather than being silently ignored.

This approach converts a potentially complicated longitudinal treatment history into a prespecified binary Week 48 outcome. That standardization is useful for a confirmatory comparison because the classification rule is specified before interpreting the treatment difference.

11. Risk Difference Versus Relative Effect Measures

The SAILING primary analysis is reported using a difference in percentage, normalized here as a risk difference. This choice has a direct clinical interpretation: it expresses the absolute separation between the two response proportions.

MeasureQuestion it answersSAILING primary analysis
Risk differenceHow many percentage points apart are the two response proportions?Yes
Risk ratioHow many times as large is the response probability in one group relative to the other?Not reported in the ClinicalTrials.gov record
Odds ratioHow do the response odds compare?Not reported in the ClinicalTrials.gov record
Hazard ratioHow do instantaneous event rates compare over time?Not the primary endpoint method

The distinction matters because a risk difference is naturally interpreted in percentage points. It should not be converted into a hazard ratio, odds ratio, or relative risk without the underlying data and an appropriate statistical model.

12. Safety

The ClinicalTrials.gov record reports serious adverse events by treatment arm. These data provide an arm-level safety count but do not provide a formal comparative statistical analysis of serious adverse events in the registry-reported statistical-analyses record.

Treatment armSerious adverse eventsAffected / at risk
DTG 50 mg once dailySerious adverse events73 / 357
RAL 400 mg twice dailySerious adverse events46 / 362

These counts should not be treated as another primary efficacy endpoint. They describe the number affected and the number at risk in each registry-reported safety population. The trial data do not provide a confidence interval or P-value for a between-arm serious-adverse-event comparison, so none is inferred here.

Safety interpretation: the denominators in the registry-reported serious-adverse-event data are 357 and 362, respectively, rather than the overall enrollment of 724. That distinction is important when reading safety percentages or attempting any comparative calculation.

13. What the Primary Result Does — and Does Not — Establish

What the estimate says

The adjusted treatment difference was 7.4 percentage points for the Week 48 binary virologic endpoint, with a two-sided 95% confidence interval of 0.7 to 14.2%.

What it does not say

The estimate does not mean that every participant experienced a 7.4 percentage-point improvement. It is a population-level comparison of the proportion meeting the registered response criterion.

What the confidence interval says

The confidence interval quantifies uncertainty around the adjusted treatment difference under the specified analysis. Its lower bound is especially important because the non-inferiority margin is defined on the same difference scale.

What the P-value says

The reported P = 0.030 is a hypothesis-testing quantity. It is not a measure of the magnitude of the treatment effect and should not be substituted for the estimated difference or its confidence interval.

Why the endpoint algorithm matters

Because the outcome uses the MSDF/snapshot framework, the response proportions reflect prespecified classifications of missing data, switching, and discontinuation. The result therefore depends on the registered endpoint definition, not merely on observed laboratory measurements among participants with complete Week 48 data.

14. Limitations

15. Why This Trial Matters Statistically

SAILING is a useful teaching case because its primary analysis illustrates several core principles of modern clinical-trial statistics without requiring a time-to-event endpoint. The central statistical problem is a binary response comparison embedded in a stratified, randomized, blinded, non-inferiority design.

ConceptHow it appears in SAILING
RandomizationThe trial uses randomized allocation in a parallel-group phase 3 design.
BlindingThe registry specifies quadruple masking.
Binary endpointWeek 48 HIV-1 RNA <50 c/mL is analyzed as a responder/nonresponder outcome.
Risk differenceThe primary effect measure is the difference in percentage between DTG and RAL.
Stratified analysisThe treatment comparison is adjusted for prespecified baseline stratification factors.
Cochran-Mantel-Haenszel methodThe posted primary statistical method is the CMH test.
Non-inferiorityThe lower bound of a two-sided 95% CI is compared with a −12% margin.
Intention-to-treat principleThe analysis population is modified from the randomized population and explicitly defined by exposure and site exclusion.
Confidence intervalsThe 95% CI quantifies uncertainty around the adjusted difference.
P-valuesThe reported P = 0.030 addresses hypothesis testing rather than effect magnitude.

16. Related Tutorials

Learn more about the methods used in this trial:

17. Related Statistical Calculators

18. Sources

Continue through Clinical Biostats

Use the trial's statistical concepts as a starting point for deeper tutorials and practical statistical calculators.

19. Record Summary

SAILING provides a clear example of how a randomized clinical trial can evaluate a binary Week 48 endpoint within a non-inferiority framework. The key statistical elements are the modified intent-to-treat exposed population, the Cochran-Mantel-Haenszel analysis, adjustment for baseline stratification factors, the risk-difference effect measure, and the prespecified −12% non-inferiority margin.

The reported adjusted difference was 7.4 percentage points, with a two-sided 95% confidence interval of 0.7 to 14.2% and P = 0.030. The lower confidence bound is the critical quantity for the stated non-inferiority criterion because it remains above the prespecified −12% margin. The interpretation should remain tied to the registered endpoint definition, the mITT-E analysis population, and the specific stratified statistical framework rather than extending the result to unreported outcomes.

Clinical Biostats methodology: A trial-results page should not merely repeat a registry result. The goal is to explain the statistical structure of the comparison, distinguish effect size from uncertainty and hypothesis testing, and show how the endpoint definition and non-inferiority framework determine the interpretation.