This page separates reported trial results from statistical interpretation. Numerical trial results on this page are restricted to the ClinicalTrials.gov record and the linked publication identifier in the trial data.
1. Trial at a Glance
SAILING was a randomized, parallel phase 3 trial comparing GSK1349572 with raltegravir, each administered with an investigator-selected background regimen, in antiretroviral-experienced, integrase inhibitor-naive adults with HIV infection. The registry reports a binary primary endpoint at Week 48 and a prespecified non-inferiority framework based on the difference in the percentage of participants achieving HIV-1 RNA <50 c/mL.
| Feature | SAILING |
|---|---|
| Trial name | SAILING |
| Phase | Phase 3 |
| Therapeutic area | Infectious Disease |
| Condition | Infection, Human Immunodeficiency Virus; HIV Infections |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 724 |
| Primary endpoint | Percentage of participants with HIV-1 RNA <50 copies/mL at Week 48 |
| Primary endpoint type | Binary |
| Primary statistical method | Cochran-Mantel-Haenszel test |
| Effect measure | Risk difference / difference in percentage |
| Hypothesis type | Non-inferiority or equivalence |
| Trial status | Completed |
| Start | October 26, 2010 |
| Primary completion | February 4, 2013 |
| Lead sponsor | ViiV Healthcare |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT01231516 |
2. Clinical Question
The primary question was whether GSK1349572 50 mg once daily, given with an investigator-selected background regimen, was non-inferior to raltegravir 400 mg twice daily with an investigator-selected background regimen with respect to the percentage of participants achieving HIV-1 RNA <50 copies/mL at Week 48.
Population
Antiretroviral-experienced, integrase inhibitor-naive adults with HIV infection.
Intervention
GSK1349572 50 mg once daily, with an investigator-selected background regimen.
Comparator
Raltegravir 400 mg twice daily, with an investigator-selected background regimen.
Primary question
Is the Week 48 virologic response with GSK1349572 non-inferior to that with raltegravir under the prespecified statistical criterion?
3. Trial Design
GSK1349572 50 mg once daily
- GSK1349572 50 mg once daily
- Investigator-selected background regimen
- GSK1349572 placebo was among the registered interventions
Raltegravir 400 mg twice daily
- Raltegravir 400 mg twice daily
- Investigator-selected background regimen
- Raltegravir placebo was among the registered interventions
The randomized parallel-group structure is important statistically because treatment assignment creates the basis for comparing the Week 48 binary outcome. The quadruple-masked design is also relevant: masking can reduce the potential for knowledge of treatment assignment to influence trial conduct and outcome assessment.
Trial chronology
Trial start
The registered study start date was October 26, 2010.
Primary completion
The registered primary completion date was February 4, 2013.
Results posted
The ClinicalTrials.gov record reports that results were posted, including one formal statistical analysis for the primary endpoint.
4. Endpoints
The registry identifies one primary endpoint. It is a binary virologic outcome assessed at Week 48.
| Endpoint | Registry definition / time frame | Endpoint type |
|---|---|---|
| Percentage of Participants With HIV-1 Ribonucleic Acid (RNA) <50 Copies/Milliliter (c/mL) at Week 48 | At Week 48 | Binary |
The registered definition specifies that the percentage was assessed using the Missing, Switch or Discontinuation = Failure (MSDF) approach, codified by the FDA “snapshot” algorithm. Under the registry definition, participants without HIV-1 RNA at Week 48 were treated as nonresponders, as were participants who switched treatment or discontinued under the circumstances specified by the algorithm.
5. Analysis Population and Treatment Comparison
| Analysis feature | Registry information |
|---|---|
| Analysis population | Modified Intent-To-Treat Exposed (mITT-E) Population |
| Population principle | Randomized participants who received at least one dose of investigational product, with the stated site exclusion |
| Groups compared | DTG 50 mg OD vs RAL 400 mg BID |
| Primary endpoint | HIV-1 RNA <50 c/mL at Week 48 |
| Statistical method | Cochran-Mantel-Haenszel |
| Effect measure | Difference in percentage / risk difference |
| Adjustment | Baseline stratification factors |
The analysis population deserves attention because the phrase intent-to-treat does not automatically mean that every randomized participant is present in every analysis population. Here, the registry specifically reports an mITT-E population defined by exposure to investigational product and a stated site exclusion. The interpretation of the numerical result therefore belongs to that analysis population rather than being silently generalized to every enrolled participant.
6. Statistical Methodology
Binary responder analysis
The primary outcome is binary: a participant either meets the HIV-1 RNA <50 c/mL criterion at Week 48 under the registered MSDF/snapshot framework or is classified as a nonresponder under that framework. The natural statistical object is therefore a difference between two response proportions.
A positive risk difference means that the observed response percentage is higher in the GSK1349572 group than in the raltegravir group. The estimate is expressed in percentage points when the component probabilities are expressed as percentages.
Cochran-Mantel-Haenszel test
The posted primary analysis used the Cochran-Mantel-Haenszel (CMH) method. The analysis text states that the adjusted difference in proportion was based on the difference in percentage and adjusted for baseline stratification factors.
For a stratified binary outcome, the CMH framework allows the comparison between treatment groups to account for prespecified strata rather than treating all observations as though they belonged to one homogeneous unstratified table. This is particularly relevant when baseline characteristics used for stratification may be associated with the response outcome.
Baseline stratification factors
The registry identifies three baseline stratification factors used for the adjusted analysis:
| Factor | Strata |
|---|---|
| HIV-1 RNA | ≤50000 vs >50000 c/mL |
| Darunavir-ritonavir use without primary protease inhibitor mutations | Yes vs no |
| Phenotypic susceptibility score to background regimen | 2 vs <2 |
These factors were incorporated into the adjusted difference in proportions. The statistical purpose is not to “fix” an imbalance after the fact; rather, stratified adjustment incorporates information from the prespecified baseline structure into the comparison.
Non-inferiority framework
The trial's posted hypothesis type is non-inferiority or equivalence. The registry provides a specific non-inferiority rule: non-inferiority of DTG 50 mg and raltegravir at Week 48 can be concluded if the lower bound of a two-sided 95% confidence interval for the difference in percentages (DTG − RAL) is greater than −12%.
Prespecified non-inferiority criterion
The relevant comparison is the two-sided 95% confidence interval for the difference in Week 48 response percentages.
If non-inferiority were established, the registry states that superiority would then be tested at the nominal 5% level based on a pre-specified strategy.
This is a central feature of the SAILING analysis. In a non-inferiority trial, the question is not initially whether the treatment has a statistically significant positive difference. The question is whether the data are sufficiently compatible with a loss no larger than the prespecified non-inferiority margin.
7. Primary Result: Week 48 Virologic Response
The posted formal analysis compares DTG 50 mg once daily with RAL 400 mg twice daily for the percentage of participants with HIV-1 RNA <50 c/mL at Week 48.
Adjusted difference in Week 48 response
95% CI: 0.7 to 14.2 · P = 0.030
Effect measure: difference in percentage (DTG − RAL)
| Primary endpoint | Comparison | Estimate | 95% CI | P-value | Method |
|---|---|---|---|---|---|
| Percentage of participants with HIV-1 RNA <50 c/mL at Week 48 | DTG 50 mg OD vs RAL 400 mg BID | 7.4 percentage points | 0.7 to 14.2 | 0.030 | Cochran-Mantel-Haenszel |
The estimated adjusted difference was 7.4 percentage points, defined as the Week 48 response percentage for DTG minus the corresponding response percentage for RAL. Because the estimate is positive, the estimated response percentage was higher in the DTG group in this analysis.
The estimate does not mean that 7.4% of all participants benefited, nor does it mean that each individual participant had a 7.4% change in probability. It is a group-level difference in the proportion meeting the prespecified binary response criterion.
The two-sided 95% confidence interval extends from 0.7 to 14.2 percentage points. This interval describes uncertainty around the estimated adjusted difference under the analysis framework; it is not a range containing the treatment effect for individual participants.
The P-value of 0.030 measures the compatibility of the observed data with the statistical null hypothesis used by the CMH analysis. It is not a measure of the size or clinical importance of the effect. The size of the estimated difference is conveyed by the 7.4 percentage-point estimate and its confidence interval.
For non-inferiority, the critical issue is the lower confidence-limit comparison with the prespecified −12% margin. The reported lower limit of 0.7% is above that margin. The registry therefore provides the statistical information needed to evaluate the stated non-inferiority criterion.
Visualizing the reported difference
The chart is a visual representation of the reported numerical estimate and confidence limits; it does not reconstruct participant-level response percentages or the underlying stratified contingency tables.
8. How the Non-Inferiority Logic Works
Non-inferiority requires a different reading strategy from a conventional superiority test. Let the treatment difference be defined as:
Under this direction, increasingly negative values favor the possibility that DTG is meaningfully worse than RAL. The prespecified non-inferiority margin is −12%.
The registry states that non-inferiority can be concluded when the lower bound of the two-sided 95% confidence interval is greater than −12%. The reported interval is 0.7 to 14.2%. Its lower bound therefore lies above the non-inferiority margin.
What the margin means
The −12% margin defines the largest loss of response that the trial's non-inferiority criterion was designed to rule out, using the registry's stated difference scale.
Why the lower bound matters
The lower confidence limit represents the unfavorable end of the uncertainty interval for the DTG-minus-RAL difference. Non-inferiority depends on this limit remaining above −12%.
Why the point estimate is not enough
A positive point estimate alone does not establish non-inferiority. The confidence interval determines how much uncertainty remains about the treatment difference.
What superiority adds
The registry states that if non-inferiority were established, superiority would be tested at the nominal 5% level according to a prespecified strategy.
9. Statistical Methods Explained
Why was the Cochran-Mantel-Haenszel test used?
The primary outcome is binary and the analysis incorporated baseline stratification factors. The CMH framework is designed for comparing treatment groups across strata, producing an adjusted comparison rather than ignoring the stratified structure. In SAILING, the registry explicitly identifies CMH as the posted primary method.
What does a risk difference of 7.4 mean?
A risk difference of 7.4 percentage points means that the adjusted percentage meeting the Week 48 HIV-1 RNA <50 c/mL criterion was estimated to be 7.4 percentage points higher with DTG than with RAL. It is an absolute difference between two proportions, not a relative percentage increase and not an individual-level probability.
Why is the confidence interval central to non-inferiority?
Non-inferiority asks whether the data rule out an unacceptably unfavorable treatment difference. That requires looking at the entire uncertainty interval, especially the lower bound when the treatment difference is defined as DTG minus RAL. Here the lower bound is 0.7%, which is above the −12% margin specified by the registry.
Why does the P-value not measure effect size?
The P-value summarizes how compatible the observed data are with a specified null hypothesis under the statistical model. It does not say that the treatment effect is “3.0%” or that a P-value of 0.030 represents a 3% improvement. The effect size is the 7.4 percentage-point difference, while the confidence interval communicates uncertainty around that estimate.
Why adjust for baseline stratification factors?
The registry states that the analysis was adjusted for HIV-1 RNA category, darunavir-ritonavir use without primary protease inhibitor mutations, and phenotypic susceptibility score to the background regimen. Stratified adjustment incorporates these prespecified baseline factors into the treatment comparison and can account for differences in outcome rates across strata.
Why does the analysis population matter?
The posted analysis used the mITT-E population rather than simply labeling the result as an analysis of all 724 enrolled participants. Because eligibility for the analysis included receipt of at least one dose of investigational product and a specified site exclusion, the numerical result should be interpreted in the population actually analyzed.
10. Understanding the ClinicalTrials.gov Endpoint Algorithm
The Week 48 endpoint uses the Missing, Switch or Discontinuation = Failure approach, described in the registry as the FDA “snapshot” algorithm. This is more than a technical detail: it determines how participants with missing Week 48 measurements or treatment changes contribute to the binary outcome.
Observed response
A participant with the required HIV-1 RNA result below the registered threshold at Week 48 can contribute as a responder under the endpoint definition.
Missing measurement
The registered algorithm treats participants without HIV-1 RNA at Week 48 as nonresponders.
Switching treatment
The endpoint framework treats treatment switching according to the MSDF/snapshot failure rules described in the registry definition.
Discontinuation
Discontinuation is incorporated into the endpoint through the registry's failure classification rather than being silently ignored.
This approach converts a potentially complicated longitudinal treatment history into a prespecified binary Week 48 outcome. That standardization is useful for a confirmatory comparison because the classification rule is specified before interpreting the treatment difference.
11. Risk Difference Versus Relative Effect Measures
The SAILING primary analysis is reported using a difference in percentage, normalized here as a risk difference. This choice has a direct clinical interpretation: it expresses the absolute separation between the two response proportions.
| Measure | Question it answers | SAILING primary analysis |
|---|---|---|
| Risk difference | How many percentage points apart are the two response proportions? | Yes |
| Risk ratio | How many times as large is the response probability in one group relative to the other? | Not reported in the ClinicalTrials.gov record |
| Odds ratio | How do the response odds compare? | Not reported in the ClinicalTrials.gov record |
| Hazard ratio | How do instantaneous event rates compare over time? | Not the primary endpoint method |
The distinction matters because a risk difference is naturally interpreted in percentage points. It should not be converted into a hazard ratio, odds ratio, or relative risk without the underlying data and an appropriate statistical model.
12. Safety
The ClinicalTrials.gov record reports serious adverse events by treatment arm. These data provide an arm-level safety count but do not provide a formal comparative statistical analysis of serious adverse events in the registry-reported statistical-analyses record.
| Treatment arm | Serious adverse events | Affected / at risk |
|---|---|---|
| DTG 50 mg once daily | Serious adverse events | 73 / 357 |
| RAL 400 mg twice daily | Serious adverse events | 46 / 362 |
These counts should not be treated as another primary efficacy endpoint. They describe the number affected and the number at risk in each registry-reported safety population. The trial data do not provide a confidence interval or P-value for a between-arm serious-adverse-event comparison, so none is inferred here.
13. What the Primary Result Does — and Does Not — Establish
The adjusted treatment difference was 7.4 percentage points for the Week 48 binary virologic endpoint, with a two-sided 95% confidence interval of 0.7 to 14.2%.
The estimate does not mean that every participant experienced a 7.4 percentage-point improvement. It is a population-level comparison of the proportion meeting the registered response criterion.
The confidence interval quantifies uncertainty around the adjusted treatment difference under the specified analysis. Its lower bound is especially important because the non-inferiority margin is defined on the same difference scale.
The reported P = 0.030 is a hypothesis-testing quantity. It is not a measure of the magnitude of the treatment effect and should not be substituted for the estimated difference or its confidence interval.
Because the outcome uses the MSDF/snapshot framework, the response proportions reflect prespecified classifications of missing data, switching, and discontinuation. The result therefore depends on the registered endpoint definition, not merely on observed laboratory measurements among participants with complete Week 48 data.
14. Limitations
- mITT-E population: the formal analysis was conducted in a modified intent-to-treat exposed population, so the reported estimate should not automatically be described as an analysis of every enrolled participant.
- Non-inferiority interpretation: non-inferiority depends on the prespecified −12% margin and the confidence interval, not simply on whether a conventional P-value is below 0.05.
- Binary time-point endpoint: the primary endpoint describes response status at Week 48 and does not by itself characterize the entire trajectory of HIV-1 RNA over time.
- Snapshot classification: the MSDF algorithm classifies missing, switched, and discontinued participants according to prespecified rules. The resulting proportion therefore cannot be interpreted as a simple observed-case percentage.
- Safety comparison: serious adverse-event counts are available by arm, but the ClinicalTrials.gov record does not provide a formal between-arm statistical test or confidence interval for this safety endpoint.
- Limited posted statistical analysis: the ClinicalTrials.gov record contains one formal statistical analysis, so this page does not infer additional efficacy analyses or subgroup results that are not included in the trial data.
- No participant-level reconstruction: the registry-reported summary does not contain individual participant data or the stratified contingency tables, so additional effect measures are not calculated.
15. Why This Trial Matters Statistically
SAILING is a useful teaching case because its primary analysis illustrates several core principles of modern clinical-trial statistics without requiring a time-to-event endpoint. The central statistical problem is a binary response comparison embedded in a stratified, randomized, blinded, non-inferiority design.
| Concept | How it appears in SAILING |
|---|---|
| Randomization | The trial uses randomized allocation in a parallel-group phase 3 design. |
| Blinding | The registry specifies quadruple masking. |
| Binary endpoint | Week 48 HIV-1 RNA <50 c/mL is analyzed as a responder/nonresponder outcome. |
| Risk difference | The primary effect measure is the difference in percentage between DTG and RAL. |
| Stratified analysis | The treatment comparison is adjusted for prespecified baseline stratification factors. |
| Cochran-Mantel-Haenszel method | The posted primary statistical method is the CMH test. |
| Non-inferiority | The lower bound of a two-sided 95% CI is compared with a −12% margin. |
| Intention-to-treat principle | The analysis population is modified from the randomized population and explicitly defined by exposure and site exclusion. |
| Confidence intervals | The 95% CI quantifies uncertainty around the adjusted difference. |
| P-values | The reported P = 0.030 addresses hypothesis testing rather than effect magnitude. |
16. Related Tutorials
Learn more about the methods used in this trial:
17. Related Statistical Calculators
18. Sources
- ClinicalTrials.gov: NCT01231516 — SAILING.
- Linked publication: PubMed record PMID 40990223.
Continue through Clinical Biostats
Use the trial's statistical concepts as a starting point for deeper tutorials and practical statistical calculators.
19. Record Summary
SAILING provides a clear example of how a randomized clinical trial can evaluate a binary Week 48 endpoint within a non-inferiority framework. The key statistical elements are the modified intent-to-treat exposed population, the Cochran-Mantel-Haenszel analysis, adjustment for baseline stratification factors, the risk-difference effect measure, and the prespecified −12% non-inferiority margin.
The reported adjusted difference was 7.4 percentage points, with a two-sided 95% confidence interval of 0.7 to 14.2% and P = 0.030. The lower confidence bound is the critical quantity for the stated non-inferiority criterion because it remains above the prespecified −12% margin. The interpretation should remain tied to the registered endpoint definition, the mITT-E analysis population, and the specific stratified statistical framework rather than extending the result to unreported outcomes.