← Clinical Trials
HIV-1 Infection Phase 3 Non-Inferiority NCT01449929

FLAMINGO: Complete Statistical Analysis of Dolutegravir in HIV-1 Infection

An independent statistical review of the randomized phase 3 FLAMINGO trial comparing dolutegravir with darunavir/ritonavir, each in combination with dual nucleoside reverse transcriptase inhibitors, in ART-naive participants with HIV-1 infection.

Completed  ·  Enrollment 488  ·  Primary completion April 22, 2013
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the information reported in the ClinicalTrials.gov record for FLAMINGO.

1. Trial at a Glance

FLAMINGO was a randomized, parallel, phase 3 trial comparing dolutegravir with darunavir/ritonavir, with both treatment strategies used in combination with dual NRTIs in ART-naive subjects with HIV-1 infection. The registry reports one primary binary endpoint: the percentage of participants with plasma HIV-1 RNA below 50 copies/mL at Week 48.

488
Enrollment
488 participants
2
Arms
Parallel design
7.1
Difference
DTG − DRV + RTV
0.025
P-value
Primary analysis
FeatureFLAMINGO
PhasePhase 3
ConditionInfection, Human Immunodeficiency Virus
PopulationART-naive subjects
DesignRandomized, parallel
MaskingNone
Primary purposeTreatment
Enrollment488.0
Primary endpointPercentage of participants with plasma HIV-1 RNA <50 copies/mL at Week 48
Primary endpoint typeBinary
Primary hypothesis frameworkNon-inferiority or equivalence, followed by superiority testing if non-inferiority was established
Primary analysisCochran-Mantel-Haenszel test
Effect measureRisk difference, reported as difference in percentage
Analysis populationmITT-E Population
ClinicalTrials.govNCT01449929

2. Clinical Question

The central question was whether dolutegravir 50 mg once daily, when used with dual NRTIs, produced at least comparable Week 48 virologic response to darunavir 800 mg plus ritonavir 100 mg once daily with dual NRTIs in ART-naive subjects with HIV-1 infection, and, after the non-inferiority criterion was met, whether the observed difference supported superiority.

Population

ART-naive subjects with human immunodeficiency virus infection enrolled in a phase 3 randomized trial.

Intervention

Dolutegravir 50 mg OAD in combination with dual NRTIs.

Comparator

Darunavir 800 mg OAD plus ritonavir 100 mg OAD in combination with dual NRTIs.

Primary question

What is the difference between treatment groups in the percentage of participants with plasma HIV-1 RNA <50 c/mL at Week 48?

3. Trial Design

01
Enroll488 participants
02
Randomize2 parallel arms
03
TreatDTG or DRV + RTV
04
AssessWeek 48 endpoint
05
CompareCMH risk difference
ARM A

Dolutegravir

  • Dolutegravir 50 mg OAD
  • Used in combination with dual NRTIs
ARM B

Darunavir / Ritonavir

  • Darunavir 800 mg OAD
  • Ritonavir 100 mg OAD
  • Used in combination with dual NRTIs
Allocation
Randomized allocation to two parallel treatment groups.
Masking
None; the registry classifies the trial as unmasked.
Primary purpose
Treatment.
Enrollment period
Start: October 31, 2011. Primary completion: April 22, 2013.

Statistically, this is a useful setting for a stratified comparison of a binary endpoint. The endpoint is not a time-to-event measure: each participant is classified according to whether the Week 48 virologic criterion is met under the registry's prespecified MSDF approach.

4. Endpoints

EndpointTime frameTypeAnalysis status
Percentage of Participants With Plasma Human Immunodeficiency Virus-1 (HIV-1) Ribonucleic Acid (RNA) <50 Copies/Milliliter (c/mL) Week 48 Binary Formal statistical analysis posted

Primary endpoint definition

The registered primary endpoint was the percentage of participants with plasma HIV-1 RNA <50 copies/mL at Week 48. Assessment used Missing, Switch or Discontinuation = Failure (MSDF), codified by the FDA “snapshot” algorithm.

Under this approach, participants without HIV-1 RNA data at Week 48 were treated as nonresponders. Participants who switched their concomitant ART before Week 48 were also handled according to the prespecified snapshot framework. This matters statistically because the endpoint is not simply the observed percentage among participants with a measured Week 48 value; treatment interruptions, missing measurements, and qualifying treatment changes are incorporated into the binary responder classification.

Why the endpoint definition matters: in a binary endpoint, the way missing observations and treatment changes are classified directly affects the numerator and denominator used to estimate response. The MSDF rule therefore forms part of the estimand's operational definition rather than being an incidental data-cleaning decision.

5. Statistical Methodology

Analysis population

The posted primary analysis used the mITT-E Population. The registry identifies this as the analysis population for the Week 48 comparison of dolutegravir with darunavir/ritonavir.

Cochran-Mantel-Haenszel test

The primary comparison used a Cochran-Mantel-Haenszel (CMH) analysis. The method is appropriate for a categorical outcome when the treatment comparison is evaluated while accounting for one or more stratification factors.

The analysis was described as a stratified analysis and was adjusted for two baseline stratification factors:

The important statistical idea is that the treatment comparison is not calculated as though these baseline strata did not exist. Instead, information from the strata is combined through the CMH framework to produce an adjusted comparison.

Risk difference

The registry reports the effect measure as a difference in percentage, normalized here as a risk difference. The comparison was defined as dolutegravir minus darunavir/ritonavir.

Effect measure
Risk difference = P(response | DTG) − P(response | DRV + RTV)

A positive value means the estimated percentage responding was higher in the dolutegravir group under the specified analysis.

Non-inferiority framework

The registry identifies the hypothesis type as non-inferiority or equivalence. The prespecified non-inferiority rule was based on the lower bound of a two-sided 95% confidence interval for the difference in percentages, defined as DTG minus DRV+RTV.

Non-inferiority criterion

Lower 95% CI > −12%

Non-inferiority of DTG 50 mg and DRV+RTV at Week 48 can be concluded if the lower bound of the two-sided 95% CI for the difference in percentages is greater than −12%.

This is fundamentally different from asking whether a conventional null hypothesis produces a small p-value. The non-inferiority question is whether the data are sufficiently incompatible with a loss of more than the prespecified clinically relevant margin.

Superiority after non-inferiority

The registry states that if non-inferiority was established, superiority could be tested at the nominal 5% level based on a subsequent comparison. The posted analysis has a two-sided 95% confidence interval and a p-value of 0.025.

Sequential logic matters: the analysis should be read in the order specified by the design. First, the confidence interval is compared with the non-inferiority margin. If that criterion is met, the superiority question can then be considered under the stated testing framework. A p-value by itself does not replace the non-inferiority margin.

6. Primary Result: Week 48 Virologic Response

The registry reports a formal statistical analysis for the primary endpoint: the percentage of participants with plasma HIV-1 RNA <50 copies/mL at Week 48.

Difference in percentage

7.1

DTG 50 mg QD − DRV 800 mg + RTV 100 mg QD

95% CI: 0.9 to 13.2   ·   P = 0.025

Primary endpointAnalysis detail
EndpointPercentage of participants with plasma HIV-1 RNA <50 c/mL at Week 48
Analysis populationmITT-E Population
ComparisonDTG 50 mg QD vs DRV 800 mg + RTV 100 mg QD
MethodCochran-Mantel-Haenszel
AdjustmentBaseline plasma HIV-1 RNA and baseline background dual NRTI therapy
Effect measureDifference in percentage
Estimate7.1
95% CI0.9 to 13.2
P-value0.025
Non-inferiority margin−12%
Clinical Biostats interpretation

The estimated difference in Week 48 virologic response was 7.1 percentage points, with the comparison defined as dolutegravir minus darunavir/ritonavir. Because the estimate is positive, the estimated response percentage was higher for dolutegravir in this analysis.

The estimate does not mean that every participant had a 7.1 percentage-point improvement, nor does it establish that the treatment difference is identical in every baseline subgroup. It is a population-level estimate produced by the specified stratified analysis.

The two-sided 95% confidence interval extends from 0.9 to 13.2. For the non-inferiority question, the relevant feature is the lower bound: 0.9 is greater than the prespecified −12% margin. Thus, based on the registry's stated rule, the reported interval satisfies the criterion for non-inferiority.

The p-value of 0.025 is evidence against the corresponding null comparison under the stated testing framework; it is not a measure of how large the treatment effect is. The magnitude of the estimated difference is described by 7.1, while the confidence interval describes statistical uncertainty around that estimate.

Because this is a binary endpoint analyzed with a stratified categorical-data method, the interpretation is different from a hazard ratio from a survival analysis. There is no proportional-hazards assumption involved in this primary analysis.

7. Reading the Non-Inferiority Result

Non-inferiority trials require a different statistical reading from superiority trials. The key question is not simply whether the estimated treatment difference is positive. The question is whether the data exclude a loss greater than the prespecified non-inferiority margin.

Step 1: Define the contrast

The treatment contrast is DTG minus DRV+RTV. Positive values favor the observed percentage of responders in the dolutegravir group.

Step 2: Locate the margin

The non-inferiority margin is −12%. Values below that boundary would represent a sufficiently large disadvantage to fail the stated criterion.

Step 3: Examine the CI

The two-sided 95% CI is 0.9 to 13.2. Its lower bound is above −12%.

Step 4: Consider superiority

After non-inferiority is established, the registry states that superiority can be tested at the nominal 5% level. The posted p-value is 0.025.

The key comparison
Lower confidence bound = 0.9%  >  −12% non-inferiority margin

The non-inferiority conclusion follows from the relationship between the confidence interval and the prespecified margin, not merely from the fact that the point estimate is positive.

This distinction is especially important when teaching non-inferiority designs. A conventional two-sided confidence interval can contain zero and still support non-inferiority if its lower bound remains above the negative non-inferiority margin. In FLAMINGO, the posted interval is entirely above zero, so the reported estimate also points in the positive direction.

8. Why the Cochran-Mantel-Haenszel Method Was Used

The primary endpoint is binary: each participant is classified according to whether plasma HIV-1 RNA is below 50 c/mL at Week 48 under the MSDF algorithm. The analysis also had prespecified baseline stratification factors. The CMH framework provides a natural way to compare two treatment groups across such strata.

Conceptual structure
Stratum 1 → treatment comparison
Stratum 2 → treatment comparison
⋮
Combined adjusted treatment comparison

Rather than ignoring the baseline strata, the CMH procedure combines the within-stratum information into an adjusted overall comparison.

This is one reason the reported estimate should not be reconstructed by simply taking an unadjusted percentage difference from the raw treatment groups. The registry explicitly identifies the primary method as a Cochran-Mantel-Haenszel analysis and states that the analysis was adjusted for baseline plasma HIV-1 RNA and baseline background dual NRTI therapy.

9. Stratification and Covariate Adjustment

Two baseline factors were used in the posted stratified analysis. Both are clinically relevant to the interpretation of virologic response and were defined before the Week 48 comparison.

Baseline stratification factorCategoriesStatistical role
Baseline plasma HIV-1 RNA≤100,000 c/mL vs >100,000 c/mLAdjustment stratum in the CMH analysis
Baseline background dual NRTI therapyABC/3TC vs TDF/FTCAdjustment stratum in the CMH analysis

Stratification does not mean that separate independent treatment trials were conducted within each subgroup. Instead, the strata are incorporated into a combined treatment comparison. This can improve the relevance and precision of the treatment comparison when the stratification variables are related to outcome and were used as part of the randomized design.

Adjustment is not the same as subgroup testing: the fact that baseline plasma HIV-1 RNA and background NRTI therapy were included as stratification factors does not mean that the registry is reporting separate treatment-effect conclusions for each category. The posted primary result is the adjusted overall comparison.

10. Statistical Methods Explained

Why use a Cochran-Mantel-Haenszel test?

The primary endpoint is binary and the trial incorporated baseline stratification factors. The CMH approach provides a way to compare treatment groups while accounting for those strata. It therefore matches the structure of the endpoint and the prespecified analysis.

What does a risk difference of 7.1 mean?

The reported effect measure is the difference in percentage between the dolutegravir and darunavir/ritonavir groups. A value of 7.1 means that the estimated Week 48 response percentage was 7.1 percentage points higher for dolutegravir under the posted analysis. It is an absolute difference, not a relative percentage increase.

Why does the non-inferiority margin matter more than simply asking whether P < 0.05?

A non-inferiority trial is designed around a clinically specified boundary. Here, the lower bound of the two-sided 95% confidence interval must be greater than −12%. A p-value alone does not tell you whether the treatment difference is sufficiently far from the non-inferiority boundary.

What does the confidence interval of 0.9 to 13.2 tell us?

It describes the statistical uncertainty around the estimated difference of 7.1 under the analysis framework. It does not describe the range of individual participant responses. For the non-inferiority decision, its lower limit is particularly important because 0.9 is above −12%.

Why are the baseline HIV-1 RNA and NRTI categories included in the analysis?

The registry states that the analysis was adjusted for these baseline stratification factors. The CMH method combines treatment information across the defined strata rather than treating all participants as if the stratification structure did not exist.

What does the p-value of 0.025 mean?

The p-value quantifies the evidence against the relevant null comparison under the specified statistical framework. It does not measure the size of the treatment effect. The effect estimate is 7.1, while the 95% confidence interval is 0.9 to 13.2.

Why does MSDF matter statistically?

The MSDF snapshot algorithm classifies participants without Week 48 HIV-1 RNA data as nonresponders and specifies how treatment switches are handled. Consequently, missingness and treatment changes affect the binary endpoint classification rather than simply disappearing from the analysis.

11. Results and Statistical Interpretation

The registry provides one formal statistical analysis, corresponding to the single registered primary endpoint. It does not provide additional formal statistical analyses in the ClinicalTrials.gov record for the other 15 posted outcome measures. Accordingly, this page does not manufacture secondary efficacy estimates or p-values that are not contained in the trial data.

Result componentReported valueStatistical meaning
Point estimate7.1Estimated difference in percentage, DTG − DRV+RTV
95% CI lower bound0.9Above the −12% non-inferiority margin
95% CI upper bound13.2Upper uncertainty bound for the estimated difference
P-value0.025Reported evidence for the formal comparison
Analysis methodCochran-Mantel-HaenszelStratified categorical-data comparison
Clinical Biostats interpretation

The most important feature of the result is the alignment between the effect estimate, its confidence interval, and the prespecified non-inferiority margin. The estimated difference is positive at 7.1, and the entire reported 95% confidence interval, 0.9 to 13.2, lies above the −12% non-inferiority boundary.

That means the data are not compatible, under this analysis framework, with a treatment difference as unfavorable to dolutegravir as the prespecified non-inferiority limit. The result also has a positive point estimate, but the inferential logic should still begin with the prespecified non-inferiority question.

The result should not be interpreted as a guarantee of benefit for every participant, nor as proof that the exact 7.1-point difference would recur in another population. The confidence interval reflects uncertainty in the estimated population-level comparison.

The analysis is also conditional on the mITT-E population, the MSDF endpoint definition, and the specified baseline stratification factors. Changing any of those analytic choices could change the numerical result.

12. Safety Results

The ClinicalTrials.gov record includes serious adverse-event counts by arm. These are presented separately from the primary virologic efficacy analysis because safety events answer a different statistical question from Week 48 virologic response.

GroupSerious adverse eventsAt riskReported affected / at risk
DTG 50 mg QD3624236/242
DRV 800 mg + RTV 100 mg QD2124221/242
Extension DTG 50 mg41234/123

The ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse-event rates. Therefore, these counts should be read descriptively rather than converted into an unreported p-value, confidence interval, risk ratio, or risk difference.

Important denominator distinction: the extension DTG 50 mg group is presented separately from the two principal randomized treatment groups. Its 4/123 count should not be combined with the 36/242 DTG randomized-arm count to create an overall DTG rate without additional information about the relationship between the populations.

13. What Is and Is Not Reported

The ClinicalTrials.gov record reports 16 outcome measures and one formal statistical analysis. Only one primary endpoint has a posted estimate, confidence interval, and p-value in the ClinicalTrials.gov record.

Reported formally

Week 48 plasma HIV-1 RNA <50 c/mL, analyzed in the mITT-E population using a CMH method with a reported difference in percentage, 95% CI, and p-value.

Not formally reported here

The ClinicalTrials.gov record does not provide formal statistical analyses for the remaining posted outcome measures.

Not added

No median outcomes, subgroup estimates, additional p-values, response counts, or derived safety comparisons have been introduced from outside the ClinicalTrials.gov record.

Why this matters

A complete statistical analysis should distinguish between what the registry reports and what would merely be plausible to calculate from other sources.

14. Missing Data and the MSDF Algorithm

The Week 48 endpoint uses Missing, Switch or Discontinuation = Failure. This is a particularly important feature of the statistical definition because it converts several forms of incomplete outcome information into a prespecified binary classification.

SituationRegistry-defined treatment for the primary endpoint
No HIV-1 RNA data at Week 48Treated as a nonresponder
Switch in concomitant ART before Week 48Handled according to the prespecified snapshot/MSDF algorithm
Observed Week 48 HIV-1 RNA <50 c/mLMeets the virologic response criterion

This approach is different from simply deleting participants with missing Week 48 values. Complete-case analysis could change the population being compared and potentially introduce bias if missingness is related to treatment or outcome. The MSDF rule instead specifies in advance how missing and treatment-switch situations affect the primary binary outcome.

Interpretation caution: calling missing observations failures is itself an analytic choice. It does not prove that every missing participant would have failed virologically; rather, it defines the outcome classification used for the prespecified trial analysis.

15. Non-Inferiority vs Superiority

FLAMINGO is particularly useful for understanding why a non-inferiority design should not be reduced to a conventional superiority test.

QuestionStatistical criterion supported by the ClinicalTrials.gov record
Is dolutegravir non-inferior?Lower bound of the two-sided 95% CI must be greater than −12%.
What was the reported lower bound?0.9.
Does 0.9 exceed −12%?Yes.
Could superiority then be tested?The registry states that superiority can be tested at the nominal 5% level if non-inferiority is established.
Reported p-value0.025.

The ordering is important. Suppose a point estimate favored the new treatment but the confidence interval extended well below the non-inferiority margin. A positive point estimate alone would not establish non-inferiority. The confidence interval must demonstrate that sufficiently poor outcomes have been excluded relative to the prespecified boundary.

16. Confidence Intervals and Precision

The reported 95% confidence interval is 0.9 to 13.2 around the estimated difference of 7.1.

Primary effect estimate and confidence interval
Lower bound
0.9
Estimate
7.1
Upper bound
13.2

The confidence interval gives two pieces of information simultaneously. First, its location relative to the non-inferiority margin determines whether the non-inferiority criterion is satisfied. Second, its width communicates the precision of the estimated treatment difference.

Why the confidence interval matters

The interval of 0.9 to 13.2 does not say that individual treatment effects vary only between 0.9 and 13.2 percentage points. It describes uncertainty around the estimated population-level treatment difference under the specified analysis. It is therefore an inferential interval, not an individual-patient prediction interval.

17. P-values and Effect Size

The primary analysis reports P = 0.025. It is useful to keep the p-value separate from the effect estimate.

Effect size

The treatment difference is 7.1 percentage points. This describes the estimated magnitude and direction of the observed comparison.

Uncertainty

The 95% CI of 0.9 to 13.2 describes uncertainty around that estimated difference.

Evidence

The p-value of 0.025 quantifies evidence under the specified hypothesis-testing framework.

Non-inferiority

The decisive design feature for the non-inferiority question is the position of the confidence interval relative to −12%.

These are complementary quantities. A p-value cannot substitute for an effect estimate, and an effect estimate without an uncertainty interval is incomplete for inferential interpretation.

18. Design Features That Are Not Part of the Posted Primary Analysis

The ClinicalTrials.gov record identifies the design as randomized and parallel, with no masking, and provide a single posted primary statistical analysis. They do not support adding a separate analysis of crossover, Bayesian methods, interim efficacy monitoring, or a multiplicity-adjustment scheme beyond the non-inferiority/superiority sequence described in the statistical-analysis record.

Design topicWhat the ClinicalTrials.gov record supports
RandomizationYes — allocation is randomized.
Parallel designYes — design model is parallel.
MaskingNo masking reported.
Non-inferiorityYes — −12% margin is reported.
StratificationYes — two baseline stratification factors are identified.
Covariate adjustmentYes — analysis adjusted for the two baseline stratification factors.
Missing-data ruleYes — MSDF snapshot algorithm.
CrossoverNot reported in the ClinicalTrials.gov record.
Bayesian methodsNot reported in the ClinicalTrials.gov record.
Interim analysisNot reported in the ClinicalTrials.gov record.

19. Limitations and Interpretation Issues

20. Why This Trial Matters Statistically

FLAMINGO is a compact teaching example of how a binary clinical endpoint can be analyzed within a randomized non-inferiority framework. The statistical story is driven less by a complicated model than by the relationship among the endpoint definition, stratification factors, effect measure, confidence interval, and non-inferiority margin.

ConceptHow it appears in FLAMINGO
RandomizationParticipants were randomized to two parallel treatment strategies.
Binary endpointWeek 48 HIV-1 RNA <50 c/mL is classified as a binary response outcome.
Risk differenceThe treatment effect is reported as a difference in percentage.
Cochran-Mantel-Haenszel testThe primary treatment comparison uses a CMH method.
Stratified analysisThe analysis accounts for baseline HIV-1 RNA and background dual NRTI therapy.
Non-inferiority marginThe lower bound of the two-sided 95% CI must exceed −12%.
Confidence intervalThe reported interval is 0.9 to 13.2 around an estimate of 7.1.
P-valueThe primary analysis reports P = 0.025.
Missing-data ruleMSDF treats participants without Week 48 HIV-1 RNA data as nonresponders.
Safety analysisSerious adverse events are available descriptively by reported group.

The trial therefore illustrates a general principle in clinical biostatistics: the method cannot be interpreted independently of the estimand and decision rule. A difference in percentages, a confidence interval, and a p-value become clinically interpretable only when their direction, population, endpoint definition, stratification, and prespecified non-inferiority boundary are understood.

21. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The posted primary analysis estimated a difference of 7.1 percentage points between dolutegravir and darunavir/ritonavir, with a two-sided 95% CI of 0.9 to 13.2 and P = 0.025. The lower confidence bound is above the prespecified −12% non-inferiority margin.

Clinical interpretation

The result concerns the percentage of participants meeting the Week 48 virologic response definition. It does not by itself describe every participant's experience, long-term outcomes, or comparative effects on endpoints not formally analyzed in the ClinicalTrials.gov record.

Keeping these interpretations separate is useful. Statistical evidence addresses the uncertainty around the prespecified comparison; clinical interpretation asks what that endpoint and treatment difference mean in the context of the disease and treatment strategy. This page restricts the latter to conclusions directly supported by the ClinicalTrials.gov record.

22. Related Tutorials

Learn more about the methods used in this trial:

23. Related Calculators

24. Sources

Continue with Clinical Biostats statistical methods

Explore tutorials and calculators covering the categorical-data, confidence-interval, stratification, risk-difference, and non-inferiority concepts illustrated by FLAMINGO.

25. Record Summary

FLAMINGO provides a focused example of randomized clinical-trial inference for a binary endpoint under a non-inferiority framework. The trial randomized 488 participants in a two-arm parallel design and evaluated the percentage with plasma HIV-1 RNA <50 copies/mL at Week 48. The formal analysis used a Cochran-Mantel-Haenszel method adjusted for baseline plasma HIV-1 RNA and background dual NRTI therapy, with the treatment effect expressed as a difference in percentage.

The reported estimate was 7.1, with a two-sided 95% confidence interval of 0.9 to 13.2 and P = 0.025. Because the lower confidence bound of 0.9 is above the prespecified −12% non-inferiority margin, the posted result satisfies the registry's stated non-inferiority criterion. The registry further states that superiority could be tested at the nominal 5% level after non-inferiority was established.

The statistical lesson extends beyond this individual trial. For non-inferiority studies, the treatment effect cannot be interpreted correctly without its direction, confidence interval, analysis population, endpoint definition, stratification structure, and prespecified margin. FLAMINGO illustrates how these elements fit together in a practical randomized comparison.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. Where the ClinicalTrials.gov record provides a formal estimate and confidence interval, the analysis explains their meaning; where formal secondary analyses are not reported, it does not manufacture them.