This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the information reported in the ClinicalTrials.gov record for FLAMINGO.
1. Trial at a Glance
FLAMINGO was a randomized, parallel, phase 3 trial comparing dolutegravir with darunavir/ritonavir, with both treatment strategies used in combination with dual NRTIs in ART-naive subjects with HIV-1 infection. The registry reports one primary binary endpoint: the percentage of participants with plasma HIV-1 RNA below 50 copies/mL at Week 48.
| Feature | FLAMINGO |
|---|---|
| Phase | Phase 3 |
| Condition | Infection, Human Immunodeficiency Virus |
| Population | ART-naive subjects |
| Design | Randomized, parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 488.0 |
| Primary endpoint | Percentage of participants with plasma HIV-1 RNA <50 copies/mL at Week 48 |
| Primary endpoint type | Binary |
| Primary hypothesis framework | Non-inferiority or equivalence, followed by superiority testing if non-inferiority was established |
| Primary analysis | Cochran-Mantel-Haenszel test |
| Effect measure | Risk difference, reported as difference in percentage |
| Analysis population | mITT-E Population |
| ClinicalTrials.gov | NCT01449929 |
2. Clinical Question
The central question was whether dolutegravir 50 mg once daily, when used with dual NRTIs, produced at least comparable Week 48 virologic response to darunavir 800 mg plus ritonavir 100 mg once daily with dual NRTIs in ART-naive subjects with HIV-1 infection, and, after the non-inferiority criterion was met, whether the observed difference supported superiority.
Population
ART-naive subjects with human immunodeficiency virus infection enrolled in a phase 3 randomized trial.
Intervention
Dolutegravir 50 mg OAD in combination with dual NRTIs.
Comparator
Darunavir 800 mg OAD plus ritonavir 100 mg OAD in combination with dual NRTIs.
Primary question
What is the difference between treatment groups in the percentage of participants with plasma HIV-1 RNA <50 c/mL at Week 48?
3. Trial Design
Dolutegravir
- Dolutegravir 50 mg OAD
- Used in combination with dual NRTIs
Darunavir / Ritonavir
- Darunavir 800 mg OAD
- Ritonavir 100 mg OAD
- Used in combination with dual NRTIs
Statistically, this is a useful setting for a stratified comparison of a binary endpoint. The endpoint is not a time-to-event measure: each participant is classified according to whether the Week 48 virologic criterion is met under the registry's prespecified MSDF approach.
4. Endpoints
| Endpoint | Time frame | Type | Analysis status |
|---|---|---|---|
| Percentage of Participants With Plasma Human Immunodeficiency Virus-1 (HIV-1) Ribonucleic Acid (RNA) <50 Copies/Milliliter (c/mL) | Week 48 | Binary | Formal statistical analysis posted |
Primary endpoint definition
The registered primary endpoint was the percentage of participants with plasma HIV-1 RNA <50 copies/mL at Week 48. Assessment used Missing, Switch or Discontinuation = Failure (MSDF), codified by the FDA “snapshot” algorithm.
Under this approach, participants without HIV-1 RNA data at Week 48 were treated as nonresponders. Participants who switched their concomitant ART before Week 48 were also handled according to the prespecified snapshot framework. This matters statistically because the endpoint is not simply the observed percentage among participants with a measured Week 48 value; treatment interruptions, missing measurements, and qualifying treatment changes are incorporated into the binary responder classification.
5. Statistical Methodology
Analysis population
The posted primary analysis used the mITT-E Population. The registry identifies this as the analysis population for the Week 48 comparison of dolutegravir with darunavir/ritonavir.
Cochran-Mantel-Haenszel test
The primary comparison used a Cochran-Mantel-Haenszel (CMH) analysis. The method is appropriate for a categorical outcome when the treatment comparison is evaluated while accounting for one or more stratification factors.
The analysis was described as a stratified analysis and was adjusted for two baseline stratification factors:
- Baseline plasma HIV-1 RNA: ≤100,000 c/mL versus >100,000 c/mL.
- Baseline background dual NRTI therapy: ABC/3TC versus TDF/FTC.
The important statistical idea is that the treatment comparison is not calculated as though these baseline strata did not exist. Instead, information from the strata is combined through the CMH framework to produce an adjusted comparison.
Risk difference
The registry reports the effect measure as a difference in percentage, normalized here as a risk difference. The comparison was defined as dolutegravir minus darunavir/ritonavir.
A positive value means the estimated percentage responding was higher in the dolutegravir group under the specified analysis.
Non-inferiority framework
The registry identifies the hypothesis type as non-inferiority or equivalence. The prespecified non-inferiority rule was based on the lower bound of a two-sided 95% confidence interval for the difference in percentages, defined as DTG minus DRV+RTV.
Non-inferiority criterion
Non-inferiority of DTG 50 mg and DRV+RTV at Week 48 can be concluded if the lower bound of the two-sided 95% CI for the difference in percentages is greater than −12%.
This is fundamentally different from asking whether a conventional null hypothesis produces a small p-value. The non-inferiority question is whether the data are sufficiently incompatible with a loss of more than the prespecified clinically relevant margin.
Superiority after non-inferiority
The registry states that if non-inferiority was established, superiority could be tested at the nominal 5% level based on a subsequent comparison. The posted analysis has a two-sided 95% confidence interval and a p-value of 0.025.
6. Primary Result: Week 48 Virologic Response
The registry reports a formal statistical analysis for the primary endpoint: the percentage of participants with plasma HIV-1 RNA <50 copies/mL at Week 48.
Difference in percentage
DTG 50 mg QD − DRV 800 mg + RTV 100 mg QD
95% CI: 0.9 to 13.2 · P = 0.025
| Primary endpoint | Analysis detail |
|---|---|
| Endpoint | Percentage of participants with plasma HIV-1 RNA <50 c/mL at Week 48 |
| Analysis population | mITT-E Population |
| Comparison | DTG 50 mg QD vs DRV 800 mg + RTV 100 mg QD |
| Method | Cochran-Mantel-Haenszel |
| Adjustment | Baseline plasma HIV-1 RNA and baseline background dual NRTI therapy |
| Effect measure | Difference in percentage |
| Estimate | 7.1 |
| 95% CI | 0.9 to 13.2 |
| P-value | 0.025 |
| Non-inferiority margin | −12% |
The estimated difference in Week 48 virologic response was 7.1 percentage points, with the comparison defined as dolutegravir minus darunavir/ritonavir. Because the estimate is positive, the estimated response percentage was higher for dolutegravir in this analysis.
The estimate does not mean that every participant had a 7.1 percentage-point improvement, nor does it establish that the treatment difference is identical in every baseline subgroup. It is a population-level estimate produced by the specified stratified analysis.
The two-sided 95% confidence interval extends from 0.9 to 13.2. For the non-inferiority question, the relevant feature is the lower bound: 0.9 is greater than the prespecified −12% margin. Thus, based on the registry's stated rule, the reported interval satisfies the criterion for non-inferiority.
The p-value of 0.025 is evidence against the corresponding null comparison under the stated testing framework; it is not a measure of how large the treatment effect is. The magnitude of the estimated difference is described by 7.1, while the confidence interval describes statistical uncertainty around that estimate.
Because this is a binary endpoint analyzed with a stratified categorical-data method, the interpretation is different from a hazard ratio from a survival analysis. There is no proportional-hazards assumption involved in this primary analysis.
7. Reading the Non-Inferiority Result
Non-inferiority trials require a different statistical reading from superiority trials. The key question is not simply whether the estimated treatment difference is positive. The question is whether the data exclude a loss greater than the prespecified non-inferiority margin.
Step 1: Define the contrast
The treatment contrast is DTG minus DRV+RTV. Positive values favor the observed percentage of responders in the dolutegravir group.
Step 2: Locate the margin
The non-inferiority margin is −12%. Values below that boundary would represent a sufficiently large disadvantage to fail the stated criterion.
Step 3: Examine the CI
The two-sided 95% CI is 0.9 to 13.2. Its lower bound is above −12%.
Step 4: Consider superiority
After non-inferiority is established, the registry states that superiority can be tested at the nominal 5% level. The posted p-value is 0.025.
The non-inferiority conclusion follows from the relationship between the confidence interval and the prespecified margin, not merely from the fact that the point estimate is positive.
This distinction is especially important when teaching non-inferiority designs. A conventional two-sided confidence interval can contain zero and still support non-inferiority if its lower bound remains above the negative non-inferiority margin. In FLAMINGO, the posted interval is entirely above zero, so the reported estimate also points in the positive direction.
8. Why the Cochran-Mantel-Haenszel Method Was Used
The primary endpoint is binary: each participant is classified according to whether plasma HIV-1 RNA is below 50 c/mL at Week 48 under the MSDF algorithm. The analysis also had prespecified baseline stratification factors. The CMH framework provides a natural way to compare two treatment groups across such strata.
Stratum 2 → treatment comparison
⋮
Combined adjusted treatment comparison
Rather than ignoring the baseline strata, the CMH procedure combines the within-stratum information into an adjusted overall comparison.
This is one reason the reported estimate should not be reconstructed by simply taking an unadjusted percentage difference from the raw treatment groups. The registry explicitly identifies the primary method as a Cochran-Mantel-Haenszel analysis and states that the analysis was adjusted for baseline plasma HIV-1 RNA and baseline background dual NRTI therapy.
9. Stratification and Covariate Adjustment
Two baseline factors were used in the posted stratified analysis. Both are clinically relevant to the interpretation of virologic response and were defined before the Week 48 comparison.
| Baseline stratification factor | Categories | Statistical role |
|---|---|---|
| Baseline plasma HIV-1 RNA | ≤100,000 c/mL vs >100,000 c/mL | Adjustment stratum in the CMH analysis |
| Baseline background dual NRTI therapy | ABC/3TC vs TDF/FTC | Adjustment stratum in the CMH analysis |
Stratification does not mean that separate independent treatment trials were conducted within each subgroup. Instead, the strata are incorporated into a combined treatment comparison. This can improve the relevance and precision of the treatment comparison when the stratification variables are related to outcome and were used as part of the randomized design.
10. Statistical Methods Explained
Why use a Cochran-Mantel-Haenszel test?
The primary endpoint is binary and the trial incorporated baseline stratification factors. The CMH approach provides a way to compare treatment groups while accounting for those strata. It therefore matches the structure of the endpoint and the prespecified analysis.
What does a risk difference of 7.1 mean?
The reported effect measure is the difference in percentage between the dolutegravir and darunavir/ritonavir groups. A value of 7.1 means that the estimated Week 48 response percentage was 7.1 percentage points higher for dolutegravir under the posted analysis. It is an absolute difference, not a relative percentage increase.
Why does the non-inferiority margin matter more than simply asking whether P < 0.05?
A non-inferiority trial is designed around a clinically specified boundary. Here, the lower bound of the two-sided 95% confidence interval must be greater than −12%. A p-value alone does not tell you whether the treatment difference is sufficiently far from the non-inferiority boundary.
What does the confidence interval of 0.9 to 13.2 tell us?
It describes the statistical uncertainty around the estimated difference of 7.1 under the analysis framework. It does not describe the range of individual participant responses. For the non-inferiority decision, its lower limit is particularly important because 0.9 is above −12%.
Why are the baseline HIV-1 RNA and NRTI categories included in the analysis?
The registry states that the analysis was adjusted for these baseline stratification factors. The CMH method combines treatment information across the defined strata rather than treating all participants as if the stratification structure did not exist.
What does the p-value of 0.025 mean?
The p-value quantifies the evidence against the relevant null comparison under the specified statistical framework. It does not measure the size of the treatment effect. The effect estimate is 7.1, while the 95% confidence interval is 0.9 to 13.2.
Why does MSDF matter statistically?
The MSDF snapshot algorithm classifies participants without Week 48 HIV-1 RNA data as nonresponders and specifies how treatment switches are handled. Consequently, missingness and treatment changes affect the binary endpoint classification rather than simply disappearing from the analysis.
11. Results and Statistical Interpretation
The registry provides one formal statistical analysis, corresponding to the single registered primary endpoint. It does not provide additional formal statistical analyses in the ClinicalTrials.gov record for the other 15 posted outcome measures. Accordingly, this page does not manufacture secondary efficacy estimates or p-values that are not contained in the trial data.
| Result component | Reported value | Statistical meaning |
|---|---|---|
| Point estimate | 7.1 | Estimated difference in percentage, DTG − DRV+RTV |
| 95% CI lower bound | 0.9 | Above the −12% non-inferiority margin |
| 95% CI upper bound | 13.2 | Upper uncertainty bound for the estimated difference |
| P-value | 0.025 | Reported evidence for the formal comparison |
| Analysis method | Cochran-Mantel-Haenszel | Stratified categorical-data comparison |
The most important feature of the result is the alignment between the effect estimate, its confidence interval, and the prespecified non-inferiority margin. The estimated difference is positive at 7.1, and the entire reported 95% confidence interval, 0.9 to 13.2, lies above the −12% non-inferiority boundary.
That means the data are not compatible, under this analysis framework, with a treatment difference as unfavorable to dolutegravir as the prespecified non-inferiority limit. The result also has a positive point estimate, but the inferential logic should still begin with the prespecified non-inferiority question.
The result should not be interpreted as a guarantee of benefit for every participant, nor as proof that the exact 7.1-point difference would recur in another population. The confidence interval reflects uncertainty in the estimated population-level comparison.
The analysis is also conditional on the mITT-E population, the MSDF endpoint definition, and the specified baseline stratification factors. Changing any of those analytic choices could change the numerical result.
12. Safety Results
The ClinicalTrials.gov record includes serious adverse-event counts by arm. These are presented separately from the primary virologic efficacy analysis because safety events answer a different statistical question from Week 48 virologic response.
| Group | Serious adverse events | At risk | Reported affected / at risk |
|---|---|---|---|
| DTG 50 mg QD | 36 | 242 | 36/242 |
| DRV 800 mg + RTV 100 mg QD | 21 | 242 | 21/242 |
| Extension DTG 50 mg | 4 | 123 | 4/123 |
The ClinicalTrials.gov record does not provide a formal statistical comparison of serious adverse-event rates. Therefore, these counts should be read descriptively rather than converted into an unreported p-value, confidence interval, risk ratio, or risk difference.
13. What Is and Is Not Reported
The ClinicalTrials.gov record reports 16 outcome measures and one formal statistical analysis. Only one primary endpoint has a posted estimate, confidence interval, and p-value in the ClinicalTrials.gov record.
Reported formally
Week 48 plasma HIV-1 RNA <50 c/mL, analyzed in the mITT-E population using a CMH method with a reported difference in percentage, 95% CI, and p-value.
Not formally reported here
The ClinicalTrials.gov record does not provide formal statistical analyses for the remaining posted outcome measures.
Not added
No median outcomes, subgroup estimates, additional p-values, response counts, or derived safety comparisons have been introduced from outside the ClinicalTrials.gov record.
Why this matters
A complete statistical analysis should distinguish between what the registry reports and what would merely be plausible to calculate from other sources.
14. Missing Data and the MSDF Algorithm
The Week 48 endpoint uses Missing, Switch or Discontinuation = Failure. This is a particularly important feature of the statistical definition because it converts several forms of incomplete outcome information into a prespecified binary classification.
| Situation | Registry-defined treatment for the primary endpoint |
|---|---|
| No HIV-1 RNA data at Week 48 | Treated as a nonresponder |
| Switch in concomitant ART before Week 48 | Handled according to the prespecified snapshot/MSDF algorithm |
| Observed Week 48 HIV-1 RNA <50 c/mL | Meets the virologic response criterion |
This approach is different from simply deleting participants with missing Week 48 values. Complete-case analysis could change the population being compared and potentially introduce bias if missingness is related to treatment or outcome. The MSDF rule instead specifies in advance how missing and treatment-switch situations affect the primary binary outcome.
15. Non-Inferiority vs Superiority
FLAMINGO is particularly useful for understanding why a non-inferiority design should not be reduced to a conventional superiority test.
| Question | Statistical criterion supported by the ClinicalTrials.gov record |
|---|---|
| Is dolutegravir non-inferior? | Lower bound of the two-sided 95% CI must be greater than −12%. |
| What was the reported lower bound? | 0.9. |
| Does 0.9 exceed −12%? | Yes. |
| Could superiority then be tested? | The registry states that superiority can be tested at the nominal 5% level if non-inferiority is established. |
| Reported p-value | 0.025. |
The ordering is important. Suppose a point estimate favored the new treatment but the confidence interval extended well below the non-inferiority margin. A positive point estimate alone would not establish non-inferiority. The confidence interval must demonstrate that sufficiently poor outcomes have been excluded relative to the prespecified boundary.
16. Confidence Intervals and Precision
The reported 95% confidence interval is 0.9 to 13.2 around the estimated difference of 7.1.
The confidence interval gives two pieces of information simultaneously. First, its location relative to the non-inferiority margin determines whether the non-inferiority criterion is satisfied. Second, its width communicates the precision of the estimated treatment difference.
The interval of 0.9 to 13.2 does not say that individual treatment effects vary only between 0.9 and 13.2 percentage points. It describes uncertainty around the estimated population-level treatment difference under the specified analysis. It is therefore an inferential interval, not an individual-patient prediction interval.
17. P-values and Effect Size
The primary analysis reports P = 0.025. It is useful to keep the p-value separate from the effect estimate.
Effect size
The treatment difference is 7.1 percentage points. This describes the estimated magnitude and direction of the observed comparison.
Uncertainty
The 95% CI of 0.9 to 13.2 describes uncertainty around that estimated difference.
Evidence
The p-value of 0.025 quantifies evidence under the specified hypothesis-testing framework.
Non-inferiority
The decisive design feature for the non-inferiority question is the position of the confidence interval relative to −12%.
These are complementary quantities. A p-value cannot substitute for an effect estimate, and an effect estimate without an uncertainty interval is incomplete for inferential interpretation.
18. Design Features That Are Not Part of the Posted Primary Analysis
The ClinicalTrials.gov record identifies the design as randomized and parallel, with no masking, and provide a single posted primary statistical analysis. They do not support adding a separate analysis of crossover, Bayesian methods, interim efficacy monitoring, or a multiplicity-adjustment scheme beyond the non-inferiority/superiority sequence described in the statistical-analysis record.
| Design topic | What the ClinicalTrials.gov record supports |
|---|---|
| Randomization | Yes — allocation is randomized. |
| Parallel design | Yes — design model is parallel. |
| Masking | No masking reported. |
| Non-inferiority | Yes — −12% margin is reported. |
| Stratification | Yes — two baseline stratification factors are identified. |
| Covariate adjustment | Yes — analysis adjusted for the two baseline stratification factors. |
| Missing-data rule | Yes — MSDF snapshot algorithm. |
| Crossover | Not reported in the ClinicalTrials.gov record. |
| Bayesian methods | Not reported in the ClinicalTrials.gov record. |
| Interim analysis | Not reported in the ClinicalTrials.gov record. |
19. Limitations and Interpretation Issues
- Single formal analysis reported: although 16 outcome measures are posted, the ClinicalTrials.gov record contains only one statistical analysis with an estimate, confidence interval, and p-value. Additional formal efficacy comparisons should not be inferred.
- Binary endpoint definition: the Week 48 endpoint depends on the MSDF snapshot algorithm, including its treatment of missing Week 48 measurements and treatment changes.
- Non-inferiority interpretation: the −12% margin is central to the design. A conventional superiority interpretation alone would not capture the primary non-inferiority question.
- Analysis population: the posted formal analysis uses the mITT-E Population. The result therefore should not automatically be generalized to another analysis population.
- Stratification: the estimate is adjusted for baseline plasma HIV-1 RNA and background dual NRTI therapy. It should not be recreated as a simple unadjusted percentage difference.
- Safety comparison: serious adverse-event counts are reported, but no formal comparative safety analysis is reported. Descriptive counts should not be converted into an unreported statistical test.
- Extension population: the 4/123 serious-adverse-event count in the extension DTG group has a different denominator from the principal randomized arms and should not be combined with them without additional population information.
- Unmasked design: the registry classifies the study as having no masking. This is relevant when considering potential sources of bias, particularly for outcomes susceptible to assessment or behavior effects.
20. Why This Trial Matters Statistically
FLAMINGO is a compact teaching example of how a binary clinical endpoint can be analyzed within a randomized non-inferiority framework. The statistical story is driven less by a complicated model than by the relationship among the endpoint definition, stratification factors, effect measure, confidence interval, and non-inferiority margin.
| Concept | How it appears in FLAMINGO |
|---|---|
| Randomization | Participants were randomized to two parallel treatment strategies. |
| Binary endpoint | Week 48 HIV-1 RNA <50 c/mL is classified as a binary response outcome. |
| Risk difference | The treatment effect is reported as a difference in percentage. |
| Cochran-Mantel-Haenszel test | The primary treatment comparison uses a CMH method. |
| Stratified analysis | The analysis accounts for baseline HIV-1 RNA and background dual NRTI therapy. |
| Non-inferiority margin | The lower bound of the two-sided 95% CI must exceed −12%. |
| Confidence interval | The reported interval is 0.9 to 13.2 around an estimate of 7.1. |
| P-value | The primary analysis reports P = 0.025. |
| Missing-data rule | MSDF treats participants without Week 48 HIV-1 RNA data as nonresponders. |
| Safety analysis | Serious adverse events are available descriptively by reported group. |
The trial therefore illustrates a general principle in clinical biostatistics: the method cannot be interpreted independently of the estimand and decision rule. A difference in percentages, a confidence interval, and a p-value become clinically interpretable only when their direction, population, endpoint definition, stratification, and prespecified non-inferiority boundary are understood.
21. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The posted primary analysis estimated a difference of 7.1 percentage points between dolutegravir and darunavir/ritonavir, with a two-sided 95% CI of 0.9 to 13.2 and P = 0.025. The lower confidence bound is above the prespecified −12% non-inferiority margin.
Clinical interpretation
The result concerns the percentage of participants meeting the Week 48 virologic response definition. It does not by itself describe every participant's experience, long-term outcomes, or comparative effects on endpoints not formally analyzed in the ClinicalTrials.gov record.
Keeping these interpretations separate is useful. Statistical evidence addresses the uncertainty around the prespecified comparison; clinical interpretation asks what that endpoint and treatment difference mean in the context of the disease and treatment strategy. This page restricts the latter to conclusions directly supported by the ClinicalTrials.gov record.
22. Related Tutorials
Learn more about the methods used in this trial:
23. Related Calculators
24. Sources
- ClinicalTrials.gov: FLAMINGO, NCT01449929.
- Linked publication: PubMed record PMID 26424673.
- Linked publication: PubMed record PMID 24698485.
Continue with Clinical Biostats statistical methods
Explore tutorials and calculators covering the categorical-data, confidence-interval, stratification, risk-difference, and non-inferiority concepts illustrated by FLAMINGO.
25. Record Summary
FLAMINGO provides a focused example of randomized clinical-trial inference for a binary endpoint under a non-inferiority framework. The trial randomized 488 participants in a two-arm parallel design and evaluated the percentage with plasma HIV-1 RNA <50 copies/mL at Week 48. The formal analysis used a Cochran-Mantel-Haenszel method adjusted for baseline plasma HIV-1 RNA and background dual NRTI therapy, with the treatment effect expressed as a difference in percentage.
The reported estimate was 7.1, with a two-sided 95% confidence interval of 0.9 to 13.2 and P = 0.025. Because the lower confidence bound of 0.9 is above the prespecified −12% non-inferiority margin, the posted result satisfies the registry's stated non-inferiority criterion. The registry further states that superiority could be tested at the nominal 5% level after non-inferiority was established.
The statistical lesson extends beyond this individual trial. For non-inferiority studies, the treatment effect cannot be interpreted correctly without its direction, confidence interval, analysis population, endpoint definition, stratification structure, and prespecified margin. FLAMINGO illustrates how these elements fit together in a practical randomized comparison.