This page separates reported trial results from statistical interpretation. Numerical results are restricted to the ClinicalTrials.gov trial data posted on ClinicalTrials.gov for TBTC Study 31. The registry provides formal analyses for the primary endpoints and selected prespecified sensitivity and secondary analyses.
1. Trial at a Glance
TBTC Study 31 was a randomized, parallel, unmasked phase 3 treatment trial in tuberculosis. The study enrolled 2516 participants and compared a control regimen with two rifapentine-containing experimental regimens, with primary efficacy assessed as TB disease-free survival at 12 months and primary safety assessed by grade 3 or higher adverse events during study drug treatment.
| Feature | TBTC Study 31 |
|---|---|
| Trial name | TBTC Study 31 |
| Brief title | TBTC Study 31: Rifapentine-containing Tuberculosis Treatment Shortening Regimens |
| Phase | Phase 3 |
| Condition | Tuberculosis |
| Allocation | Randomized |
| Design | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 2516 |
| Arms | 3 |
| Lead sponsor | Centers for Disease Control and Prevention |
| Sponsor type | FED |
| Study dates | Start: 2016-01-25; primary completion: 2020-07 |
| Status | Completed |
| ClinicalTrials.gov | NCT02410772 |
2. Clinical Question
The trial evaluates whether rifapentine-containing regimens can provide an effective shorter treatment approach for tuberculosis while maintaining an acceptable safety profile. The registry describes two experimental regimens and a control regimen and specifies non-inferiority as the hypothesis type for the posted primary comparisons.
Population
Participants enrolled in a phase 3 treatment trial for tuberculosis. The primary efficacy analyses used microbiologically eligible populations, with additional assessable-population analyses.
Experimental regimens
Regimen 2: 2HPZ/2HP. Regimen 3: 2HPZM/2HPM. Both are rifapentine-containing experimental regimens.
Control regimen
Regimen 1: 2HRZE/4HR.
Primary question
Can the experimental rifapentine-containing regimens be evaluated as non-inferior to the control regimen for TB disease-free survival, while also meeting the prespecified safety objective?
3. Trial Design
2HRZE/4HR
- Control regimen.
- Compared with Regimen 3 first for the primary efficacy and safety analyses described in the registry.
- Also compared with Regimen 2 after the prespecified sequential non-inferiority framework.
2HPZ/2HP
- Rifapentine-containing experimental regimen.
- Evaluated against Regimen 1 for TB disease-free survival and safety.
- Also evaluated at 18 months in secondary analyses.
2HPZM/2HPM
- Rifapentine-containing experimental regimen including moxifloxacin.
- Evaluated against Regimen 1 for TB disease-free survival and safety.
- Also evaluated at 18 months in secondary analyses.
4. Endpoints
| Endpoint | Registry definition / wording | Time frame | Type |
|---|---|---|---|
| Primary efficacy | TB Disease-free Survival at 12M After Study Treatment Assignment Among Participants in Control Regimen, Regimen1 (2HRZE/4HR) to Experimental Regimens, Regimen3 (2HPZM/2HPM) and Regimen2 (2HPZ/2HP) (Modified Intent to Treat [MITT] Population) | Twelve months after treatment assignment | Time-to-event |
| Primary efficacy | TB Disease-free Survival at 12M After Study Treatment Assignment Among Participants in Control Regimen, Regimen1 (2HRZE/4HR) to Experimental Regimens, Regimen3 (2HPZM/2HPM) and Regimen2 (2HPZ/2HP) (Assessable Population) | Twelve months after treatment assignment | Time-to-event |
| Primary safety | Percentage Participants With Grade 3 or Higher Adverse Events During Study Drug Treatment in Control Regimen (Regimen 1 2HRZE/4HR) Compared to Experimental Regimens, Regimen 3 (2HPZM/2HPM) and Regimen 2 (2HPZ/2HP) (Safety Analysis Population) | Four months and up to 14 days after last does of after study treatment (Regimen 2 and 3) or Six months and up to 14 days | Binary |
The registry posts 12 outcome measures and 14 statistical analyses, including six analyses associated with the three primary endpoints. The primary efficacy endpoint is expressed as a time-to-event outcome, but the posted effect measure is a risk difference, reflecting the registry's unfavorable-outcome formulation for the non-inferiority comparison.
5. Statistical Methodology
Cochran-Mantel-Haenszel test
The registry identifies the Cochran-Mantel-Haenszel test as the reported method for all six primary endpoint analyses. The method is commonly used to compare categorical outcomes while accounting for stratification factors. In this trial record, the effect measure reported alongside the method is the risk difference.
Here, the registry's analysis notes describe ra and rb as the proportions of unfavorable outcomes in the control and investigational arms, respectively. Thus, the sign of the reported risk difference must be read in the context of that unfavorable-outcome definition.
Non-inferiority framework
The registry states that non-inferiority was assessed using the upper bound of the two-sided 95% confidence interval for the difference between the percentage of patients classified as having an unfavorable status on the control and investigational regimens.
The posted secondary analyses explicitly identify 6.6% as the margin to define inferiority. Under the registry's unfavorable-outcome framing, an upper confidence-limit value at or below the margin is consistent with the stated non-inferiority criterion.
This is an important distinction from conventional superiority testing. A non-inferiority analysis is not simply asking whether a treatment difference has a small P-value. It asks whether the confidence interval excludes an amount of inferiority considered too large according to the prespecified margin.
Sequential comparison strategy
The registry analysis notes state that for the primary efficacy endpoint, Regimen 1 versus Regimen 3 would be considered first. If the non-inferiority criterion was met, Regimen 1 versus Regimen 2 would then be compared. The same sequential approach is described for the primary safety endpoint.
Why the sequence matters
The order creates a prespecified decision pathway. The Regimen 1 versus Regimen 2 comparison is not described as an unconditional first comparison in the registry analysis notes.
Why the CI matters
The non-inferiority decision is tied to the upper confidence bound relative to the margin. The P-value alone does not encode the clinical margin.
Analysis populations
The efficacy results include both a modified intent-to-treat population and an assessable population. The registry describes the microbiologically eligible population as including randomized participants excluding those with no evidence of cultures positive for M. tuberculosis or with specified resistance. The assessable population additionally excludes microbiologically eligible participants without assessable outcomes under the registry's stated criteria.
Safety analyses included all randomized participants who received at least one dose of study treatment.
6. Primary Results: TB Disease-Free Survival at 12 Months
The registry posts four primary efficacy analyses: two in the modified intent-to-treat population and two in the assessable population. The estimates are reported as risk differences, with the registry's non-inferiority interpretation based on the upper bound of the two-sided 95% confidence interval.
Modified Intent-to-Treat Population: Regimen 1 vs Regimen 3
Risk difference
95% CI: -2.6 to 4.5 · P = 0.05
Two-sided 95% confidence interval; non-inferiority hypothesis.
The reported estimate is a risk difference of 1 under the registry's unfavorable-outcome formulation: the control-regimen unfavorable-outcome proportion minus the Regimen 3 unfavorable-outcome proportion is estimated at 1 percentage point.
The estimate does not mean that Regimen 3 produces a 1% improvement in every patient's tuberculosis outcome, nor does it represent a hazard ratio or a relative risk. It is an absolute difference in the proportions defined by the registry's endpoint framework.
The two-sided 95% CI of -2.6 to 4.5 describes uncertainty around that estimated difference. Because the upper bound is 4.5, it lies below the 6.6% non-inferiority margin stated in the registry. The non-inferiority conclusion should nevertheless be understood as a conclusion under the prespecified analysis framework, not as proof that the regimens are identical.
The P-value of 0.05 does not measure the size of the treatment effect. In a non-inferiority setting, the confidence interval's position relative to the prespecified margin is central to interpretation.
Modified Intent-to-Treat Population: Regimen 1 vs Regimen 3 — Additional Posted Analysis
Risk difference
95% CI: -0.6 to 6.6 · P = 0.05
Two-sided 95% confidence interval; non-inferiority hypothesis.
This additional registry-posted analysis reports a risk difference of 3, again under the unfavorable-outcome framing. Its two-sided 95% CI extends from -0.6 to 6.6.
The upper confidence bound is exactly 6.6, the margin identified in the registry's secondary analyses as the margin to define inferiority. This illustrates why a non-inferiority analysis cannot be interpreted from the point estimate alone: an estimate that is numerically small can still have a confidence interval reaching the prespecified boundary.
The P-value of 0.05 should not be substituted for the non-inferiority criterion. The relevant question is how the confidence interval relates to the prespecified margin and the direction of the unfavorable-outcome difference.
Assessable Population: Regimen 1 vs Regimen 3
Risk difference
95% CI: -1.1 to 5.1 · P = 0.05
Two-sided 95% confidence interval; non-inferiority hypothesis.
The assessable-population estimate is a risk difference of 2, with a two-sided 95% CI from -1.1 to 5.1. The upper confidence limit remains below the 6.6% non-inferiority margin reported elsewhere in the registry results.
This analysis uses a different population from the modified intent-to-treat analysis. That distinction matters because excluding participants who lack an assessable outcome can change the statistical population and therefore the precision and interpretation of the estimate.
The confidence interval addresses uncertainty around the estimated group difference; it does not describe individual patient outcomes. The P-value of 0.05 is not an effect-size measure.
Assessable Population: Regimen 1 vs Regimen 2
Risk difference
95% CI: 1.2 to 7.7 · P = 0.05
Two-sided 95% confidence interval; non-inferiority hypothesis.
The reported risk difference is 4.4, with a two-sided 95% CI from 1.2 to 7.7. Under the registry's unfavorable-outcome definition, the positive estimate indicates a higher estimated unfavorable-outcome proportion in the control-minus-experimental difference.
The important non-inferiority feature is the upper confidence limit of 7.7. The registry identifies 6.6% as the margin to define inferiority, so this confidence interval extends beyond that margin. On the stated confidence-interval criterion, this result does not establish non-inferiority of Regimen 2 in this assessable-population comparison.
The P-value of 0.05 does not override that margin-based assessment. This is one of the clearest statistical lessons of the trial: a non-inferiority conclusion depends on the prespecified margin and confidence interval, not merely on whether a P-value reaches a conventional threshold.
7. Primary Safety Results
The third primary endpoint was the percentage of participants with grade 3 or higher adverse events during study drug treatment in the safety analysis population. The registry defines the safety population as all randomized participants who received at least one dose of study treatment.
| Comparison | Effect measure | Estimate | 95% CI | P-value | Hypothesis |
|---|---|---|---|---|---|
| Regimen 1 vs Regimen 3 | Risk difference | -0.6 | -4.3 to 3.2 | 0.05 | Non-inferiority |
| Regimen 1 vs Regimen 2 | Risk difference | -5.1 | -8.7 to -1.5 | 0.05 | Non-inferiority |
Regimen 1 vs Regimen 3
Grade 3 or higher adverse events
95% CI: -4.3 to 3.2 · P = 0.05
The risk difference is -0.6 under the registry's control-minus-experimental framing for the safety endpoint. The two-sided 95% CI ranges from -4.3 to 3.2.
The estimate is an absolute difference in the percentage of participants experiencing grade 3 or higher adverse events; it is not a relative risk and does not describe the severity or clinical importance of individual adverse events.
The confidence interval provides the precision of the estimated difference. Because the upper bound is 3.2, it remains below the 6.6% margin described in the registry's non-inferiority framework. The P-value of 0.05 does not itself establish the magnitude or clinical importance of the safety difference.
Regimen 1 vs Regimen 2
Grade 3 or higher adverse events
95% CI: -8.7 to -1.5 · P = 0.05
The reported estimate is -5.1, with a two-sided 95% CI from -8.7 to -1.5. Under the registry's stated direction, the negative estimate means the unfavorable-event percentage was lower in the control-minus-experimental difference.
The entire confidence interval is below zero, so the interval is consistent with a lower unfavorable-event proportion for the experimental regimen under this risk-difference definition. More importantly for the non-inferiority framework, the upper bound of -1.5 is below the 6.6% margin.
The result should still be read as a group-level safety comparison rather than evidence that every participant had fewer adverse events. The P-value of 0.05 is not a measure of clinical effect size.
8. Serious Adverse Events by Arm
The ClinicalTrials.gov record also report serious adverse events by treatment arm. These counts are presented as affected participants over participants at risk.
| Regimen | Serious adverse events | Affected / at risk |
|---|---|---|
| Regimen 1 (2HRZE/4HR) | Serious adverse events | 56/825 |
| Regimen 2 (2HPZ/2HP) | Serious adverse events | 39/835 |
| Regimen 3 (2HPMZ/2HPM) | Serious adverse events | 37/846 |
These are descriptive arm-level counts and denominators reported by the registry data. They should not be substituted for the formal primary safety analysis, which used grade 3 or higher adverse events and reported risk differences with confidence intervals.
9. Secondary Endpoint Results: TB Disease-Free Survival at 18 Months
The registry posts four secondary analyses of TB disease-free survival at eighteen months. These results use risk differences and identify 6.6% (δ = 0.066) as the margin to define inferiority. The registry does not report a normalized statistical method for these secondary analyses.
| Population | Comparison | Estimate | 95% CI | Hypothesis |
|---|---|---|---|---|
| Microbiologically eligible | Regimen 1 vs Regimen 3 | 1.09 | -2.45 to 4.63 | Non-inferiority |
| Microbiologically eligible | Regimen 1 vs Regimen 2 | 4.12 | 0.45 to 7.79 | Non-inferiority |
| Assessable | Regimen 1 vs Regimen 3 | 1.05 | -2.01 to 4.11 | Non-inferiority |
| Assessable | Regimen 1 vs Regimen 2 | 3.66 | 0.42 to 6.90 | Non-inferiority |
For Regimen 3, the upper confidence limits are 4.63 in the microbiologically eligible analysis and 4.11 in the assessable analysis, both below the 6.6% margin. For Regimen 2, the upper limits are 7.79 and 6.90, respectively, both above the stated margin. These observations describe how the posted confidence intervals relate to the registry's margin; they do not convert the secondary analyses into superiority tests.
10. Prespecified Sensitivity Analyses
The registry reports sensitivity analyses designed to examine how conclusions could change under two extreme assumptions concerning losses to follow-up and non-tuberculosis deaths.
| Assumption | Population | Comparison | Estimate | 95% CI |
|---|---|---|---|---|
| All losses to follow-up and non-tuberculosis deaths unfavorable | Microbiologically eligible | Regimen 1 vs Regimen 2 | 3.0 | -0.6 to 6.6 |
| All losses to follow-up and non-tuberculosis deaths unfavorable | Microbiologically eligible | Regimen 1 vs Regimen 3 | 2.29 | -1.12 to 5.70 |
| All losses to follow-up and non-tuberculosis deaths favorable | Assessable | Regimen 1 vs Regimen 2 | 5.9 | 2.8 to 8.9 |
| All losses to follow-up and non-tuberculosis deaths favorable | Assessable | Regimen 1 vs Regimen 3 | 3.4 | 0.5 to 6.3 |
Why these sensitivity analyses matter
Missing outcome information can matter particularly in a non-inferiority trial. If missing observations are handled under assumptions that systematically favor one regimen, the estimated treatment difference can move closer to or farther from the non-inferiority boundary. The two posted sensitivity scenarios therefore provide a way to examine the stability of the result under different extreme classifications.
11. Statistical Methods Explained
Why was a risk difference used?
The registry reports risk difference as the effect measure even though the primary efficacy endpoint is classified as time-to-event. The analysis notes frame the comparison in terms of the proportion of participants with an unfavorable outcome. A risk difference therefore expresses the absolute difference between those proportions under the registry's endpoint definition.
Why is non-inferiority judged against the margin rather than the P-value?
Non-inferiority asks whether the data are compatible with an unacceptable degree of inferiority. Here, the registry specifies a 6.6% margin and states that the upper bound of the two-sided 95% confidence interval is used for the assessment. A P-value of 0.05 does not tell the reader whether the confidence interval crosses that clinically defined boundary.
What does the 6.6% margin mean?
The registry's secondary analyses explicitly identify 6.6% (δ = 0.066) as the margin to define inferiority. It is a prespecified tolerance for the difference in unfavorable-outcome proportions, not a claim that the two treatments have exactly the same effect.
Why are both modified intent-to-treat and assessable populations shown?
They answer closely related questions using different analysis populations. The modified intent-to-treat analysis is based on the microbiologically eligible randomized population described by the registry, whereas the assessable analysis excludes additional participants without an assessable outcome under the registry's stated criteria. Comparing both helps show how population definition affects the estimated risk difference and its precision.
Why is the order of Regimen 3 and Regimen 2 comparisons important?
The registry analysis notes specify that Regimen 1 versus Regimen 3 is considered first for the primary efficacy endpoint. If the non-inferiority criterion is met, Regimen 1 versus Regimen 2 is then compared. The same sequential structure is described for the primary safety endpoint. This means the analysis should be read as a prespecified decision sequence rather than as two unrelated pairwise tests.
What does a negative risk difference mean here?
The answer depends on the endpoint's direction. The registry's non-inferiority comment defines the comparison using the percentage of participants with an unfavorable status in the control regimen minus the corresponding percentage in the investigational regimen. Consequently, a negative value indicates that the control-minus-experimental difference is below zero under that definition. It does not automatically mean "better" without first establishing which outcome is being differenced.
How should the 12-month disease-free survival endpoint be understood?
It is registered as a time-to-event endpoint with a time frame of twelve months after treatment assignment, but the posted effect measure is a risk difference based on unfavorable outcome classification. The registry therefore does not provide a hazard ratio in the statistical analyses posted on ClinicalTrials.gov. A Kaplan-Meier approach would ordinarily be a natural way to estimate a conventional time-to-event survival function, but the posted primary analyses themselves are identified as Cochran-Mantel-Haenszel analyses with risk differences.
12. Interpreting Confidence Intervals in This Trial
The confidence intervals are central to understanding TBTC Study 31 because the trial uses a non-inferiority framework. Three distinct questions should be kept separate:
Where is the point estimate?
The estimate summarizes the observed group difference according to the registry's risk-difference definition.
How precise is it?
The confidence interval describes uncertainty around the estimate. Wider intervals indicate less precision than narrower intervals.
Does it cross the NI boundary?
For the stated framework, the upper confidence limit is compared with the 6.6% margin.
What does it not say?
A confidence interval is not a range containing the effects experienced by individual patients, and it is not a probability distribution for the treatment effect.
Reading the results in this order keeps the non-inferiority logic visible rather than allowing the P-value to dominate the interpretation.
13. Randomization and What It Contributes Statistically
The study was randomized, which is the fundamental design feature supporting a causal comparison between the treatment assignments. Randomization does not guarantee identical groups in every characteristic; rather, it provides a principled basis for comparing outcomes between assigned regimens under the trial protocol.
The three-arm parallel structure also permits two experimental-versus-control comparisons within the same trial. The registry's analysis notes impose a specific sequence on those comparisons, beginning with Regimen 1 versus Regimen 3.
Participants were assigned to one of three study regimens.
Participants remained associated with their randomized regimen rather than following a crossover design in the registry description.
The registry identifies the study as unmasked.
The trial evaluates treatment regimens for tuberculosis.
14. Missing Data and Sensitivity Analysis
The ClinicalTrials.gov record does not describe a conventional imputation method such as multiple imputation. Instead, the posted prespecified sensitivity analyses explicitly assign all losses to follow-up and non-tuberculosis deaths as either unfavorable or favorable.
This is statistically important because non-inferiority studies can be sensitive to assumptions about missing outcomes. An analysis that treats every missing observation as unfavorable may produce a different risk difference from one that treats those observations as favorable. The ClinicalTrials.gov record demonstrate that the investigators examined both directions rather than relying on a single missing-data assumption.
| Sensitivity scenario | Why it is informative |
|---|---|
| All losses to follow-up and non-tuberculosis deaths unfavorable | Tests the effect of assigning missing or specified non-TB-death outcomes to the unfavorable category. |
| All losses to follow-up and non-tuberculosis deaths favorable | Tests the opposite extreme classification. |
15. Multiplicity and Sequential Testing
The registry identifies three primary endpoints and six primary-endpoint analyses. The primary efficacy endpoint is assessed in more than one population and against more than one experimental regimen, while the safety endpoint is also evaluated against both experimental regimens.
The ClinicalTrials.gov record specifically state that the comparison of Regimen 1 versus Regimen 3 is considered first and that, if the non-inferiority criterion is met, Regimen 1 versus Regimen 2 is then compared. This hierarchy is an important part of the statistical design because it defines how the pairwise comparisons are ordered.
| Endpoint family | Population / comparison | Posted analysis |
|---|---|---|
| Primary efficacy | MITT, Regimen 1 vs Regimen 3 | Risk difference; 95% CI; non-inferiority |
| Primary efficacy | MITT, Regimen 1 vs Regimen 3 — additional posted analysis | Risk difference; 95% CI; non-inferiority |
| Primary efficacy | Assessable, Regimen 1 vs Regimen 3 | Risk difference; 95% CI; non-inferiority |
| Primary efficacy | Assessable, Regimen 1 vs Regimen 2 | Risk difference; 95% CI; non-inferiority |
| Primary safety | Safety population, Regimen 1 vs Regimen 3 | Risk difference; 95% CI; non-inferiority |
| Primary safety | Safety population, Regimen 1 vs Regimen 2 | Risk difference; 95% CI; non-inferiority |
The ClinicalTrials.gov record does not provide an alpha-spending procedure or Bayesian method. Accordingly, no such methods are attributed to TBTC Study 31 here.
16. Kaplan-Meier Estimation and Time-to-Event Endpoints
TB disease-free survival is registered as a time-to-event endpoint. In general, time-to-event analysis is useful when participants can have different lengths of follow-up and when the timing of an event matters, because participants who have not experienced the event by their last observation can be censored rather than simply discarded.
Here, di represents events at time ti, while ni represents participants at risk immediately before that time.
However, the posted primary statistical analyses in the ClinicalTrials.gov record is not reported as Kaplan-Meier or Cox analyses. They are reported as Cochran-Mantel-Haenszel analyses with risk difference as the effect measure. This distinction should be preserved rather than replacing the registry's actual method with a more familiar survival-analysis label.
17. Design Limitations
- No masking: the registry states that placebos were not used and that neither participants nor site staff were blinded to treatment assignment. This creates the potential for knowledge of treatment assignment to affect behavior, reporting, assessment, or other trial processes.
- HIV coinfection representation: only 8% of participants were HIV-coinfected, limiting the power to compare regimens in this important population.
- Analysis-population dependence: the primary efficacy results differ in their defined populations, including modified intent-to-treat and assessable populations. Non-inferiority interpretation should therefore remain tied to the prespecified population used for each analysis.
- Time-to-event endpoint versus posted effect measure: TB disease-free survival is classified as a time-to-event endpoint, while the posted analyses use risk difference. Readers should not interpret these results as hazard ratios.
- Non-inferiority margin: the 6.6% margin is central to interpretation. A confidence interval that reaches beyond the margin cannot be treated as demonstrating non-inferiority merely because the point estimate is relatively small.
- Unmasked design: because treatment assignment was not masked, observed outcomes may be influenced by knowledge of assignment in ways that a blinded design might reduce.
- Missing outcome assumptions: the sensitivity analyses demonstrate that assumptions concerning losses to follow-up and non-tuberculosis deaths can change the estimated risk difference.
18. Why This Trial Matters Statistically
TBTC Study 31 is a useful statistical teaching case because it combines a three-arm randomized design with a non-inferiority objective, multiple analysis populations, risk differences, confidence-interval-based decision rules, sequential experimental-versus-control comparisons, and prespecified sensitivity analyses.
| Concept | How it appears in TBTC Study 31 |
|---|---|
| Randomization | Three-arm randomized parallel phase 3 design. |
| Non-inferiority | Primary comparisons are explicitly identified as non-inferiority analyses. |
| Risk difference | The registry's primary effect measure. |
| Confidence interval | The upper bound of the two-sided 95% CI is central to the non-inferiority assessment. |
| Cochran-Mantel-Haenszel test | Reported statistical method for all six primary endpoint analyses. |
| Sequential comparisons | Regimen 1 vs Regimen 3 is considered first, followed by Regimen 1 vs Regimen 2 if the stated criterion is met. |
| Analysis populations | Modified intent-to-treat, assessable, microbiologically eligible, and safety populations appear in the posted analyses. |
| Missing-data sensitivity | Extreme favorable and unfavorable assumptions are explicitly evaluated. |
| Safety analysis | Grade 3 or higher adverse events are evaluated as a primary endpoint. |
| Time-to-event endpoint | TB disease-free survival is registered at 12 and 18 months. |
19. A Practical Guide to Reading the Results
A reader can interpret the primary results without reducing the trial to a single number by following a structured sequence.
Identify the estimand direction
Determine whether the risk difference is defined as control minus experimental and whether the underlying event is favorable or unfavorable.
Read the point estimate
The estimate summarizes the observed absolute group difference in the registry's chosen scale.
Read the full confidence interval
Do not infer precision from the point estimate alone. The interval shows how uncertain the estimate is.
Compare the upper bound with 6.6%
For the stated non-inferiority framework, the upper confidence limit is compared with the prespecified margin.
Check the analysis population
A result from the modified intent-to-treat population and a result from the assessable population are not interchangeable.
Examine sensitivity analyses
Ask whether extreme assumptions about losses to follow-up and non-tuberculosis deaths materially alter the relationship with the non-inferiority margin.
20. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The registry reports risk differences with two-sided 95% confidence intervals and identifies non-inferiority as the hypothesis type. Several posted upper confidence limits are below the 6.6% margin, while some Regimen 2 comparisons have upper limits above it.
Clinical interpretation
The statistical evidence should be interpreted in the context of the actual treatment regimens, the trial population, the unmasked design, the limited representation of HIV-coinfected participants, and the prespecified definition of unfavorable outcomes.
The statistical analysis can establish what the reported estimates and confidence intervals say under the stated framework. It cannot by itself eliminate the design limitations or extend the findings beyond the population represented by the trial.
21. Important Interpretation Issues
- Non-inferiority is not equivalence: demonstrating that an upper confidence limit lies within a non-inferiority margin does not establish that two regimens have identical effects.
- Non-inferiority is not superiority: the fact that an interval lies on one side of zero is a different question from whether an experimental regimen is superior under a prespecified superiority hypothesis.
- P-values are not effect sizes: the reported P-value of 0.05 for the primary analyses does not tell the reader how large or clinically important the risk difference is.
- The margin is part of the conclusion: the same point estimate can lead to a different non-inferiority interpretation if the prespecified margin changes.
- Population definition matters: modified intent-to-treat and assessable populations can produce different estimates and confidence intervals.
- Time-to-event terminology should be used carefully: the registry labels TB disease-free survival as time-to-event, but the posted effect measure is risk difference rather than a hazard ratio.
- Safety and efficacy are distinct: disease-free survival and adverse-event outcomes answer different questions and should not be combined into a single numerical measure.
22. Related Statistical Tutorials
Learn more about the methods used in this trial:
23. Related Statistical Calculators
24. Sources
- ClinicalTrials.gov: TBTC Study 31 — NCT02410772.
- PubMed: PMID 39477924.
- PubMed: PMID 39012226.
- PubMed: PMID 38462673.
- PubMed: PMID 36790881.
- PubMed: PMID 36041016.
Continue through the Clinical Biostats statistical tutorials
Use the trial's endpoints and methods as a practical route into non-inferiority design, confidence intervals, risk differences, categorical-data methods, and time-to-event analysis.
25. Record Summary
TBTC Study 31 provides a detailed example of how a randomized three-arm trial can be evaluated through a non-inferiority framework. The ClinicalTrials.gov record uses the Cochran-Mantel-Haenszel method and risk difference, with two-sided 95% confidence intervals and a stated 6.6% margin to define inferiority. The primary efficacy analyses include modified intent-to-treat and assessable populations, while the primary safety analyses use randomized participants who received at least one dose of study treatment.
The most important statistical lesson is that the point estimate, confidence interval, non-inferiority margin, analysis population, and prespecified comparison sequence must be interpreted together. For Regimen 3, the posted primary confidence intervals have upper bounds below the stated 6.6% margin. For Regimen 2, the assessable primary efficacy comparison has an upper confidence limit of 7.7, above that margin. The secondary 18-month analyses and sensitivity analyses provide additional perspectives on how the estimates behave under different populations and assumptions.