This page separates reported trial results from statistical interpretation. Trial-specific numerical facts are restricted to the ClinicalTrials.gov record. Where the registry does not report a formal statistical comparison, that distinction is preserved.
1. Trial at a Glance
REMoxTB was a randomized, parallel-group, quadruple-masked phase 3 treatment trial evaluating two moxifloxacin-containing treatment-shortening regimens against a control regimen in patients with pulmonary tuberculosis.
| Feature | REMoxTB |
|---|---|
| Trial name | REMoxTB |
| Brief title | Controlled Comparison of Two Moxifloxacin Containing Treatment Shortening Regimens in Pulmonary Tuberculosis |
| Phase | Phase 3 |
| Status | Completed |
| Therapeutic area | Infectious Disease |
| Condition | Pulmonary Tuberculosis |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 1931 |
| Arms | 3 |
| Primary endpoints | 2 |
| Outcome measures posted | 9 |
| Statistical analyses posted | 2 |
| Primary-Endpoint Analyses | 2 |
| Lead sponsor | Global Alliance for TB Drug Development |
| Start | 2008-01 |
| Primary completion | 2013-10 |
2. Clinical Question
The central statistical question was whether either of two moxifloxacin-containing treatment-shortening regimens could achieve a sufficiently similar rate of unfavorable outcomes to the control regimen to satisfy a prespecified non-inferiority criterion.
Population
The registered condition was pulmonary tuberculosis.
Interventions
The trial evaluated moxifloxacin-containing regimens. The registered intervention list includes Moxifloxacin, Ethambutol, Isoniazid, Pyrazinamide, and Rifampicin.
Comparator
The control regimen was Regimen 1 — 2EHRZ/4HR.
Primary question
Could the treatment-shortening regimens remain within the prespecified non-inferiority boundary for the unfavorable-outcome endpoint?
3. Trial Design
2EHRZ/4HR
- Control regimen
- Used as the reference group for the posted primary analyses
2MHRZ/2MHR
- Moxifloxacin-containing regimen
- Compared with Regimen 1 in the posted analysis
2EMRZ/2MR
- Moxifloxacin-containing regimen
- Compared with Regimen 1 in the posted analysis
4. Endpoints
| Endpoint | Registry definition | Time frame | Type |
|---|---|---|---|
| Combined Failure of Bacteriological Cure and Relapse Within One Year of Completion of Therapy as Defined by Culture Using Solid Media (Lowenstein-Jensen - LJ) | The primary efficacy outcome was the proportion of patients who had bacteriologically or clinically defined failure or relapse within 18 months after randomization, described as a composite unfavorable outcome. Culture-negative status was defined as two negative-culture results at different visits without an intervening positive result. | 18 months (within one year of completion of therapy) | Binary |
| Number of Patients With Grade 3 or 4 Adverse Events (Using a Modified DAIDS National Institute of Allergy and Infectious Diseases Scale of Adverse Event Reporting) | The number of participants includes all patients who had at least one grade 3 or 4 adverse event. | 18 months (within one year of completion of therapy) | Binary |
The first primary endpoint is the endpoint for which the registry supplies two formal statistical analyses. The second primary endpoint is a binary safety endpoint, but the ClinicalTrials.gov record does not provide a corresponding formal statistical analysis for it.
5. Statistical Methodology
Per-protocol analysis
The posted efficacy analyses were performed in the per-protocol population. This is important in a non-inferiority trial because departures from the assigned regimen can make two treatment strategies appear more similar than they would be under strict adherence, while excluding protocol deviations can also change the population being compared.
For REMoxTB, the registry explicitly identifies per-protocol analysis as an analysis concept and reports an adjusted difference in the proportion of patients with an unfavorable outcome.
Risk difference and adjusted difference in proportions
The principal effect measure is an absolute difference in proportions rather than a relative risk or odds ratio. Conceptually, if \(p_T\) denotes the unfavorable-outcome proportion in a moxifloxacin-containing regimen and \(p_C\) denotes the corresponding proportion in the control group, a risk difference can be written as:
A positive value means the unfavorable-outcome proportion is higher in the treatment group than in the control group. For this trial, the registry reports an adjusted difference rather than supplying the underlying arm-specific proportions.
The non-inferiority margin
The registry defines non-inferiority using a 6 percentage-point boundary. Specifically, non-inferiority was defined as a between-group difference of less than 6 percentage points in the upper boundary of the two-sided 97.5% Wald confidence interval for the difference in the proportion of patients with an unfavorable outcome.
Non-inferiority decision rule
The posted analysis uses the upper boundary of the two-sided 97.5% Wald confidence interval.
Because the outcome is unfavorable, a sufficiently small positive difference is compatible with non-inferiority; a confidence interval extending beyond 6 percentage points is not.
Confidence intervals
A confidence interval provides a range of values compatible with the statistical model and observed data under the specified confidence procedure. In a non-inferiority trial, the direction and upper boundary of the interval are particularly important because the question is whether the treatment could be worse than the control by more than the prespecified clinically relevant amount.
Wald confidence interval
The registry describes the non-inferiority assessment as using a two-sided 97.5% Wald confidence interval. A Wald interval is constructed from an estimated effect and its estimated standard error. The resulting interval is then compared with the prespecified non-inferiority boundary.
What is not reported in the ClinicalTrials.gov record
The ClinicalTrials.gov record identifies the formal analysis as a per-protocol analysis and identify the effect measures and confidence intervals, but the normalized statistical method is listed as not reported. Accordingly, this page does not assign a specific regression model, covariance adjustment, stratification scheme, imputation method, or other analytic technique that is not explicitly present in the trial data.
6. Results
The registry contains two formal statistical analyses, both for the first primary efficacy endpoint. Each compares one moxifloxacin-containing regimen with the control regimen in the per-protocol population.
Primary Efficacy Result: Regimen 2 vs Control
Adjusted difference in unfavorable-outcome proportions
Two-sided 97.5% CI: 1.7 to 10.5 percentage points
Analysis population: Per protocol
Comparison: Regimen 1 — 2EHRZ/4HR vs Regimen 2 — 2MHRZ/2MHR
The reported estimate is an adjusted difference from control in the proportion of participants with failure or relapse. The estimate of 6.1 percentage points is positive, indicating that the estimated unfavorable-outcome proportion was higher for Regimen 2 than for the control regimen under the reported analysis.
What the estimate means: The reported adjusted difference was 6.1 percentage points in the unfavorable direction for Regimen 2 relative to control. Because the effect is expressed as an absolute difference, it describes a difference in proportions rather than a relative hazard or odds.
What it does not mean: It does not mean that every patient had a 6.1 percentage-point increase in personal risk, nor does it provide the underlying event proportions in the two groups. Those arm-specific proportions are not reported in the ClinicalTrials.gov record used for this page.
What the confidence interval says: The two-sided 97.5% confidence interval extends from 1.7 to 10.5 percentage points. Its upper boundary is above the prespecified 6 percentage-point non-inferiority margin.
Non-inferiority implication: Under the registry's stated rule, non-inferiority requires the upper boundary to be less than 6 percentage points. An upper boundary of 10.5 therefore does not satisfy that criterion.
Why a p-value is not the effect size: No p-value is reported for this statistical analysis in the ClinicalTrials.gov record. In general, a p-value addresses evidence against a specified null hypothesis; it does not quantify the magnitude of the treatment difference or replace the confidence interval in a non-inferiority assessment.
Analysis-population caution: This result is explicitly based on the per-protocol population. It should therefore not be described as an intention-to-treat estimate.
Primary Efficacy Result: Regimen 3 vs Control
Adjusted difference from control
Two-sided 97.5% CI: 6.7 to 16.1 percentage points
Analysis population: Per protocol
Comparison: Regimen 1 — 2EHRZ/4HR vs Regimen 3 — 2EMRZ/2MR
The reported adjusted difference from control was 11.4 percentage points for the rate of unfavorable outcome. The confidence interval ranges from 6.7 to 16.1 percentage points, placing the entire interval above the 6 percentage-point non-inferiority margin.
What the estimate means: The estimated unfavorable-outcome proportion for Regimen 3 was 11.4 percentage points higher than the control group's proportion under the reported adjusted comparison.
What it does not mean: It is not a relative increase of 11.4%, and it is not a statement that an individual patient's probability of failure or relapse increased by exactly 11.4 percentage points. It is a group-level adjusted difference in proportions.
What the confidence interval says: The two-sided 97.5% confidence interval is 6.7 to 16.1 percentage points. The entire interval lies above the 6 percentage-point boundary.
Non-inferiority implication: Because the upper boundary is 16.1 percentage points and the lower boundary is already 6.7 percentage points, the reported interval does not satisfy the registry's non-inferiority criterion of an upper boundary below 6 percentage points.
Why a p-value is not the effect size: No p-value is reported for this analysis in the ClinicalTrials.gov record. A p-value, when available, would not substitute for the estimated difference and its confidence interval. For a non-inferiority question, the location of the confidence interval relative to the prespecified margin is central.
Analysis-population caution: As with the Regimen 2 comparison, this result comes from the per-protocol population rather than an intention-to-treat analysis.
| Primary efficacy comparison | Analysis population | Effect measure | Estimate | Two-sided 97.5% CI | Non-inferiority margin |
|---|---|---|---|---|---|
| Regimen 2 vs Control | Per protocol | Adjusted difference in proportions | 6.1 percentage points | 1.7 to 10.5 | 6 percentage points |
| Regimen 3 vs Control | Per protocol | Adjusted difference from control | 11.4 percentage points | 6.7 to 16.1 | 6 percentage points |
7. Interpreting the Non-Inferiority Framework
Non-inferiority trials are designed around a different question from conventional superiority trials. Instead of asking whether a new treatment is better than control, the analysis asks whether the new treatment is not unacceptably worse than control according to a prespecified margin.
The registry wording defines the boundary using the difference in the proportion of patients with an unfavorable outcome. Because higher unfavorable-outcome rates are worse, the relevant concern is whether the treatment could be more than 6 percentage points worse than control.
Regimen 2
The estimate was 6.1 percentage points, with a two-sided 97.5% CI of 1.7 to 10.5. The upper boundary exceeds 6 percentage points.
Regimen 3
The estimate was 11.4 percentage points, with a two-sided 97.5% CI of 6.7 to 16.1. The entire interval is above 6 percentage points.
This illustrates why a non-inferiority conclusion should not be based simply on whether a confidence interval includes zero. A treatment could have a confidence interval that includes zero and still fail non-inferiority if its upper boundary extends beyond the permitted margin. Conversely, the non-inferiority question is specifically about whether the plausible range crosses the prespecified unacceptable-worse threshold.
8. Why the Per-Protocol Population Matters
The posted primary efficacy analyses use the per-protocol population. That choice is especially important for interpreting non-inferiority studies.
Protocol adherence
A per-protocol analysis focuses on participants who met the protocol-defined requirements for the analysis. The ClinicalTrials.gov record identifies this population but do not provide the registry's detailed eligibility criteria for inclusion in it.
Non-inferiority concern
In a non-inferiority setting, treatment-group similarity can sometimes be produced by departures from assigned treatment. For that reason, analysis population definitions deserve particular scrutiny.
The important point for reading this REMoxTB result is that the reported 6.1 and 11.4 percentage-point estimates are per-protocol estimates. They should not be silently relabeled as ITT results or interpreted as though both populations produced the same estimate.
9. Safety Results
The second registered primary endpoint was the number of patients with grade 3 or 4 adverse events during the 18-month time frame. The ClinicalTrials.gov record does not include a formal comparison for that endpoint.
the ClinicalTrials.gov record does, however, report serious adverse events by arm. These are a different safety measure from the registered grade 3 or 4 adverse-event endpoint and should not be substituted for it.
| Regimen | Serious adverse events | Affected / at risk |
|---|---|---|
| Regimen 1 — 2EHRZ/4HR (Control Regimen) | Serious adverse events | 38/639 |
| Regimen 2 — 2MHRZ/2MHR | Serious adverse events | 46/655 |
| Regimen 3 — 2EMRZ/2MR | Serious adverse events | 40/636 |
These figures provide arm-level information about serious adverse events, but they do not constitute the formal analysis of the registered grade 3 or 4 adverse-event endpoint. No risk difference, confidence interval, p-value, or hypothesis-test result for the safety endpoint is included in the statistical analyses posted on ClinicalTrials.gov.
10. Statistical Methods Explained
Why was a non-inferiority design used?
The trial's registered hypothesis type was non-inferiority or equivalence. In a treatment-shortening comparison, the statistical question can be framed around whether a new regimen remains sufficiently close to the control on the unfavorable-outcome endpoint. The ClinicalTrials.gov record establishes the non-inferiority framework but do not state the clinical or operational rationale used to select the treatment duration.
Why is the margin more important than zero?
For superiority testing, zero is often the key reference value for an absolute difference. Non-inferiority uses a different reference: the prespecified amount of worse outcome considered unacceptable. In REMoxTB, that boundary was 6 percentage points. Therefore, the relevant question is whether the upper boundary of the two-sided 97.5% confidence interval is below 6, not merely whether the interval includes zero.
What does an adjusted difference of 6.1 percentage points mean?
It means the reported adjusted estimate of the unfavorable-outcome proportion was 6.1 percentage points higher for Regimen 2 than for the control regimen in the per-protocol analysis. It is an absolute effect measure. It should not be converted into a relative risk or odds ratio without the underlying arm-specific proportions.
Why does the confidence interval extend beyond the point estimate?
The point estimate is a single estimate from the observed data. The confidence interval represents uncertainty around that estimate. For Regimen 2, the interval ranges from 1.7 to 10.5 percentage points; for Regimen 3, it ranges from 6.7 to 16.1 percentage points. The upper boundary is particularly important because the non-inferiority rule concerns the possibility that treatment is worse than control by more than 6 percentage points.
Why use a per-protocol analysis?
A per-protocol analysis evaluates the treatment comparison among participants meeting the protocol-defined criteria for that analysis. In non-inferiority trials, adherence and protocol deviations can affect the ability to distinguish genuine similarity from similarity created by departures from treatment. The registry explicitly identifies the primary efficacy analyses as per-protocol.
Why can't the serious-adverse-event counts be used for the second primary endpoint?
The second primary endpoint is specifically the number of patients with grade 3 or 4 adverse events using a modified DAIDS NIAID scale. The ClinicalTrials.gov record instead identify serious adverse events. Because those endpoint definitions are different, replacing one with the other would change the outcome being analyzed.
Why is the absence of a p-value important?
The statistical analyses posted on ClinicalTrials.gov do not report p-values. That is not a reason to infer one. More importantly, for the stated non-inferiority criterion, the confidence interval's upper boundary provides the direct comparison with the 6 percentage-point margin. A p-value, even if available, would not describe the magnitude of the estimated difference or its precision.
11. What the Risk Difference Does — and Does Not — Mean
The risk difference expresses an absolute difference in proportions. In this trial, the reported adjusted differences were 6.1 percentage points for Regimen 2 versus control and 11.4 percentage points for Regimen 3 versus control.
These values do not mean that the probability of an unfavorable outcome for every individual patient changed by exactly those amounts. They summarize the difference between treatment groups under the reported statistical analysis.
The 97.5% confidence intervals quantify uncertainty around the estimated differences under the specified Wald procedure. For Regimen 2, the interval is 1.7 to 10.5 percentage points. For Regimen 3, it is 6.7 to 16.1 percentage points.
The registry's rule is based on the upper boundary of the confidence interval. The Regimen 2 upper boundary is 10.5 percentage points, while the Regimen 3 upper boundary is 16.1 percentage points. Both are above the 6 percentage-point boundary specified in the registry analysis description.
12. Results in Statistical Context
| Feature | Regimen 2 vs Control | Regimen 3 vs Control |
|---|---|---|
| Analysis population | Per protocol | Per protocol |
| Effect measure | Adjusted difference in proportions | Adjusted difference from control |
| Estimate | 6.1 percentage points | 11.4 percentage points |
| Two-sided CI | 97.5% | 97.5% |
| CI lower boundary | 1.7 | 6.7 |
| CI upper boundary | 10.5 | 16.1 |
| Non-inferiority margin | 6 percentage points | 6 percentage points |
| Upper-bound comparison | 10.5 exceeds 6 | 16.1 exceeds 6 |
| Formal p-value reported | No | No |
The two comparisons therefore tell a consistent statistical story with respect to the prespecified non-inferiority boundary: neither confidence interval has an upper boundary below 6 percentage points. The second comparison is additionally separated from the margin at its lower confidence boundary, because its lower boundary is 6.7 percentage points.
13. What Is Not Reported in the Supplied Registry Data
A careful statistical analysis also requires knowing what the available record does not establish. The registry-reported REMoxTB data do not provide numerical baseline characteristics, subgroup estimates, median event times, hazard ratios, odds ratios, arm-specific proportions for the primary efficacy endpoint, formal p-values for the posted primary analyses, or a formal statistical comparison for the registered grade 3 or 4 adverse-event endpoint.
| Topic | Status in the ClinicalTrials.gov record |
|---|---|
| Primary efficacy estimates | Reported |
| Confidence intervals | Reported |
| Non-inferiority margin | Reported as 6 percentage points |
| Analysis population | Per protocol |
| Normalized statistical method | Not reported |
| Formal p-values | Not reported |
| Arm-specific primary efficacy proportions | Not reported |
| Grade 3 or 4 AE formal comparison | Not reported |
| Serious adverse events by arm | Reported |
| Subgroup results | Not reported |
| Bayesian methods | Not reported |
| Interim analysis details | Not reported |
| Missing-data or imputation method | Not reported |
| Stratification factors | Not reported |
This distinction prevents a common error in trial interpretation: treating a method as though it was used simply because it is conventional for a particular endpoint. The registry does not state a statistical method for this analysis, so this page does not assign an unreported model to the REMoxTB analysis.
14. Limitations
- Per-protocol efficacy population: the posted primary efficacy analyses were conducted in the per-protocol population. This is an important consideration for non-inferiority interpretation.
- Incomplete method specification: the ClinicalTrials.gov record identifies the effect measures and confidence intervals but lists the normalized statistical method as not reported.
- No formal p-values reported: the statistical-analysis records provide estimates and confidence intervals but no p-values.
- No underlying arm-specific efficacy proportions: the ClinicalTrials.gov record provides differences from control but do not provide the component proportions from which those differences arise.
- Safety endpoint distinction: serious adverse-event counts are available, but they are not equivalent to the registered grade 3 or 4 adverse-event endpoint.
- No subgroup information: subgroup estimates are not reported, so treatment-effect heterogeneity cannot be evaluated from this record.
- No missing-data method reported: the ClinicalTrials.gov record does not identify an imputation or missing-data strategy.
- No interim-analysis information: the ClinicalTrials.gov record does not specify whether or how interim statistical monitoring was performed.
- No Bayesian methodology reported: there is no basis in the ClinicalTrials.gov record for describing the analysis as Bayesian.
15. Why This Trial Matters Statistically
REMoxTB is a useful teaching example because its primary analysis illustrates one of the most important distinctions in clinical-trial statistics: non-inferiority is not simply a test of whether two treatments are statistically different. The interpretation depends on the direction of the effect, the confidence interval, the prespecified margin, and the analysis population.
| Concept | How it appears in REMoxTB |
|---|---|
| Randomization | Randomized allocation across 3 parallel treatment arms |
| Masking | Quadruple masking |
| Non-inferiority | Primary efficacy comparisons evaluated against a 6 percentage-point boundary |
| Risk difference | Adjusted difference in proportions / adjusted difference from control |
| Confidence interval | Two-sided 97.5% Wald intervals |
| Per-protocol analysis | Primary efficacy analyses were based on the per-protocol population |
| Binary endpoint | Primary efficacy outcome is an unfavorable-outcome proportion |
| Composite outcome | Failure and relapse are combined in the primary efficacy endpoint |
| Safety analysis | Serious adverse events are reported by treatment arm |
| Endpoint-definition discipline | Serious adverse events should not be substituted for grade 3 or 4 adverse events |
16. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The posted per-protocol analyses estimated unfavorable-outcome differences of 6.1 and 11.4 percentage points for the two moxifloxacin-containing regimens versus control. Their two-sided 97.5% confidence-interval upper boundaries were 10.5 and 16.1 percentage points, respectively, compared with a 6 percentage-point non-inferiority boundary.
Clinical interpretation
The registry data establish the trial's comparison framework and the reported unfavorable-outcome estimates, but they do not provide enough additional clinical detail on this page to quantify other dimensions of treatment benefit, such as subgroup-specific effects or time-to-event outcomes.
The statistical conclusion should therefore remain tied to the endpoint and analysis actually reported. A non-inferiority analysis is not a general statement about every possible outcome of a treatment regimen; it is a conclusion about whether the prespecified endpoint satisfies the prespecified margin under the stated analysis.
17. Design Features That Affect Interpretation
Three-arm comparison
The trial contains a control regimen and two moxifloxacin-containing regimens. Each posted primary analysis compares one experimental regimen with the same control regimen.
Quadruple masking
The registry classifies the study as quadruple-masked. Masking is a design feature intended to reduce opportunities for knowledge of treatment assignment to influence trial conduct or assessment.
Binary primary efficacy endpoint
The primary efficacy outcome is represented as the proportion of participants experiencing the composite unfavorable outcome.
Non-inferiority boundary
The 6 percentage-point boundary converts a broad clinical question into an explicit statistical decision rule.
18. A Worked Reading of the Two Primary Analyses
Step 1: Identify the estimand
The reported estimand is an adjusted difference in the proportion of participants with an unfavorable outcome, comparing each moxifloxacin-containing regimen with the control regimen.
Step 2: Identify the analysis population
The registry identifies the analysis population as per protocol. This population label must remain attached to the result because it affects how the estimate should be interpreted.
Step 3: Read the point estimate
For Regimen 2, the estimate is 6.1 percentage points. For Regimen 3, the estimate is 11.4 percentage points. Both estimates are positive differences in the unfavorable-outcome direction.
Step 4: Read the confidence interval
The Regimen 2 interval is 1.7 to 10.5 percentage points. The Regimen 3 interval is 6.7 to 16.1 percentage points. The intervals communicate considerably more than the point estimates alone because they show the range of effect sizes compatible with the analysis.
Step 5: Compare the upper boundary with the margin
The non-inferiority boundary is 6 percentage points. Regimen 2 has an upper boundary of 10.5; Regimen 3 has an upper boundary of 16.1. Neither upper boundary is below 6.
The key statistical comparison
For both posted efficacy analyses, the upper confidence-limit value exceeds the registry's 6 percentage-point non-inferiority boundary.
This is the core of the reported non-inferiority analysis. No additional p-value, hazard ratio, or reconstructed event rate is necessary to state what the ClinicalTrials.gov record reports.
19. Related Tutorials
Learn more about the methods used in this trial:
20. Related Calculators
21. Sources
- ClinicalTrials.gov: REMoxTB — NCT00864383.
- Linked publication: PubMed record — PMID 25196020.
- Linked publication: PubMed record — PMID 38468320.
- Linked publication: PubMed record — PMID 37986887.
- Linked publication: PubMed record — PMID 29779774.
- Linked publication: PubMed record — PMID 26847437.
Continue through the Clinical Biostats statistical pathway
Explore the statistical concepts that appear in randomized non-inferiority trials, from confidence intervals and risk differences to protocol-based analysis populations.
22. Record Summary
REMoxTB provides a clear example of how a randomized phase 3 non-inferiority trial can be interpreted through the relationship among the effect estimate, confidence interval, analysis population, and non-inferiority margin. The registry analyses report adjusted unfavorable-outcome differences of 6.1 percentage points for Regimen 2 versus control and 11.4 percentage points for Regimen 3 versus control, with two-sided 97.5% confidence intervals of 1.7 to 10.5 and 6.7 to 16.1 percentage points, respectively. The registry defines non-inferiority using an upper confidence boundary below 6 percentage points, so the reported upper boundaries of 10.5 and 16.1 are central to interpreting the posted analyses.
The page also illustrates why endpoint definitions must remain precise. The registered second primary endpoint concerns grade 3 or 4 adverse events, whereas the ClinicalTrials.gov record reports serious adverse events of 38/639, 46/655, and 40/636 across the three regimens. Those measures should not be conflated. Likewise, the absence of reported p-values, detailed normalized statistical methods, subgroup analyses, or missing-data methods means those features should not be reconstructed from convention or assumed from the endpoint type.