This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. The numerical trial results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
INBUILD was a randomized, double-blind, parallel phase 3 trial evaluating 150 mg nintedanib versus placebo in patients with progressive fibrosing interstitial lung disease. The registry reports 663 enrolled participants, two arms, two registered primary endpoints, 20 posted outcome measures, and 18 posted statistical analyses.
| Feature | INBUILD |
|---|---|
| Trial name | INBUILD |
| Brief title | Efficacy and Safety of Nintedanib in Patients With Progressive Fibrosing Interstitial Lung Disease (PF-ILD) |
| Phase | Phase 3 |
| Status | Completed |
| Condition | Lung Diseases, Interstitial |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Double |
| Primary purpose | Treatment |
| Enrollment | 663 |
| Interventions | Nintedanib; Placebo |
| Lead sponsor | Boehringer Ingelheim |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT02999178 |
2. Clinical Question
The central statistical question was whether 150 mg nintedanib differed from placebo in the annual rate of decline in forced vital capacity (FVC) over the planned 52-week assessment period, both in the overall population and in participants with an HRCT fibrotic pattern described as UIP-like fibrotic pattern only.
Population
Randomized participants with progressive fibrosing interstitial lung disease, analyzed either in the overall population or in the HRCT UIP-like fibrotic-pattern subgroup specified by the registered endpoint.
Intervention
150 mg nintedanib.
Comparator
Placebo.
Primary question
Does nintedanib change the modeled annual rate of FVC decline compared with placebo over 52 weeks?
The ClinicalTrials.gov record also contain secondary analyses covering K-BILD total score, L-PF symptom domains, acute ILD exacerbation or death, death, progression or death, and categorical FVC decline thresholds. These endpoints use different statistical estimands, so they should not be treated as interchangeable measures of the same treatment effect.
3. Trial Design
150 mg Nintedanib
- Drug intervention.
- Compared with placebo under randomized allocation.
- Primary analysis population required receipt of at least 1 dose of trial medication.
Placebo
- Placebo intervention.
- Compared with 150 mg nintedanib.
- Primary analysis population required receipt of at least 1 dose of trial medication.
4. Endpoints
The registry identifies two primary endpoints. Both concern annual FVC decline and were formally analyzed using mixed-model methodology.
| Primary endpoint | Registry time frame | Statistical approach |
|---|---|---|
| Annual Rate of Decline in Forced Vital Capacity - Overall Population | Baseline, 2, 4, 6, 12, 24, 36, 52 weeks after first drug intake (planned post-baseline visits) | Mixed Models Analysis; mixed-effects model |
| Annual Rate of Decline in Forced Vital Capacity - Participants With HRCT Fibrotic Pattern=UIP-like Fibrotic Pattern Only | Baseline, 2, 4, 6, 12, 24, 36, 52 weeks after first drug intake (planned post-baseline visits) | Mixed Models Analysis; mixed-effects model |
What FVC represents
The registry defines forced vital capacity as the volume of air, measured in milliliter, that can be forcibly exhaled from the lungs after taking the deepest breath possible. The primary estimand is not a single FVC measurement at week 52. It is the annual rate of decline estimated from repeated measurements over the 52-week period.
Secondary endpoint families
| Endpoint family | Examples in the posted analyses | Method |
|---|---|---|
| Patient-reported outcomes | K-BILD total score; L-PF Symptoms Dyspnea; L-PF Symptoms Cough | MMRM |
| Time-to-event | Acute ILD exacerbation or death; death; progression or death | Log-rank test with Cox model for HR |
| Binary FVC outcomes | Relative decline greater than 10%; relative decline greater than 5% | Logistic regression |
For several secondary endpoints, the registry explicitly states that no formal hypotheses were tested. Those estimates are therefore best interpreted as descriptive or supportive analyses rather than as independent confirmatory tests.
5. Statistical Methodology
Mixed-effects model for the primary FVC endpoint
The primary analysis used a mixed-effects model for repeated FVC measurements. The registry states that the decrease in FVC was assumed to be linear within each participant over 52 weeks. Participant-specific intercepts and slopes were assumed to be normally distributed with an unstructured covariance matrix, and the within-participant error was assumed to be independent and normally distributed with mean 0 and a common variance.
The model therefore uses the entire planned sequence of FVC measurements rather than reducing each participant to a single observed value. This is statistically important because the endpoint is defined as a rate of decline. A model of the longitudinal trajectory can use the pattern of repeated observations to estimate the treatment difference in slope.
The INBUILD registry analysis extends this longitudinal idea with treatment and other model components. The key estimand is the difference between treatment-group rates of FVC decline over the analyzed period.
Kenward-Roger approximation
The registry reports use of the Kenward-Roger approximation to estimate denominator degrees of freedom and adjust the covariance structure used for inference. This is a small-sample refinement commonly used with mixed models because the effective degrees of freedom are not necessarily obvious from the nominal number of observations.
MMRM for repeated patient-reported outcomes
K-BILD and L-PF symptom-domain outcomes were analyzed using mixed model repeated measures. These models incorporate repeated observations over visits and include participant-level random effects. The analyses posted on ClinicalTrials.gov specify baseline terms, visit, treatment-by-visit interactions, and baseline-by-visit interactions, with the exact fixed-effect structure varying by endpoint.
Log-rank test and Cox proportional-hazards model
Time to acute ILD exacerbation or death, time to death, and time to progression or death were evaluated with log-rank methods. Where a hazard ratio was reported, the registry states that a Cox proportional-hazards model was used to estimate the hazard ratio as Nintedanib divided by Placebo. Some analyses were stratified by HRCT fibrotic pattern.
An HR below 1 indicates a lower estimated instantaneous event rate in the nintedanib group under the Cox model. It does not directly equal an absolute risk reduction, a probability of benefit, or a percentage of patients who experience an event.
Logistic regression for categorical FVC decline
The relative decline endpoints at week 52 were analyzed with logistic regression. The models included baseline FVC percent predicted as a continuous covariate, with HRCT fibrotic pattern included as a binary covariate where specified. The reported effect measure is an adjusted odds ratio calculated as Nintedanib divided by Placebo.
6. Primary Results: Annual Rate of FVC Decline — Overall Population
The first primary endpoint was the annual rate of decline in FVC in the overall population. The analysis included all randomized participants in the overall population who received at least 1 dose of trial medication.
Adjusted mean difference in annual FVC decline
95% CI: 65.42 to 148.50 mL/year · P < .0001
150 mg Nintedanib vs Placebo; mixed-effects model.
| Primary endpoint | Nintedanib | Placebo | Adjusted mean difference | 95% CI | P-value |
|---|---|---|---|---|---|
| Annual Rate of Decline in FVC - Overall Population | 150 mg Nintedanib vs Placebo | 106.96 mL/year | 65.42 to 148.50 | <.0001 | |
The adjusted mean difference of 106.96 mL/year is the modeled difference in annual FVC decline between nintedanib and placebo. Because the endpoint concerns decline, a positive difference in the Nintedanib-minus-Placebo comparison corresponds to a less negative FVC trajectory for the nintedanib group under the model.
The 95% confidence interval of 65.42 to 148.50 mL/year describes the statistical uncertainty around the estimated between-group difference under the model. It does not mean that individual patients' FVC changes must fall inside this interval.
The P-value of <.0001 addresses the statistical evidence against the null comparison specified by the analysis. It does not measure the size or clinical importance of the treatment effect. The effect estimate and its confidence interval are needed to understand magnitude and precision.
The analysis also relies on the stated longitudinal assumptions, including a linear decline within each participant over 52 weeks and the specified distributional and covariance assumptions. As with any longitudinal model, departures from those assumptions can affect the estimated slope difference.
Why a longitudinal model is appropriate here
FVC was planned to be measured repeatedly at baseline and at 2, 4, 6, 12, 24, 36, and 52 weeks after first drug intake. A slope-based analysis uses these repeated measurements to characterize the trajectory rather than comparing only two observations. This makes the statistical estimand directly aligned with the registered endpoint: annual rate of decline.
7. Primary Results: Annual Rate of FVC Decline — UIP-like Fibrotic Pattern
The second primary endpoint restricted the analysis to participants with HRCT fibrotic pattern described as UIP-like fibrotic pattern only. The analysis population consisted of all randomized participants meeting that pattern definition who received at least 1 dose of trial medication.
Adjusted mean difference in annual FVC decline
95% CI: 70.81 to 185.59 mL/year · P < .0001
150 mg Nintedanib vs Placebo; mixed-effects model.
| Primary endpoint | Comparison | Adjusted mean difference | 95% CI | P-value |
|---|---|---|---|---|
| Annual Rate of Decline in FVC - Participants With HRCT Fibrotic Pattern=UIP-like Fibrotic Pattern Only | 150 mg Nintedanib vs Placebo | 128.20 mL/year | 70.81 to 185.59 | <.0001 |
The estimated adjusted mean difference was 128.20 mL/year, with the comparison defined as Nintedanib minus Placebo. For an endpoint measuring decline, the positive difference corresponds to a less negative modeled FVC trajectory in the nintedanib group.
The 95% CI of 70.81 to 185.59 mL/year provides the precision interval reported by the registry. It indicates that the estimated treatment difference is not represented by a single exact number; uncertainty remains around the model-based estimate.
The P-value of <.0001 provides evidence against the relevant null comparison but is not a measure of how large or clinically meaningful the effect is. The magnitude and confidence interval are more informative for understanding the estimated difference.
This endpoint is also a prespecified subgroup-defined primary endpoint. Its interpretation should remain tied to the population specified by the registry rather than being generalized automatically to every possible fibrotic pattern.
8. Comparing the Two Primary FVC Analyses
| Primary endpoint | Estimate | 95% CI | P-value | Interpretation of direction |
|---|---|---|---|---|
| Overall population | 106.96 mL/year | 65.42 to 148.50 | <.0001 | Higher modeled annual FVC rate relative to placebo; because the endpoint is decline, this corresponds to less decline. |
| UIP-like fibrotic pattern only | 128.20 mL/year | 70.81 to 185.59 | <.0001 | Higher modeled annual FVC rate relative to placebo; because the endpoint is decline, this corresponds to less decline. |
The two estimates address different populations. The overall-population estimate uses all randomized participants in the defined overall population who received at least one dose, whereas the second estimate is restricted to the specified UIP-like HRCT fibrotic pattern. The difference between the numerical estimates should therefore not be treated as a formal test that treatment effects differ between these populations. A comparison of effect estimates across populations requires an appropriate interaction or heterogeneity analysis.
9. Secondary Results: K-BILD Total Score
The registry reports absolute change from baseline in K-BILD total score at week 52 using MMRM. These analyses were not associated with formal hypothesis testing in the ClinicalTrials.gov record.
| Population | Adjusted mean difference | 95% CI | P-value |
|---|---|---|---|
| Overall population | 1.34 | -0.31 to 2.98 | 0.1115 |
| HRCT UIP-like fibrotic pattern only | 1.53 | -0.68 to 3.74 | 0.1747 |
Overall population
The adjusted mean difference was 1.34 units, calculated as Nintedanib minus Placebo, with a 95% CI of -0.31 to 2.98 and P=0.1115. The MMRM included baseline K-BILD total score, visit, treatment-by-visit and baseline-by-visit interactions, with participant as a random effect.
UIP-like fibrotic pattern only
The adjusted mean difference was 1.53 units, with a 95% CI of -0.68 to 3.74 and P=0.1747. The registry explicitly states that no formal hypotheses were tested for this endpoint.
10. Secondary Results: Acute ILD Exacerbation or Death
Time to first acute ILD exacerbation or death over 52 weeks was analyzed with a log-rank method. The hazard ratio was estimated with a Cox proportional-hazards model. In the overall population, the analysis was stratified by HRCT fibrotic pattern.
| Population | HR | 95% CI | P-value | Analysis |
|---|---|---|---|---|
| Overall population | 0.80 | 0.48 to 1.34 | 0.3948 | Log-rank; stratified Cox model |
| HRCT UIP-like fibrotic pattern only | 0.67 | 0.36 to 1.24 | 0.1985 | Log-rank; Cox model |
The time frame was from first drug intake until the date of first acute ILD exacerbation, death, or last contact date, up to 372 days. No formal hypotheses were tested for these secondary analyses according to the ClinicalTrials.gov record.
An HR of 0.80 means that the estimated instantaneous event rate in the nintedanib group was 0.80 times the corresponding rate in the placebo group under the Cox model. This is a relative model-based measure; it is not an 80% probability of avoiding an event.
The 95% CI of 0.48 to 1.34 includes 1.00, indicating substantial statistical uncertainty about the direction and magnitude of the relative hazard difference in this analysis.
The UIP-like estimate of 0.67 has a 95% CI of 0.36 to 1.24. Again, the interval includes 1.00. The numerical difference between the two subgroup-specific estimates does not by itself establish that the treatment effects differ.
11. Secondary Results: Time to Death
| Population | HR | 95% CI | P-value | Analysis |
|---|---|---|---|---|
| Overall population | 0.94 | 0.47 to 1.86 | 0.8544 | Log-rank; stratified Cox model |
| HRCT UIP-like fibrotic pattern only | 0.68 | 0.32 to 1.47 | 0.3291 | Log-rank; Cox model |
The endpoint was time to death over 52 weeks, measured from first drug intake until death or last contact date, up to 372 days. The overall-population analysis was stratified by HRCT fibrotic pattern.
The overall-population HR of 0.94 is accompanied by a wide 95% CI of 0.47 to 1.86. The width of this interval shows that the reported estimate has limited precision for this time-to-death analysis. The P-value of 0.8544 does not quantify the size of the effect; it addresses the statistical evidence against the relevant null comparison.
The UIP-like estimate of 0.68 has a 95% CI of 0.32 to 1.47. The interval includes 1.00, so the point estimate should not be read as establishing a definitive relative hazard reduction.
12. Secondary Results: Time to Progression or Death
Overall population
95% CI: 0.49 to 0.85 · P = 0.0017
Time to progression or death over 52 weeks; stratified by HRCT fibrotic pattern.
| Population | HR | 95% CI | P-value |
|---|---|---|---|
| Overall population | 0.65 | 0.49 to 0.85 | 0.0017 |
| HRCT UIP-like fibrotic pattern only | 0.64 | 0.45 to 0.89 | 0.0081 |
The registry reports a log-rank analysis with a Cox proportional-hazards model. The overall-population model was stratified by HRCT fibrotic pattern. The hazard ratio was calculated as Nintedanib divided by Placebo.
An HR of 0.65 means that the estimated instantaneous rate of progression or death was 0.65 times the corresponding rate in the placebo group under the fitted Cox model. In relative terms, 1 - 0.65 = 0.35, so the point estimate corresponds to an approximately 35% lower estimated hazard. This is a direct interpretation of the hazard-ratio point estimate, not a statement that 35% of patients avoided progression or death.
The 95% CI of 0.49 to 0.85 quantifies uncertainty around the estimated hazard ratio. Because the entire reported interval is below 1.00, the estimated relative hazard is below 1 throughout the confidence interval.
The P-value of 0.0017 provides evidence against the relevant null comparison. It does not say that the treatment effect is 0.0017 in size, nor does it tell us how many additional days an individual patient will remain progression-free.
UIP-like fibrotic pattern only
For participants with the HRCT UIP-like fibrotic pattern only, the reported hazard ratio was 0.64, with a 95% CI of 0.45 to 0.89 and P=0.0081.
The point estimates of 0.65 and 0.64 are numerically close, but the ClinicalTrials.gov record does not provide a formal interaction test between these populations. Therefore, the results support describing the estimates separately rather than claiming that one population had a statistically different treatment effect.
13. Secondary Results: Relative FVC Decline Thresholds
The registry also reports binary endpoints based on whether participants experienced a relative decline from baseline in FVC percent predicted greater than 10% or greater than 5% at week 52. These were analyzed with logistic regression and adjusted odds ratios.
| Endpoint | Population | Adjusted OR | 95% CI | P-value |
|---|---|---|---|---|
| Relative decline >10% | Overall | 0.70 | 0.52 to 0.96 | Not stated as formal hypothesis |
| Relative decline >10% | HRCT UIP-like pattern only | 0.63 | 0.43 to 0.94 | Not stated as formal hypothesis |
| Relative decline >5% | Overall | 0.50 | 0.36 to 0.68 | Not stated as formal hypothesis |
| Relative decline >5% | HRCT UIP-like pattern only | 0.46 | 0.31 to 0.69 | Not stated as formal hypothesis |
For the overall-population >10% and >5% endpoints, the logistic regression model included baseline FVC percent predicted as a continuous covariate and HRCT fibrotic pattern as a binary covariate. For the corresponding UIP-like analyses, the model included baseline FVC percent predicted.
How to interpret an odds ratio
An adjusted odds ratio below 1 indicates lower modeled odds of the specified binary outcome in the nintedanib group relative to placebo. It is important not to call an odds ratio a risk ratio unless the two measures happen to coincide under special circumstances.
Logistic regression models the odds of the specified binary outcome. An odds ratio of 0.50 therefore represents half the modeled odds, not necessarily half the probability of the outcome.
For example, the overall-population adjusted OR of 0.50 for a relative FVC decline greater than 5% corresponds to half the modeled odds under the specified logistic regression model. The registry does not provide the corresponding arm-specific event percentages in the ClinicalTrials.gov record, so an absolute risk difference should not be inferred from the odds ratio alone.
14. Secondary Results: L-PF Dyspnea Domain
| Population | Adjusted mean difference | 95% CI | P-value |
|---|---|---|---|
| Overall population | -3.53 | -6.14 to -0.92 | Not formally tested |
| HRCT UIP-like fibrotic pattern only | -4.18 | -7.48 to -0.88 | Not formally tested |
The endpoint was absolute change from baseline in the Living With Pulmonary Fibrosis (L-PF) Symptoms Dyspnea Domain Score at week 52. The overall-population MMRM included baseline, HRCT fibrotic pattern, visit, treatment-by-visit interaction, baseline-by-visit interaction, and participant as a random effect. The UIP-like analysis used baseline, visit, treatment-by-visit interaction, baseline-by-visit interaction, and participant as a random effect.
The adjusted mean difference was calculated as Nintedanib minus Placebo. Therefore, the negative point estimates indicate a lower adjusted change in the nintedanib group relative to placebo on the reported score scale. The clinical direction of a score change should be interpreted using the scoring convention of the instrument; the statistical comparison itself is the reported between-group difference.
15. Secondary Results: L-PF Cough Domain
| Population | Adjusted mean difference | 95% CI | P-value |
|---|---|---|---|
| Overall population | -6.09 | -9.65 to -2.53 | Not formally tested |
| HRCT UIP-like fibrotic pattern only | -7.28 | -11.86 to -2.71 | Not formally tested |
These analyses used MMRM and assessed absolute change from baseline in the L-PF Symptoms Cough Domain Score at week 52. The overall-population model included baseline, HRCT fibrotic pattern, visit, treatment-by-visit interaction, baseline-by-visit interaction, and participant as a random effect. The UIP-like analysis included baseline, visit, treatment-by-visit interaction, baseline-by-visit interaction, and participant as a random effect.
The confidence intervals for both cough analyses are entirely below zero. That describes the direction and precision of the estimated adjusted mean difference under the model. It does not, by itself, establish a clinically meaningful threshold because the ClinicalTrials.gov record does not provide a minimally clinically important difference for this endpoint.
The absence of formal hypothesis testing is also important. These results should not be treated as though each were a separately powered confirmatory endpoint simply because a confidence interval and model-based estimate are available.
16. Statistical Methods Explained
Why was a mixed-effects model used for the primary FVC endpoint?
FVC was measured repeatedly over time, and the registered endpoint was an annual rate of decline. A mixed-effects model is designed for longitudinal observations because measurements from the same participant are correlated rather than independent. The model can therefore estimate a treatment-group difference in trajectories while accounting for participant-level variation in intercepts and slopes.
What does an adjusted mean difference of 106.96 mL/year mean?
It is the modeled difference in annual FVC decline between nintedanib and placebo, after applying the specified longitudinal model. Because the comparison is Nintedanib minus Placebo and the endpoint is a rate of decline, a positive difference means the estimated FVC trajectory is higher in the nintedanib group than in the placebo group.
It does not mean that every patient loses exactly 106.96 mL less FVC per year. It is a population-level model estimate of the difference in average trajectory.
Why use MMRM for K-BILD and L-PF scores?
These endpoints were assessed repeatedly at multiple visits. MMRM permits the analysis to use the longitudinal outcome structure, including baseline and visit effects and treatment-by-visit interactions. Participant-level random effects account for repeated observations within individuals.
What does an HR of 0.65 mean?
An HR of 0.65 means that, under the Cox proportional-hazards model, the estimated instantaneous rate of the event is 0.65 times the corresponding rate in the comparator group. The complementary expression, 1 − 0.65 = 0.35, corresponds to an approximately 35% lower estimated hazard.
This does not mean a 35% absolute reduction in the probability of progression or death, nor does it mean that every patient experiences the same proportional change.
Why is the confidence interval essential?
A point estimate is only one summary of the observed data. The confidence interval provides a range representing statistical uncertainty under the model and sampling framework. For example, the progression-or-death HR of 0.65 has a 95% CI from 0.49 to 0.85, so the estimate should be understood as uncertain rather than as exactly 0.65.
Why should odds ratios not be interpreted as risk ratios?
Logistic regression models odds rather than probabilities. An odds ratio of 0.50 means that the modeled odds are half as large. The corresponding probability reduction depends on the baseline probability, so the same odds ratio can correspond to different absolute risk differences in different settings.
Why does a P-value not measure treatment-effect size?
The P-value summarizes how compatible the observed data are with a specified null hypothesis under the statistical model. It depends on both the size of the estimated effect and the amount of information available. A small P-value therefore does not automatically mean a large clinical effect, while a larger P-value does not prove that the treatment effects are identical.
17. Time-to-Event Analysis in INBUILD
Three families of time-to-event outcomes appear in the ClinicalTrials.gov record: acute ILD exacerbation or death, death, and progression or death. Each begins at first drug intake and uses the date of the relevant event or the last contact date for follow-up, with the specified analyses extending up to 372 days.
| Endpoint | Overall HR | 95% CI | P-value | Model detail |
|---|---|---|---|---|
| Acute ILD exacerbation or death | 0.80 | 0.48 to 1.34 | 0.3948 | Stratified by HRCT fibrotic pattern |
| Death | 0.94 | 0.47 to 1.86 | 0.8544 | Stratified by HRCT fibrotic pattern |
| Progression or death | 0.65 | 0.49 to 0.85 | 0.0017 | Stratified by HRCT fibrotic pattern |
Why censoring matters
Time-to-event analyses can include participants whose event has not occurred by the end of their available follow-up. Such observations are censored rather than treated as if the event never occurred. The statistical contribution of a censored participant therefore depends on how long that participant was known to be event-free.
This is one reason a time-to-event analysis cannot be reconstructed correctly from a simple count of events without the corresponding event times and censoring information.
The proportional-hazards assumption
The Cox model used to estimate the reported hazard ratios is a proportional-hazards model. Conceptually, the model assumes that the relative hazard between treatment groups can be represented by a common hazard ratio over the modeled time scale. If that relationship changes substantially over time, a single HR can become a less complete description of the treatment effect.
18. Covariate Adjustment and Stratification
Several secondary analyses explicitly incorporate baseline covariates. The logistic regression analyses adjust for baseline FVC percent predicted and, where specified, HRCT fibrotic pattern. The overall-population time-to-event analyses were stratified by HRCT fibrotic pattern.
Covariate adjustment
Logistic regression uses baseline FVC percent predicted as a continuous covariate and, for specified analyses, HRCT fibrotic pattern as a binary covariate.
Stratified survival analysis
For selected time-to-event analyses, the Cox model was stratified by HRCT fibrotic pattern so that the baseline hazard could differ across the specified pattern strata.
Adjustment and stratification serve different statistical purposes. Covariate adjustment models the association between the outcome and specified baseline variables, while stratification permits different baseline hazards across strata without estimating a single coefficient for the stratification factor in the same way as an ordinary covariate.
19. Analysis Populations
The ClinicalTrials.gov record explicitly define the populations used for each endpoint. The primary FVC analyses used randomized participants who received at least 1 dose of trial medication. Secondary repeated-measures analyses additionally required participants to contribute to the model evaluation.
| Population concept | Registry definition in the analyses posted on ClinicalTrials.gov |
|---|---|
| Overall primary FVC population | All randomized participants in the overall population who received at least 1 dose of trial medication. |
| UIP-like primary FVC population | All randomized participants with HRCT fibrotic pattern=UIP-like fibrotic pattern only who received at least 1 dose of trial medication. |
| Repeated-measures secondary analyses | Randomized participants who received at least 1 dose and contributed to the model evaluation. |
| Time-to-event analyses | Randomized participants in the specified population who received at least 1 dose of trial medication. |
This distinction is statistically important. Randomization establishes the treatment comparison, but the analyses posted on ClinicalTrials.gov do not simply use an undifferentiated set of all enrolled participants. Each endpoint has a specified analysis population, and interpretation should remain tied to that population.
20. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm. The reported counts are based on affected participants over the corresponding at-risk populations.
| Safety measure | Nintedanib | Placebo |
|---|---|---|
| Serious adverse events, affected / at risk | 147 / 332 | 164 / 331 |
These data provide arm-specific serious-adverse-event counts and denominators. They should be interpreted separately from the efficacy analyses because safety is an exposure-related outcome, whereas the primary efficacy endpoint is a randomized longitudinal comparison of FVC trajectories.
21. Missing Data and Repeated Measurements
Longitudinal clinical-trial data inevitably involve measurements that may not be available at every planned visit. The registry description gives important information about the modeling framework, including repeated measurements, participant-level random effects, the assumed error structure, and the model evaluation population. It does not provide a separate detailed missing-data or imputation strategy in the registry-reported statistical-analysis text.
That distinction matters because a mixed model is not synonymous with a particular imputation method. The model defines how observed repeated measurements contribute to estimation; the assumptions concerning why observations are missing and whether the model adequately represents the observed-data process remain part of the broader statistical analysis.
22. Multiplicity and Hypothesis Testing
The ClinicalTrials.gov record identifies the two FVC endpoints as primary and classify their hypothesis type as superiority. Several secondary analyses are explicitly labeled Other / not stated, with the registry comment that no formal hypotheses were tested.
| Analysis family | Hypothesis designation | Interpretation |
|---|---|---|
| Overall annual FVC decline | Superiority | Primary formal comparison |
| UIP-like annual FVC decline | Superiority | Primary formal comparison |
| K-BILD | Other / not stated | No formal hypotheses were tested |
| Acute ILD exacerbation or death | Other / not stated | No formal hypotheses were tested |
| Death | Other / not stated | No formal hypotheses were tested |
| Progression or death | Other / not stated | No formal hypotheses were tested |
| Binary FVC decline thresholds | Other / not stated | No formal hypotheses were tested |
| L-PF symptom domains | Other / not stated | No formal hypotheses were tested |
Consequently, the secondary P-values should not be treated as if each represented a separately powered confirmatory claim. The distinction between primary superiority testing and secondary analyses without formal hypothesis testing is central to a statistically disciplined reading of the results.
23. Interim Analysis, Non-Inferiority, Crossover, and Bayesian Methods
Non-inferiority margin
The ClinicalTrials.gov record does not report a non-inferiority margin. The primary hypotheses are described as superiority.
Crossover
The ClinicalTrials.gov record does not report a crossover design or crossover analysis.
Factorial design
The registered design model is parallel, not factorial.
Bayesian methods
No Bayesian method is reported in the statistical analyses posted on ClinicalTrials.gov. The posted methods are frequentist mixed models, MMRM, log-rank testing, and logistic regression.
The ClinicalTrials.gov record also does not identify an interim analysis in the ClinicalTrials.gov record. Accordingly, no interim alpha-spending or stopping-boundary interpretation is attributed to INBUILD on this page.
24. Statistical Methods Explained in More Depth
Why does the primary endpoint use a rate instead of a week-52 difference?
The registered endpoint is an annual rate of decline. The trial planned repeated FVC measurements through 52 weeks, allowing the analysis to characterize the trajectory over time. This is conceptually different from simply subtracting baseline FVC from the week-52 measurement.
A rate-based estimand can make better use of the repeated measurements when the scientific question concerns the progression of lung function rather than only its value at one final visit.
What does the unstructured covariance assumption mean?
The registry-reported primary-analysis description states that participant-specific intercepts and slopes were assumed to have an unstructured covariance matrix. In practical terms, the model does not impose a simple fixed relationship between the variability of participant-specific intercepts, participant-specific slopes, and their covariance. The covariance structure is therefore flexible within the model framework.
Why use the Kenward-Roger approximation?
Mixed models do not always have straightforward denominator degrees of freedom for fixed-effect tests. The Kenward-Roger approximation provides an adjusted inference procedure that accounts for estimation of the covariance parameters. The registry specifically identifies it as part of the primary FVC analysis.
Why include treatment-by-visit interaction in MMRM?
A treatment-by-visit interaction allows the difference between treatment groups to vary across visits. This is useful for repeated-measures endpoints because the treatment contrast does not have to be assumed constant at every scheduled time point.
What does a negative adjusted mean difference mean for L-PF scores?
The registry defines the adjusted mean difference as Nintedanib minus Placebo. A negative estimate therefore means the modeled change is numerically lower in the nintedanib group. Whether that represents improvement or worsening depends on the direction of the specific instrument's scoring scale; the statistical sign alone should not be translated into a clinical direction without knowing the instrument's scoring convention.
Why use logistic regression for the FVC threshold endpoints?
The outcome is binary: participants either meet the specified relative-decline threshold or they do not. Logistic regression is designed for this structure and permits adjustment for baseline FVC percent predicted and, for specified analyses, HRCT fibrotic pattern.
25. Limitations and Interpretation Issues
- Longitudinal-model assumptions: the primary FVC analysis assumes that the decrease in FVC is linear within each participant over 52 weeks and specifies normality and covariance assumptions for the model components.
- Analysis population: the primary analyses are defined among randomized participants who received at least 1 dose, so the numerical estimates should not be described as estimates based on every enrolled participant regardless of treatment exposure.
- Subgroup interpretation: the UIP-like fibrotic-pattern endpoint is population-specific. A numerical difference between the overall and UIP-like estimates is not a formal test of heterogeneity.
- Time-to-event assumptions: Cox hazard ratios rely on the proportional-hazards model. The ClinicalTrials.gov record does not report a diagnostic assessment of that assumption.
- Secondary hypothesis testing: the registry states that no formal hypotheses were tested for the reported secondary analyses. Their P-values should therefore not be interpreted as independent confirmatory findings.
- Odds ratios: logistic-regression odds ratios should not be translated directly into risk ratios or absolute probability differences.
- Confidence intervals: a confidence interval quantifies statistical uncertainty around the estimated effect; it does not describe the distribution of individual patient responses.
- Missing-data strategy: the ClinicalTrials.gov record describes the longitudinal models but do not provide a separate detailed imputation strategy for the primary FVC endpoint.
- Clinical meaning of patient-reported scores: the ClinicalTrials.gov record provides adjusted mean differences but do not provide a minimally clinically important difference for the K-BILD or L-PF domains.
- Safety comparison: the ClinicalTrials.gov record consists of serious-adverse-event counts and denominators; it does not provide a formal comparative statistical test.
26. Why This Trial Matters Statistically
INBUILD is a useful teaching example because several common clinical-trial methods appear in one trial record. The primary endpoint is longitudinal and continuous, while the secondary outcomes span repeated patient-reported measures, time-to-event outcomes, and binary clinical thresholds.
| Statistical concept | How it appears in INBUILD |
|---|---|
| Randomization | Randomized allocation to nintedanib or placebo. |
| Double masking | The trial is registered as double-masked. |
| Longitudinal analysis | Repeated FVC measurements are modeled through 52 weeks. |
| Mixed-effects model | Primary annual FVC decline is analyzed with a mixed-effects model. |
| Kenward-Roger approximation | Used to estimate denominator degrees of freedom and adjust inference in the primary FVC analysis. |
| MMRM | Used for K-BILD and L-PF repeated-measures outcomes. |
| Hazard ratio | Used for acute ILD exacerbation or death, death, and progression or death. |
| Log-rank test | Used for the reported time-to-event comparisons. |
| Cox model | Used to estimate hazard ratios for time-to-event outcomes. |
| Logistic regression | Used for relative FVC decline thresholds greater than 10% and greater than 5%. |
| Covariate adjustment | Baseline FVC percent predicted and, for specified analyses, HRCT fibrotic pattern enter the logistic models. |
| Stratified analysis | Selected time-to-event analyses are stratified by HRCT fibrotic pattern. |
| Confidence intervals | Reported for the primary and secondary effect estimates. |
| Secondary endpoint interpretation | Many secondary analyses explicitly state that no formal hypotheses were tested. |
27. What the Primary FVC Estimate Does — and Does Not — Mean
The overall-population adjusted mean difference of 106.96 mL/year summarizes the modeled difference in annual FVC decline between nintedanib and placebo. It is a population-level treatment comparison derived from repeated measurements.
It does not mean that every individual participant experienced exactly 106.96 mL/year less decline. Individual trajectories vary, and the estimate summarizes the modeled average treatment difference.
The 95% CI of 65.42 to 148.50 mL/year expresses uncertainty around the estimated difference. A single point estimate should therefore not be presented without its corresponding uncertainty interval when describing the statistical result.
The P-value of <.0001 indicates strong statistical evidence against the relevant null comparison, but it does not tell us whether the effect is large, small, or clinically important. The estimated magnitude and its confidence interval answer the effect-size question more directly.
28. Relative Effects and Absolute Clinical Measures
The INBUILD results illustrate why statistical interpretation should not rely on a single effect measure. The primary endpoints use an adjusted mean difference in an annual FVC trajectory, while the time-to-event outcomes use hazard ratios and the categorical FVC outcomes use odds ratios.
| Effect measure | INBUILD example | Core question answered |
|---|---|---|
| Adjusted mean difference | 106.96 mL/year for overall annual FVC decline | How different are the modeled mean longitudinal rates between groups? |
| Hazard ratio | 0.65 for progression or death in the overall population | How do the estimated instantaneous event rates compare under the Cox model? |
| Odds ratio | 0.50 for relative FVC decline greater than 5% in the overall population | How do the modeled odds of the binary outcome compare? |
These measures should not be converted into one another without the necessary underlying information. In particular, an odds ratio cannot be treated as a risk ratio, and a hazard ratio cannot be interpreted as an absolute difference in event probability.
29. Trial Timeline
Trial start
The registered trial start date was January 17, 2017.
Primary completion
The registered primary completion date was April 23, 2019.
Registry status
The trial is registered as completed, with results posted on ClinicalTrials.gov.
30. Overall Statistical Reading of the Evidence
The primary FVC analyses provide two adjusted mean differences favoring a higher modeled annual FVC trajectory in the nintedanib group relative to placebo: 106.96 mL/year in the overall population and 128.20 mL/year in participants with the specified UIP-like fibrotic pattern. Both confidence intervals are entirely above zero, and both reported P-values are <.0001.
The secondary results are more heterogeneous in statistical form. Progression or death produced hazard ratios of 0.65 overall and 0.64 in the UIP-like population, while acute ILD exacerbation or death and time to death produced wider confidence intervals. Binary FVC decline analyses produced adjusted odds ratios below 1, and several repeated patient-reported outcomes produced adjusted mean differences with their corresponding confidence intervals.
The most important statistical distinction is between the primary superiority analyses and the many secondary analyses for which the registry states that no formal hypotheses were tested. A statistically disciplined summary therefore preserves the hierarchy of the trial rather than treating every reported P-value as equivalent evidence.
31. Related Tutorials
Learn more about the methods used in this trial:
32. Related Statistical Calculators
33. Sources
- ClinicalTrials.gov: INBUILD — NCT02999178.
- PubMed record: PMID 42084658.
- PubMed record: PMID 37751022.
- PubMed record: PMID 36894966.
- PubMed record: PMID 36642509.
- PubMed record: PMID 35927208.
Continue with the underlying statistical methods
Explore the statistical concepts used in INBUILD through Clinical Biostats tutorials, calculators, and clinical-trial analyses.
34. Record Summary
INBUILD provides a broad statistical teaching example centered on a longitudinal continuous endpoint. Its two primary analyses estimate annual FVC decline using mixed-effects models, with repeated measurements through 52 weeks and explicit assumptions concerning participant-specific intercepts and slopes, covariance, and within-participant error. The registry also reports MMRM analyses for patient-reported outcomes, Cox and log-rank methods for time-to-event endpoints, and logistic regression for binary FVC decline thresholds.
The primary overall-population estimate was an adjusted mean difference of 106.96 mL/year with a 95% CI of 65.42 to 148.50 and P<.0001. In the specified UIP-like fibrotic-pattern population, the corresponding estimate was 128.20 mL/year with a 95% CI of 70.81 to 185.59 and P<.0001. Secondary endpoints provide additional measures of disease progression, patient-reported outcomes, and categorical FVC decline, but the registry states that no formal hypotheses were tested for these analyses.
The statistical story is therefore best understood through the combination of longitudinal effect estimates, confidence intervals, time-to-event hazard ratios, adjusted odds ratios, analysis populations, and model assumptions. Each answers a different statistical question, and none should be treated as a substitute for the others.