← Clinical Trials
Progressive Fibrosing ILD Phase 3 Completed NCT02999178

INBUILD: Complete Statistical Analysis of Nintedanib in Progressive Fibrosing Interstitial Lung Disease

An independent statistical review of the randomized phase 3 INBUILD trial evaluating nintedanib versus placebo in patients with progressive fibrosing interstitial lung disease, with emphasis on longitudinal FVC analysis, repeated-measures models, time-to-event methods, logistic regression, and interpretation of uncertainty.

Trial start: 2017-01-17  ·  Primary completion: 2019-04-23  ·  Enrollment: 663
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. The numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

INBUILD was a randomized, double-blind, parallel phase 3 trial evaluating 150 mg nintedanib versus placebo in patients with progressive fibrosing interstitial lung disease. The registry reports 663 enrolled participants, two arms, two registered primary endpoints, 20 posted outcome measures, and 18 posted statistical analyses.

663
Enrolled
Phase 3 trial
2
Arms
Nintedanib vs placebo
2
Primary endpoints
Both formally analyzed
18
Analyses posted
20 outcome measures
FeatureINBUILD
Trial nameINBUILD
Brief titleEfficacy and Safety of Nintedanib in Patients With Progressive Fibrosing Interstitial Lung Disease (PF-ILD)
PhasePhase 3
StatusCompleted
ConditionLung Diseases, Interstitial
AllocationRandomized
Design modelParallel
MaskingDouble
Primary purposeTreatment
Enrollment663
InterventionsNintedanib; Placebo
Lead sponsorBoehringer Ingelheim
Sponsor typeIndustry
ClinicalTrials.govNCT02999178

2. Clinical Question

The central statistical question was whether 150 mg nintedanib differed from placebo in the annual rate of decline in forced vital capacity (FVC) over the planned 52-week assessment period, both in the overall population and in participants with an HRCT fibrotic pattern described as UIP-like fibrotic pattern only.

Population

Randomized participants with progressive fibrosing interstitial lung disease, analyzed either in the overall population or in the HRCT UIP-like fibrotic-pattern subgroup specified by the registered endpoint.

Intervention

150 mg nintedanib.

Comparator

Placebo.

Primary question

Does nintedanib change the modeled annual rate of FVC decline compared with placebo over 52 weeks?

The ClinicalTrials.gov record also contain secondary analyses covering K-BILD total score, L-PF symptom domains, acute ILD exacerbation or death, death, progression or death, and categorical FVC decline thresholds. These endpoints use different statistical estimands, so they should not be treated as interchangeable measures of the same treatment effect.

3. Trial Design

01
Randomize663 enrolled
02
Two armsNintedanib or placebo
03
Repeated FVCBaseline through 52 weeks
04
Secondary outcomesPatient-reported and event-based
05
Statistical analysisMixed models, survival, logistic regression
INTERVENTION

150 mg Nintedanib

  • Drug intervention.
  • Compared with placebo under randomized allocation.
  • Primary analysis population required receipt of at least 1 dose of trial medication.
CONTROL

Placebo

  • Placebo intervention.
  • Compared with 150 mg nintedanib.
  • Primary analysis population required receipt of at least 1 dose of trial medication.
Allocation
Randomized, parallel-group design.
Masking
Double-masked trial.
Primary purpose
Treatment.
Results availability
Results posted; 18 statistical analyses were reported for 20 outcome measures.

4. Endpoints

The registry identifies two primary endpoints. Both concern annual FVC decline and were formally analyzed using mixed-model methodology.

Primary endpointRegistry time frameStatistical approach
Annual Rate of Decline in Forced Vital Capacity - Overall Population Baseline, 2, 4, 6, 12, 24, 36, 52 weeks after first drug intake (planned post-baseline visits) Mixed Models Analysis; mixed-effects model
Annual Rate of Decline in Forced Vital Capacity - Participants With HRCT Fibrotic Pattern=UIP-like Fibrotic Pattern Only Baseline, 2, 4, 6, 12, 24, 36, 52 weeks after first drug intake (planned post-baseline visits) Mixed Models Analysis; mixed-effects model

What FVC represents

The registry defines forced vital capacity as the volume of air, measured in milliliter, that can be forcibly exhaled from the lungs after taking the deepest breath possible. The primary estimand is not a single FVC measurement at week 52. It is the annual rate of decline estimated from repeated measurements over the 52-week period.

Secondary endpoint families

Endpoint familyExamples in the posted analysesMethod
Patient-reported outcomesK-BILD total score; L-PF Symptoms Dyspnea; L-PF Symptoms CoughMMRM
Time-to-eventAcute ILD exacerbation or death; death; progression or deathLog-rank test with Cox model for HR
Binary FVC outcomesRelative decline greater than 10%; relative decline greater than 5%Logistic regression

For several secondary endpoints, the registry explicitly states that no formal hypotheses were tested. Those estimates are therefore best interpreted as descriptive or supportive analyses rather than as independent confirmatory tests.

5. Statistical Methodology

Mixed-effects model for the primary FVC endpoint

The primary analysis used a mixed-effects model for repeated FVC measurements. The registry states that the decrease in FVC was assumed to be linear within each participant over 52 weeks. Participant-specific intercepts and slopes were assumed to be normally distributed with an unstructured covariance matrix, and the within-participant error was assumed to be independent and normally distributed with mean 0 and a common variance.

The model therefore uses the entire planned sequence of FVC measurements rather than reducing each participant to a single observed value. This is statistically important because the endpoint is defined as a rate of decline. A model of the longitudinal trajectory can use the pattern of repeated observations to estimate the treatment difference in slope.

Conceptual model
FVCij = intercepti + slopei × timej + errorij

The INBUILD registry analysis extends this longitudinal idea with treatment and other model components. The key estimand is the difference between treatment-group rates of FVC decline over the analyzed period.

Kenward-Roger approximation

The registry reports use of the Kenward-Roger approximation to estimate denominator degrees of freedom and adjust the covariance structure used for inference. This is a small-sample refinement commonly used with mixed models because the effective degrees of freedom are not necessarily obvious from the nominal number of observations.

MMRM for repeated patient-reported outcomes

K-BILD and L-PF symptom-domain outcomes were analyzed using mixed model repeated measures. These models incorporate repeated observations over visits and include participant-level random effects. The analyses posted on ClinicalTrials.gov specify baseline terms, visit, treatment-by-visit interactions, and baseline-by-visit interactions, with the exact fixed-effect structure varying by endpoint.

Log-rank test and Cox proportional-hazards model

Time to acute ILD exacerbation or death, time to death, and time to progression or death were evaluated with log-rank methods. Where a hazard ratio was reported, the registry states that a Cox proportional-hazards model was used to estimate the hazard ratio as Nintedanib divided by Placebo. Some analyses were stratified by HRCT fibrotic pattern.

Hazard-ratio interpretation
HR = hazard in Nintedanib group / hazard in Placebo group

An HR below 1 indicates a lower estimated instantaneous event rate in the nintedanib group under the Cox model. It does not directly equal an absolute risk reduction, a probability of benefit, or a percentage of patients who experience an event.

Logistic regression for categorical FVC decline

The relative decline endpoints at week 52 were analyzed with logistic regression. The models included baseline FVC percent predicted as a continuous covariate, with HRCT fibrotic pattern included as a binary covariate where specified. The reported effect measure is an adjusted odds ratio calculated as Nintedanib divided by Placebo.

6. Primary Results: Annual Rate of FVC Decline — Overall Population

The first primary endpoint was the annual rate of decline in FVC in the overall population. The analysis included all randomized participants in the overall population who received at least 1 dose of trial medication.

Adjusted mean difference in annual FVC decline

106.96 mL/year

95% CI: 65.42 to 148.50 mL/year   ·   P < .0001

150 mg Nintedanib vs Placebo; mixed-effects model.

Primary endpointNintedanibPlaceboAdjusted mean difference95% CIP-value
Annual Rate of Decline in FVC - Overall Population 150 mg Nintedanib vs Placebo 106.96 mL/year 65.42 to 148.50 <.0001
Clinical Biostats interpretation

The adjusted mean difference of 106.96 mL/year is the modeled difference in annual FVC decline between nintedanib and placebo. Because the endpoint concerns decline, a positive difference in the Nintedanib-minus-Placebo comparison corresponds to a less negative FVC trajectory for the nintedanib group under the model.

The 95% confidence interval of 65.42 to 148.50 mL/year describes the statistical uncertainty around the estimated between-group difference under the model. It does not mean that individual patients' FVC changes must fall inside this interval.

The P-value of <.0001 addresses the statistical evidence against the null comparison specified by the analysis. It does not measure the size or clinical importance of the treatment effect. The effect estimate and its confidence interval are needed to understand magnitude and precision.

The analysis also relies on the stated longitudinal assumptions, including a linear decline within each participant over 52 weeks and the specified distributional and covariance assumptions. As with any longitudinal model, departures from those assumptions can affect the estimated slope difference.

Why a longitudinal model is appropriate here

FVC was planned to be measured repeatedly at baseline and at 2, 4, 6, 12, 24, 36, and 52 weeks after first drug intake. A slope-based analysis uses these repeated measurements to characterize the trajectory rather than comparing only two observations. This makes the statistical estimand directly aligned with the registered endpoint: annual rate of decline.

7. Primary Results: Annual Rate of FVC Decline — UIP-like Fibrotic Pattern

The second primary endpoint restricted the analysis to participants with HRCT fibrotic pattern described as UIP-like fibrotic pattern only. The analysis population consisted of all randomized participants meeting that pattern definition who received at least 1 dose of trial medication.

Adjusted mean difference in annual FVC decline

128.20 mL/year

95% CI: 70.81 to 185.59 mL/year   ·   P < .0001

150 mg Nintedanib vs Placebo; mixed-effects model.

Primary endpointComparisonAdjusted mean difference95% CIP-value
Annual Rate of Decline in FVC - Participants With HRCT Fibrotic Pattern=UIP-like Fibrotic Pattern Only 150 mg Nintedanib vs Placebo 128.20 mL/year 70.81 to 185.59 <.0001
Clinical Biostats interpretation

The estimated adjusted mean difference was 128.20 mL/year, with the comparison defined as Nintedanib minus Placebo. For an endpoint measuring decline, the positive difference corresponds to a less negative modeled FVC trajectory in the nintedanib group.

The 95% CI of 70.81 to 185.59 mL/year provides the precision interval reported by the registry. It indicates that the estimated treatment difference is not represented by a single exact number; uncertainty remains around the model-based estimate.

The P-value of <.0001 provides evidence against the relevant null comparison but is not a measure of how large or clinically meaningful the effect is. The magnitude and confidence interval are more informative for understanding the estimated difference.

This endpoint is also a prespecified subgroup-defined primary endpoint. Its interpretation should remain tied to the population specified by the registry rather than being generalized automatically to every possible fibrotic pattern.

8. Comparing the Two Primary FVC Analyses

Primary endpointEstimate95% CIP-valueInterpretation of direction
Overall population 106.96 mL/year 65.42 to 148.50 <.0001 Higher modeled annual FVC rate relative to placebo; because the endpoint is decline, this corresponds to less decline.
UIP-like fibrotic pattern only 128.20 mL/year 70.81 to 185.59 <.0001 Higher modeled annual FVC rate relative to placebo; because the endpoint is decline, this corresponds to less decline.

The two estimates address different populations. The overall-population estimate uses all randomized participants in the defined overall population who received at least one dose, whereas the second estimate is restricted to the specified UIP-like HRCT fibrotic pattern. The difference between the numerical estimates should therefore not be treated as a formal test that treatment effects differ between these populations. A comparison of effect estimates across populations requires an appropriate interaction or heterogeneity analysis.

Important distinction: the ClinicalTrials.gov record provides two primary endpoint estimates, but they do not provide a formal interaction test comparing the treatment effect between the overall population and the UIP-like fibrotic-pattern population. Therefore, the numerical difference between 106.96 and 128.20 mL/year should not be described as evidence of treatment-effect heterogeneity.

9. Secondary Results: K-BILD Total Score

The registry reports absolute change from baseline in K-BILD total score at week 52 using MMRM. These analyses were not associated with formal hypothesis testing in the ClinicalTrials.gov record.

PopulationAdjusted mean difference95% CIP-value
Overall population 1.34 -0.31 to 2.98 0.1115
HRCT UIP-like fibrotic pattern only 1.53 -0.68 to 3.74 0.1747

Overall population

The adjusted mean difference was 1.34 units, calculated as Nintedanib minus Placebo, with a 95% CI of -0.31 to 2.98 and P=0.1115. The MMRM included baseline K-BILD total score, visit, treatment-by-visit and baseline-by-visit interactions, with participant as a random effect.

UIP-like fibrotic pattern only

The adjusted mean difference was 1.53 units, with a 95% CI of -0.68 to 3.74 and P=0.1747. The registry explicitly states that no formal hypotheses were tested for this endpoint.

How to read these results: the confidence intervals cross zero in both analyses. The registry's explicit statement that no formal hypotheses were tested is important: the reported P-values should not be treated as if these were independent confirmatory tests within the trial's primary error-control framework.

10. Secondary Results: Acute ILD Exacerbation or Death

Time to first acute ILD exacerbation or death over 52 weeks was analyzed with a log-rank method. The hazard ratio was estimated with a Cox proportional-hazards model. In the overall population, the analysis was stratified by HRCT fibrotic pattern.

PopulationHR95% CIP-valueAnalysis
Overall population 0.80 0.48 to 1.34 0.3948 Log-rank; stratified Cox model
HRCT UIP-like fibrotic pattern only 0.67 0.36 to 1.24 0.1985 Log-rank; Cox model

The time frame was from first drug intake until the date of first acute ILD exacerbation, death, or last contact date, up to 372 days. No formal hypotheses were tested for these secondary analyses according to the ClinicalTrials.gov record.

Reading the hazard ratios

An HR of 0.80 means that the estimated instantaneous event rate in the nintedanib group was 0.80 times the corresponding rate in the placebo group under the Cox model. This is a relative model-based measure; it is not an 80% probability of avoiding an event.

The 95% CI of 0.48 to 1.34 includes 1.00, indicating substantial statistical uncertainty about the direction and magnitude of the relative hazard difference in this analysis.

The UIP-like estimate of 0.67 has a 95% CI of 0.36 to 1.24. Again, the interval includes 1.00. The numerical difference between the two subgroup-specific estimates does not by itself establish that the treatment effects differ.

11. Secondary Results: Time to Death

PopulationHR95% CIP-valueAnalysis
Overall population 0.94 0.47 to 1.86 0.8544 Log-rank; stratified Cox model
HRCT UIP-like fibrotic pattern only 0.68 0.32 to 1.47 0.3291 Log-rank; Cox model

The endpoint was time to death over 52 weeks, measured from first drug intake until death or last contact date, up to 372 days. The overall-population analysis was stratified by HRCT fibrotic pattern.

Why the confidence interval matters

The overall-population HR of 0.94 is accompanied by a wide 95% CI of 0.47 to 1.86. The width of this interval shows that the reported estimate has limited precision for this time-to-death analysis. The P-value of 0.8544 does not quantify the size of the effect; it addresses the statistical evidence against the relevant null comparison.

The UIP-like estimate of 0.68 has a 95% CI of 0.32 to 1.47. The interval includes 1.00, so the point estimate should not be read as establishing a definitive relative hazard reduction.

12. Secondary Results: Time to Progression or Death

Overall population

HR 0.65

95% CI: 0.49 to 0.85   ·   P = 0.0017

Time to progression or death over 52 weeks; stratified by HRCT fibrotic pattern.

PopulationHR95% CIP-value
Overall population 0.65 0.49 to 0.85 0.0017
HRCT UIP-like fibrotic pattern only 0.64 0.45 to 0.89 0.0081

The registry reports a log-rank analysis with a Cox proportional-hazards model. The overall-population model was stratified by HRCT fibrotic pattern. The hazard ratio was calculated as Nintedanib divided by Placebo.

Clinical Biostats interpretation

An HR of 0.65 means that the estimated instantaneous rate of progression or death was 0.65 times the corresponding rate in the placebo group under the fitted Cox model. In relative terms, 1 - 0.65 = 0.35, so the point estimate corresponds to an approximately 35% lower estimated hazard. This is a direct interpretation of the hazard-ratio point estimate, not a statement that 35% of patients avoided progression or death.

The 95% CI of 0.49 to 0.85 quantifies uncertainty around the estimated hazard ratio. Because the entire reported interval is below 1.00, the estimated relative hazard is below 1 throughout the confidence interval.

The P-value of 0.0017 provides evidence against the relevant null comparison. It does not say that the treatment effect is 0.0017 in size, nor does it tell us how many additional days an individual patient will remain progression-free.

UIP-like fibrotic pattern only

For participants with the HRCT UIP-like fibrotic pattern only, the reported hazard ratio was 0.64, with a 95% CI of 0.45 to 0.89 and P=0.0081.

The point estimates of 0.65 and 0.64 are numerically close, but the ClinicalTrials.gov record does not provide a formal interaction test between these populations. Therefore, the results support describing the estimates separately rather than claiming that one population had a statistically different treatment effect.

13. Secondary Results: Relative FVC Decline Thresholds

The registry also reports binary endpoints based on whether participants experienced a relative decline from baseline in FVC percent predicted greater than 10% or greater than 5% at week 52. These were analyzed with logistic regression and adjusted odds ratios.

EndpointPopulationAdjusted OR95% CIP-value
Relative decline >10% Overall 0.70 0.52 to 0.96 Not stated as formal hypothesis
Relative decline >10% HRCT UIP-like pattern only 0.63 0.43 to 0.94 Not stated as formal hypothesis
Relative decline >5% Overall 0.50 0.36 to 0.68 Not stated as formal hypothesis
Relative decline >5% HRCT UIP-like pattern only 0.46 0.31 to 0.69 Not stated as formal hypothesis

For the overall-population >10% and >5% endpoints, the logistic regression model included baseline FVC percent predicted as a continuous covariate and HRCT fibrotic pattern as a binary covariate. For the corresponding UIP-like analyses, the model included baseline FVC percent predicted.

How to interpret an odds ratio

An adjusted odds ratio below 1 indicates lower modeled odds of the specified binary outcome in the nintedanib group relative to placebo. It is important not to call an odds ratio a risk ratio unless the two measures happen to coincide under special circumstances.

Odds are not probabilities
Odds = p / (1 − p)

Logistic regression models the odds of the specified binary outcome. An odds ratio of 0.50 therefore represents half the modeled odds, not necessarily half the probability of the outcome.

For example, the overall-population adjusted OR of 0.50 for a relative FVC decline greater than 5% corresponds to half the modeled odds under the specified logistic regression model. The registry does not provide the corresponding arm-specific event percentages in the ClinicalTrials.gov record, so an absolute risk difference should not be inferred from the odds ratio alone.

14. Secondary Results: L-PF Dyspnea Domain

PopulationAdjusted mean difference95% CIP-value
Overall population -3.53 -6.14 to -0.92 Not formally tested
HRCT UIP-like fibrotic pattern only -4.18 -7.48 to -0.88 Not formally tested

The endpoint was absolute change from baseline in the Living With Pulmonary Fibrosis (L-PF) Symptoms Dyspnea Domain Score at week 52. The overall-population MMRM included baseline, HRCT fibrotic pattern, visit, treatment-by-visit interaction, baseline-by-visit interaction, and participant as a random effect. The UIP-like analysis used baseline, visit, treatment-by-visit interaction, baseline-by-visit interaction, and participant as a random effect.

The adjusted mean difference was calculated as Nintedanib minus Placebo. Therefore, the negative point estimates indicate a lower adjusted change in the nintedanib group relative to placebo on the reported score scale. The clinical direction of a score change should be interpreted using the scoring convention of the instrument; the statistical comparison itself is the reported between-group difference.

15. Secondary Results: L-PF Cough Domain

PopulationAdjusted mean difference95% CIP-value
Overall population -6.09 -9.65 to -2.53 Not formally tested
HRCT UIP-like fibrotic pattern only -7.28 -11.86 to -2.71 Not formally tested

These analyses used MMRM and assessed absolute change from baseline in the L-PF Symptoms Cough Domain Score at week 52. The overall-population model included baseline, HRCT fibrotic pattern, visit, treatment-by-visit interaction, baseline-by-visit interaction, and participant as a random effect. The UIP-like analysis included baseline, visit, treatment-by-visit interaction, baseline-by-visit interaction, and participant as a random effect.

Reading the repeated-measures results

The confidence intervals for both cough analyses are entirely below zero. That describes the direction and precision of the estimated adjusted mean difference under the model. It does not, by itself, establish a clinically meaningful threshold because the ClinicalTrials.gov record does not provide a minimally clinically important difference for this endpoint.

The absence of formal hypothesis testing is also important. These results should not be treated as though each were a separately powered confirmatory endpoint simply because a confidence interval and model-based estimate are available.

16. Statistical Methods Explained

Why was a mixed-effects model used for the primary FVC endpoint?

FVC was measured repeatedly over time, and the registered endpoint was an annual rate of decline. A mixed-effects model is designed for longitudinal observations because measurements from the same participant are correlated rather than independent. The model can therefore estimate a treatment-group difference in trajectories while accounting for participant-level variation in intercepts and slopes.

What does an adjusted mean difference of 106.96 mL/year mean?

It is the modeled difference in annual FVC decline between nintedanib and placebo, after applying the specified longitudinal model. Because the comparison is Nintedanib minus Placebo and the endpoint is a rate of decline, a positive difference means the estimated FVC trajectory is higher in the nintedanib group than in the placebo group.

It does not mean that every patient loses exactly 106.96 mL less FVC per year. It is a population-level model estimate of the difference in average trajectory.

Why use MMRM for K-BILD and L-PF scores?

These endpoints were assessed repeatedly at multiple visits. MMRM permits the analysis to use the longitudinal outcome structure, including baseline and visit effects and treatment-by-visit interactions. Participant-level random effects account for repeated observations within individuals.

What does an HR of 0.65 mean?

An HR of 0.65 means that, under the Cox proportional-hazards model, the estimated instantaneous rate of the event is 0.65 times the corresponding rate in the comparator group. The complementary expression, 1 − 0.65 = 0.35, corresponds to an approximately 35% lower estimated hazard.

This does not mean a 35% absolute reduction in the probability of progression or death, nor does it mean that every patient experiences the same proportional change.

Why is the confidence interval essential?

A point estimate is only one summary of the observed data. The confidence interval provides a range representing statistical uncertainty under the model and sampling framework. For example, the progression-or-death HR of 0.65 has a 95% CI from 0.49 to 0.85, so the estimate should be understood as uncertain rather than as exactly 0.65.

Why should odds ratios not be interpreted as risk ratios?

Logistic regression models odds rather than probabilities. An odds ratio of 0.50 means that the modeled odds are half as large. The corresponding probability reduction depends on the baseline probability, so the same odds ratio can correspond to different absolute risk differences in different settings.

Why does a P-value not measure treatment-effect size?

The P-value summarizes how compatible the observed data are with a specified null hypothesis under the statistical model. It depends on both the size of the estimated effect and the amount of information available. A small P-value therefore does not automatically mean a large clinical effect, while a larger P-value does not prove that the treatment effects are identical.

17. Time-to-Event Analysis in INBUILD

Three families of time-to-event outcomes appear in the ClinicalTrials.gov record: acute ILD exacerbation or death, death, and progression or death. Each begins at first drug intake and uses the date of the relevant event or the last contact date for follow-up, with the specified analyses extending up to 372 days.

EndpointOverall HR95% CIP-valueModel detail
Acute ILD exacerbation or death 0.80 0.48 to 1.34 0.3948 Stratified by HRCT fibrotic pattern
Death 0.94 0.47 to 1.86 0.8544 Stratified by HRCT fibrotic pattern
Progression or death 0.65 0.49 to 0.85 0.0017 Stratified by HRCT fibrotic pattern

Why censoring matters

Time-to-event analyses can include participants whose event has not occurred by the end of their available follow-up. Such observations are censored rather than treated as if the event never occurred. The statistical contribution of a censored participant therefore depends on how long that participant was known to be event-free.

This is one reason a time-to-event analysis cannot be reconstructed correctly from a simple count of events without the corresponding event times and censoring information.

The proportional-hazards assumption

The Cox model used to estimate the reported hazard ratios is a proportional-hazards model. Conceptually, the model assumes that the relative hazard between treatment groups can be represented by a common hazard ratio over the modeled time scale. If that relationship changes substantially over time, a single HR can become a less complete description of the treatment effect.

Interpretive caution: the ClinicalTrials.gov record reports Cox proportional-hazards analyses, but they do not provide a diagnostic assessment of proportional hazards. The reported HRs should therefore be described as model-based summary measures rather than as complete descriptions of the entire event-time distributions.

18. Covariate Adjustment and Stratification

Several secondary analyses explicitly incorporate baseline covariates. The logistic regression analyses adjust for baseline FVC percent predicted and, where specified, HRCT fibrotic pattern. The overall-population time-to-event analyses were stratified by HRCT fibrotic pattern.

Covariate adjustment

Logistic regression uses baseline FVC percent predicted as a continuous covariate and, for specified analyses, HRCT fibrotic pattern as a binary covariate.

Stratified survival analysis

For selected time-to-event analyses, the Cox model was stratified by HRCT fibrotic pattern so that the baseline hazard could differ across the specified pattern strata.

Adjustment and stratification serve different statistical purposes. Covariate adjustment models the association between the outcome and specified baseline variables, while stratification permits different baseline hazards across strata without estimating a single coefficient for the stratification factor in the same way as an ordinary covariate.

19. Analysis Populations

The ClinicalTrials.gov record explicitly define the populations used for each endpoint. The primary FVC analyses used randomized participants who received at least 1 dose of trial medication. Secondary repeated-measures analyses additionally required participants to contribute to the model evaluation.

Population conceptRegistry definition in the analyses posted on ClinicalTrials.gov
Overall primary FVC population All randomized participants in the overall population who received at least 1 dose of trial medication.
UIP-like primary FVC population All randomized participants with HRCT fibrotic pattern=UIP-like fibrotic pattern only who received at least 1 dose of trial medication.
Repeated-measures secondary analyses Randomized participants who received at least 1 dose and contributed to the model evaluation.
Time-to-event analyses Randomized participants in the specified population who received at least 1 dose of trial medication.

This distinction is statistically important. Randomization establishes the treatment comparison, but the analyses posted on ClinicalTrials.gov do not simply use an undifferentiated set of all enrolled participants. Each endpoint has a specified analysis population, and interpretation should remain tied to that population.

20. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm. The reported counts are based on affected participants over the corresponding at-risk populations.

Safety measureNintedanibPlacebo
Serious adverse events, affected / at risk 147 / 332 164 / 331

These data provide arm-specific serious-adverse-event counts and denominators. They should be interpreted separately from the efficacy analyses because safety is an exposure-related outcome, whereas the primary efficacy endpoint is a randomized longitudinal comparison of FVC trajectories.

Safety interpretation: the ClinicalTrials.gov record does not provide a statistical hypothesis test for the serious-adverse-event comparison. The safest presentation is therefore the reported affected/at-risk counts rather than an inferred P-value or an unreported comparative measure.

21. Missing Data and Repeated Measurements

Longitudinal clinical-trial data inevitably involve measurements that may not be available at every planned visit. The registry description gives important information about the modeling framework, including repeated measurements, participant-level random effects, the assumed error structure, and the model evaluation population. It does not provide a separate detailed missing-data or imputation strategy in the registry-reported statistical-analysis text.

That distinction matters because a mixed model is not synonymous with a particular imputation method. The model defines how observed repeated measurements contribute to estimation; the assumptions concerning why observations are missing and whether the model adequately represents the observed-data process remain part of the broader statistical analysis.

Do not infer an imputation method: because the ClinicalTrials.gov record does not specify a separate imputation procedure for the primary FVC analysis, this page does not attribute one to the trial.

22. Multiplicity and Hypothesis Testing

The ClinicalTrials.gov record identifies the two FVC endpoints as primary and classify their hypothesis type as superiority. Several secondary analyses are explicitly labeled Other / not stated, with the registry comment that no formal hypotheses were tested.

Analysis familyHypothesis designationInterpretation
Overall annual FVC decline Superiority Primary formal comparison
UIP-like annual FVC decline Superiority Primary formal comparison
K-BILD Other / not stated No formal hypotheses were tested
Acute ILD exacerbation or death Other / not stated No formal hypotheses were tested
Death Other / not stated No formal hypotheses were tested
Progression or death Other / not stated No formal hypotheses were tested
Binary FVC decline thresholds Other / not stated No formal hypotheses were tested
L-PF symptom domains Other / not stated No formal hypotheses were tested

Consequently, the secondary P-values should not be treated as if each represented a separately powered confirmatory claim. The distinction between primary superiority testing and secondary analyses without formal hypothesis testing is central to a statistically disciplined reading of the results.

23. Interim Analysis, Non-Inferiority, Crossover, and Bayesian Methods

Non-inferiority margin

The ClinicalTrials.gov record does not report a non-inferiority margin. The primary hypotheses are described as superiority.

Crossover

The ClinicalTrials.gov record does not report a crossover design or crossover analysis.

Factorial design

The registered design model is parallel, not factorial.

Bayesian methods

No Bayesian method is reported in the statistical analyses posted on ClinicalTrials.gov. The posted methods are frequentist mixed models, MMRM, log-rank testing, and logistic regression.

The ClinicalTrials.gov record also does not identify an interim analysis in the ClinicalTrials.gov record. Accordingly, no interim alpha-spending or stopping-boundary interpretation is attributed to INBUILD on this page.

24. Statistical Methods Explained in More Depth

Why does the primary endpoint use a rate instead of a week-52 difference?

The registered endpoint is an annual rate of decline. The trial planned repeated FVC measurements through 52 weeks, allowing the analysis to characterize the trajectory over time. This is conceptually different from simply subtracting baseline FVC from the week-52 measurement.

A rate-based estimand can make better use of the repeated measurements when the scientific question concerns the progression of lung function rather than only its value at one final visit.

What does the unstructured covariance assumption mean?

The registry-reported primary-analysis description states that participant-specific intercepts and slopes were assumed to have an unstructured covariance matrix. In practical terms, the model does not impose a simple fixed relationship between the variability of participant-specific intercepts, participant-specific slopes, and their covariance. The covariance structure is therefore flexible within the model framework.

Why use the Kenward-Roger approximation?

Mixed models do not always have straightforward denominator degrees of freedom for fixed-effect tests. The Kenward-Roger approximation provides an adjusted inference procedure that accounts for estimation of the covariance parameters. The registry specifically identifies it as part of the primary FVC analysis.

Why include treatment-by-visit interaction in MMRM?

A treatment-by-visit interaction allows the difference between treatment groups to vary across visits. This is useful for repeated-measures endpoints because the treatment contrast does not have to be assumed constant at every scheduled time point.

What does a negative adjusted mean difference mean for L-PF scores?

The registry defines the adjusted mean difference as Nintedanib minus Placebo. A negative estimate therefore means the modeled change is numerically lower in the nintedanib group. Whether that represents improvement or worsening depends on the direction of the specific instrument's scoring scale; the statistical sign alone should not be translated into a clinical direction without knowing the instrument's scoring convention.

Why use logistic regression for the FVC threshold endpoints?

The outcome is binary: participants either meet the specified relative-decline threshold or they do not. Logistic regression is designed for this structure and permits adjustment for baseline FVC percent predicted and, for specified analyses, HRCT fibrotic pattern.

25. Limitations and Interpretation Issues

26. Why This Trial Matters Statistically

INBUILD is a useful teaching example because several common clinical-trial methods appear in one trial record. The primary endpoint is longitudinal and continuous, while the secondary outcomes span repeated patient-reported measures, time-to-event outcomes, and binary clinical thresholds.

Statistical conceptHow it appears in INBUILD
RandomizationRandomized allocation to nintedanib or placebo.
Double maskingThe trial is registered as double-masked.
Longitudinal analysisRepeated FVC measurements are modeled through 52 weeks.
Mixed-effects modelPrimary annual FVC decline is analyzed with a mixed-effects model.
Kenward-Roger approximationUsed to estimate denominator degrees of freedom and adjust inference in the primary FVC analysis.
MMRMUsed for K-BILD and L-PF repeated-measures outcomes.
Hazard ratioUsed for acute ILD exacerbation or death, death, and progression or death.
Log-rank testUsed for the reported time-to-event comparisons.
Cox modelUsed to estimate hazard ratios for time-to-event outcomes.
Logistic regressionUsed for relative FVC decline thresholds greater than 10% and greater than 5%.
Covariate adjustmentBaseline FVC percent predicted and, for specified analyses, HRCT fibrotic pattern enter the logistic models.
Stratified analysisSelected time-to-event analyses are stratified by HRCT fibrotic pattern.
Confidence intervalsReported for the primary and secondary effect estimates.
Secondary endpoint interpretationMany secondary analyses explicitly state that no formal hypotheses were tested.

27. What the Primary FVC Estimate Does — and Does Not — Mean

The estimate

The overall-population adjusted mean difference of 106.96 mL/year summarizes the modeled difference in annual FVC decline between nintedanib and placebo. It is a population-level treatment comparison derived from repeated measurements.

What it does not mean

It does not mean that every individual participant experienced exactly 106.96 mL/year less decline. Individual trajectories vary, and the estimate summarizes the modeled average treatment difference.

Why the confidence interval matters

The 95% CI of 65.42 to 148.50 mL/year expresses uncertainty around the estimated difference. A single point estimate should therefore not be presented without its corresponding uncertainty interval when describing the statistical result.

Why the P-value is secondary to the effect estimate

The P-value of <.0001 indicates strong statistical evidence against the relevant null comparison, but it does not tell us whether the effect is large, small, or clinically important. The estimated magnitude and its confidence interval answer the effect-size question more directly.

28. Relative Effects and Absolute Clinical Measures

The INBUILD results illustrate why statistical interpretation should not rely on a single effect measure. The primary endpoints use an adjusted mean difference in an annual FVC trajectory, while the time-to-event outcomes use hazard ratios and the categorical FVC outcomes use odds ratios.

Effect measureINBUILD exampleCore question answered
Adjusted mean difference 106.96 mL/year for overall annual FVC decline How different are the modeled mean longitudinal rates between groups?
Hazard ratio 0.65 for progression or death in the overall population How do the estimated instantaneous event rates compare under the Cox model?
Odds ratio 0.50 for relative FVC decline greater than 5% in the overall population How do the modeled odds of the binary outcome compare?

These measures should not be converted into one another without the necessary underlying information. In particular, an odds ratio cannot be treated as a risk ratio, and a hazard ratio cannot be interpreted as an absolute difference in event probability.

29. Trial Timeline

2017-01-17

Trial start

The registered trial start date was January 17, 2017.

2019-04-23

Primary completion

The registered primary completion date was April 23, 2019.

Completed

Registry status

The trial is registered as completed, with results posted on ClinicalTrials.gov.

30. Overall Statistical Reading of the Evidence

The primary FVC analyses provide two adjusted mean differences favoring a higher modeled annual FVC trajectory in the nintedanib group relative to placebo: 106.96 mL/year in the overall population and 128.20 mL/year in participants with the specified UIP-like fibrotic pattern. Both confidence intervals are entirely above zero, and both reported P-values are <.0001.

The secondary results are more heterogeneous in statistical form. Progression or death produced hazard ratios of 0.65 overall and 0.64 in the UIP-like population, while acute ILD exacerbation or death and time to death produced wider confidence intervals. Binary FVC decline analyses produced adjusted odds ratios below 1, and several repeated patient-reported outcomes produced adjusted mean differences with their corresponding confidence intervals.

The most important statistical distinction is between the primary superiority analyses and the many secondary analyses for which the registry states that no formal hypotheses were tested. A statistically disciplined summary therefore preserves the hierarchy of the trial rather than treating every reported P-value as equivalent evidence.

Clinical Biostats methodology: the purpose of this page is to explain the statistical structure of the trial, not merely reproduce a list of numbers. Effect estimates, confidence intervals, P-values, analysis populations, model assumptions, and endpoint hierarchy should be read together.

31. Related Tutorials

Learn more about the methods used in this trial:

32. Related Statistical Calculators

33. Sources

Continue with the underlying statistical methods

Explore the statistical concepts used in INBUILD through Clinical Biostats tutorials, calculators, and clinical-trial analyses.

34. Record Summary

INBUILD provides a broad statistical teaching example centered on a longitudinal continuous endpoint. Its two primary analyses estimate annual FVC decline using mixed-effects models, with repeated measurements through 52 weeks and explicit assumptions concerning participant-specific intercepts and slopes, covariance, and within-participant error. The registry also reports MMRM analyses for patient-reported outcomes, Cox and log-rank methods for time-to-event endpoints, and logistic regression for binary FVC decline thresholds.

The primary overall-population estimate was an adjusted mean difference of 106.96 mL/year with a 95% CI of 65.42 to 148.50 and P<.0001. In the specified UIP-like fibrotic-pattern population, the corresponding estimate was 128.20 mL/year with a 95% CI of 70.81 to 185.59 and P<.0001. Secondary endpoints provide additional measures of disease progression, patient-reported outcomes, and categorical FVC decline, but the registry states that no formal hypotheses were tested for these analyses.

The statistical story is therefore best understood through the combination of longitudinal effect estimates, confidence intervals, time-to-event hazard ratios, adjusted odds ratios, analysis populations, and model assumptions. Each answers a different statistical question, and none should be treated as a substitute for the others.