← Clinical Trials
Idiopathic Pulmonary Fibrosis Phase 3 Time-to-Event Analysis NCT00768300

ARTEMIS-IPF: Complete Statistical Analysis of Ambrisentan in Idiopathic Pulmonary Fibrosis

An independent statistical review of the randomized, double-blind, placebo-controlled ARTEMIS-IPF trial evaluating the safety and effectiveness of ambrisentan in idiopathic pulmonary fibrosis, with emphasis on its time-to-death-or-disease-progression endpoint and prespecified secondary analyses.

Phase 3  ·  Randomized  ·  Double-blind  ·  Parallel design  ·  Enrollment 494
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. View NCT00768300 on ClinicalTrials.gov.

1. Trial at a Glance

ARTEMIS-IPF was a phase 3 randomized, double-blind, parallel trial evaluating ambrisentan versus placebo in participants with idiopathic pulmonary fibrosis. The registry reports 494 enrolled participants, one registered primary time-to-event endpoint, nine posted outcome measures, and five posted statistical analyses.

494
Enrolled
ClinicalTrials.gov enrollment
2
Arms
Ambrisentan vs placebo
1.74
Primary HR
95% CI 1.14–2.66
0.010
Primary P-value
Two-sided
FeatureARTEMIS-IPF
Trial nameARTEMIS-IPF
PhasePhase 3
ConditionIdiopathic Pulmonary Fibrosis
DesignRandomized, double-blind, parallel
AllocationRandomized
Primary purposeTreatment
Enrollment494
InterventionsAmbrisentan; placebo
Primary endpoint typeTime-to-event
Primary hypothesis typeSuperiority
Trial statusTerminated
Start2008-12
Primary completion2011-02
Lead sponsorGilead Sciences
Sponsor typeIndustry
ClinicalTrials.govNCT00768300

2. Clinical Question

The primary statistical question was whether the randomized comparison of ambrisentan versus placebo differed with respect to time to death or disease progression in idiopathic pulmonary fibrosis. The registry classified the primary hypothesis as superiority.

Population

Participants enrolled in a phase 3 trial for idiopathic pulmonary fibrosis.

Intervention

Ambrisentan.

Comparator

Placebo.

Primary question

Does randomized assignment to ambrisentan versus placebo change the time to death or disease progression?

This question is inherently longitudinal. A participant who has not experienced death or the defined progression event by the end of available follow-up does not simply become an ordinary non-event; the participant contributes information up to the point of censoring. That structure is why survival-analysis methods are central to the primary analysis.

3. Trial Design

01
Randomize494 enrolled
02
Two armsAmbrisentan vs placebo
03
Double-blindMasked treatment assignment
04
FollowDeath / IPF progression
05
AnalyzeTime-to-event comparison
INTERVENTION

Ambrisentan

  • Drug intervention.
  • Participants were randomized to the ambrisentan arm.
  • Primary comparison was time to death or IPF progression.
CONTROL

Placebo

  • Placebo drug intervention.
  • Participants were randomized to the placebo arm.
  • Primary comparison was time to death or IPF progression.
Allocation
Randomized allocation to parallel treatment groups.
Masking
Double-blind.
Purpose
Treatment.
Phase
Phase 3.

The design has an important statistical implication: randomization establishes the framework for comparing outcomes between the two assigned groups, while double blinding reduces the potential for knowledge of treatment assignment to influence treatment administration, outcome assessment, or participant behavior.

4. Endpoints

EndpointTime frameTypeRegistry definition / analysis
Time to Death or Disease (IPF) Progression Up to 48 months Time-to-event The median time to death or disease progression was based on Kaplan-Meier estimates of pooling over strata, and was defined as the first occurrence of either 1) a decrease of ≥ 10% in FVC (L) and a decrease of ≥ 5% in diffuse lung capacity for carbon monoxide (DLCO) (ml/min/mmHg), or 2) a decrease of ≥ 5% in FVC (L) and a decrease of ≥ 15% in DLCO (ml/min/mmHg).
Change in FVC % Predicted at Week 48 Baseline and Week 48 Binary in the posted analysis record Change in FVC % predicted. Participants in the Full Analysis Set with evaluable change data were analyzed.
Change in DLCO % Predicted at Week 48 Baseline and Week 48 Binary in the posted analysis record Change in DLCO % predicted. Participants in the Full Analysis Set with evaluable change data were analyzed.
Change in 6MWT at Week 48 Baseline and Week 48 Continuous Change in 6MWT, with outcome unit reported as meters. Participants in the Full Analysis Set with evaluable change data were analyzed.
Change in Dyspnea Score at Week 48 as Assessed by the Transitional Dyspnea Index (TDI) Baseline and Week 48 Continuous Change in dyspnea score, with outcome unit reported as units on a scale. Participants in the Full Analysis Set with evaluable change data were analyzed.

The primary endpoint is a composite time-to-event endpoint. The event is not limited to death: disease progression, according to the registry's specified pulmonary-function criteria, can occur first and therefore determine the event time.

Endpoint interpretation: Because the primary endpoint combines death and IPF progression, its hazard ratio should be interpreted as the relative effect on the composite endpoint. It should not be described as a hazard ratio for mortality alone.

5. Analysis Population and Stratification

The primary statistical analysis used the Full Analysis Set, defined in the registry analysis record as participants who were randomized and treated. The secondary analyses used participants in the Full Analysis Set with evaluable change data.

Analysis populationDefinition / role
Full Analysis Set Participants who were randomized and treated; used for the primary time-to-event analysis.
Full Analysis Set with evaluable change data Participants with evaluable change data; used for the posted Week 48 secondary analyses.

The primary hazard ratio was based on a stratified Cox proportional-hazards model. The registry states that the strata were defined by baseline presence of pulmonary hypertension and whether a surgical lung biopsy was performed with definite or probable UIP based on core pathology review.

Stratum 1

Baseline presence of pulmonary hypertension.

Stratum 2

Whether a surgical lung biopsy was performed with definite or probable UIP based on core pathology review.

Stratification allows the primary comparison to account for these prespecified baseline factors rather than treating every participant as belonging to a single homogeneous risk set. It does not mean that the treatment effect is estimated separately and independently within every stratum; the posted analysis reports a stratified Cox model producing an overall hazard-ratio estimate.

6. Statistical Methodology

Kaplan-Meier estimation

The registry states that the median time to death or disease progression was based on Kaplan-Meier estimates of pooling over strata. Kaplan-Meier estimation is designed for time-to-event data in which some participants may not have experienced the event by the time their follow-up ends.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di is the number of events at event time ti, and ni is the number at risk immediately before that time.

The key point is that Kaplan-Meier estimation uses the timing of events and censoring rather than reducing every participant to a simple yes/no event indicator.

Log-rank test

The posted primary analysis used a log-rank test. The log-rank test compares the observed and expected event patterns between randomized groups over follow-up. For this trial, it was used to test the superiority hypothesis for the time-to-death-or-IPF-progression endpoint.

Stratified Cox proportional-hazards model

The hazard ratio was based on a stratified Cox proportional-hazards model using the registry-specified strata. The Cox model provides a relative treatment-effect estimate while accounting for the time dimension of the endpoint.

Conceptual interpretation
HR = estimated hazard in ambrisentan group ÷ estimated hazard in placebo group

For this trial, the reported HR compares the estimated hazard of the composite endpoint of death or IPF progression between the randomized groups.

Van Elteren test

The secondary FVC, DLCO, and 6MWT analyses used the Van Elteren test. This is a stratified nonparametric procedure. Its use is consistent with comparing treatment groups while accounting for stratification, without requiring the outcome distribution to satisfy the assumptions of a conventional parametric comparison.

Hodges-Lehmann estimation

The registry states that the point estimates and 95% confidence intervals for the FVC, DLCO, and 6MWT analyses were based on the Hodges-Lehmann estimate of treatment effect. The Hodges-Lehmann approach provides a nonparametric estimate of a location-type treatment difference that is naturally paired with rank-based tests such as the Van Elteren procedure.

Wilcoxon / Mann-Whitney analysis

The TDI dyspnea analysis used the Wilcoxon (Mann-Whitney) method. This is another rank-based nonparametric approach for comparing two groups. The registry again reports a Hodges-Lehmann estimate and 95% confidence interval for the treatment effect.

7. Primary Result: Time to Death or Disease (IPF) Progression

The primary endpoint was evaluated over up to 48 months. The Full Analysis Set consisted of participants who were randomized and treated. The groups compared were ambrisentan versus placebo.

Hazard ratio for death or IPF progression

1.74

95% CI: 1.14–2.66   ·   P = 0.010

Two-sided confidence interval · Log-rank analysis · Superiority hypothesis

Primary endpointAnalysis populationMethodEffect estimate95% CIP-value
Time to Death or Disease (IPF) Progression Full Analysis Set: participants who were randomized and treated Log-rank test; hazard ratio based on stratified Cox proportional-hazards model HR 1.74 1.14–2.66 0.010
Clinical Biostats interpretation

The reported HR of 1.74 means that, under the stratified Cox model used for the primary analysis, the estimated instantaneous rate of experiencing the composite endpoint was 1.74 times as high in the ambrisentan group as in the placebo group. Expressed as a relative difference, 1.74 corresponds to a 74% higher estimated hazard for the composite endpoint under this model.

This does not mean that 74% of participants experienced the event, that 74% more participants necessarily died, or that each individual participant had exactly a 74% increase in risk. The endpoint combines death with a defined measure of IPF progression, so the HR is specifically about that composite event.

The 95% confidence interval of 1.14–2.66 describes uncertainty around the estimated hazard ratio under the statistical model and sampling framework. It does not describe the range of individual patient responses. Because the interval is entirely above 1, the estimated treatment-group hazard is above the comparator hazard throughout the reported interval.

The P-value of 0.010 addresses the statistical evidence against the null hypothesis used for the superiority comparison; it is not a measure of the size or clinical importance of the treatment effect. The magnitude of the association is communicated by the HR and its confidence interval.

The Cox interpretation also depends on the proportional-hazards framework. A single HR summarizes the relative hazard over follow-up under that model; it should not automatically be translated into a constant relative risk at every individual time point.

Why the primary result is a survival-analysis result

A conventional two-group comparison of proportions would discard the timing information in this endpoint. The primary analysis instead uses the entire follow-up structure: when participants experience death or IPF progression, when they are censored, and the risk set present at successive event times.

The combination of the log-rank test and a stratified Cox model is therefore coherent with the endpoint structure. The log-rank test addresses the group comparison, while the Cox model supplies the reported hazard-ratio effect measure and confidence interval.

8. Secondary Result: Change in FVC % Predicted at Week 48

The registry reports a Week 48 comparison of change in FVC % predicted using the Van Elteren test. Participants in the Full Analysis Set with evaluable change data were analyzed.

Hodges-Lehmann treatment-effect estimate

4.29

95% CI: -0.805–9.376   ·   P = 0.086

Two-sided confidence interval · Van Elteren test · Superiority hypothesis

EndpointMethodEstimate95% CIP-value
Change in FVC % Predicted at Week 48 Van Elteren test 4.29 -0.805–9.376 0.086
Clinical Biostats interpretation

The reported point estimate is 4.29, with a 95% confidence interval from -0.805 to 9.376. The registry identifies this as a Hodges-Lehmann estimate of treatment effect for percent change from baseline.

The confidence interval spans zero. Thus, the interval includes treatment effects in either direction relative to zero under the reported estimation framework. The P-value of 0.086 is a measure of evidence for the specified superiority comparison, not a measure of how large or clinically meaningful the point estimate is.

This secondary endpoint is also structurally different from the primary endpoint. FVC change at a fixed Week 48 time point does not contain the same event-time information as the composite time-to-event endpoint.

9. Secondary Result: Change in DLCO % Predicted at Week 48

The Week 48 DLCO analysis used the same general nonparametric framework: participants in the Full Analysis Set with evaluable change data were analyzed with the Van Elteren test.

Hodges-Lehmann treatment-effect estimate

2.85

95% CI: -2.20–7.90   ·   P = 0.250

Two-sided confidence interval · Van Elteren test · Superiority hypothesis

EndpointMethodEstimate95% CIP-value
Change in DLCO % Predicted at Week 48 Van Elteren test 2.85 -2.20–7.90 0.250
Clinical Biostats interpretation

The estimated treatment effect is 2.85, with a 95% confidence interval of -2.20 to 7.90. The interval crosses zero, so the reported uncertainty range includes both a negative and positive treatment difference.

The P-value of 0.250 is not an effect-size measure. A larger P-value does not establish that the two treatments are identical, just as a smaller P-value would not by itself establish clinical importance. The confidence interval is essential for understanding the range of effects compatible with the reported analysis.

10. Secondary Result: Change in 6MWT at Week 48

The registry reports a Week 48 change in 6-minute walk test (6MWT), analyzed with the Van Elteren test among participants in the Full Analysis Set with evaluable change data.

Hodges-Lehmann treatment-effect estimate

16.00 meters

95% CI: -5.00–37.00   ·   P = 0.150

Two-sided confidence interval · Van Elteren test · Superiority hypothesis

EndpointMethodEstimate95% CIP-value
Change in 6MWT at Week 48 Van Elteren test 16.00 -5.00–37.00 0.150
Clinical Biostats interpretation

The reported point estimate is 16.00 meters. The 95% confidence interval extends from -5.00 to 37.00, which includes zero. The interval therefore reflects uncertainty spanning a possible difference in either direction under the reported estimation framework.

The P-value of 0.150 should not be interpreted as the probability that the treatment effect is zero. It instead describes the evidence against the specified null hypothesis within the statistical testing framework.

The registry analysis notes that the point estimate and confidence interval were based on the Hodges-Lehmann estimate of treatment effect for percent change from baseline, even though the outcome unit is reported as meters. The registry wording is retained here rather than attempting to reinterpret or recompute the estimand.

11. Secondary Result: Change in Dyspnea Score at Week 48

Dyspnea was assessed using the Transitional Dyspnea Index (TDI). The Week 48 analysis used the Wilcoxon (Mann-Whitney) method among participants in the Full Analysis Set with evaluable change data.

Hodges-Lehmann treatment-effect estimate

0.50

95% CI: 0.00–1.00   ·   P = 0.793

Two-sided confidence interval · Wilcoxon (Mann-Whitney) test · Superiority hypothesis

EndpointMethodEstimate95% CIP-value
Change in Dyspnea Score at Week 48 as Assessed by the Transitional Dyspnea Index (TDI) Wilcoxon (Mann-Whitney) 0.50 0.00–1.00 0.793
Clinical Biostats interpretation

The reported treatment-effect estimate is 0.50 on the TDI scale, with a 95% confidence interval from 0.00 to 1.00. The registry identifies the estimate and confidence interval as being based on the Hodges-Lehmann estimate of treatment effect.

The P-value of 0.793 indicates limited statistical evidence against the null hypothesis in this analysis. It does not measure the size of the estimated treatment effect, and it should not be converted into a probability that the treatment has no effect.

The endpoint is analyzed with a rank-based method rather than a conventional mean-comparison model. That distinction matters when explaining what the reported point estimate represents.

12. Secondary Results at a Glance

EndpointMethodEstimate95% CIP-value
Change in FVC % Predicted at Week 48Van Elteren4.29-0.805–9.3760.086
Change in DLCO % Predicted at Week 48Van Elteren2.85-2.20–7.900.250
Change in 6MWT at Week 48Van Elteren16.00-5.00–37.000.150
Change in Dyspnea Score at Week 48, TDIWilcoxon (Mann-Whitney)0.500.00–1.000.793

These secondary endpoints illustrate why a clinical trial should not be summarized using a single P-value. The trial contains a time-to-event endpoint, pulmonary-function measures, exercise capacity, and a dyspnea measure, each with a different estimand and statistical method.

13. Safety: Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by arm as affected participants over the reported at-risk population.

ArmParticipants with serious adverse eventsAt riskReported proportion
Ambrisentan 73 329 73/329
Placebo 25 163 25/163
Serious adverse events: affected / at risk
Ambrisentan
73/329
Placebo
25/163
Safety interpretation: The denominators shown here are the at-risk denominators reported in the ClinicalTrials.gov record. They should not be replaced with inferred randomized-arm sample sizes. The serious-adverse-event figures are descriptive safety data and are not the same estimand as the primary time-to-event efficacy analysis.

The distinction between efficacy and safety populations is important. The primary efficacy analysis was defined using participants who were randomized and treated, whereas the ClinicalTrials.gov record is reported as affected participants over the corresponding at-risk populations. A serious-adverse-event proportion does not account for the timing of events in the same way as a time-to-event analysis.

14. Statistical Methods Explained

Why was a log-rank test used for the primary endpoint?

The primary outcome is a time-to-event endpoint measured over up to 48 months. The log-rank test is designed to compare the event-time distributions of two groups while incorporating the timing of events and censoring. A simple comparison of the proportion experiencing an event would discard much of that information.

What does an HR of 1.74 mean?

An HR of 1.74 means that the estimated instantaneous hazard of the composite endpoint was 1.74 times as high in the ambrisentan group as in the placebo group under the fitted stratified Cox model. It is a relative hazard measure, not a probability and not a statement that every participant has a 74% higher individual risk.

Why does the confidence interval matter?

The point estimate alone does not communicate statistical precision. The primary 95% CI, 1.14–2.66, describes uncertainty around the HR estimate under the specified model and sampling framework. It is therefore more informative than the HR alone.

Why does the P-value not measure effect size?

The P-value quantifies the compatibility of the observed data with a specified null hypothesis within the testing framework. It does not tell the reader how large the treatment difference is. The HR and confidence interval provide the effect-size information for the primary endpoint.

What is the Van Elteren test doing?

The Van Elteren test is a stratified nonparametric comparison. In ARTEMIS-IPF it was used for the Week 48 FVC, DLCO, and 6MWT analyses. The accompanying Hodges-Lehmann estimates provide a nonparametric treatment-effect estimate rather than requiring a conventional normal-theory mean comparison.

Why use the Wilcoxon / Mann-Whitney test for TDI?

The TDI change analysis used a rank-based Wilcoxon (Mann-Whitney) comparison. Rank-based methods can be useful when the outcome distribution does not warrant reliance on a particular parametric distributional assumption. The registry paired this analysis with a Hodges-Lehmann treatment-effect estimate and 95% confidence interval.

Why is the analysis stratified?

The primary Cox analysis was stratified by baseline presence of pulmonary hypertension and by whether a surgical lung biopsy was performed with definite or probable UIP based on core pathology review. Stratification incorporates these factors into the survival comparison without requiring the analysis to assume a common baseline hazard across the strata.

15. Understanding the Primary Endpoint More Deeply

The primary endpoint combines two clinically distinct types of events: death and a prespecified definition of IPF disease progression. This design has statistical advantages and interpretive consequences.

Composite endpoint

The first qualifying occurrence determines the event time. Death and progression therefore enter the same primary time-to-event analysis.

Event timing

The analysis retains information about when an event occurs rather than only whether it eventually occurred.

Censoring

Participants without an observed primary event contribute follow-up information until the relevant censoring time.

Model-based effect

The reported HR comes from a stratified Cox proportional-hazards model rather than directly from a simple event-rate ratio.

Composite endpoints can be statistically efficient because multiple types of events contribute to the analysis. At the same time, the components can have different clinical meanings. A hazard ratio for the composite should therefore not be described as though it were a mortality-only effect estimate.

16. What the Hazard Ratio Does — and Does Not — Mean

Primary statistical interpretation

The primary HR of 1.74 indicates a higher estimated hazard of the composite endpoint in the ambrisentan group than in the placebo group under the stratified Cox model. Numerically, 1.74 corresponds to a 74% higher estimated hazard relative to the placebo hazard.

It does not mean that 74% of participants experienced death or progression, that 74% more participants necessarily experienced the event, or that the treatment changes every participant's individual probability by exactly 74%.

Why the confidence interval matters

The 95% CI of 1.14–2.66 gives a measure of uncertainty around the estimated hazard ratio. The entire interval lies above 1, so the reported uncertainty range remains on the same side of the null value for the HR.

Why the P-value matters differently

The P-value of 0.010 is evidence from the prespecified statistical comparison against its null hypothesis. It is not a percentage probability that the treatment is effective or ineffective, and it should never replace the effect estimate and confidence interval when communicating the magnitude of a result.

Proportional-hazards caution

The Cox model expresses the treatment comparison through a hazard ratio. That interpretation relies on the model's proportional-hazards framework. Even when an overall HR is reported, the HR should not automatically be interpreted as a fixed relative difference in risk at every point in time.

17. Interpreting the Week 48 Nonparametric Results

The secondary analyses provide a useful contrast with the primary survival analysis. FVC, DLCO, 6MWT, and TDI are evaluated at a defined follow-up point rather than through a time-to-first-event framework.

EndpointEstimator / testWhat the reported estimate represents
FVC % predicted Hodges-Lehmann / Van Elteren Nonparametric treatment-effect estimate for change from baseline.
DLCO % predicted Hodges-Lehmann / Van Elteren Nonparametric treatment-effect estimate for change from baseline.
6MWT Hodges-Lehmann / Van Elteren Registry-reported point estimate of treatment effect; outcome unit is meters.
TDI Hodges-Lehmann / Wilcoxon (Mann-Whitney) Nonparametric treatment-effect estimate for change in dyspnea score.

Three of the four Week 48 secondary estimates have confidence intervals that cross zero. The TDI interval reaches zero at its lower boundary. These confidence intervals should be read directly rather than translating P-values into categorical statements about whether a treatment "works."

18. Multiplicity and Secondary Endpoints

The ClinicalTrials.gov record identifies one registered primary endpoint and several secondary analyses. The primary hypothesis type is superiority, and the posted statistical methods include log-rank, Van Elteren, and Wilcoxon / Mann-Whitney procedures.

Because multiple outcomes are reported, the secondary P-values should be interpreted in the context of their secondary status. A set of secondary analyses can generate several statistical tests, and the nominal P-value for any one test does not, by itself, establish control of the familywise type I error across all secondary outcomes.

Analysis familyRole in the ClinicalTrials.gov recordInterpretive focus
Time to Death or Disease (IPF) Progression Primary endpoint Primary superiority comparison using log-rank testing and a stratified Cox HR.
FVC % predicted Secondary Week 48 nonparametric treatment-effect estimate.
DLCO % predicted Secondary Week 48 nonparametric treatment-effect estimate.
6MWT Secondary Week 48 nonparametric treatment-effect estimate.
TDI dyspnea score Secondary Week 48 rank-based treatment comparison.

The ClinicalTrials.gov record does not provide an alpha-allocation scheme, multiplicity-adjustment procedure, or detailed interim-analysis plan. Those features therefore should not be inferred from the presence of multiple reported endpoints.

19. Missing Data, Censoring, and Analysis Populations

The primary endpoint is a time-to-event outcome, so censoring is intrinsic to the analysis. A participant who has not experienced death or the defined progression event during observed follow-up contributes information until the censoring time.

The secondary analyses are different. The registry specifically states that participants in the Full Analysis Set with evaluable change data were analyzed. This means the secondary results are conditional on evaluability of the relevant Week 48 change measure.

Primary endpoint

Time-to-event analysis accounts for event timing and censoring.

Secondary endpoints

Analyses were conducted among Full Analysis Set participants with evaluable change data.

Imputation

The ClinicalTrials.gov record does not specify a missing-data imputation method.

Interpretive caution

The available analysis record should not be expanded into an assumed imputation strategy that is not reported in the ClinicalTrials.gov record.

This distinction is important because an analysis based on evaluable Week 48 change data answers a somewhat different statistical question from a pure ITT time-to-event analysis. The ClinicalTrials.gov record does not provide the number of participants with evaluable data for each secondary endpoint, so no such denominator should be inferred.

20. Blinding and Randomization as Statistical Protections

ARTEMIS-IPF was randomized and double-blind. These design features are not merely administrative details; they support the validity of the statistical comparison.

Randomization
Randomized allocation helps balance measured and unmeasured prognostic factors in expectation, providing the foundation for the between-group comparison.
Double blinding
Masking treatment assignment can reduce differential behavior, treatment decisions, and assessment-related influences.
Parallel design
Participants remain in their assigned parallel treatment groups rather than serving as their own controls.
Superiority hypothesis
The registered hypothesis type is superiority rather than non-inferiority or equivalence.

Randomization does not guarantee identical baseline characteristics in every finite sample. Its principal statistical value is that treatment assignment is randomized, allowing the treatment groups to serve as the basis for a causal comparison under the trial design and assumptions.

21. Limitations

22. Why This Trial Matters Statistically

ARTEMIS-IPF is a useful teaching example because the registry record brings together several core concepts in clinical-trial biostatistics without relying on a single statistical technique.

ConceptHow it appears in ARTEMIS-IPF
RandomizationParticipants were randomized to ambrisentan or placebo in a parallel phase 3 design.
BlindingThe trial was double-blind.
Time-to-event analysisThe primary endpoint was time to death or IPF progression over up to 48 months.
Kaplan-Meier estimationThe registry states that median time to the primary event was based on Kaplan-Meier estimates pooling over strata.
Log-rank testThe primary statistical analysis used a log-rank test.
Hazard ratioThe primary treatment effect was reported as HR 1.74 with a 95% CI of 1.14–2.66.
Stratified Cox modelThe HR was based on a stratified Cox proportional-hazards model.
Stratified analysisStrata included baseline pulmonary hypertension and surgical lung biopsy/UIP status.
Nonparametric testingFVC, DLCO, and 6MWT used the Van Elteren test; TDI used Wilcoxon / Mann-Whitney.
Hodges-Lehmann estimationThe registry used Hodges-Lehmann treatment-effect estimates for the reported Week 48 secondary analyses.
Confidence intervalsBoth the primary hazard ratio and secondary treatment-effect estimates were reported with 95% CIs.
P-valuesThe registry reports P-values for the primary and secondary statistical comparisons.
Analysis populationsThe primary analysis used the Full Analysis Set of randomized and treated participants; secondary analyses used evaluable change data.

The most important statistical lesson is that endpoint structure should drive method selection. The primary outcome requires survival-analysis methods because time and censoring matter. The Week 48 outcomes are analyzed using rank-based methods because the registry specifies Van Elteren or Wilcoxon procedures and Hodges-Lehmann effect estimation.

23. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue through Clinical Biostats

Use the trial's statistical methods as a starting point for deeper study of survival analysis, confidence intervals, nonparametric tests, and clinical-trial methodology.

26. Record Summary

ARTEMIS-IPF provides a compact but statistically rich example of randomized clinical-trial analysis. Its primary endpoint is a time-to-event composite analyzed with Kaplan-Meier estimation, a log-rank test, and a stratified Cox proportional-hazards model. The reported primary HR of 1.74 has a 95% CI of 1.14–2.66 and a P-value of 0.010. The secondary analyses demonstrate a different statistical pathway, using Van Elteren tests and Hodges-Lehmann estimates for FVC, DLCO, and 6MWT, and a Wilcoxon / Mann-Whitney analysis with a Hodges-Lehmann estimate for TDI.

The statistical interpretation therefore depends on keeping several distinctions clear: composite versus component endpoints, time-to-event versus fixed-time outcomes, hazard ratios versus absolute measures, effect estimates versus P-values, and primary versus secondary analyses. These distinctions are essential for reading the trial without reducing its evidence to a single number.

Clinical Biostats methodology: This page separates reported registry results from statistical explanation. Numerical results, endpoint definitions, analysis populations, and statistical methods are restricted to the registry-reported ARTEMIS-IPF trial data.