This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. View NCT00768300 on ClinicalTrials.gov.
1. Trial at a Glance
ARTEMIS-IPF was a phase 3 randomized, double-blind, parallel trial evaluating ambrisentan versus placebo in participants with idiopathic pulmonary fibrosis. The registry reports 494 enrolled participants, one registered primary time-to-event endpoint, nine posted outcome measures, and five posted statistical analyses.
| Feature | ARTEMIS-IPF |
|---|---|
| Trial name | ARTEMIS-IPF |
| Phase | Phase 3 |
| Condition | Idiopathic Pulmonary Fibrosis |
| Design | Randomized, double-blind, parallel |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 494 |
| Interventions | Ambrisentan; placebo |
| Primary endpoint type | Time-to-event |
| Primary hypothesis type | Superiority |
| Trial status | Terminated |
| Start | 2008-12 |
| Primary completion | 2011-02 |
| Lead sponsor | Gilead Sciences |
| Sponsor type | Industry |
| ClinicalTrials.gov | NCT00768300 |
2. Clinical Question
The primary statistical question was whether the randomized comparison of ambrisentan versus placebo differed with respect to time to death or disease progression in idiopathic pulmonary fibrosis. The registry classified the primary hypothesis as superiority.
Population
Participants enrolled in a phase 3 trial for idiopathic pulmonary fibrosis.
Intervention
Ambrisentan.
Comparator
Placebo.
Primary question
Does randomized assignment to ambrisentan versus placebo change the time to death or disease progression?
This question is inherently longitudinal. A participant who has not experienced death or the defined progression event by the end of available follow-up does not simply become an ordinary non-event; the participant contributes information up to the point of censoring. That structure is why survival-analysis methods are central to the primary analysis.
3. Trial Design
Ambrisentan
- Drug intervention.
- Participants were randomized to the ambrisentan arm.
- Primary comparison was time to death or IPF progression.
Placebo
- Placebo drug intervention.
- Participants were randomized to the placebo arm.
- Primary comparison was time to death or IPF progression.
The design has an important statistical implication: randomization establishes the framework for comparing outcomes between the two assigned groups, while double blinding reduces the potential for knowledge of treatment assignment to influence treatment administration, outcome assessment, or participant behavior.
4. Endpoints
| Endpoint | Time frame | Type | Registry definition / analysis |
|---|---|---|---|
| Time to Death or Disease (IPF) Progression | Up to 48 months | Time-to-event | The median time to death or disease progression was based on Kaplan-Meier estimates of pooling over strata, and was defined as the first occurrence of either 1) a decrease of ≥ 10% in FVC (L) and a decrease of ≥ 5% in diffuse lung capacity for carbon monoxide (DLCO) (ml/min/mmHg), or 2) a decrease of ≥ 5% in FVC (L) and a decrease of ≥ 15% in DLCO (ml/min/mmHg). |
| Change in FVC % Predicted at Week 48 | Baseline and Week 48 | Binary in the posted analysis record | Change in FVC % predicted. Participants in the Full Analysis Set with evaluable change data were analyzed. |
| Change in DLCO % Predicted at Week 48 | Baseline and Week 48 | Binary in the posted analysis record | Change in DLCO % predicted. Participants in the Full Analysis Set with evaluable change data were analyzed. |
| Change in 6MWT at Week 48 | Baseline and Week 48 | Continuous | Change in 6MWT, with outcome unit reported as meters. Participants in the Full Analysis Set with evaluable change data were analyzed. |
| Change in Dyspnea Score at Week 48 as Assessed by the Transitional Dyspnea Index (TDI) | Baseline and Week 48 | Continuous | Change in dyspnea score, with outcome unit reported as units on a scale. Participants in the Full Analysis Set with evaluable change data were analyzed. |
The primary endpoint is a composite time-to-event endpoint. The event is not limited to death: disease progression, according to the registry's specified pulmonary-function criteria, can occur first and therefore determine the event time.
5. Analysis Population and Stratification
The primary statistical analysis used the Full Analysis Set, defined in the registry analysis record as participants who were randomized and treated. The secondary analyses used participants in the Full Analysis Set with evaluable change data.
| Analysis population | Definition / role |
|---|---|
| Full Analysis Set | Participants who were randomized and treated; used for the primary time-to-event analysis. |
| Full Analysis Set with evaluable change data | Participants with evaluable change data; used for the posted Week 48 secondary analyses. |
The primary hazard ratio was based on a stratified Cox proportional-hazards model. The registry states that the strata were defined by baseline presence of pulmonary hypertension and whether a surgical lung biopsy was performed with definite or probable UIP based on core pathology review.
Stratum 1
Baseline presence of pulmonary hypertension.
Stratum 2
Whether a surgical lung biopsy was performed with definite or probable UIP based on core pathology review.
Stratification allows the primary comparison to account for these prespecified baseline factors rather than treating every participant as belonging to a single homogeneous risk set. It does not mean that the treatment effect is estimated separately and independently within every stratum; the posted analysis reports a stratified Cox model producing an overall hazard-ratio estimate.
6. Statistical Methodology
Kaplan-Meier estimation
The registry states that the median time to death or disease progression was based on Kaplan-Meier estimates of pooling over strata. Kaplan-Meier estimation is designed for time-to-event data in which some participants may not have experienced the event by the time their follow-up ends.
Here, di is the number of events at event time ti, and ni is the number at risk immediately before that time.
The key point is that Kaplan-Meier estimation uses the timing of events and censoring rather than reducing every participant to a simple yes/no event indicator.
Log-rank test
The posted primary analysis used a log-rank test. The log-rank test compares the observed and expected event patterns between randomized groups over follow-up. For this trial, it was used to test the superiority hypothesis for the time-to-death-or-IPF-progression endpoint.
Stratified Cox proportional-hazards model
The hazard ratio was based on a stratified Cox proportional-hazards model using the registry-specified strata. The Cox model provides a relative treatment-effect estimate while accounting for the time dimension of the endpoint.
For this trial, the reported HR compares the estimated hazard of the composite endpoint of death or IPF progression between the randomized groups.
Van Elteren test
The secondary FVC, DLCO, and 6MWT analyses used the Van Elteren test. This is a stratified nonparametric procedure. Its use is consistent with comparing treatment groups while accounting for stratification, without requiring the outcome distribution to satisfy the assumptions of a conventional parametric comparison.
Hodges-Lehmann estimation
The registry states that the point estimates and 95% confidence intervals for the FVC, DLCO, and 6MWT analyses were based on the Hodges-Lehmann estimate of treatment effect. The Hodges-Lehmann approach provides a nonparametric estimate of a location-type treatment difference that is naturally paired with rank-based tests such as the Van Elteren procedure.
Wilcoxon / Mann-Whitney analysis
The TDI dyspnea analysis used the Wilcoxon (Mann-Whitney) method. This is another rank-based nonparametric approach for comparing two groups. The registry again reports a Hodges-Lehmann estimate and 95% confidence interval for the treatment effect.
7. Primary Result: Time to Death or Disease (IPF) Progression
The primary endpoint was evaluated over up to 48 months. The Full Analysis Set consisted of participants who were randomized and treated. The groups compared were ambrisentan versus placebo.
Hazard ratio for death or IPF progression
95% CI: 1.14–2.66 · P = 0.010
Two-sided confidence interval · Log-rank analysis · Superiority hypothesis
| Primary endpoint | Analysis population | Method | Effect estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Time to Death or Disease (IPF) Progression | Full Analysis Set: participants who were randomized and treated | Log-rank test; hazard ratio based on stratified Cox proportional-hazards model | HR 1.74 | 1.14–2.66 | 0.010 |
The reported HR of 1.74 means that, under the stratified Cox model used for the primary analysis, the estimated instantaneous rate of experiencing the composite endpoint was 1.74 times as high in the ambrisentan group as in the placebo group. Expressed as a relative difference, 1.74 corresponds to a 74% higher estimated hazard for the composite endpoint under this model.
This does not mean that 74% of participants experienced the event, that 74% more participants necessarily died, or that each individual participant had exactly a 74% increase in risk. The endpoint combines death with a defined measure of IPF progression, so the HR is specifically about that composite event.
The 95% confidence interval of 1.14–2.66 describes uncertainty around the estimated hazard ratio under the statistical model and sampling framework. It does not describe the range of individual patient responses. Because the interval is entirely above 1, the estimated treatment-group hazard is above the comparator hazard throughout the reported interval.
The P-value of 0.010 addresses the statistical evidence against the null hypothesis used for the superiority comparison; it is not a measure of the size or clinical importance of the treatment effect. The magnitude of the association is communicated by the HR and its confidence interval.
The Cox interpretation also depends on the proportional-hazards framework. A single HR summarizes the relative hazard over follow-up under that model; it should not automatically be translated into a constant relative risk at every individual time point.
Why the primary result is a survival-analysis result
A conventional two-group comparison of proportions would discard the timing information in this endpoint. The primary analysis instead uses the entire follow-up structure: when participants experience death or IPF progression, when they are censored, and the risk set present at successive event times.
The combination of the log-rank test and a stratified Cox model is therefore coherent with the endpoint structure. The log-rank test addresses the group comparison, while the Cox model supplies the reported hazard-ratio effect measure and confidence interval.
8. Secondary Result: Change in FVC % Predicted at Week 48
The registry reports a Week 48 comparison of change in FVC % predicted using the Van Elteren test. Participants in the Full Analysis Set with evaluable change data were analyzed.
Hodges-Lehmann treatment-effect estimate
95% CI: -0.805–9.376 · P = 0.086
Two-sided confidence interval · Van Elteren test · Superiority hypothesis
| Endpoint | Method | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Change in FVC % Predicted at Week 48 | Van Elteren test | 4.29 | -0.805–9.376 | 0.086 |
The reported point estimate is 4.29, with a 95% confidence interval from -0.805 to 9.376. The registry identifies this as a Hodges-Lehmann estimate of treatment effect for percent change from baseline.
The confidence interval spans zero. Thus, the interval includes treatment effects in either direction relative to zero under the reported estimation framework. The P-value of 0.086 is a measure of evidence for the specified superiority comparison, not a measure of how large or clinically meaningful the point estimate is.
This secondary endpoint is also structurally different from the primary endpoint. FVC change at a fixed Week 48 time point does not contain the same event-time information as the composite time-to-event endpoint.
9. Secondary Result: Change in DLCO % Predicted at Week 48
The Week 48 DLCO analysis used the same general nonparametric framework: participants in the Full Analysis Set with evaluable change data were analyzed with the Van Elteren test.
Hodges-Lehmann treatment-effect estimate
95% CI: -2.20–7.90 · P = 0.250
Two-sided confidence interval · Van Elteren test · Superiority hypothesis
| Endpoint | Method | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Change in DLCO % Predicted at Week 48 | Van Elteren test | 2.85 | -2.20–7.90 | 0.250 |
The estimated treatment effect is 2.85, with a 95% confidence interval of -2.20 to 7.90. The interval crosses zero, so the reported uncertainty range includes both a negative and positive treatment difference.
The P-value of 0.250 is not an effect-size measure. A larger P-value does not establish that the two treatments are identical, just as a smaller P-value would not by itself establish clinical importance. The confidence interval is essential for understanding the range of effects compatible with the reported analysis.
10. Secondary Result: Change in 6MWT at Week 48
The registry reports a Week 48 change in 6-minute walk test (6MWT), analyzed with the Van Elteren test among participants in the Full Analysis Set with evaluable change data.
Hodges-Lehmann treatment-effect estimate
95% CI: -5.00–37.00 · P = 0.150
Two-sided confidence interval · Van Elteren test · Superiority hypothesis
| Endpoint | Method | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Change in 6MWT at Week 48 | Van Elteren test | 16.00 | -5.00–37.00 | 0.150 |
The reported point estimate is 16.00 meters. The 95% confidence interval extends from -5.00 to 37.00, which includes zero. The interval therefore reflects uncertainty spanning a possible difference in either direction under the reported estimation framework.
The P-value of 0.150 should not be interpreted as the probability that the treatment effect is zero. It instead describes the evidence against the specified null hypothesis within the statistical testing framework.
The registry analysis notes that the point estimate and confidence interval were based on the Hodges-Lehmann estimate of treatment effect for percent change from baseline, even though the outcome unit is reported as meters. The registry wording is retained here rather than attempting to reinterpret or recompute the estimand.
11. Secondary Result: Change in Dyspnea Score at Week 48
Dyspnea was assessed using the Transitional Dyspnea Index (TDI). The Week 48 analysis used the Wilcoxon (Mann-Whitney) method among participants in the Full Analysis Set with evaluable change data.
Hodges-Lehmann treatment-effect estimate
95% CI: 0.00–1.00 · P = 0.793
Two-sided confidence interval · Wilcoxon (Mann-Whitney) test · Superiority hypothesis
| Endpoint | Method | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Change in Dyspnea Score at Week 48 as Assessed by the Transitional Dyspnea Index (TDI) | Wilcoxon (Mann-Whitney) | 0.50 | 0.00–1.00 | 0.793 |
The reported treatment-effect estimate is 0.50 on the TDI scale, with a 95% confidence interval from 0.00 to 1.00. The registry identifies the estimate and confidence interval as being based on the Hodges-Lehmann estimate of treatment effect.
The P-value of 0.793 indicates limited statistical evidence against the null hypothesis in this analysis. It does not measure the size of the estimated treatment effect, and it should not be converted into a probability that the treatment has no effect.
The endpoint is analyzed with a rank-based method rather than a conventional mean-comparison model. That distinction matters when explaining what the reported point estimate represents.
12. Secondary Results at a Glance
| Endpoint | Method | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Change in FVC % Predicted at Week 48 | Van Elteren | 4.29 | -0.805–9.376 | 0.086 |
| Change in DLCO % Predicted at Week 48 | Van Elteren | 2.85 | -2.20–7.90 | 0.250 |
| Change in 6MWT at Week 48 | Van Elteren | 16.00 | -5.00–37.00 | 0.150 |
| Change in Dyspnea Score at Week 48, TDI | Wilcoxon (Mann-Whitney) | 0.50 | 0.00–1.00 | 0.793 |
These secondary endpoints illustrate why a clinical trial should not be summarized using a single P-value. The trial contains a time-to-event endpoint, pulmonary-function measures, exercise capacity, and a dyspnea measure, each with a different estimand and statistical method.
13. Safety: Serious Adverse Events
The ClinicalTrials.gov record reports serious adverse events by arm as affected participants over the reported at-risk population.
| Arm | Participants with serious adverse events | At risk | Reported proportion |
|---|---|---|---|
| Ambrisentan | 73 | 329 | 73/329 |
| Placebo | 25 | 163 | 25/163 |
The distinction between efficacy and safety populations is important. The primary efficacy analysis was defined using participants who were randomized and treated, whereas the ClinicalTrials.gov record is reported as affected participants over the corresponding at-risk populations. A serious-adverse-event proportion does not account for the timing of events in the same way as a time-to-event analysis.
14. Statistical Methods Explained
Why was a log-rank test used for the primary endpoint?
The primary outcome is a time-to-event endpoint measured over up to 48 months. The log-rank test is designed to compare the event-time distributions of two groups while incorporating the timing of events and censoring. A simple comparison of the proportion experiencing an event would discard much of that information.
What does an HR of 1.74 mean?
An HR of 1.74 means that the estimated instantaneous hazard of the composite endpoint was 1.74 times as high in the ambrisentan group as in the placebo group under the fitted stratified Cox model. It is a relative hazard measure, not a probability and not a statement that every participant has a 74% higher individual risk.
Why does the confidence interval matter?
The point estimate alone does not communicate statistical precision. The primary 95% CI, 1.14–2.66, describes uncertainty around the HR estimate under the specified model and sampling framework. It is therefore more informative than the HR alone.
Why does the P-value not measure effect size?
The P-value quantifies the compatibility of the observed data with a specified null hypothesis within the testing framework. It does not tell the reader how large the treatment difference is. The HR and confidence interval provide the effect-size information for the primary endpoint.
What is the Van Elteren test doing?
The Van Elteren test is a stratified nonparametric comparison. In ARTEMIS-IPF it was used for the Week 48 FVC, DLCO, and 6MWT analyses. The accompanying Hodges-Lehmann estimates provide a nonparametric treatment-effect estimate rather than requiring a conventional normal-theory mean comparison.
Why use the Wilcoxon / Mann-Whitney test for TDI?
The TDI change analysis used a rank-based Wilcoxon (Mann-Whitney) comparison. Rank-based methods can be useful when the outcome distribution does not warrant reliance on a particular parametric distributional assumption. The registry paired this analysis with a Hodges-Lehmann treatment-effect estimate and 95% confidence interval.
Why is the analysis stratified?
The primary Cox analysis was stratified by baseline presence of pulmonary hypertension and by whether a surgical lung biopsy was performed with definite or probable UIP based on core pathology review. Stratification incorporates these factors into the survival comparison without requiring the analysis to assume a common baseline hazard across the strata.
15. Understanding the Primary Endpoint More Deeply
The primary endpoint combines two clinically distinct types of events: death and a prespecified definition of IPF disease progression. This design has statistical advantages and interpretive consequences.
Composite endpoint
The first qualifying occurrence determines the event time. Death and progression therefore enter the same primary time-to-event analysis.
Event timing
The analysis retains information about when an event occurs rather than only whether it eventually occurred.
Censoring
Participants without an observed primary event contribute follow-up information until the relevant censoring time.
Model-based effect
The reported HR comes from a stratified Cox proportional-hazards model rather than directly from a simple event-rate ratio.
Composite endpoints can be statistically efficient because multiple types of events contribute to the analysis. At the same time, the components can have different clinical meanings. A hazard ratio for the composite should therefore not be described as though it were a mortality-only effect estimate.
16. What the Hazard Ratio Does — and Does Not — Mean
The primary HR of 1.74 indicates a higher estimated hazard of the composite endpoint in the ambrisentan group than in the placebo group under the stratified Cox model. Numerically, 1.74 corresponds to a 74% higher estimated hazard relative to the placebo hazard.
It does not mean that 74% of participants experienced death or progression, that 74% more participants necessarily experienced the event, or that the treatment changes every participant's individual probability by exactly 74%.
The 95% CI of 1.14–2.66 gives a measure of uncertainty around the estimated hazard ratio. The entire interval lies above 1, so the reported uncertainty range remains on the same side of the null value for the HR.
The P-value of 0.010 is evidence from the prespecified statistical comparison against its null hypothesis. It is not a percentage probability that the treatment is effective or ineffective, and it should never replace the effect estimate and confidence interval when communicating the magnitude of a result.
The Cox model expresses the treatment comparison through a hazard ratio. That interpretation relies on the model's proportional-hazards framework. Even when an overall HR is reported, the HR should not automatically be interpreted as a fixed relative difference in risk at every point in time.
17. Interpreting the Week 48 Nonparametric Results
The secondary analyses provide a useful contrast with the primary survival analysis. FVC, DLCO, 6MWT, and TDI are evaluated at a defined follow-up point rather than through a time-to-first-event framework.
| Endpoint | Estimator / test | What the reported estimate represents |
|---|---|---|
| FVC % predicted | Hodges-Lehmann / Van Elteren | Nonparametric treatment-effect estimate for change from baseline. |
| DLCO % predicted | Hodges-Lehmann / Van Elteren | Nonparametric treatment-effect estimate for change from baseline. |
| 6MWT | Hodges-Lehmann / Van Elteren | Registry-reported point estimate of treatment effect; outcome unit is meters. |
| TDI | Hodges-Lehmann / Wilcoxon (Mann-Whitney) | Nonparametric treatment-effect estimate for change in dyspnea score. |
Three of the four Week 48 secondary estimates have confidence intervals that cross zero. The TDI interval reaches zero at its lower boundary. These confidence intervals should be read directly rather than translating P-values into categorical statements about whether a treatment "works."
18. Multiplicity and Secondary Endpoints
The ClinicalTrials.gov record identifies one registered primary endpoint and several secondary analyses. The primary hypothesis type is superiority, and the posted statistical methods include log-rank, Van Elteren, and Wilcoxon / Mann-Whitney procedures.
Because multiple outcomes are reported, the secondary P-values should be interpreted in the context of their secondary status. A set of secondary analyses can generate several statistical tests, and the nominal P-value for any one test does not, by itself, establish control of the familywise type I error across all secondary outcomes.
| Analysis family | Role in the ClinicalTrials.gov record | Interpretive focus |
|---|---|---|
| Time to Death or Disease (IPF) Progression | Primary endpoint | Primary superiority comparison using log-rank testing and a stratified Cox HR. |
| FVC % predicted | Secondary | Week 48 nonparametric treatment-effect estimate. |
| DLCO % predicted | Secondary | Week 48 nonparametric treatment-effect estimate. |
| 6MWT | Secondary | Week 48 nonparametric treatment-effect estimate. |
| TDI dyspnea score | Secondary | Week 48 rank-based treatment comparison. |
The ClinicalTrials.gov record does not provide an alpha-allocation scheme, multiplicity-adjustment procedure, or detailed interim-analysis plan. Those features therefore should not be inferred from the presence of multiple reported endpoints.
19. Missing Data, Censoring, and Analysis Populations
The primary endpoint is a time-to-event outcome, so censoring is intrinsic to the analysis. A participant who has not experienced death or the defined progression event during observed follow-up contributes information until the censoring time.
The secondary analyses are different. The registry specifically states that participants in the Full Analysis Set with evaluable change data were analyzed. This means the secondary results are conditional on evaluability of the relevant Week 48 change measure.
Primary endpoint
Time-to-event analysis accounts for event timing and censoring.
Secondary endpoints
Analyses were conducted among Full Analysis Set participants with evaluable change data.
Imputation
The ClinicalTrials.gov record does not specify a missing-data imputation method.
Interpretive caution
The available analysis record should not be expanded into an assumed imputation strategy that is not reported in the ClinicalTrials.gov record.
This distinction is important because an analysis based on evaluable Week 48 change data answers a somewhat different statistical question from a pure ITT time-to-event analysis. The ClinicalTrials.gov record does not provide the number of participants with evaluable data for each secondary endpoint, so no such denominator should be inferred.
20. Blinding and Randomization as Statistical Protections
ARTEMIS-IPF was randomized and double-blind. These design features are not merely administrative details; they support the validity of the statistical comparison.
Randomization does not guarantee identical baseline characteristics in every finite sample. Its principal statistical value is that treatment assignment is randomized, allowing the treatment groups to serve as the basis for a causal comparison under the trial design and assumptions.
21. Limitations
- Composite primary endpoint: the primary result combines death and IPF progression. The HR should therefore not be interpreted as a mortality-only effect.
- Proportional-hazards assumption: the HR is a model-based summary from a stratified Cox model and should not automatically be treated as a constant relative effect at every time point.
- Secondary endpoint evaluability: the Week 48 analyses were conducted among Full Analysis Set participants with evaluable change data, and the ClinicalTrials.gov record does not provide the number evaluable for each endpoint.
- Multiplicity: several secondary outcomes were analyzed, but the ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure or familywise error-control strategy.
- Missing-data methods: the ClinicalTrials.gov record does not specify an imputation method for missing secondary endpoint measurements.
- Safety denominators: serious adverse events are reported using the registry-reported affected/at-risk counts. These should not be replaced by inferred denominators.
- Endpoint interpretation: fixed-time pulmonary-function and functional outcomes answer different questions from the primary time-to-event endpoint and should not be treated as interchangeable measures.
- Registry scope: the analysis presented here is restricted to the numbers, endpoint definitions, methods, and results reported in the ClinicalTrials.gov record. Additional publication-level details are not introduced.
22. Why This Trial Matters Statistically
ARTEMIS-IPF is a useful teaching example because the registry record brings together several core concepts in clinical-trial biostatistics without relying on a single statistical technique.
| Concept | How it appears in ARTEMIS-IPF |
|---|---|
| Randomization | Participants were randomized to ambrisentan or placebo in a parallel phase 3 design. |
| Blinding | The trial was double-blind. |
| Time-to-event analysis | The primary endpoint was time to death or IPF progression over up to 48 months. |
| Kaplan-Meier estimation | The registry states that median time to the primary event was based on Kaplan-Meier estimates pooling over strata. |
| Log-rank test | The primary statistical analysis used a log-rank test. |
| Hazard ratio | The primary treatment effect was reported as HR 1.74 with a 95% CI of 1.14–2.66. |
| Stratified Cox model | The HR was based on a stratified Cox proportional-hazards model. |
| Stratified analysis | Strata included baseline pulmonary hypertension and surgical lung biopsy/UIP status. |
| Nonparametric testing | FVC, DLCO, and 6MWT used the Van Elteren test; TDI used Wilcoxon / Mann-Whitney. |
| Hodges-Lehmann estimation | The registry used Hodges-Lehmann treatment-effect estimates for the reported Week 48 secondary analyses. |
| Confidence intervals | Both the primary hazard ratio and secondary treatment-effect estimates were reported with 95% CIs. |
| P-values | The registry reports P-values for the primary and secondary statistical comparisons. |
| Analysis populations | The primary analysis used the Full Analysis Set of randomized and treated participants; secondary analyses used evaluable change data. |
The most important statistical lesson is that endpoint structure should drive method selection. The primary outcome requires survival-analysis methods because time and censoring matter. The Week 48 outcomes are analyzed using rank-based methods because the registry specifies Van Elteren or Wilcoxon procedures and Hodges-Lehmann effect estimation.
23. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Sources
- ClinicalTrials.gov: NCT00768300 — ARTEMIS-IPF.
- PubMed: PMID 23648946.
- PubMed: PMID 24717624.
- PubMed: PMID 24177001.
Continue through Clinical Biostats
Use the trial's statistical methods as a starting point for deeper study of survival analysis, confidence intervals, nonparametric tests, and clinical-trial methodology.
26. Record Summary
ARTEMIS-IPF provides a compact but statistically rich example of randomized clinical-trial analysis. Its primary endpoint is a time-to-event composite analyzed with Kaplan-Meier estimation, a log-rank test, and a stratified Cox proportional-hazards model. The reported primary HR of 1.74 has a 95% CI of 1.14–2.66 and a P-value of 0.010. The secondary analyses demonstrate a different statistical pathway, using Van Elteren tests and Hodges-Lehmann estimates for FVC, DLCO, and 6MWT, and a Wilcoxon / Mann-Whitney analysis with a Hodges-Lehmann estimate for TDI.
The statistical interpretation therefore depends on keeping several distinctions clear: composite versus component endpoints, time-to-event versus fixed-time outcomes, hazard ratios versus absolute measures, effect estimates versus P-values, and primary versus secondary analyses. These distinctions are essential for reading the trial without reducing its evidence to a single number.