This page separates reported trial results from statistical interpretation. Trial-specific numerical results and design facts are restricted to the ClinicalTrials.gov record for NCT01106014. The registry provides the official trial record.
1. Trial at a Glance
GRIPHON was a completed phase 3 randomized, parallel-group trial evaluating selexipag versus placebo in pulmonary arterial hypertension. The trial enrolled 1156 patients and used quadruple masking. Its registered primary endpoint was a time-to-event outcome defined from randomization through 7 days after the last study-drug intake.
| Feature | GRIPHON |
|---|---|
| Phase | Phase 3 |
| Condition | Pulmonary Arterial Hypertension |
| Design | Randomized, parallel |
| Masking | Quadruple |
| Primary purpose | Treatment |
| Enrollment | 1156 |
| Interventions | Selexipag and placebo |
| Primary endpoint type | Time-to-event |
| Results posted | Yes |
| Statistical analyses posted | 3 |
| Lead sponsor | Actelion |
| Trial status | Completed |
2. Clinical Question
The central statistical question was whether the randomized comparison of selexipag versus placebo differed with respect to the time from randomization to the first morbidity event or death from all causes, within the registered follow-up window.
Population
Patients enrolled in the phase 3 trial for pulmonary arterial hypertension.
Intervention
Selexipag.
Comparator
Placebo.
Primary question
Does selexipag alter the time to the first morbidity event or death from all causes compared with placebo?
3. Trial Design
Selexipag
- Drug intervention
- Randomized parallel-group assignment
- Compared with placebo
Placebo
- Placebo comparator
- Randomized parallel-group assignment
- Compared with selexipag
4. Endpoints
| Endpoint | Time frame | Type |
|---|---|---|
| Time From Randomization to the First Morbidity Event or Death (All Causes) up to 7 Days After the Last Study Drug Intake | Up to 7 days after end of double-blind treatment (maximum: 4.3 years) | Time-to-event |
| Change From Baseline to Week 26 in 6-minute Walk Distance (6MWD) at Trough | Week 26 | Continuous |
| Absence of Worsening From Baseline to Week 26 in Modified NYHA/WHO Functional Class (WHO FC) | Week 26 | Binary |
Primary endpoint definition
The registered primary endpoint was the time from randomization to the first morbidity event or death from all causes up to 7 days after the last study drug intake. The registry definition states that the endpoint was analyzed with the Kaplan-Meier method, with event-free Kaplan-Meier estimates at different time points.
5. Analysis Populations and Statistical Framework
| Analysis | Population | Role |
|---|---|---|
| Primary endpoint | Full analysis set | All randomized patients evaluated according to the study drug to which they were randomized; the registry describes this as the intention-to-treat analysis. |
| 6MWD | Full analysis set | Continuous secondary endpoint analyzed using ANCOVA. |
| WHO functional class | Full analysis set, excluding baseline WHO FC IV patients | Binary secondary endpoint analyzed using the Cochran-Mantel-Haenszel test. |
The primary endpoint analysis is explicitly anchored to randomized treatment assignment. This is important because an intention-to-treat analysis preserves the treatment comparison created by randomization rather than redefining treatment groups according to subsequent treatment exposure.
6. Statistical Methodology
Kaplan-Meier estimation
The registry definition for the primary endpoint states that the endpoint was analyzed using the Kaplan-Meier method, including event-free Kaplan-Meier estimates at different time points. This is appropriate for a time-to-event endpoint because patients can have different follow-up durations and some patients may not experience the event during observation.
Here, di represents events occurring at time ti, while ni represents the number at risk immediately before that time.
Log-rank test
The formal primary comparison was a log-rank test. The registry analysis notes specify that the primary analysis was performed on the Full Analysis Set using a one-sided unstratified log-rank test.
The log-rank test compares the observed and expected pattern of events between treatment groups across follow-up. It is therefore a global comparison of the time-to-event experience rather than a comparison of a single fixed-time percentage.
Hazard ratio
The effect measure reported for the primary endpoint was a hazard ratio. The estimate was 0.6, with a two-sided 99% confidence interval of 0.46 to 0.78.
A hazard ratio is a relative time-to-event measure. It is not a probability, a percentage of patients who benefit, or an absolute difference in event-free survival.
ANCOVA for 6-minute walk distance
The change from baseline to Week 26 in 6-minute walk distance was analyzed using ANCOVA. The registry analysis notes specify a non-parametric ANCOVA with 6MWD as a covariate at baseline.
Covariate adjustment can improve precision by accounting for baseline measurement differences. Here, the baseline 6MWD value was incorporated into the analysis rather than comparing Week 26 values without adjustment.
Missing-data imputation for 6MWD
The registry analysis specifies prespecified imputation rules for missing Week 26 6MWD values. If a patient was unable to walk at Week 26, 0 meter was imputed. If that rule did not apply, the second lowest observed 6MWD value, 10 meters, at Week 26 was imputed. Missing values were imputed for 21.6% of subjects.
Cochran-Mantel-Haenszel analysis
The binary WHO functional-class endpoint was analyzed using the Cochran-Mantel-Haenszel test. The reported effect measure was an odds ratio, with the estimate expressed as 1.161 and a two-sided 99% confidence interval of 0.811 to 1.664.
7. Primary Result: Time to First Morbidity Event or Death
The primary endpoint was analyzed in the Full Analysis Set according to randomized treatment assignment. The reported method was a log-rank test, specifically described in the analysis notes as a one-sided unstratified log-rank test.
Hazard ratio for first morbidity event or death
99% CI: 0.46–0.78 · P < 0.0001
Hypothesis type: Superiority
| Primary endpoint | Selexipag vs placebo |
|---|---|
| Effect measure | Hazard ratio |
| Estimate | 0.6 |
| Confidence interval | Two-sided 99% CI: 0.46–0.78 |
| P-value | <0.0001 |
| Hypothesis | Superiority |
| Analysis population | Full analysis set / intention-to-treat |
| Statistical method | One-sided unstratified log-rank test |
The reported hazard ratio of 0.6 means that the estimated instantaneous rate of experiencing the primary event, under the time-to-event comparison, was approximately 40% lower with selexipag than with placebo. That is a relative hazard interpretation; it is not the same as saying that 40% of patients avoided an event, or that every individual patient had a 40% reduction in risk.
The two-sided 99% confidence interval of 0.46 to 0.78 describes the statistical uncertainty around the reported hazard-ratio estimate under the analysis framework. It is an interval for the treatment-effect parameter, not a range containing the outcomes that individual patients might experience.
The P-value <0.0001 addresses the evidence against the null hypothesis under the specified testing framework. It does not measure the magnitude of the treatment effect. The magnitude is described by the hazard ratio and its confidence interval.
The interpretation also depends on the time-to-event framework, including censoring and the assumptions underlying a hazard-ratio summary. A single hazard ratio is most naturally interpreted as a relative event-rate comparison over follow-up; it should not automatically be translated into an absolute risk reduction at a particular time point.
Finally, the analysis notes distinguish the one-sided hypothesis test from the reported two-sided 99% confidence interval. Those are related but not identical pieces of statistical reporting and should not be treated as interchangeable.
8. Secondary Result: Change in 6-Minute Walk Distance
The first posted secondary analysis evaluated the change from baseline to Week 26 in 6-minute walk distance at trough. The analysis used the Full Analysis Set and a non-parametric ANCOVA with baseline 6MWD as a covariate.
Median difference in final values
Two-sided 99% CI: 1–24 · P = 0.0027
Hypothesis type: Superiority
| Endpoint | Result |
|---|---|
| Outcome | Change from baseline to Week 26 in 6-minute walk distance at trough |
| Effect measure | Median Difference (Final Values) |
| Estimate | 12 |
| Confidence interval | Two-sided 99% CI: 1–24 |
| P-value | 0.0027 |
| Analysis population | Full analysis set |
| Method | ANCOVA |
The reported estimate of 12 is the median difference in final values between the randomized treatment groups under the specified non-parametric ANCOVA analysis. It describes a difference in the analyzed 6-minute walk-distance outcome; it does not mean that every patient improved by exactly 12 meters.
The two-sided 99% confidence interval of 1 to 24 communicates the precision of the estimated location shift under the stated analysis. It does not represent the range of individual patient changes.
The P-value of 0.0027 describes evidence against the null hypothesis within the superiority testing framework. It does not quantify clinical importance or the magnitude of individual benefit.
Interpretation should also account for the fact that 21.6% of subjects had missing values that were imputed according to prespecified rules. The imputation strategy is therefore part of the statistical result rather than a minor data-cleaning detail.
9. Secondary Result: Absence of Worsening in WHO Functional Class
The second posted secondary analysis examined the absence of worsening from baseline to Week 26 in modified NYHA/WHO Functional Class. Patients with WHO FC IV at baseline were excluded because they could not shift to a worse category.
Odds ratio for absence of worsening
Two-sided 99% CI: 0.811–1.664 · P = 0.2843
Hypothesis type: Non-inferiority or equivalence
| Endpoint | Result |
|---|---|
| Outcome | Absence of worsening from baseline to Week 26 in modified NYHA/WHO Functional Class |
| Effect measure | Odds ratio |
| Estimate | 1.161 |
| Confidence interval | Two-sided 99% CI: 0.811–1.664 |
| P-value | 0.2843 |
| Analysis population | Full analysis set; baseline WHO FC IV patients excluded |
| Method | Cochran-Mantel-Haenszel test |
| Hypothesis | Non-inferiority or equivalence |
An odds ratio of 1.161 means that the estimated odds of remaining free from worsening were 1.161 times the corresponding odds in the comparator group under the reported analysis. An odds ratio is not a risk ratio and should not be interpreted as a 16.1% increase in the probability of avoiding worsening.
The two-sided 99% confidence interval of 0.811 to 1.664 spans 1. This indicates substantial statistical uncertainty about the relative odds under this analysis. The interval is more informative about precision than the point estimate alone.
The P-value of 0.2843 is not a measure of effect size. It summarizes evidence against the relevant null hypothesis under the stated testing framework.
Importantly, the registry labels the hypothesis as non-inferiority or equivalence and states that it was assumed that the probabilities for absence of worsening at Week 26 were the same for both treatment groups. The ClinicalTrials.gov record does not provide a numerical non-inferiority margin. Therefore, a formal non-inferiority conclusion cannot be reconstructed from the reported odds ratio and P-value alone.
10. How the Three Posted Analyses Fit Together
| Endpoint | Data type | Method | Effect measure | Hypothesis |
|---|---|---|---|---|
| Time to first morbidity event or death | Time-to-event | Log-rank | Hazard ratio | Superiority |
| Change in 6MWD at Week 26 | Continuous | ANCOVA | Median difference | Superiority |
| Absence of worsening in WHO FC at Week 26 | Binary | Cochran-Mantel-Haenszel | Odds ratio | Non-inferiority or equivalence |
This combination illustrates why statistical analysis must be matched to endpoint structure. A time-to-event endpoint requires methods that account for follow-up and censoring. A continuous Week 26 endpoint can use covariate-adjusted regression. A binary functional-class outcome can be analyzed through contingency-table methods such as the Cochran-Mantel-Haenszel procedure.
The effect measures should likewise not be interchanged. A hazard ratio, median difference, and odds ratio answer different statistical questions and have different interpretations.
11. Serious Adverse Events
The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm using affected patients and the corresponding at-risk denominator.
| Safety measure | Selexipag | Placebo |
|---|---|---|
| Serious adverse events | 252 / 575 | 272 / 577 |
The affected/at-risk counts should be read as 252/575 and 272/577, rather than converted into a newly calculated percentage here. The ClinicalTrials.gov record does not provide a formal statistical comparison for this safety measure, so this page does not infer a P-value, confidence interval, risk ratio, or odds ratio from the counts.
12. Statistical Methods Explained
Why was a log-rank test used for the primary endpoint?
The primary endpoint records the time until the first morbidity event or death. Unlike a simple binary endpoint measured at one fixed time, this outcome uses information about when events occurred and can accommodate patients whose observation ends before an event occurs. The log-rank test is designed for comparing time-to-event distributions between randomized groups.
What does a hazard ratio of 0.6 mean?
A hazard ratio of 0.6 represents an estimated instantaneous event rate that is 60% of the comparator rate under the time-to-event analysis. Equivalently, it corresponds to a 40% lower estimated hazard relative to the comparator. It is not a statement that 40% of patients avoided an event.
Why is the confidence interval 99% rather than 95%?
The primary and secondary results posted on ClinicalTrials.gov for GRIPHON report two-sided 99% confidence intervals. The confidence level is part of the prespecified statistical reporting framework. A 99% interval is wider than a 95% interval calculated from the same underlying information, reflecting a higher confidence level.
Why was ANCOVA used for 6-minute walk distance?
The Week 26 endpoint is continuous and the registry specifies ANCOVA with baseline 6MWD as a covariate. Covariate adjustment can account for baseline measurements when estimating the treatment-group difference at follow-up. The registry further specifies that the analysis was non-parametric.
Why does missing-data imputation matter for 6MWD?
Because the Week 26 outcome was not available for every subject, the analysis needed explicit rules for assigning values to missing observations. The registry specifies 0 meters when a patient was unable to walk at Week 26 and otherwise uses the second lowest observed Week 26 value, 10 meters. The stated 21.6% imputation rate means those rules affect a meaningful portion of the analyzed data.
What does an odds ratio of 1.161 mean?
An odds ratio of 1.161 indicates that the estimated odds of absence of worsening were 1.161 times those in the comparator group. Odds are not probabilities, so this cannot be translated directly into a 16.1% higher probability. The accompanying 99% confidence interval, 0.811 to 1.664, also shows that the estimate is imprecise enough to include an odds ratio of 1.
Why can't the non-inferiority result be judged from the P-value alone?
Non-inferiority requires comparison of the estimated treatment effect and its confidence interval with a prespecified non-inferiority margin. The ClinicalTrials.gov record identifies the WHO functional-class analysis as non-inferiority or equivalence and state the equal-probability assumption, but they do not provide a numerical margin. Therefore, the formal non-inferiority decision cannot be reconstructed from the reported P-value alone.
13. One-Sided Testing and Two-Sided Confidence Intervals
The primary analysis notes contain an important statistical detail: the primary analysis was performed using a one-sided unstratified log-rank test, while the reported hazard-ratio confidence interval is a two-sided 99% confidence interval.
One-sided hypothesis test
A one-sided test evaluates evidence in a prespecified direction. Here, that direction is relevant to the superiority hypothesis for the primary endpoint.
Two-sided confidence interval
A two-sided 99% confidence interval expresses uncertainty around the hazard-ratio estimate on both sides of the estimate.
These two reporting components answer related but different questions. The P-value describes the evidence under the specified hypothesis test, while the confidence interval describes the precision of the estimated treatment effect. Neither should be treated as a substitute for the other.
14. Stratified Analysis and the Cochran-Mantel-Haenszel Framework
The registry-reported analysis metadata identifies stratified analysis as an additional concept associated with the primary endpoint. However, the specific primary analysis notes state that the primary analysis itself was an unstratified log-rank test. The page therefore does not infer particular randomization strata or insert a stratification factor that is not reported in the ClinicalTrials.gov record.
The WHO functional-class analysis used the Cochran-Mantel-Haenszel test. This family of methods is useful when a binary outcome is evaluated while accounting for categorical strata. The registry names the procedure but does not specify the individual stratification variables used in that analysis.
15. Missing Data and Imputation
Missing-data handling is explicitly documented for the 6-minute walk distance analysis. The registry states that missing values were imputed for 21.6% of subjects.
| Situation | Imputation rule |
|---|---|
| Patient unable to walk at Week 26 | 0 meter imputed |
| Rule 1 did not apply | Second lowest observed 6MWD value at Week 26, 10 meters, imputed |
| Subjects with imputed values | 21.6% |
These rules have a direct statistical consequence: the analysis does not simply discard all subjects without an observed Week 26 value. Instead, specified values enter the final-value analysis. This makes the imputation assumptions part of the interpretation of the reported median difference and its confidence interval.
The ClinicalTrials.gov record does not describe a separate missing-data or imputation strategy for the primary time-to-event endpoint. Accordingly, no additional imputation procedure is attributed to that analysis.
16. Non-Inferiority and Equivalence: What the Registry Supports
The WHO functional-class analysis is explicitly labeled with a hypothesis type of non-inferiority or equivalence. The registry comment states that it was assumed that the probabilities for absence of worsening in WHO functional class at Week 26 were the same for both treatment groups.
A formal non-inferiority conclusion requires a prespecified margin defining how much loss of effect would still be considered acceptable. The registry-reported GRIPHON data do not provide that numerical margin.
Consequently, the reported odds ratio of 1.161, 99% CI 0.811–1.664, and P-value 0.2843 can be reported exactly, but the ClinicalTrials.gov record is insufficient to reconstruct the complete non-inferiority decision rule.
17. Multiplicity and Multiple Analyses
The registry reports 3 statistical analyses: one primary-endpoint analysis and two secondary analyses. The ClinicalTrials.gov record identifies different hypothesis types across these analyses, including superiority and non-inferiority or equivalence.
| Analysis | Role | Hypothesis | Reported P-value |
|---|---|---|---|
| First morbidity event or death | Primary | Superiority | <0.0001 |
| 6MWD change at Week 26 | Secondary | Superiority | 0.0027 |
| Absence of WHO FC worsening | Secondary | Non-inferiority or equivalence | 0.2843 |
The ClinicalTrials.gov record does not describe a multiplicity-adjustment procedure or an endpoint hierarchy beyond identifying the primary and secondary roles. Therefore, the individual P-values should not be reinterpreted as if an unreported multiplicity strategy had been applied.
18. Results Summary
| Endpoint | Estimate | 99% CI | P-value | Method |
|---|---|---|---|---|
| Time to first morbidity event or death | HR 0.6 | 0.46–0.78 | <0.0001 | Log-rank |
| Change from baseline to Week 26 in 6MWD | Median difference 12 | 1–24 | 0.0027 | ANCOVA |
| Absence of WHO FC worsening at Week 26 | OR 1.161 | 0.811–1.664 | 0.2843 | Cochran-Mantel-Haenszel |
The results illustrate three distinct statistical estimands. The primary analysis estimates a relative difference in time-to-event hazards. The 6MWD analysis estimates a difference in the location of the Week 26 outcome after baseline adjustment. The WHO functional-class analysis estimates a relative difference in odds for a binary outcome.
19. What the Primary Hazard Ratio Does — and Does Not — Mean
The primary hazard ratio of 0.6 corresponds to an estimated instantaneous event rate approximately 40% lower in the selexipag group than in the placebo group under the reported time-to-event analysis.
The hazard ratio does not mean that 40% of patients avoided morbidity or death, that individual patients had exactly a 40% reduction in risk, or that the absolute probability of an event was reduced by 40 percentage points.
The two-sided 99% confidence interval of 0.46 to 0.78 quantifies uncertainty around the estimated hazard ratio. It does not describe the distribution of treatment effects across individual patients.
The reported P < 0.0001 provides evidence against the null hypothesis under the specified one-sided log-rank framework. A P-value is not an effect-size measure and should be interpreted alongside the hazard ratio and confidence interval.
20. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The primary randomized time-to-event comparison produced a hazard ratio of 0.6 with a two-sided 99% confidence interval of 0.46–0.78 and a one-sided log-rank P-value below 0.0001. The Week 26 6MWD analysis reported a median difference of 12, while the WHO functional-class analysis reported an odds ratio of 1.161.
Clinical interpretation
The statistical results describe differences in time to the composite primary event, walking distance, and functional-class worsening. Translating these estimates into clinical importance requires considering the endpoint definitions, absolute event experience, uncertainty, missing-data assumptions, and the complete clinical context.
21. Important Limitations and Interpretation Issues
- Primary endpoint complexity: the primary endpoint combines the first morbidity event with death from all causes. A composite endpoint should be interpreted according to its complete registered definition rather than treating the hazard ratio as a mortality-only effect.
- Time-to-event interpretation: the hazard ratio is a relative time-to-event measure and should not be converted directly into an absolute risk reduction.
- One-sided testing: the primary analysis used a one-sided unstratified log-rank test, while the reported confidence interval was two-sided at 99%. These components should be interpreted according to their respective statistical roles.
- Non-inferiority margin unavailable: the WHO functional-class analysis is labeled non-inferiority or equivalence, but the ClinicalTrials.gov record does not state a numerical non-inferiority margin.
- Missing data: 21.6% of subjects had imputed Week 26 6MWD values. The specified imputation rules therefore materially form part of the secondary endpoint analysis.
- Baseline WHO FC IV exclusion: patients with WHO FC IV at baseline were excluded from the absence-of-worsening analysis because they could not shift to a worse category.
- Multiplicity: the registry provides three posted statistical analyses but does not provide a multiplicity-adjustment strategy in the ClinicalTrials.gov record. Individual P-values should therefore not be assigned an unreported confirmatory error-control framework.
- Stratification details: stratified analysis is identified as an analysis concept, but the ClinicalTrials.gov record does not specify the individual strata. The primary analysis is explicitly described as unstratified.
- Safety inference: serious-adverse-event counts are reported by arm, but no formal safety comparison is reported. A new comparative P-value or effect estimate should not be reconstructed from the counts alone.
22. Why This Trial Matters Statistically
GRIPHON is a useful teaching case because the registry presents several different statistical structures within one randomized clinical trial. The primary endpoint is time-to-event, while the posted secondary analyses include both continuous and binary outcomes.
| Concept | How it appears in GRIPHON |
|---|---|
| Randomization | Randomized parallel-group phase 3 design |
| Blinding | Quadruple masking |
| Intention-to-treat analysis | Primary endpoint analyzed using the Full Analysis Set according to randomized treatment |
| Kaplan-Meier estimation | Used for the primary time-to-event endpoint |
| Log-rank testing | Primary comparison between selexipag and placebo |
| Hazard ratio | Primary effect measure, estimate 0.6 |
| Confidence interval | Two-sided 99% CI for the primary hazard ratio |
| ANCOVA | Analysis of Week 26 6MWD with baseline 6MWD as a covariate |
| Missing-data imputation | Explicit rules for missing Week 26 6MWD; 21.6% of subjects imputed |
| Cochran-Mantel-Haenszel test | Binary WHO functional-class analysis |
| Odds ratio | Effect measure for absence of WHO FC worsening |
| Non-inferiority / equivalence | Hypothesis type reported for the WHO functional-class analysis |
| One-sided testing | Primary analysis used a one-sided unstratified log-rank test |
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Statistical Calculators
25. Sources
- ClinicalTrials.gov: NCT01106014 — GRIPHON.
- Linked publication: PubMed record — PMID 26699168.
- Linked publication: PubMed record — PMID 39222737.
- Linked publication: PubMed record — PMID 34727317.
- Linked publication: PubMed record — PMID 33545163.
- Linked publication: PubMed record — PMID 30982349.
Continue through the Clinical Biostats statistical pathway
Explore the statistical concepts behind randomized trials, survival endpoints, covariate adjustment, categorical analyses, confidence intervals, and missing-data methods.
26. Record Summary
GRIPHON provides a compact example of how different clinical-trial endpoints require different statistical tools. The primary endpoint is a time-to-event outcome analyzed in the Full Analysis Set with a one-sided unstratified log-rank test and reported using a hazard ratio. The two secondary analyses use different structures: a non-parametric ANCOVA for Week 26 6-minute walk distance and a Cochran-Mantel-Haenszel analysis for absence of worsening in WHO functional class.
The primary hazard ratio was 0.6, with a two-sided 99% confidence interval of 0.46–0.78 and a reported P-value of <0.0001. The 6MWD analysis reported a median difference of 12, with a two-sided 99% confidence interval of 1–24 and P = 0.0027. The WHO functional-class analysis reported an odds ratio of 1.161, with a two-sided 99% confidence interval of 0.811–1.664 and P = 0.2843.
The most important statistical lesson is that these estimates should not be collapsed into a single measure of "benefit." Hazard ratios, median differences, and odds ratios describe different aspects of the randomized comparison. Their interpretation depends on the endpoint definition, analysis population, missing-data rules, hypothesis framework, and confidence intervals.