← Clinical Trials
Pulmonary Arterial Hypertension Phase 3 Randomized NCT01106014

GRIPHON: Complete Statistical Analysis of Selexipag in Pulmonary Arterial Hypertension

An independent statistical review of the randomized phase 3 GRIPHON trial evaluating selexipag versus placebo in patients with pulmonary arterial hypertension, with emphasis on the primary time-to-event endpoint, secondary efficacy analyses, and the methods used to quantify treatment effects.

Trial status: Completed  ·  Enrollment: 1156  ·  Sponsor: Actelion
Scope of this record

This page separates reported trial results from statistical interpretation. Trial-specific numerical results and design facts are restricted to the ClinicalTrials.gov record for NCT01106014. The registry provides the official trial record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

GRIPHON was a completed phase 3 randomized, parallel-group trial evaluating selexipag versus placebo in pulmonary arterial hypertension. The trial enrolled 1156 patients and used quadruple masking. Its registered primary endpoint was a time-to-event outcome defined from randomization through 7 days after the last study-drug intake.

1156
Enrollment
Randomized trial
2
Arms
Selexipag vs placebo
0.6
Primary HR
99% CI 0.46–0.78
<0.0001
Primary P-value
Superiority analysis
FeatureGRIPHON
PhasePhase 3
ConditionPulmonary Arterial Hypertension
DesignRandomized, parallel
MaskingQuadruple
Primary purposeTreatment
Enrollment1156
InterventionsSelexipag and placebo
Primary endpoint typeTime-to-event
Results postedYes
Statistical analyses posted3
Lead sponsorActelion
Trial statusCompleted

2. Clinical Question

The central statistical question was whether the randomized comparison of selexipag versus placebo differed with respect to the time from randomization to the first morbidity event or death from all causes, within the registered follow-up window.

Population

Patients enrolled in the phase 3 trial for pulmonary arterial hypertension.

Intervention

Selexipag.

Comparator

Placebo.

Primary question

Does selexipag alter the time to the first morbidity event or death from all causes compared with placebo?

3. Trial Design

01
Randomize1156 enrolled
02
Parallel armsSelexipag vs placebo
03
Quadruple maskingMasked trial design
04
Follow-upTime-to-event endpoint
05
AnalyzeLog-rank comparison
ARM A

Selexipag

  • Drug intervention
  • Randomized parallel-group assignment
  • Compared with placebo
ARM B

Placebo

  • Placebo comparator
  • Randomized parallel-group assignment
  • Compared with selexipag
Allocation
Randomized
Design model
Parallel
Masking
Quadruple
Primary purpose
Treatment

4. Endpoints

EndpointTime frameType
Time From Randomization to the First Morbidity Event or Death (All Causes) up to 7 Days After the Last Study Drug Intake Up to 7 days after end of double-blind treatment (maximum: 4.3 years) Time-to-event
Change From Baseline to Week 26 in 6-minute Walk Distance (6MWD) at Trough Week 26 Continuous
Absence of Worsening From Baseline to Week 26 in Modified NYHA/WHO Functional Class (WHO FC) Week 26 Binary

Primary endpoint definition

The registered primary endpoint was the time from randomization to the first morbidity event or death from all causes up to 7 days after the last study drug intake. The registry definition states that the endpoint was analyzed with the Kaplan-Meier method, with event-free Kaplan-Meier estimates at different time points.

5. Analysis Populations and Statistical Framework

AnalysisPopulationRole
Primary endpoint Full analysis set All randomized patients evaluated according to the study drug to which they were randomized; the registry describes this as the intention-to-treat analysis.
6MWD Full analysis set Continuous secondary endpoint analyzed using ANCOVA.
WHO functional class Full analysis set, excluding baseline WHO FC IV patients Binary secondary endpoint analyzed using the Cochran-Mantel-Haenszel test.

The primary endpoint analysis is explicitly anchored to randomized treatment assignment. This is important because an intention-to-treat analysis preserves the treatment comparison created by randomization rather than redefining treatment groups according to subsequent treatment exposure.

6. Statistical Methodology

Kaplan-Meier estimation

The registry definition for the primary endpoint states that the endpoint was analyzed using the Kaplan-Meier method, including event-free Kaplan-Meier estimates at different time points. This is appropriate for a time-to-event endpoint because patients can have different follow-up durations and some patients may not experience the event during observation.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents events occurring at time ti, while ni represents the number at risk immediately before that time.

Log-rank test

The formal primary comparison was a log-rank test. The registry analysis notes specify that the primary analysis was performed on the Full Analysis Set using a one-sided unstratified log-rank test.

The log-rank test compares the observed and expected pattern of events between treatment groups across follow-up. It is therefore a global comparison of the time-to-event experience rather than a comparison of a single fixed-time percentage.

Hazard ratio

The effect measure reported for the primary endpoint was a hazard ratio. The estimate was 0.6, with a two-sided 99% confidence interval of 0.46 to 0.78.

Interpretation of a hazard ratio
HR < 1  →  lower estimated instantaneous event rate in the selexipag group

A hazard ratio is a relative time-to-event measure. It is not a probability, a percentage of patients who benefit, or an absolute difference in event-free survival.

ANCOVA for 6-minute walk distance

The change from baseline to Week 26 in 6-minute walk distance was analyzed using ANCOVA. The registry analysis notes specify a non-parametric ANCOVA with 6MWD as a covariate at baseline.

Covariate adjustment can improve precision by accounting for baseline measurement differences. Here, the baseline 6MWD value was incorporated into the analysis rather than comparing Week 26 values without adjustment.

Missing-data imputation for 6MWD

The registry analysis specifies prespecified imputation rules for missing Week 26 6MWD values. If a patient was unable to walk at Week 26, 0 meter was imputed. If that rule did not apply, the second lowest observed 6MWD value, 10 meters, at Week 26 was imputed. Missing values were imputed for 21.6% of subjects.

Why this matters: imputation is part of the estimand's operational definition. A result from an ANCOVA is not independent of how missing Week 26 outcomes are handled. The relatively explicit registry rules allow the reader to understand how missing observations entered this particular analysis.

Cochran-Mantel-Haenszel analysis

The binary WHO functional-class endpoint was analyzed using the Cochran-Mantel-Haenszel test. The reported effect measure was an odds ratio, with the estimate expressed as 1.161 and a two-sided 99% confidence interval of 0.811 to 1.664.

7. Primary Result: Time to First Morbidity Event or Death

The primary endpoint was analyzed in the Full Analysis Set according to randomized treatment assignment. The reported method was a log-rank test, specifically described in the analysis notes as a one-sided unstratified log-rank test.

Hazard ratio for first morbidity event or death

0.6

99% CI: 0.46–0.78   ·   P < 0.0001

Hypothesis type: Superiority

Primary endpointSelexipag vs placebo
Effect measureHazard ratio
Estimate0.6
Confidence intervalTwo-sided 99% CI: 0.46–0.78
P-value<0.0001
HypothesisSuperiority
Analysis populationFull analysis set / intention-to-treat
Statistical methodOne-sided unstratified log-rank test
Clinical Biostats interpretation

The reported hazard ratio of 0.6 means that the estimated instantaneous rate of experiencing the primary event, under the time-to-event comparison, was approximately 40% lower with selexipag than with placebo. That is a relative hazard interpretation; it is not the same as saying that 40% of patients avoided an event, or that every individual patient had a 40% reduction in risk.

The two-sided 99% confidence interval of 0.46 to 0.78 describes the statistical uncertainty around the reported hazard-ratio estimate under the analysis framework. It is an interval for the treatment-effect parameter, not a range containing the outcomes that individual patients might experience.

The P-value <0.0001 addresses the evidence against the null hypothesis under the specified testing framework. It does not measure the magnitude of the treatment effect. The magnitude is described by the hazard ratio and its confidence interval.

The interpretation also depends on the time-to-event framework, including censoring and the assumptions underlying a hazard-ratio summary. A single hazard ratio is most naturally interpreted as a relative event-rate comparison over follow-up; it should not automatically be translated into an absolute risk reduction at a particular time point.

Finally, the analysis notes distinguish the one-sided hypothesis test from the reported two-sided 99% confidence interval. Those are related but not identical pieces of statistical reporting and should not be treated as interchangeable.

8. Secondary Result: Change in 6-Minute Walk Distance

The first posted secondary analysis evaluated the change from baseline to Week 26 in 6-minute walk distance at trough. The analysis used the Full Analysis Set and a non-parametric ANCOVA with baseline 6MWD as a covariate.

Median difference in final values

12 meters

Two-sided 99% CI: 1–24   ·   P = 0.0027

Hypothesis type: Superiority

EndpointResult
OutcomeChange from baseline to Week 26 in 6-minute walk distance at trough
Effect measureMedian Difference (Final Values)
Estimate12
Confidence intervalTwo-sided 99% CI: 1–24
P-value0.0027
Analysis populationFull analysis set
MethodANCOVA
Clinical Biostats interpretation

The reported estimate of 12 is the median difference in final values between the randomized treatment groups under the specified non-parametric ANCOVA analysis. It describes a difference in the analyzed 6-minute walk-distance outcome; it does not mean that every patient improved by exactly 12 meters.

The two-sided 99% confidence interval of 1 to 24 communicates the precision of the estimated location shift under the stated analysis. It does not represent the range of individual patient changes.

The P-value of 0.0027 describes evidence against the null hypothesis within the superiority testing framework. It does not quantify clinical importance or the magnitude of individual benefit.

Interpretation should also account for the fact that 21.6% of subjects had missing values that were imputed according to prespecified rules. The imputation strategy is therefore part of the statistical result rather than a minor data-cleaning detail.

9. Secondary Result: Absence of Worsening in WHO Functional Class

The second posted secondary analysis examined the absence of worsening from baseline to Week 26 in modified NYHA/WHO Functional Class. Patients with WHO FC IV at baseline were excluded because they could not shift to a worse category.

Odds ratio for absence of worsening

1.161

Two-sided 99% CI: 0.811–1.664   ·   P = 0.2843

Hypothesis type: Non-inferiority or equivalence

EndpointResult
OutcomeAbsence of worsening from baseline to Week 26 in modified NYHA/WHO Functional Class
Effect measureOdds ratio
Estimate1.161
Confidence intervalTwo-sided 99% CI: 0.811–1.664
P-value0.2843
Analysis populationFull analysis set; baseline WHO FC IV patients excluded
MethodCochran-Mantel-Haenszel test
HypothesisNon-inferiority or equivalence
Clinical Biostats interpretation

An odds ratio of 1.161 means that the estimated odds of remaining free from worsening were 1.161 times the corresponding odds in the comparator group under the reported analysis. An odds ratio is not a risk ratio and should not be interpreted as a 16.1% increase in the probability of avoiding worsening.

The two-sided 99% confidence interval of 0.811 to 1.664 spans 1. This indicates substantial statistical uncertainty about the relative odds under this analysis. The interval is more informative about precision than the point estimate alone.

The P-value of 0.2843 is not a measure of effect size. It summarizes evidence against the relevant null hypothesis under the stated testing framework.

Importantly, the registry labels the hypothesis as non-inferiority or equivalence and states that it was assumed that the probabilities for absence of worsening at Week 26 were the same for both treatment groups. The ClinicalTrials.gov record does not provide a numerical non-inferiority margin. Therefore, a formal non-inferiority conclusion cannot be reconstructed from the reported odds ratio and P-value alone.

10. How the Three Posted Analyses Fit Together

EndpointData typeMethodEffect measureHypothesis
Time to first morbidity event or death Time-to-event Log-rank Hazard ratio Superiority
Change in 6MWD at Week 26 Continuous ANCOVA Median difference Superiority
Absence of worsening in WHO FC at Week 26 Binary Cochran-Mantel-Haenszel Odds ratio Non-inferiority or equivalence

This combination illustrates why statistical analysis must be matched to endpoint structure. A time-to-event endpoint requires methods that account for follow-up and censoring. A continuous Week 26 endpoint can use covariate-adjusted regression. A binary functional-class outcome can be analyzed through contingency-table methods such as the Cochran-Mantel-Haenszel procedure.

The effect measures should likewise not be interchanged. A hazard ratio, median difference, and odds ratio answer different statistical questions and have different interpretations.

11. Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm using affected patients and the corresponding at-risk denominator.

Safety measureSelexipagPlacebo
Serious adverse events252 / 575272 / 577
Serious adverse events: affected patients
Selexipag
252
Placebo
272

The affected/at-risk counts should be read as 252/575 and 272/577, rather than converted into a newly calculated percentage here. The ClinicalTrials.gov record does not provide a formal statistical comparison for this safety measure, so this page does not infer a P-value, confidence interval, risk ratio, or odds ratio from the counts.

Safety interpretation: the serious-adverse-event counts describe an observed safety category by arm. They do not, by themselves, establish a causal difference between treatment groups or provide the complete safety profile of the trial.

12. Statistical Methods Explained

Why was a log-rank test used for the primary endpoint?

The primary endpoint records the time until the first morbidity event or death. Unlike a simple binary endpoint measured at one fixed time, this outcome uses information about when events occurred and can accommodate patients whose observation ends before an event occurs. The log-rank test is designed for comparing time-to-event distributions between randomized groups.

What does a hazard ratio of 0.6 mean?

A hazard ratio of 0.6 represents an estimated instantaneous event rate that is 60% of the comparator rate under the time-to-event analysis. Equivalently, it corresponds to a 40% lower estimated hazard relative to the comparator. It is not a statement that 40% of patients avoided an event.

Why is the confidence interval 99% rather than 95%?

The primary and secondary results posted on ClinicalTrials.gov for GRIPHON report two-sided 99% confidence intervals. The confidence level is part of the prespecified statistical reporting framework. A 99% interval is wider than a 95% interval calculated from the same underlying information, reflecting a higher confidence level.

Why was ANCOVA used for 6-minute walk distance?

The Week 26 endpoint is continuous and the registry specifies ANCOVA with baseline 6MWD as a covariate. Covariate adjustment can account for baseline measurements when estimating the treatment-group difference at follow-up. The registry further specifies that the analysis was non-parametric.

Why does missing-data imputation matter for 6MWD?

Because the Week 26 outcome was not available for every subject, the analysis needed explicit rules for assigning values to missing observations. The registry specifies 0 meters when a patient was unable to walk at Week 26 and otherwise uses the second lowest observed Week 26 value, 10 meters. The stated 21.6% imputation rate means those rules affect a meaningful portion of the analyzed data.

What does an odds ratio of 1.161 mean?

An odds ratio of 1.161 indicates that the estimated odds of absence of worsening were 1.161 times those in the comparator group. Odds are not probabilities, so this cannot be translated directly into a 16.1% higher probability. The accompanying 99% confidence interval, 0.811 to 1.664, also shows that the estimate is imprecise enough to include an odds ratio of 1.

Why can't the non-inferiority result be judged from the P-value alone?

Non-inferiority requires comparison of the estimated treatment effect and its confidence interval with a prespecified non-inferiority margin. The ClinicalTrials.gov record identifies the WHO functional-class analysis as non-inferiority or equivalence and state the equal-probability assumption, but they do not provide a numerical margin. Therefore, the formal non-inferiority decision cannot be reconstructed from the reported P-value alone.

13. One-Sided Testing and Two-Sided Confidence Intervals

The primary analysis notes contain an important statistical detail: the primary analysis was performed using a one-sided unstratified log-rank test, while the reported hazard-ratio confidence interval is a two-sided 99% confidence interval.

One-sided hypothesis test

A one-sided test evaluates evidence in a prespecified direction. Here, that direction is relevant to the superiority hypothesis for the primary endpoint.

Two-sided confidence interval

A two-sided 99% confidence interval expresses uncertainty around the hazard-ratio estimate on both sides of the estimate.

These two reporting components answer related but different questions. The P-value describes the evidence under the specified hypothesis test, while the confidence interval describes the precision of the estimated treatment effect. Neither should be treated as a substitute for the other.

14. Stratified Analysis and the Cochran-Mantel-Haenszel Framework

The registry-reported analysis metadata identifies stratified analysis as an additional concept associated with the primary endpoint. However, the specific primary analysis notes state that the primary analysis itself was an unstratified log-rank test. The page therefore does not infer particular randomization strata or insert a stratification factor that is not reported in the ClinicalTrials.gov record.

The WHO functional-class analysis used the Cochran-Mantel-Haenszel test. This family of methods is useful when a binary outcome is evaluated while accounting for categorical strata. The registry names the procedure but does not specify the individual stratification variables used in that analysis.

Methodological discipline: statistical methods should be reported at the level supported by the registry. Knowing that a method is stratified does not justify reconstructing the strata when their identities are not contained in the ClinicalTrials.gov record.

15. Missing Data and Imputation

Missing-data handling is explicitly documented for the 6-minute walk distance analysis. The registry states that missing values were imputed for 21.6% of subjects.

SituationImputation rule
Patient unable to walk at Week 260 meter imputed
Rule 1 did not applySecond lowest observed 6MWD value at Week 26, 10 meters, imputed
Subjects with imputed values21.6%

These rules have a direct statistical consequence: the analysis does not simply discard all subjects without an observed Week 26 value. Instead, specified values enter the final-value analysis. This makes the imputation assumptions part of the interpretation of the reported median difference and its confidence interval.

The ClinicalTrials.gov record does not describe a separate missing-data or imputation strategy for the primary time-to-event endpoint. Accordingly, no additional imputation procedure is attributed to that analysis.

16. Non-Inferiority and Equivalence: What the Registry Supports

The WHO functional-class analysis is explicitly labeled with a hypothesis type of non-inferiority or equivalence. The registry comment states that it was assumed that the probabilities for absence of worsening in WHO functional class at Week 26 were the same for both treatment groups.

General non-inferiority logic
Treatment effect + confidence interval   compared with   prespecified NI margin

A formal non-inferiority conclusion requires a prespecified margin defining how much loss of effect would still be considered acceptable. The registry-reported GRIPHON data do not provide that numerical margin.

Consequently, the reported odds ratio of 1.161, 99% CI 0.811–1.664, and P-value 0.2843 can be reported exactly, but the ClinicalTrials.gov record is insufficient to reconstruct the complete non-inferiority decision rule.

17. Multiplicity and Multiple Analyses

The registry reports 3 statistical analyses: one primary-endpoint analysis and two secondary analyses. The ClinicalTrials.gov record identifies different hypothesis types across these analyses, including superiority and non-inferiority or equivalence.

AnalysisRoleHypothesisReported P-value
First morbidity event or deathPrimarySuperiority<0.0001
6MWD change at Week 26SecondarySuperiority0.0027
Absence of WHO FC worseningSecondaryNon-inferiority or equivalence0.2843

The ClinicalTrials.gov record does not describe a multiplicity-adjustment procedure or an endpoint hierarchy beyond identifying the primary and secondary roles. Therefore, the individual P-values should not be reinterpreted as if an unreported multiplicity strategy had been applied.

Important distinction: three posted statistical analyses do not automatically imply three independent confirmatory tests. The inferential status of each result depends on the prespecified protocol and statistical analysis plan, including any alpha allocation or multiplicity procedure. Those details are not reported here.

18. Results Summary

EndpointEstimate99% CIP-valueMethod
Time to first morbidity event or death HR 0.6 0.46–0.78 <0.0001 Log-rank
Change from baseline to Week 26 in 6MWD Median difference 12 1–24 0.0027 ANCOVA
Absence of WHO FC worsening at Week 26 OR 1.161 0.811–1.664 0.2843 Cochran-Mantel-Haenszel

The results illustrate three distinct statistical estimands. The primary analysis estimates a relative difference in time-to-event hazards. The 6MWD analysis estimates a difference in the location of the Week 26 outcome after baseline adjustment. The WHO functional-class analysis estimates a relative difference in odds for a binary outcome.

19. What the Primary Hazard Ratio Does — and Does Not — Mean

Relative effect

The primary hazard ratio of 0.6 corresponds to an estimated instantaneous event rate approximately 40% lower in the selexipag group than in the placebo group under the reported time-to-event analysis.

What it does not mean

The hazard ratio does not mean that 40% of patients avoided morbidity or death, that individual patients had exactly a 40% reduction in risk, or that the absolute probability of an event was reduced by 40 percentage points.

Precision

The two-sided 99% confidence interval of 0.46 to 0.78 quantifies uncertainty around the estimated hazard ratio. It does not describe the distribution of treatment effects across individual patients.

Statistical significance

The reported P < 0.0001 provides evidence against the null hypothesis under the specified one-sided log-rank framework. A P-value is not an effect-size measure and should be interpreted alongside the hazard ratio and confidence interval.

20. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The primary randomized time-to-event comparison produced a hazard ratio of 0.6 with a two-sided 99% confidence interval of 0.46–0.78 and a one-sided log-rank P-value below 0.0001. The Week 26 6MWD analysis reported a median difference of 12, while the WHO functional-class analysis reported an odds ratio of 1.161.

Clinical interpretation

The statistical results describe differences in time to the composite primary event, walking distance, and functional-class worsening. Translating these estimates into clinical importance requires considering the endpoint definitions, absolute event experience, uncertainty, missing-data assumptions, and the complete clinical context.

21. Important Limitations and Interpretation Issues

22. Why This Trial Matters Statistically

GRIPHON is a useful teaching case because the registry presents several different statistical structures within one randomized clinical trial. The primary endpoint is time-to-event, while the posted secondary analyses include both continuous and binary outcomes.

ConceptHow it appears in GRIPHON
RandomizationRandomized parallel-group phase 3 design
BlindingQuadruple masking
Intention-to-treat analysisPrimary endpoint analyzed using the Full Analysis Set according to randomized treatment
Kaplan-Meier estimationUsed for the primary time-to-event endpoint
Log-rank testingPrimary comparison between selexipag and placebo
Hazard ratioPrimary effect measure, estimate 0.6
Confidence intervalTwo-sided 99% CI for the primary hazard ratio
ANCOVAAnalysis of Week 26 6MWD with baseline 6MWD as a covariate
Missing-data imputationExplicit rules for missing Week 26 6MWD; 21.6% of subjects imputed
Cochran-Mantel-Haenszel testBinary WHO functional-class analysis
Odds ratioEffect measure for absence of WHO FC worsening
Non-inferiority / equivalenceHypothesis type reported for the WHO functional-class analysis
One-sided testingPrimary analysis used a one-sided unstratified log-rank test

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical concepts behind randomized trials, survival endpoints, covariate adjustment, categorical analyses, confidence intervals, and missing-data methods.

26. Record Summary

GRIPHON provides a compact example of how different clinical-trial endpoints require different statistical tools. The primary endpoint is a time-to-event outcome analyzed in the Full Analysis Set with a one-sided unstratified log-rank test and reported using a hazard ratio. The two secondary analyses use different structures: a non-parametric ANCOVA for Week 26 6-minute walk distance and a Cochran-Mantel-Haenszel analysis for absence of worsening in WHO functional class.

The primary hazard ratio was 0.6, with a two-sided 99% confidence interval of 0.46–0.78 and a reported P-value of <0.0001. The 6MWD analysis reported a median difference of 12, with a two-sided 99% confidence interval of 1–24 and P = 0.0027. The WHO functional-class analysis reported an odds ratio of 1.161, with a two-sided 99% confidence interval of 0.811–1.664 and P = 0.2843.

The most important statistical lesson is that these estimates should not be collapsed into a single measure of "benefit." Hazard ratios, median differences, and odds ratios describe different aspects of the randomized comparison. Their interpretation depends on the endpoint definition, analysis population, missing-data rules, hypothesis framework, and confidence intervals.

Clinical Biostats methodology: A trial-results page should not merely repeat the registry. The goal is to reconstruct the statistical story of the trial while clearly separating reported evidence from educational interpretation and avoiding unsupported assumptions about unreported design details.