← Clinical Trials
Diabetic Nephropathy Phase 3 Time-to-Event Analysis NCT01858532

SONAR: Complete Statistical Analysis of Atrasentan in Diabetic Nephropathy

An independent statistical review of the randomized phase 3 SONAR trial evaluating atrasentan versus placebo in diabetic nephropathy, with emphasis on the composite renal endpoint, responder-set analysis, hazard ratios, confidence intervals, and stratified log-rank testing.

Trial start: 2013-05-17  ·  Primary completion: 2018-03-29  ·  Sponsor: AbbVie
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

SONAR was a randomized, parallel-group, phase 3 trial in diabetic nephropathy comparing atrasentan with placebo. The registered primary endpoint was a time-to-event composite renal endpoint, analyzed in the Intent-to-Treat (ITT) Responder Set consisting of participants who achieved at least 30% reduction in their albumin to creatinine ratio (UACR) during the Enrichment Period.

5107
Enrollment
Participants
2
Arms
Atrasentan vs placebo
0.654
Primary HR
95% CI 0.488–0.878
0.005
Primary Cox P
Two-sided 95% CI
FeatureSONAR
Trial nameSONAR
Brief titleStudy Of Diabetic Nephropathy With Atrasentan
PhasePhase 3
ConditionDiabetic Nephropathy
DesignRandomized, parallel
MaskingQuadruple
AllocationRandomized
Primary purposeTreatment
Enrollment5107
InterventionsAtrasentan; placebo
Primary endpoint typeTime-to-event
ClinicalTrials.govNCT01858532
Trial statusTerminated

2. Clinical Question

The central statistical question was whether atrasentan was associated with a different time to the first occurrence of a component of the composite renal endpoint than placebo among randomized participants who met the prespecified responder criterion during the Enrichment Period.

Population

Participants with diabetic nephropathy enrolled in the phase 3 SONAR trial. The primary efficacy analysis was conducted in the ITT Responder Set: participants who achieved at least 30% reduction in UACR during the Enrichment Period.

Intervention

Atrasentan during the trial's randomized treatment comparison.

Comparator

Placebo.

Primary question

How did time to the first occurrence of a component of the composite renal endpoint compare between randomized atrasentan and placebo responders?

3. Trial Design

01
Enroll5107 participants
02
EnrichmentUACR response assessed
03
RandomizeAtrasentan or placebo
04
ObserveTime-to-event follow-up
05
AnalyzeCox + stratified log-rank
Allocation
Randomized allocation in a parallel-group phase 3 design.
Masking
Quadruple masking was registered for the trial.
Primary purpose
Treatment.
Analysis framework
Time-to-event treatment comparison using a Cox proportional-hazards model and a stratified log-rank test.
RANDOMIZED ARM

Atrasentan

  • Atrasentan was the active intervention.
  • Primary comparisons were made against placebo.
  • The primary analysis used the ITT Responder Set as randomized.
RANDOMIZED ARM

Placebo

  • Placebo was the randomized comparator.
  • The primary comparison was performed against randomized atrasentan.
  • The same ITT Responder Set framework was used for the primary analysis.

The enrichment feature is statistically important. The primary analysis was not simply a comparison of every enrolled participant: it was conducted among participants who achieved at least 30% reduction in UACR during the Enrichment Period, while retaining their randomized treatment assignment for the subsequent comparison. This creates a clearly defined analysis population and preserves the randomized comparison within that responder set.

4. Endpoints

Primary Endpoint

EndpointDefinition / time frameAnalysis
Time to the First Occurrence of a Component of the Composite Renal Endpoint in the Intent-to-Treat (ITT) Responder Set (as Randomized) From randomization to individual end of observation, up to 53 months. The time to the first occurrence of a component of the composite renal endpoint was defined as doubling of serum creatinine, confirmed by a 30-day serum creatinine measurement, or onset of end stage renal disease: eGFR less than 15 ml/min/1.73 m2 confirmed by a 90-day eGFR measurement, receiving chronic dialysis, renal transplantation, or renal death. Cox proportional-hazards model and stratified log-rank test

The registry's registry-reported endpoint definition ends after the listed renal-death component. The analysis on this page uses only those components and timing details contained in the ClinicalTrials.gov record.

Secondary Endpoints

Secondary endpointTime frameAnalysis method
Time to a 50% Estimated Glomerular Filtration Rate Reduction in the ITT Responder Set (as Randomized) From randomization to individual end of observation, up to 53 months Cox proportional-hazards model; stratified log-rank test
Time to Cardio-renal Composite Endpoint in the ITT Responder Set (as Randomized) From randomization to individual end of observation, up to 53 months Cox proportional-hazards model; stratified log-rank test
Time to First Occurrence of a Component of Composite Renal Endpoint for All Randomized Participants (Pooled) From randomization to individual end of observation, up to 53 months Cox proportional-hazards model
Time to the Cardiovascular Composite Endpoint in the ITT Responder Set (as Randomized) From randomization to individual end of observation, up to 53 months Cox proportional-hazards model; stratified log-rank test

5. Analysis Population and Estimand

The primary efficacy analysis used the Intent-to-Treat (ITT) Responder Set. The registry defines these participants as those who achieved at least 30% reduction in their UACR during the Enrichment Period. They were then compared according to randomized treatment assignment: ITT Responder Atrasentan versus ITT Responder Placebo.

Primary analysis population
ITT Responder Set = participants achieving ≥30% UACR reduction during the Enrichment Period

The treatment comparison remains based on the randomized assignment within this responder set. This is different from an as-treated comparison, which would classify participants according to treatment actually received.

This distinction matters because the responder criterion was determined during the Enrichment Period. The resulting primary analysis therefore answers a narrower question than a comparison of all 5107 enrolled participants: it estimates the randomized treatment difference among the prespecified population that met the UACR-response criterion.

6. Results: Primary Renal Endpoint

ClinicalTrials.gov reports two formal primary analyses for the same registered time-to-event endpoint. One used a Cox proportional-hazards regression model with prespecified covariate adjustment and reported a hazard ratio with a two-sided 95% confidence interval. The other used a stratified log-rank test and reported a P-value.

Cox Proportional-Hazards Analysis

Hazard ratio for the composite renal endpoint

0.654

95% CI: 0.488–0.878   ·   P = 0.005

Analysis population: ITT Responder Set, as randomized.

MeasureReported result
Effect measureHazard Ratio (HR)
Estimate0.654
95% confidence interval0.488–0.878
Confidence intervalTwo-sided
P-value=0.005
ModelCox proportional-hazards model
AdjustmentPrespecified covariates were adjusted for in the Cox regression model.
Clinical Biostats interpretation

An HR of 0.654 means that, under the Cox model, the estimated instantaneous rate of experiencing the first qualifying component of the composite renal endpoint was 34.6% lower for randomized atrasentan responders relative to randomized placebo responders over the analyzed follow-up. This is a relative hazard measure, not a statement that 34.6% of patients avoided the endpoint.

The 95% CI of 0.488–0.878 describes the precision of the estimated hazard ratio under the model. It excludes 1, so the reported interval is compatible with a lower hazard for atrasentan than placebo under the specified model. It does not establish that every patient benefits, nor does it provide an absolute risk reduction.

The reported P = 0.005 addresses the statistical evidence against the null hypothesis represented by the analysis; it is not a measure of the size or clinical importance of the treatment effect. The P-value should therefore be interpreted alongside the hazard ratio and confidence interval.

Because this is a Cox analysis, the hazard-ratio interpretation depends on the model and its proportional-hazards framework. The ClinicalTrials.gov record does not report a separate assessment of the proportional-hazards assumption, so the HR should not be interpreted as a guaranteed constant risk ratio at every time point.

Stratified Log-Rank Analysis

Stratified log-rank treatment comparison

P = 0.029

Primary endpoint  ·  ITT Responder Set  ·  Atrasentan vs placebo

The second primary analysis used a stratified log-rank test for the treatment comparison and reported P = 0.029. Unlike the Cox analysis, this result does not itself provide a treatment-effect estimate such as a hazard ratio or a confidence interval.

Clinical Biostats interpretation

The stratified log-rank P-value provides a hypothesis-test result for the difference between the treatment-group time-to-event experience under the specified stratified analysis. It does not quantify the magnitude of that difference. The HR of 0.654 and its 95% CI come from the separate Cox model.

A P-value of 0.029 should not be read as a 2.9% probability that the treatment effect is real, nor as a 97.1% probability that atrasentan is beneficial. It measures compatibility with the null hypothesis under the statistical test, given the analysis framework.

The combination of the log-rank result and the Cox estimate is informative because the two methods address related aspects of the same time-to-event comparison: the stratified log-rank test supplies the formal comparison, while the Cox model supplies an estimated relative effect and its precision.

As with any time-to-event analysis, censoring and the assumptions underlying the statistical methods affect interpretation. The ClinicalTrials.gov record does not provide enough information to independently reconstruct the censoring process or a Kaplan-Meier curve.

Educational note: a Kaplan-Meier curve is not reconstructed here because the ClinicalTrials.gov record contains summary analysis results rather than participant-level event and censoring times. A valid reconstruction requires the underlying event/censoring information or sufficiently detailed source data.

7. Secondary Endpoint Results

50% eGFR Reduction

Time to a 50% eGFR reduction

HR 0.779

95% CI: 0.573–1.060   ·   P = 0.112

The Cox model estimated a hazard ratio of 0.779 for time to a 50% estimated glomerular filtration rate reduction in the ITT Responder Set. The two-sided 95% CI was 0.573–1.060, with P = 0.112.

The registry also reports a stratified log-rank analysis for this endpoint with P = 0.289.

Clinical Biostats interpretation

The HR of 0.779 corresponds to a model-estimated 22.1% lower instantaneous event rate for atrasentan relative to placebo. The 95% CI includes 1, extending from 0.573 to 1.060, so the estimate is imprecise enough that the data are compatible with effects on either side of the conventional HR = 1 reference under this model.

The P-value of 0.112 is not a measure of effect size. The HR and confidence interval provide the information about estimated magnitude and precision; the P-value addresses the hypothesis test.

Cardio-renal Composite Endpoint

Time to cardio-renal composite endpoint

HR 0.801

95% CI: 0.642–0.999   ·   P = 0.049

For the cardio-renal composite endpoint in the ITT Responder Set, the prespecified-covariate-adjusted Cox model produced an HR of 0.801 with a two-sided 95% CI of 0.642–0.999 and P = 0.049.

The corresponding stratified log-rank analysis reported P = 0.089.

Clinical Biostats interpretation

An HR of 0.801 represents an estimated 19.9% lower instantaneous event rate under the Cox model for atrasentan versus placebo in the responder set. The upper confidence-limit value of 0.999 is very close to 1, illustrating why the point estimate alone should not be treated as a complete description of the uncertainty.

The Cox-model P-value is 0.049, while the separate stratified log-rank P-value is 0.089. These are results from different statistical procedures, so they should not be substituted for one another or treated as if they were two independent treatment effects. The registry does not provide additional information here that would justify a broader multiplicity interpretation.

Composite Renal Endpoint in All Randomized Participants

All randomized participants, pooled

HR 0.72

95% CI: 0.58–0.89   ·   P = 0.002

In the pooled ITT population of participants randomized to atrasentan or placebo during the Double-Blind Treatment Period, the Cox model estimated an HR of 0.72, with a 95% CI of 0.58–0.89 and P = 0.002.

Clinical Biostats interpretation

An HR of 0.72 corresponds to a 28% lower estimated instantaneous event rate for atrasentan versus placebo under the Cox model. This analysis differs importantly from the primary responder-set analysis because its population is described as all randomized participants rather than only participants who achieved the UACR-response criterion during enrichment.

Because the analysis populations differ, the HR 0.72 should not be presented as interchangeable with the primary HR 0.654. They answer related but distinct statistical questions.

Cardiovascular Composite Endpoint

Time to cardiovascular composite endpoint

HR 0.884

95% CI: 0.643–1.215   ·   P = 0.447

The Cox model for the cardiovascular composite endpoint in the ITT Responder Set produced an HR of 0.884, with a two-sided 95% CI of 0.643–1.215 and P = 0.447.

The corresponding stratified log-rank analysis reported P = 0.446.

Clinical Biostats interpretation

The HR of 0.884 corresponds to a 11.6% lower estimated instantaneous event rate under the Cox model, but the 95% CI extends from 0.643 to 1.215 and therefore includes the null value of 1. The ClinicalTrials.gov record therefore do not provide a precise estimate of the direction or magnitude of the cardiovascular composite effect.

The closely corresponding Cox and stratified log-rank P-values, 0.447 and 0.446, are results from different analyses of the same endpoint. Neither P-value is an effect-size measure.

8. Secondary Endpoint Summary

EndpointPopulationEffect / test95% CIP-value
50% eGFR reduction ITT Responder Set HR 0.779 0.573–1.060 0.112
Cardio-renal composite ITT Responder Set HR 0.801 0.642–0.999 0.049
Composite renal endpoint, all randomized participants All randomized participants HR 0.72 0.58–0.89 0.002
Cardiovascular composite ITT Responder Set HR 0.884 0.643–1.215 0.447
EndpointStratified log-rank P-value
Primary composite renal endpoint=0.029
50% eGFR reduction=0.289
Cardio-renal composite0.089
Cardiovascular composite0.446

The registry reports nine statistical analyses in total, including two formal primary-endpoint analyses and seven secondary analyses. The ClinicalTrials.gov record therefore support a broad statistical description of the registered time-to-event results, but they do not support treating every reported P-value as an independently confirmatory test.

9. Safety

The ClinicalTrials.gov record reports serious adverse-event counts for three labeled analysis contexts. Because the denominators and labels differ, these figures should be presented exactly as reported rather than converted into percentages or treated as directly comparable treatment-group rates.

Safety contextSerious adverse eventsAffected / at risk
Enrichment AtrasentanSerious adverse events292/5107
Double-Blind AtrasentanSerious adverse events685/1829
Double-Blind PlaceboSerious adverse events626/1830
Safety interpretation: these counts are not interchangeable denominators. In particular, the ClinicalTrials.gov record labels one figure as "Enrichment Atrasentan" and the other two as "Double-Blind" treatment groups. The data does not provide a single common safety analysis population that would justify collapsing these figures into one direct randomized comparison.

10. Statistical Methodology

Cox Proportional-Hazards Model

The primary Cox analysis used a proportional-hazards regression model with prespecified covariate adjustment to estimate the hazard ratio of atrasentan to placebo and its 95% confidence interval. The same model family was used for several secondary time-to-event endpoints.

Conceptual Cox model
h(t | X) = h0(t) exp(β1Xtreatment + β2X2 + … + βpXp)

For a binary treatment indicator, the treatment hazard ratio is represented by exp(β1). The registry's analysis notes specify that prespecified covariates were adjusted for.

The Cox model is particularly useful for time-to-event outcomes because it uses the ordering and timing of observed events while accommodating right censoring. Its treatment effect is expressed as a hazard ratio rather than a difference in proportions at one fixed time.

Hazard Ratio

Interpretation of the reported effect measure
HR = hAtrasentan(t) / hPlacebo(t)

An HR below 1 indicates a lower estimated instantaneous event rate for atrasentan relative to placebo under the fitted Cox model.

The hazard ratio is not a relative risk. It does not mean that the same percentage of participants avoided the event, and it does not directly provide an absolute risk difference. For SONAR, the HR should be interpreted in conjunction with its confidence interval and with the definition of the analyzed population.

Stratified Log-Rank Test

The registered treatment comparison for the primary endpoint also used a stratified log-rank test. The ClinicalTrials.gov record identifies the method as a stratified analysis but does not provide the specific stratification variables. Therefore, no particular stratification factors are asserted here.

The stratified log-rank approach compares the time-to-event experience between randomized treatment groups while incorporating the registered stratified-analysis framework. It produces a P-value for the treatment comparison but, by itself, does not produce the hazard-ratio estimate shown in the Cox analysis.

Intention-to-Treat Analysis

The registry labels the primary analysis population as an Intent-to-Treat (ITT) Responder Set and explicitly states that the comparison was made "as Randomized." This is important: after the responder criterion is applied, participants remain classified according to their randomized atrasentan or placebo assignment rather than being reassigned according to subsequent exposure.

The term ITT therefore describes the treatment-assignment principle within the specified responder set. It should not be confused with a claim that all 5107 enrolled participants entered the primary efficacy analysis.

Time-to-Event Analysis and Censoring

The primary endpoint was followed from randomization to each participant's individual end of observation, up to 53 months. Time-to-event methods allow participants who do not experience the endpoint during observed follow-up to contribute information until their observation ends.

The ClinicalTrials.gov record does not provide individual censoring times, numbers censored, or a censoring table. Consequently, this page does not attempt to reconstruct survival probabilities or a Kaplan-Meier curve.

11. Statistical Methods Explained

Why was a Cox proportional-hazards model used?

The primary endpoint is explicitly a time-to-event outcome, with follow-up beginning at randomization and continuing until the individual end of observation. A Cox model is suited to this structure because it estimates a relative hazard while allowing participants to have different observed follow-up times and providing a hazard-ratio effect measure.

What does HR 0.654 mean?

Within the specified ITT Responder Set and Cox model, HR 0.654 means the estimated instantaneous rate of the first qualifying composite renal event was 34.6% lower for atrasentan than placebo. It does not mean that 34.6% fewer participants experienced the endpoint, and it is not an absolute risk reduction.

Why is the confidence interval important?

The 95% CI of 0.488–0.878 shows the statistical precision around the reported HR 0.654 under the fitted model. It is substantially more informative than the point estimate alone because it shows a range of effect estimates compatible with the data under the stated confidence procedure. It also indicates that the interval excludes HR = 1.

Why are the Cox P-value and log-rank P-value different?

The Cox regression and stratified log-rank test are different statistical procedures. The Cox analysis estimates a hazard ratio and its confidence interval, while the log-rank analysis supplies a treatment-comparison test. For the primary endpoint, the registry reports P = 0.005 from the Cox analysis and P = 0.029 from the stratified log-rank analysis. These values should not be treated as two estimates of the same quantity.

Why does the responder population matter?

The primary efficacy analysis was restricted to participants who achieved at least 30% reduction in UACR during the Enrichment Period. Therefore, the primary HR describes the randomized treatment comparison within this prespecified responder population. It is not automatically generalizable to every participant represented by the overall enrollment figure of 5107.

Why should the HR not be interpreted as a fixed risk ratio?

A hazard ratio compares instantaneous event rates under a Cox model. It is not the ratio of cumulative event probabilities at a particular time. Its interpretation also relies on the proportional-hazards framework. The ClinicalTrials.gov record does not report a separate proportional-hazards diagnostic, so no stronger claim about the constancy of the HR over time should be made.

12. Reading the P-values Carefully

SONAR illustrates why statistical interpretation should not reduce a clinical trial to a list of P-values. The primary Cox result combines three distinct pieces of information:

The stratified log-rank analysis provides a separate treatment-comparison P-value of 0.029. The secondary endpoints then provide a mixture of hazard-ratio estimates, confidence intervals, and log-rank P-values.

Effect size

The HR describes the estimated relative event rate under the Cox model.

Precision

The confidence interval shows uncertainty around the estimated hazard ratio.

Evidence against the null

The P-value quantifies the result of a specified hypothesis test; it does not measure treatment benefit.

Clinical meaning

Clinical importance cannot be inferred from a P-value alone and requires consideration of the endpoint, magnitude, uncertainty, and population.

13. Comparing the Analysis Populations

One of the most important statistical distinctions in the registry-reported SONAR results is between the responder-set analyses and the analysis reported for all randomized participants.

AnalysisPopulationReported HR95% CIP-value
Primary composite renal endpoint ITT Responder Set 0.654 0.488–0.878 =0.005
Composite renal endpoint All randomized participants 0.72 0.58–0.89 0.002

These two estimates should not be interpreted as conflicting versions of one statistic. The first is explicitly restricted to participants who achieved the UACR-response criterion during enrichment; the second is reported for all randomized participants. Different analysis populations can produce different estimates even when they concern closely related endpoints.

The analysis-population distinction is especially important when communicating results because the headline enrollment number, 5107, should not be silently substituted for the population actually used in the primary efficacy analysis.

14. What the Results Do and Do Not Establish

The reported primary Cox analysis provides evidence about the relative hazard of the specified composite renal endpoint in the ITT Responder Set under the prespecified-covariate-adjusted Cox model. The HR is below 1 and its 95% confidence interval excludes 1.

That statistical result does not by itself establish an absolute reduction in the probability of renal events, a reduction in mortality, or a particular patient-level probability of benefit. Those would require different quantities or additional reported analyses.

Similarly, the secondary results should be interpreted endpoint by endpoint. For example, the 50% eGFR-reduction analysis reports HR 0.779 with a 95% CI of 0.573–1.060 and P = 0.112, whereas the cardiovascular composite analysis reports HR 0.884 with a 95% CI of 0.643–1.215 and P = 0.447. Neither result should be reduced to a binary "worked" or "failed" label without considering the estimate and its uncertainty.

15. Limitations

Data discipline: the ClinicalTrials.gov record does not report baseline characteristics, subgroup estimates, median time-to-event values, Kaplan-Meier survival probabilities, or a complete event-count table. Those quantities are intentionally not added here from outside sources.

16. Why This Trial Matters Statistically

SONAR is particularly useful for teaching clinical trial statistics because its primary analysis combines several important ideas in one design: randomization, enrichment, an ITT analysis principle, time-to-event endpoints, Cox regression, stratified log-rank testing, covariate adjustment, hazard ratios, and confidence intervals.

The enrichment framework also creates an important interpretive distinction. A conventional randomized trial can often be summarized as a treatment comparison across the randomized population. Here, the primary endpoint was analyzed in a responder set defined by a reduction in UACR during the Enrichment Period. That means the statistical question is conditional on entry into that prespecified responder population.

The trial also demonstrates why multiple representations of an endpoint can be useful. The Cox model provides an interpretable effect estimate and confidence interval, while the stratified log-rank test provides a complementary formal treatment comparison. Looking at both prevents the analysis from becoming dependent on a single numerical summary.

Finally, the difference between the responder-set HR of 0.654 and the all-randomized HR of 0.72 is a practical reminder that analysis population is part of the result. A hazard ratio cannot be interpreted correctly without knowing which participants generated it.

17. Trial Timeline

2013-05-17

Trial start

The registered SONAR trial began on 2013-05-17.

Phase 3

Randomized parallel-group study

The trial used randomized allocation, a parallel design, and quadruple masking to compare atrasentan with placebo in diabetic nephropathy.

Enrichment Period

Responder-set definition

The primary analysis population was defined by achieving at least 30% reduction in UACR during the Enrichment Period, after which participants were analyzed according to randomized assignment.

2018-03-29

Primary completion

The registered primary completion date was 2018-03-29.

18. Bottom-Line Statistical Reading

Primary Cox result

HR 0.654

95% CI 0.488–0.878   ·   P = 0.005

Primary composite renal endpoint in the ITT Responder Set, as randomized.

The principal reported analysis estimates a lower hazard of the first qualifying composite renal endpoint for atrasentan relative to placebo among the ITT Responder Set, with HR 0.654 and a two-sided 95% CI of 0.488–0.878. The corresponding stratified log-rank test reports P = 0.029.

The most important qualification is the analysis population: the primary estimate concerns participants who achieved at least 30% UACR reduction during the Enrichment Period. Secondary analyses include both responder-set endpoints and an analysis of all randomized participants, and their effect estimates should be interpreted in their respective populations rather than pooled into one overall statistic.

Statistical takeaway: SONAR demonstrates how the meaning of a clinical-trial result depends on more than the P-value. The estimand, analysis population, endpoint definition, time-to-event method, effect measure, confidence interval, and hypothesis test all need to be read together.

Explore the statistical methods behind SONAR

Use the related tutorials and calculators to build a deeper understanding of time-to-event analysis, hazard ratios, confidence intervals, and randomized clinical-trial methodology.