← Clinical Trials
Renal Cell Carcinoma Phase 3 Non-Inferiority NCT00720941

COMPARZ: Complete Statistical Analysis of Pazopanib in Renal Cell Carcinoma

An independent statistical review of the randomized phase 3 COMPARZ trial comparing pazopanib with sunitinib in the treatment of locally advanced and/or metastatic renal cell carcinoma, with particular focus on its progression-free survival endpoint and non-inferiority analysis.

COMPARZ  ·  Phase 3  ·  Randomized parallel design  ·  1110 enrolled
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical trial results are restricted to the information posted in the ClinicalTrials.gov record. The registry reports one formal statistical analysis for the primary progression-free survival endpoint.

Registry source: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. View the COMPARZ ClinicalTrials.gov record.

1. Trial at a Glance

COMPARZ was a completed randomized phase 3 trial comparing pazopanib with sunitinib for locally advanced and/or metastatic renal cell carcinoma. The registered primary endpoint was progression-free survival (PFS), a time-to-event outcome analyzed using a Cox proportional-hazards model in the intention-to-treat population.

1110
Enrolled
Phase 3
2
Arms
Parallel design
1.0466
PFS HR
Pazopanib vs sunitinib
0.8982–1.2195
95% CI
Two-sided
FeatureCOMPARZ
Trial nameCOMPARZ
PhasePhase 3
PopulationLocally advanced and/or metastatic renal cell carcinoma
DesignRandomized, parallel-group
MaskingNone
Primary purposeTreatment
Enrollment1110
InterventionsPazopanib and sunitinib
Primary endpointProgression-free Survival (PFS)
Primary endpoint typeTime-to-event
Primary statistical methodCox proportional-hazards model
Hypothesis frameworkNon-inferiority or equivalence
ClinicalTrials.govNCT00720941

2. Clinical Question

The registered trial compares pazopanib with sunitinib in participants with locally advanced and/or metastatic renal cell carcinoma. The primary statistical question is whether the progression-free survival experience with pazopanib is sufficiently close to that with sunitinib to satisfy the prespecified non-inferiority criterion.

Population

Participants with locally advanced and/or metastatic renal cell carcinoma.

Intervention

Pazopanib 800 mg.

Comparator

Sunitinib 50 mg.

Primary question

Does the hazard of progression or death with pazopanib remain within the prespecified non-inferiority boundary relative to sunitinib?

3. Trial Design

01
Randomize1110 enrolled
02
Two armsPazopanib vs sunitinib
03
FollowTime-to-event outcome
04
Assess PFSProgression or death
05
ModelCox proportional hazards
INTERVENTION

Pazopanib 800 mg

  • Pazopanib was the randomized intervention.
  • The reported serious adverse-event count was 242 affected participants among 554 at risk.
COMPARATOR

Sunitinib 50 mg

  • Sunitinib was the randomized comparator.
  • The reported serious adverse-event count was 227 affected participants among 548 at risk.
Allocation
Randomized allocation.
Design model
Parallel-group design.
Masking
None.
Primary purpose
Treatment.

Trial timing

2008-08-14

Study start

The registered study start date was August 14, 2008.

2012-05-21

Primary completion

The registered primary completion date was May 21, 2012.

4. Endpoints

The ClinicalTrials.gov record identifies one registered primary endpoint: progression-free survival. The endpoint is a time-to-event outcome, so the analysis must account for both the timing of events and participants who have not experienced an event by the end of their observed follow-up.

EndpointRegistry definition / time frameStatistical role
Progression-free Survival (PFS) From randomization until the earliest date of disease progression or date of death from any cause, assessed up to approximately 39 months Primary endpoint

How PFS is defined in the registry

PFS was defined as the interval between the date of randomization and the earliest date of progressive disease (PD), as defined by the Independent Review Committee (IRC), or death due to any cause. The IRC defined PD according to Response Evaluation Criteria in Solid Tumors (RECIST), Version 1.

Thus, the endpoint combines two possible events: disease progression and death from any cause. A participant does not need to experience both events for PFS to occur; the first qualifying event determines the PFS event time.

Endpoint interpretation: because PFS is measured from randomization, its comparison is naturally aligned with the randomized treatment assignment. Participants who have neither progressed nor died contribute follow-up information until they are censored according to the applicable analysis rules.

5. Statistical Methodology

Primary analysis population: intention-to-treat

The posted PFS analysis used the Intent-to-Treat (ITT) Population. The analysis was based on the assigned randomized treatment, not on the actual treatment received or not received. This is an important feature of the analysis because the purpose of randomization is to create comparable treatment groups at the point of assignment.

Cox proportional-hazards model

The primary PFS comparison used a Cox proportional-hazards model. The effect measure was the hazard ratio (HR), comparing pazopanib 800 mg with sunitinib 50 mg.

Primary model
HR = hazard of PFS event with pazopanib ÷ hazard of PFS event with sunitinib

The posted analysis reports an HR of 1.0466. Values above 1 correspond to a higher estimated hazard in the numerator group, while values below 1 correspond to a lower estimated hazard.

Covariate adjustment

The HR was estimated by the Cox regression model using treatment stratification factors as covariates. The model was adjusted for Karnofsky Performance Scale scores, prior nephrectomy, and baseline levels of lactate dehydrogenase, categorized as ≤1.5xULN and >1.5xULN.

This adjustment means that the reported HR is not simply an unadjusted ratio of event rates. It is a model-based estimate incorporating the specified baseline variables. Covariate adjustment can improve precision when the selected variables explain some of the variation in the endpoint, while the ITT framework retains treatment assignment as the basis for the comparison.

Stratified analysis

The analysis text identifies stratified analysis as an additional statistical concept and states that the treatment stratification factors were used as covariates in the Cox regression. This is distinct from merely comparing two crude event proportions: the time-to-event model incorporates the timing of events and the specified adjustment structure.

6. Results: Progression-Free Survival

The registry contains one formal statistical analysis, and it is for the primary endpoint of progression-free survival. The analysis compares pazopanib 800 mg with sunitinib 50 mg in the ITT population.

Hazard ratio for progression or death

1.0466

95% CI: 0.8982–1.2195   ·   Two-sided 95% CI

Non-inferiority criterion: the upper limit of the 95% CI must be <1.25.

Primary endpointPazopanib 800 mg vs Sunitinib 50 mg
OutcomeProgression-free Survival (PFS)
Analysis populationIntent-to-Treat (ITT)
Effect measureHazard Ratio
Estimate1.0466
95% CI0.8982–1.2195
CI typeTwo-sided
ModelCox proportional-hazards model
Hypothesis typeNon-inferiority or equivalence
Non-inferiority boundaryUpper 95% CI limit <1.25
P-valueNot reported in the ClinicalTrials.gov statistical analysis
Clinical Biostats interpretation

The estimated HR of 1.0466 means that the fitted Cox model estimated the instantaneous hazard of progression or death in the pazopanib group at approximately 1.0466 times that in the sunitinib group. In relative terms, the estimate is close to 1, so the point estimate itself indicates relatively similar estimated hazards.

It does not mean that 1.0466 times as many patients progressed or died, nor does it describe an absolute difference in the probability of progression-free survival. A hazard ratio is a time-to-event measure and depends on the model and observed follow-up.

The 95% confidence interval of 0.8982–1.2195 describes the uncertainty around the estimated hazard ratio under the specified statistical framework. It includes values below 1, values close to 1, and values above 1. The interval is therefore important for interpreting the precision of the treatment-effect estimate rather than relying on the point estimate alone.

The non-inferiority question is especially important here. The registry states that non-inferiority is defined as excluding a difference of greater than 25% in the hazards, with the upper limit of the 95% confidence interval required to be <1.25. The reported upper confidence limit is 1.2195, which is below that prespecified boundary. Under the criterion stated in the registry, the confidence interval therefore satisfies the stated non-inferiority rule.

The registry does not provide a formal p-value for this analysis in the ClinicalTrials.gov record. A p-value should not be reconstructed from the HR and confidence interval because the requested reporting standard is to use the posted values exactly. In addition, a p-value would not replace the non-inferiority margin as the central criterion for this analysis.

Reading the confidence interval against the non-inferiority margin

Non-inferiority logic
95% CI upper limit = 1.2195  <  1.25 non-inferiority boundary

The key comparison is between the upper confidence limit and the prespecified margin. The registry defines non-inferiority as excluding a hazard difference greater than 25%, operationalized here by requiring the upper limit of the 95% CI to be below 1.25.

This illustrates why non-inferiority trials should not be interpreted simply by asking whether an HR is statistically different from 1. The scientific question is different: can sufficiently unfavorable differences be excluded? The confidence interval supplies the information needed to compare the observed uncertainty with the prespecified clinically relevant boundary.

7. Understanding the Non-Inferiority Framework

The COMPARZ analysis is explicitly described as using a non-inferiority or equivalence hypothesis type, with a non-inferiority criterion based on the hazard ratio. The registry states that non-inferiority is defined as excluding a difference of greater than 25% in the hazards.

What the margin represents

The 1.25 boundary represents the largest unfavorable hazard ratio that the registry's stated criterion allows before non-inferiority would no longer be established.

What must be excluded

The upper confidence limit must remain below 1.25. It is not enough for the point estimate alone to be below the boundary.

Why the CI is central

The confidence interval represents uncertainty around the estimated hazard ratio and therefore determines whether sufficiently unfavorable effects can be excluded.

Why HR = 1 is different

A test against 1 asks about equality of hazards. The non-inferiority question instead asks whether an unacceptably large disadvantage can be ruled out.

For a superiority trial, an effect estimate near 1 can be difficult to distinguish from a meaningful treatment difference when the confidence interval is wide. In a non-inferiority trial, the same confidence interval can be acceptable if its unfavorable extreme remains inside the prespecified margin. That is the essential statistical distinction illustrated by this analysis.

Non-inferiority caution: satisfying a non-inferiority margin does not mean that the two treatments are identical in every respect. It means that, under the prespecified endpoint, analysis population, model, and confidence-interval rule, the data exclude a difference worse than the specified margin. Equivalence is a separate and generally stronger claim unless the protocol specifies an equivalence framework with corresponding bounds.

8. What the Hazard Ratio Does — and Does Not — Mean

Statistical interpretation

An HR of 1.0466 means that the fitted model estimates the instantaneous hazard of progression or death in the pazopanib group relative to the sunitinib group at 1.0466. Because the estimate is close to 1, the estimated relative hazards are close in the model's treatment comparison.

It does not mean that 4.66% more patients experienced progression or death. The HR is not an absolute risk difference, a risk ratio at a fixed time, or a difference in median PFS.

Why the confidence interval matters

The 95% CI of 0.8982–1.2195 shows that the point estimate should not be interpreted in isolation. The interval permits a range of plausible relative hazard values under the model. For the non-inferiority question, its upper endpoint is particularly important because the registry's margin is expressed on the hazard-ratio scale.

Why the p-value is not the main non-inferiority criterion

The registry analysis does not report a p-value. More importantly, the stated non-inferiority rule is based on whether the upper limit of the confidence interval is below 1.25. A p-value for testing HR = 1 would address a different hypothesis and would not by itself establish non-inferiority.

The proportional-hazards assumption

The Cox model is built around a proportional-hazards framework. Conceptually, the model treats the relative hazard between treatment groups as having a stable multiplicative structure over the analyzed time scale. If hazards vary substantially in relative terms over time, a single HR can become an incomplete description of the treatment contrast.

The ClinicalTrials.gov record identifies the Cox proportional-hazards model but do not provide a diagnostic assessment of the proportional-hazards assumption. Therefore, the HR should be understood as the model-based summary reported by the registry rather than as proof that the hazard ratio was constant at every point in time.

9. Intention-to-Treat Analysis

The primary PFS analysis was conducted in the Intent-to-Treat (ITT) Population. The registry explicitly states that the analysis was based on the assigned randomized treatment, not on the actual treatment received or not received.

PrincipleApplication in COMPARZ
Assignment basisParticipants were analyzed according to assigned randomized treatment.
Primary endpointProgression-free Survival.
Analysis populationIntent-to-Treat population.
ModelCox proportional-hazards model.
Effect measureHazard ratio.

The ITT principle matters because changing the analysis population according to treatment actually received can compromise the balance created by randomization. In a randomized trial, treatment assignment is the foundation of the causal comparison. An ITT analysis maintains that foundation even when actual exposure differs from assignment.

10. Covariate Adjustment and Stratified Analysis

The primary Cox model was adjusted for the treatment stratification factors identified in the analysis text. Specifically, the HR was adjusted for Karnofsky Performance Scale scores, prior nephrectomy, and baseline lactate dehydrogenase with categories of ≤1.5xULN and >1.5xULN.

Adjustment variableRegistered analysis detail
Karnofsky Performance ScaleIncluded as a model adjustment variable.
Prior nephrectomyIncluded as a model adjustment variable.
Baseline lactate dehydrogenase≤1.5xULN vs >1.5xULN.

Why adjust for baseline factors?

Covariate adjustment serves a different purpose from randomization. Randomization creates the treatment comparison; adjustment can account for prespecified baseline characteristics within the statistical model. The resulting HR is therefore a conditional, model-based estimate rather than a simple descriptive comparison.

An important practical point is that adjustment does not turn a randomized trial into an observational study. Treatment remains assigned randomly. Instead, the Cox model uses additional information to estimate the treatment effect more efficiently under the specified model.

11. Statistical Methods Explained

Why was a Cox proportional-hazards model used?

PFS is a time-to-event endpoint. Participants can experience progression or death at different times, while others may remain event-free at the end of their observed follow-up. A Cox model is designed to compare event hazards while incorporating the timing of events and censoring rather than reducing the outcome to a simple yes/no proportion.

What does an HR of 1.0466 mean?

It means that the fitted model estimated the instantaneous hazard of progression or death with pazopanib at 1.0466 times the hazard with sunitinib. An HR above 1 favors the comparator in terms of the direction of the event hazard, but the point estimate alone is not sufficient for interpreting the non-inferiority question.

Why is 1.25 more important than 1 for the non-inferiority question?

The value 1 represents equal hazards. The non-inferiority margin is different: it represents the largest unfavorable difference that the trial is designed to rule out. Here the registry states that the upper 95% confidence limit must be below 1.25. Therefore, the relevant question is whether the confidence interval extends beyond 1.25, not simply whether it crosses 1.

Why does the confidence interval determine the non-inferiority conclusion?

The point estimate is only one estimate of the treatment effect. The confidence interval describes uncertainty around it. A non-inferiority analysis must establish that the data are sufficiently precise to exclude an unacceptably large disadvantage. In this record, the upper confidence limit is 1.2195, which is below the 1.25 boundary specified in the registry.

Why was ITT analysis used?

Analyzing participants according to randomized assignment preserves the comparison created by randomization. The registry explicitly states that the PFS analysis was based on assigned treatment rather than actual treatment received or not received. This makes the treatment assignment the stable reference point for the primary efficacy comparison.

What does covariate adjustment add to the Cox model?

The model incorporates the specified baseline variables—Karnofsky Performance Scale scores, prior nephrectomy, and baseline lactate dehydrogenase category—when estimating the treatment hazard ratio. This means the reported 1.0466 is an adjusted model-based estimate rather than an unadjusted comparison of crude event frequencies.

Why is censoring important in PFS?

Not every participant will necessarily experience progression or death during the period in which their outcome is observed. Time-to-event methods allow such participants to contribute information up to their censoring time. This is one reason a Kaplan-Meier or Cox framework is preferable to simply calculating the percentage of participants who experienced an event.

12. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants among those at risk. These figures should be reported separately from the PFS analysis because safety and efficacy answer different questions and may use different analysis populations.

Safety measurePazopanib 800 mgSunitinib 50 mg
Serious adverse events242/554227/548
Numerator242 affected participants227 affected participants
Denominator554 at risk548 at risk
Serious adverse events: affected participants
Pazopanib 800 mg
242
Sunitinib 50 mg
227

The registry supplies the affected and at-risk counts but does not provide a formal statistical comparison of serious adverse events in the ClinicalTrials.gov record. Accordingly, these figures should be treated as descriptive safety information rather than as evidence of a statistically tested difference between treatment groups.

Safety interpretation caution: the denominators for the reported serious-adverse-event data are 554 and 548, whereas the overall trial enrollment is 1110. The ClinicalTrials.gov record does not state that these denominators are the same analysis population used for the primary PFS model. The safety counts should therefore not be silently substituted for the ITT efficacy population.

13. Results Reporting: What Is and Is Not Available

The ClinicalTrials.gov record reports 19 outcome measures and 1 statistical analysis. The single posted formal statistical analysis is for the primary PFS endpoint.

Evidence categoryWhat the ClinicalTrials.gov record supports
Primary endpointFormal statistical analysis reported for PFS.
Effect estimateHR 1.0466.
Confidence interval95% two-sided CI 0.8982–1.2195.
Formal modelCox proportional-hazards model.
Analysis populationIntent-to-Treat population.
Non-inferiority criterionUpper 95% CI limit must be <1.25.
Formal p-valueNot reported in the registry-reported statistical analysis.
Secondary endpoint estimatesNot included in the ClinicalTrials.gov record.
Median PFSNot included in the ClinicalTrials.gov record.
Subgroup estimatesNot included in the ClinicalTrials.gov record.
Kaplan-Meier estimatesNot included in the ClinicalTrials.gov record.

This distinction is important for reproducibility. A complete trial record may contain many outcome measures, but the presence of an outcome measure does not mean that a formal comparative statistical estimate has been posted for it. The analysis presented here therefore does not manufacture additional efficacy estimates from information outside the ClinicalTrials.gov record.

14. Censoring and Time-to-Event Interpretation

PFS is not simply a binary endpoint. Its definition uses the time from randomization to the first qualifying event—progression or death. Participants without either event at the relevant end of observation are handled through censoring within the time-to-event framework.

Conceptual survival framework
PFS time = min(time to PD, time to death)

The endpoint therefore records the earliest qualifying event after randomization. A participant who has neither event during observed follow-up contributes information up to the censoring time.

This structure explains why the Cox model is appropriate for the endpoint. Two participants can both be classified as progression-free at a particular administrative cutoff while having contributed very different amounts of follow-up. Time-to-event analysis preserves that information rather than treating them as identical observations.

The ClinicalTrials.gov record does not describe a separate missing-data or imputation strategy. PFS is instead defined through its event and censoring framework. No additional imputation method should be attributed to the trial without supporting registry data.

15. Statistical Interpretation of the Primary Result

The central numerical result is the HR of 1.0466 with a two-sided 95% CI of 0.8982–1.2195. The estimate is close to 1, while the confidence interval extends on both sides of 1. The non-inferiority analysis, however, is governed by a different reference value: 1.25.

Point estimate

1.0466 is the fitted estimate of the relative hazard of progression or death for pazopanib versus sunitinib.

Precision

The 95% CI of 0.8982–1.2195 describes uncertainty around the estimated relative hazard.

Equality reference

HR = 1 represents equal hazards, but equality is not the same hypothesis as non-inferiority.

Non-inferiority reference

The upper confidence limit of 1.2195 is below the prespecified 1.25 boundary.

In practical statistical terms, the result demonstrates the importance of distinguishing effect estimation from hypothesis definition. The HR provides the estimated treatment contrast; the confidence interval describes uncertainty; and the non-inferiority margin defines what degree of unfavorable difference the trial was designed to exclude.

It is also important not to convert the HR into an absolute clinical statement. Without additional reported survival probabilities, event counts for PFS, or median PFS values, the ClinicalTrials.gov record does not support an absolute estimate of how many additional participants remained progression-free at a particular time point.

16. What the Statistical Analysis Can Support

QuestionSupported interpretation
What was the estimated relative PFS hazard?HR 1.0466 for pazopanib 800 mg versus sunitinib 50 mg.
How precise was the estimate?Two-sided 95% CI 0.8982–1.2195.
Was the stated non-inferiority boundary exceeded?No. The upper CI limit of 1.2195 is below 1.25.
What analysis population was used?Intent-to-Treat population.
What model was used?Cox proportional-hazards model.
Was the model adjusted?Yes. Adjustment included Karnofsky Performance Scale scores, prior nephrectomy, and baseline lactate dehydrogenase category.
Was a formal p-value reported?Not in the registry-reported statistical analysis.
Can an absolute PFS difference be calculated from the ClinicalTrials.gov record?No. The ClinicalTrials.gov record does not include the necessary time-specific survival estimates or median PFS values.

17. Limitations

18. Why This Trial Matters Statistically

COMPARZ is a useful teaching example because its primary result illustrates a statistical situation that is frequently misunderstood: a trial can be designed around non-inferiority rather than superiority, and the key inferential comparison is therefore between a confidence interval and a prespecified margin rather than simply between an estimate and the null value of 1.

ConceptHow it appears in COMPARZ
RandomizationThe study used randomized allocation in a parallel-group phase 3 design.
Intention-to-treat analysisThe primary PFS analysis was based on assigned randomized treatment.
Time-to-event endpointPFS was defined from randomization to progression or death.
Cox proportional-hazards modelThe primary PFS HR was estimated with a Cox model.
Hazard ratioThe treatment effect was summarized as HR 1.0466.
Confidence intervalThe two-sided 95% CI was 0.8982–1.2195.
Covariate adjustmentThe model adjusted for Karnofsky Performance Scale, prior nephrectomy, and baseline LDH category.
Stratified analysisStratification factors were identified as part of the analysis framework.
Non-inferiority designThe upper 95% CI limit was required to be below 1.25.
Safety analysisSerious adverse events were reported descriptively by treatment arm.

The trial therefore provides a compact example of how a modern time-to-event analysis moves from a clinical endpoint definition to a statistical model, from the model to an effect estimate, and finally from the confidence interval to a decision rule specified before looking at the observed result.

19. A Step-by-Step Reading of the Primary Analysis

01
Define PFSProgression or death
02
Preserve ITTAssigned treatment
03
Fit Cox modelAdjusted analysis
04
Estimate HR1.0466
05
Check margin1.2195 < 1.25

Step 1 — Define the endpoint. PFS begins at randomization and ends at the earliest qualifying progression or death event.

Step 2 — Preserve treatment assignment. The analysis uses the ITT population, with participants analyzed according to randomized treatment assignment.

Step 3 — Estimate the treatment effect. A Cox proportional-hazards model is used, incorporating the specified adjustment variables.

Step 4 — Quantify the effect. The fitted model produces an HR of 1.0466 with a two-sided 95% CI of 0.8982–1.2195.

Step 5 — Apply the non-inferiority rule. The registry specifies that the upper 95% CI must be below 1.25. The reported upper limit is 1.2195, which is below that threshold.

This sequence is more informative than simply labeling the result "positive" or "negative." It shows exactly how the design hypothesis, endpoint, analysis population, statistical model, effect measure, confidence interval, and margin fit together.

20. Related Tutorials

Learn more about the statistical methods used in this trial:

21. Related Calculators

22. Sources

The numerical statistical results presented on this page are restricted to the ClinicalTrials.gov record. The linked PubMed records are provided as trial-related source references; no additional numerical results from those publications are incorporated into the analysis above.

Continue through the Clinical Biostats statistical pathway

Use the methods illustrated by COMPARZ to explore survival analysis, confidence intervals, non-inferiority design, Cox regression, and clinical-trial analysis in greater depth.

23. Record Summary

COMPARZ provides a clear example of a randomized phase 3 non-inferiority analysis using a time-to-event endpoint. The primary endpoint was progression-free survival, defined from randomization to the earliest date of disease progression or death from any cause. The analysis was performed in the ITT population using a Cox proportional-hazards model with adjustment for Karnofsky Performance Scale scores, prior nephrectomy, and baseline lactate dehydrogenase category.

The reported HR of 1.0466 had a two-sided 95% confidence interval of 0.8982–1.2195. The registry defines non-inferiority as excluding a difference of greater than 25% in the hazards, with the upper limit of the 95% confidence interval required to be <1.25. Because the reported upper confidence limit is 1.2195, the stated non-inferiority criterion is satisfied under the registry's specified rule.

The statistical lesson is broader than the numerical result. In a non-inferiority trial, the interpretation depends on the relationship between the entire confidence interval and the prespecified margin. The HR communicates the estimated relative hazard, the confidence interval communicates uncertainty, and the margin defines the magnitude of unfavorable effect that the study is designed to exclude.

Clinical Biostats methodology: A trial-results page should distinguish the reported evidence from statistical interpretation. For COMPARZ, that means preserving the registry's PFS definition, ITT population, adjusted Cox model, hazard ratio, confidence interval, and non-inferiority margin without introducing unreported medians, event counts, p-values, subgroup estimates, or other results.