← Clinical Trials
Kidney Neoplasms Phase 3 Disease-Free Survival NCT00375674

S-TRAC: Complete Statistical Analysis of Sunitinib in High-Risk Recurrent Renal Cell Cancer

An independent statistical analysis of the randomized phase 3 S-TRAC trial comparing sunitinib with placebo in patients at high risk of recurrent renal cell cancer, focusing on disease-free survival, overall survival, Cox proportional-hazards modeling, and the interpretation of time-to-event results.

Trial status: COMPLETED  ·  Enrollment: 674  ·  Sponsor: Pfizer
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

S-TRAC was a randomized, parallel-group, triple-masked phase 3 clinical trial comparing sunitinib with placebo in patients at high risk of recurrent renal cell cancer. The registry reports 674 enrolled participants, two treatment arms, two registered primary endpoints, and formal statistical analyses based on Cox proportional-hazards models and the log-rank test.

674
Enrollment
2 treatment arms
3
Phase
Randomized trial
0.761
BICR DFS HR
95% CI 0.594–0.975
0.929
OS HR
95% CI 0.670–1.289
FeatureS-TRAC
Trial nameS-TRAC
PhasePhase 3
ConditionKidney Neoplasms
DesignRandomized, parallel-group, triple-masked
AllocationRandomized
Primary purposeTreatment
Enrollment674
InterventionsSunitinib malate and placebo
Primary endpoints2 registered endpoints
Primary endpoint typesBinary; time-to-event
Results postedYes
Statistical analyses posted3
ClinicalTrials.govNCT00375674

2. Clinical Question

The central statistical question was whether assignment to sunitinib, compared with placebo, was associated with a different disease-free survival experience in participants at high risk of recurrent renal cell cancer. The registry contains two formal primary endpoint analyses: disease-free survival assessed by blinded independent central review and disease-free survival assessed by the investigator, with the latter explicitly described as stratified by the University of California Los Angeles Integrated Staging System (UISS) High Risk Group in the intent-to-treat population.

Population

Participants with kidney neoplasms who were enrolled in the phase 3 S-TRAC trial and randomized to one of two treatment arms.

Intervention

Sunitinib malate.

Comparator

Placebo.

Primary question

Does sunitinib produce a different disease-free survival outcome from placebo under the prespecified superiority framework?

3. Trial Design

01
Enroll674 participants
02
Randomize2 treatment arms
03
MaskTriple masking
04
Assess DFSCentral review / investigator
05
AnalyzeCox model / log-rank
Allocation
Randomized allocation in a parallel-group design.
Masking
Triple masking.
Primary purpose
Treatment.
Hypothesis type
Superiority.
ARM 1

Sunitinib

  • Sunitinib malate
  • Randomized treatment assignment
ARM 2

Placebo

  • Placebo
  • Randomized comparator assignment

The registry records a start date of August 1, 2007 and a primary completion date of April 7, 2016. The trial is listed as completed. The sponsor is Pfizer, classified in the registry as industry.

Design principle: randomization creates the basis for comparing outcomes according to assigned treatment rather than according to treatment received. The primary efficacy analyses reported here use the intent-to-treat population, preserving that randomized comparison.

4. Endpoints

Registered primary endpointDefinitionTime frame
Disease-free Survival (DFS) — Assessed by Blinded Independent Central Review DFS was defined as the time interval, in years, from the date of randomization to the first date of recurrence or occurrence of a secondary malignancy or death. Recurrence refers to relapse of the primary tumor in-situ or at metastatic sites. Every 12 weeks during the first 3 years and every 6 months after that unless the participant had withdrawn consent.
DFS — Assessed by the Investigator, stratified by UISS High Risk Group — Intent to Treat Population DFS was defined as the time interval, in years, from the date of randomization to the first date of recurrence or occurrence of a secondary malignancy or death. Recurrence refers to relapse of the primary tumor in-situ or at metastatic sites. Every 12 weeks during the first 3 years and every 6 months after that unless the participant had withdrawn consent.

The registry therefore defines DFS as a time-to-event endpoint. Participants who have not experienced the defined event at the time of their last informative assessment can contribute follow-up without necessarily contributing an observed event. That distinction is fundamental to why survival-analysis methods are used rather than ordinary comparisons of proportions.

5. Analysis Populations and Statistical Structure

The formal primary analyses use the intent-to-treat population. The registry defines this population as including all participants who were randomized, with study-drug assignment designated according to initial randomization, regardless of whether participants received study drug or received treatment thereafter.

Analysis featureRegistry-supported description
Primary efficacy populationIntent-to-treat population
Group comparisonSunitinib vs placebo
Primary modelCox proportional-hazards model
StratificationUISS High Risk Group
Primary effect measureHazard ratio
Hypothesis typeSuperiority

This structure is important because the treatment comparison is not simply a comparison of the percentage of participants who eventually experienced recurrence. The timing of recurrence, secondary malignancy, or death is incorporated into the analysis.

6. Statistical Methodology

Kaplan-Meier estimation

A time-to-event endpoint such as DFS is commonly summarized using the Kaplan-Meier estimator. It estimates the probability of remaining event-free over time while allowing participants to be censored when their follow-up ends without an observed event.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at time ti, while ni represents the number at risk immediately before that time.

The registry's formal analyses do not list Kaplan-Meier estimation as a method; they identify Cox proportional-hazards modeling for the primary DFS analyses and a log-rank test for the secondary overall-survival analysis. Kaplan-Meier estimation is nevertheless the natural descriptive framework for displaying a time-to-event endpoint such as DFS; it should not be confused with the inferential model used to obtain the reported hazard ratios.

Cox proportional-hazards model

The two primary DFS analyses use a Cox proportional-hazards model. The reported model is stratified by the UISS High Risk Group. The model estimates the relative hazard associated with the randomized treatment comparison while accounting for the specified stratification factor.

Hazard-ratio interpretation
HR = hazard in sunitinib group ÷ hazard in placebo group

An HR below 1 indicates a lower estimated instantaneous event rate in the sunitinib group under the fitted model. It does not directly equal a percentage reduction in the probability of an event by a fixed date.

Log-rank test

The registry reports a log-rank test for the secondary overall-survival analysis. The log-rank test is designed for comparing survival distributions between randomized groups while accounting for the timing of observed events and censoring.

Stratified analysis

Both primary DFS analyses are described as stratified analyses. The BICR DFS analysis specifically states that the Cox proportional-hazards model was stratified by the UISS High Risk Group. Stratification allows the comparison to account for a prespecified categorical factor without requiring that factor to have one common regression coefficient in the ordinary unstratified sense.

Intention-to-treat analysis

The ITT population includes all randomized participants and preserves their initial treatment assignment. This matters because post-randomization treatment changes do not redefine the original randomized comparison. For efficacy, the resulting estimate therefore answers a question about the effect associated with treatment assignment rather than a narrower question about adherence among participants who remained on assigned treatment.

7. Primary Results: Disease-Free Survival by Blinded Independent Central Review

The first formal primary analysis evaluated disease-free survival as assessed by blinded independent central review. The registry reports a superiority analysis using a Cox proportional-hazards model stratified by UISS High Risk Group in the ITT population.

DFS hazard ratio

0.761

95% CI: 0.594–0.975   ·   P = 0.030

Two-sided 95% confidence interval · Sunitinib vs placebo

Primary endpointSunitinib vs placebo95% CIP-valueModel
Disease-free Survival — Blinded Independent Central Review HR 0.761 0.594–0.975 0.030 Cox proportional-hazards model, stratified by UISS High Risk Group
Clinical Biostats interpretation

The reported HR of 0.761 means that, under the fitted Cox model, the estimated instantaneous rate of the DFS event in the sunitinib group was about 76.1% of that in the placebo group. Expressed as a simple relative-hazard interpretation, this corresponds to an estimated 23.9% lower hazard for the sunitinib group relative to placebo.

The HR does not mean that 23.9% fewer participants experienced recurrence, secondary malignancy, or death. It also does not mean that every participant had exactly a 23.9% reduction in individual risk. A hazard ratio is a model-based relative measure of event rates over follow-up.

The two-sided 95% confidence interval of 0.594–0.975 describes the statistical uncertainty around the estimated hazard ratio under the model and sampling framework. The interval is relatively broad compared with the point estimate, so the point estimate should not be treated as exact.

The P-value of 0.030 measures the compatibility of the observed data with the specified null hypothesis under the analysis framework. It is not a measure of effect size, clinical importance, or the probability that the treatment works. The magnitude of the effect is better conveyed by the hazard ratio and its confidence interval.

Because the analysis uses a Cox proportional-hazards model, interpretation also depends on the proportional-hazards framework. A single HR summarizes the relative hazard under that model; it is not itself a complete description of how treatment effects may vary over time.

Reading the confidence interval

The lower confidence-limit value of 0.594 and upper confidence-limit value of 0.975 provide a range of model-compatible hazard-ratio values under the stated 95% confidence framework. The interval remains below 1, while approaching 1 at its upper boundary. That combination is reflected in the reported two-sided P-value of 0.030.

Important distinction: the registry reports a hazard ratio and P-value, not an absolute DFS probability at a particular time point. Therefore, the result should not be converted into an absolute risk reduction or number needed to treat from the ClinicalTrials.gov record.

8. Primary Results: Investigator-Assessed Disease-Free Survival

The second formal primary analysis evaluated DFS according to investigator assessment and explicitly identifies the UISS High Risk Group as the stratification factor. The analysis was performed in the ITT population using the Cox proportional-hazards model.

Investigator-assessed DFS hazard ratio

0.811

95% CI: 0.643–1.023   ·   P = 0.077

Two-sided 95% confidence interval · Sunitinib vs placebo

Primary endpointSunitinib vs placebo95% CIP-valueModel
DFS — Investigator assessment, stratified by UISS High Risk Group HR 0.811 0.643–1.023 0.077 Cox proportional-hazards model, stratified by UISS High Risk Group
Clinical Biostats interpretation

The HR of 0.811 corresponds to an estimated hazard approximately 81.1% as large in the sunitinib group as in the placebo group under the fitted model, or approximately a 18.9% lower estimated hazard.

Again, this is not the same as saying that 18.9% fewer participants had an event. The hazard ratio describes a relative event-rate measure under the Cox model, not an absolute probability difference.

The 95% confidence interval is 0.643–1.023. Unlike the BICR analysis, this interval extends slightly above 1. That means the confidence interval includes values corresponding to a higher estimated hazard as well as values corresponding to a lower estimated hazard for sunitinib.

The reported P-value of 0.077 is a hypothesis-testing quantity, not an effect-size measure. It should not be interpreted as a 7.7% probability that the treatment effect is due to chance, nor as a direct probability that the null hypothesis is true.

The difference between the two primary DFS estimates is statistically informative: the BICR analysis gives HR 0.761, whereas the investigator-assessed analysis gives HR 0.811. They are related analyses of the same randomized comparison but use different assessment sources. The ClinicalTrials.gov record does not provide enough information to attribute the numerical difference to any particular source of disagreement.

9. Comparing the Two Primary DFS Analyses

The two primary analyses address the same broad endpoint but differ in how recurrence and other DFS events were assessed. The first uses blinded independent central review; the second uses investigator assessment and is explicitly identified as stratified by the UISS High Risk Group.

FeatureBICR DFSInvestigator DFS
Endpoint rolePrimaryPrimary
AssessmentBlinded Independent Central ReviewInvestigator
PopulationITTITT
ComparisonSunitinib vs placeboSunitinib vs placebo
ModelCox proportional-hazards modelCox proportional-hazards model
StratificationUISS High Risk GroupUISS High Risk Group
HR0.7610.811
95% CI0.594–0.9750.643–1.023
P-value0.0300.077

Both estimates are below 1, so both point estimates favor a lower estimated DFS-event hazard in the sunitinib group. The estimates should nevertheless be interpreted separately rather than averaged or combined. They arise from distinct endpoint-assessment approaches, and the registry does not provide the information required to construct a pooled estimate.

Do not treat the P-values as competing measures of effect size. The BICR analysis has P = 0.030 and the investigator analysis has P = 0.077, but the corresponding effect estimates and confidence intervals provide the more informative description of magnitude and precision. A P-value alone cannot tell how large or clinically important an effect is.

10. Secondary Result: Overall Survival

The registry also reports a formal statistical analysis for overall survival, stratified by UISS High Risk Group in the ITT population. The analysis method is recorded as the log-rank test, with the reported hazard ratio based on a Cox proportional-hazards model stratified by UISS High Risk Group.

Overall survival hazard ratio

0.929

95% CI: 0.670–1.289   ·   P = 0.661

Two-sided 95% confidence interval · Sunitinib vs placebo

EndpointSunitinib vs placebo95% CIP-valueReported method
Overall Survival — stratified by UISS High Risk Group, ITT HR 0.929 0.670–1.289 0.661 Log-rank test; HR based on stratified Cox proportional-hazards model
Clinical Biostats interpretation

An OS HR of 0.929 means that the estimated instantaneous rate of death in the sunitinib group was 92.9% of the corresponding rate in the placebo group under the reported Cox model. As a simple relative-hazard transformation, this corresponds to an estimated 7.1% lower hazard for sunitinib.

The confidence interval of 0.670–1.289 is substantially wider than the point estimate alone suggests. It spans values below and above 1, so the data are compatible with a range of relative hazard values under the stated confidence framework.

The P-value of 0.661 is not an estimate of the probability that the treatment has no effect. It is a measure of the data's compatibility with the specified null hypothesis under the test framework. It also does not establish that the two treatments have identical survival outcomes.

The result should therefore be described using the estimate and confidence interval rather than reducing it to a statement based solely on whether the P-value crosses a conventional threshold.

11. Primary and Secondary Results in One View

EndpointRoleHR95% CIP-valueAnalysis
DFS — BICRPrimary0.7610.594–0.9750.030Stratified Cox model
DFS — InvestigatorPrimary0.8110.643–1.0230.077Stratified Cox model
Overall SurvivalSecondary0.9290.670–1.2890.661Log-rank; HR from stratified Cox model

This table illustrates why clinical-trial interpretation should preserve the full statistical record. The three hazard ratios are all different, their confidence intervals have different widths, and their P-values answer different inferential questions. A single summary statement would discard information about endpoint definition, assessment method, and statistical precision.

12. Serious Adverse Events

The registry reports serious adverse events by treatment arm as affected participants divided by participants at risk.

ArmAffected / at riskSimple observed fraction
Sunitinib67 / 30667/306
Placebo52 / 30452/304

The ClinicalTrials.gov record provides the serious-adverse-event counts and denominators but do not provide a formal statistical comparison for this safety measure. Accordingly, the figures should be read descriptively rather than converted into a treatment-effect estimate or hypothesis-test conclusion.

Safety interpretation: efficacy and safety answer different statistical questions. The primary efficacy analyses use the ITT population and time-to-event methodology, whereas the registry-reported serious-adverse-event information is presented as affected participants over those at risk. These quantities should not be combined into a single numerical benefit-risk measure.

13. Statistical Methods Explained

Why was a Cox proportional-hazards model used?

DFS and OS are time-to-event outcomes. A Cox model can use both the occurrence of an event and the amount of follow-up contributed by participants who have not yet experienced the event. It also produces a hazard ratio that summarizes the relative event rate between randomized groups under the model.

What does an HR of 0.761 mean?

An HR of 0.761 means that the estimated instantaneous event rate in the sunitinib group was 0.761 times that of the placebo group under the fitted Cox model. Equivalently, the estimated hazard was 23.9% lower because 1 − 0.761 = 0.239. That calculation describes the relative hazard; it does not provide an absolute probability difference.

Why is the confidence interval important?

The point estimate is only one estimate from the observed data. The 95% confidence interval communicates statistical precision. For the BICR DFS analysis, the interval is 0.594–0.975; for investigator-assessed DFS, it is 0.643–1.023. The widths and locations of these intervals provide information that the P-values alone cannot provide.

Why does the P-value not measure effect size?

A P-value is tied to a hypothesis-testing framework. It reflects how compatible the observed data are with the specified null hypothesis under that framework. It is influenced by both the estimated effect and the amount of information in the analysis. Effect magnitude is better described using the hazard ratio together with its confidence interval and, when available, absolute event probabilities.

Why was the analysis stratified by UISS High Risk Group?

The registry explicitly states that the primary Cox analyses were stratified by the UISS High Risk Group. Stratification allows the model to account for this factor when comparing treatment groups. It does not mean that the UISS factor itself is assigned a single ordinary regression coefficient in the same way as an unstratified covariate.

What does intention-to-treat add to the analysis?

ITT keeps participants in the treatment group to which they were originally randomized. This maintains the randomized comparison and avoids redefining the efficacy population based on subsequent treatment exposure. The registry specifically states that the ITT population includes all randomized participants regardless of whether they received study drug or received treatment thereafter.

Why should DFS and OS not be interpreted as the same endpoint?

DFS and OS capture different clinical events. The registered DFS definition includes recurrence, occurrence of a secondary malignancy, or death, beginning at randomization. OS instead concerns survival. A treatment effect on DFS therefore cannot automatically be translated into an equivalent treatment effect on OS.

14. Interpreting Hazard Ratios Without Overstating Them

Example: BICR DFS

The HR of 0.761 is a relative measure of the event hazard under the Cox model. The corresponding 23.9% reduction is a transformation of that hazard ratio, not an absolute 23.9-percentage-point improvement in disease-free survival.

Example: Investigator DFS

The HR of 0.811 corresponds to an estimated 18.9% lower hazard under the same type of model. Its confidence interval, 0.643–1.023, indicates more uncertainty than the point estimate alone reveals.

Example: Overall survival

The OS HR of 0.929 corresponds to an estimated 7.1% lower hazard under the reported Cox model. The 95% CI of 0.670–1.289 is broad enough to include hazard ratios above 1, so the point estimate should not be treated as a precise estimate of the underlying treatment effect.

These examples illustrate why hazard ratios should be reported with their confidence intervals. The HR supplies a compact measure of relative effect, while the confidence interval shows how much statistical uncertainty surrounds that estimate.

15. What the Results Do — and Do Not — Establish

What is directly reported

The registry reports two primary DFS analyses, both using stratified Cox proportional-hazards models in the ITT population, with HRs of 0.761 and 0.811 and their associated confidence intervals and P-values.

What the point estimates mean

Each HR below 1 corresponds to a lower estimated instantaneous event rate for sunitinib under the fitted model.

What the confidence intervals add

They quantify uncertainty around each HR. The investigator-assessed DFS interval extends above 1, while the BICR DFS interval remains below 1.

What cannot be inferred

The ClinicalTrials.gov record does not support an absolute DFS difference, median DFS, subgroup-specific treatment effects, or an individual patient's probability of benefit.

16. Assessment of the Time-to-Event Framework

The S-TRAC primary endpoints are well suited to time-to-event analysis because the registered DFS definition explicitly incorporates the interval from randomization until the first occurrence of a qualifying event. This structure preserves information about when events occur rather than reducing every participant to a simple yes/no status at one arbitrary time.

There are three statistical layers to keep distinct:

  1. Endpoint definition: DFS specifies which events count and establishes randomization as time zero.
  2. Analysis population: the ITT population preserves the randomized treatment assignment.
  3. Effect model: the stratified Cox model estimates the relative hazard associated with treatment assignment.

The registry's reported analysis therefore represents more than a simple comparison of recurrence percentages. It is a model-based analysis of event timing using all available follow-up consistent with the endpoint and censoring framework.

17. Limitations

18. Why This Trial Matters Statistically

S-TRAC is a useful teaching example because the registry data bring together several central concepts in clinical-trial biostatistics without requiring a complicated endpoint structure. The primary question is time-to-event based, the efficacy population is ITT, the primary model is a stratified Cox proportional-hazards model, and the effect measure is a hazard ratio.

ConceptHow it appears in S-TRAC
RandomizationParticipants were randomized to sunitinib or placebo.
Parallel-group designThe trial used a parallel design with 2 arms.
Triple maskingThe registry identifies the trial as triple-masked.
Intention-to-treat analysisPrimary analyses included all randomized participants according to initial assignment.
Time-to-event endpointDFS was defined from randomization to recurrence, secondary malignancy, or death.
Hazard ratioDFS and OS treatment effects were reported as hazard ratios.
Cox proportional-hazards modelUsed for both formal primary DFS analyses and to derive the OS HR.
Stratified analysisPrimary Cox analyses were stratified by UISS High Risk Group.
Log-rank testReported method for the secondary OS analysis.
Confidence intervals95% two-sided intervals accompany all three reported HR estimates.
Superiority testingThe registered hypothesis type is superiority.

19. Statistical Interpretation of the Three Reported Hazard Ratios

Putting the three formal estimates side by side shows an important lesson in clinical-trial interpretation: an effect estimate is never complete without its endpoint definition, analysis population, statistical model, confidence interval, and assessment context.

Reported hazard ratios
BICR DFS
0.761
Investigator DFS
0.811
Overall survival
0.929

The visual comparison is descriptive only. It does not imply that one endpoint should be prioritized over another or that the numerical differences between HRs constitute a formal test of heterogeneity. Different endpoints measure different events and can naturally produce different effect estimates.

20. Clinical Biostats Interpretation Framework

Start with the endpoint

Before reading the HR, identify exactly what constitutes an event. For S-TRAC DFS, the registry includes recurrence, secondary malignancy, or death.

Then identify the population

The primary analyses use the ITT population, maintaining the original randomized assignment.

Then identify the model

The primary DFS analyses use stratified Cox proportional-hazards models with UISS High Risk Group as the stratification factor.

Finally read HR, CI, and P-value together

The HR describes relative magnitude, the CI describes uncertainty, and the P-value addresses compatibility with the specified null hypothesis.

This sequence prevents a common statistical error: beginning with the P-value and then working backward to construct a narrative. A more informative approach begins with the clinical endpoint, identifies the population and analysis method, and then interprets the effect estimate together with its uncertainty.

21. Related Tutorials

Learn more about the methods used in this trial:

22. Related Calculators

23. Sources

Continue through Clinical Biostats

Build from this trial's time-to-event methods into tutorials and statistical calculators covering survival analysis, hazard ratios, confidence intervals, and clinical-trial methodology.

24. Record Summary

S-TRAC provides a clear example of a randomized phase 3 time-to-event analysis. The trial enrolled 674 participants and compared sunitinib with placebo using a randomized, parallel-group, triple-masked design. The two registered primary endpoints were disease-free survival assessed by blinded independent central review and investigator-assessed disease-free survival, with both primary analyses conducted in the ITT population using Cox proportional-hazards models stratified by UISS High Risk Group.

The BICR DFS analysis reported an HR of 0.761 with a two-sided 95% CI of 0.594–0.975 and P = 0.030. The investigator-assessed DFS analysis reported an HR of 0.811 with a two-sided 95% CI of 0.643–1.023 and P = 0.077. The secondary OS analysis reported an HR of 0.929 with a two-sided 95% CI of 0.670–1.289 and P = 0.661.

The statistical lesson is broader than any individual P-value. The treatment effect must be interpreted in the context of the endpoint definition, randomized analysis population, assessment method, stratification factor, Cox model, confidence interval, and hypothesis-testing framework. In particular, a hazard ratio is a relative model-based measure rather than an absolute probability difference, and its interpretation is strongest when accompanied by its confidence interval and the full definition of the event being analyzed.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. For S-TRAC, the ClinicalTrials.gov record supports a focused analysis of randomized time-to-event methodology, Cox proportional-hazards modeling, stratification, hazard ratios, confidence intervals, and the interpretation of primary and secondary endpoint results without adding unsupported clinical or numerical claims.