This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
S-TRAC was a randomized, parallel-group, triple-masked phase 3 clinical trial comparing sunitinib with placebo in patients at high risk of recurrent renal cell cancer. The registry reports 674 enrolled participants, two treatment arms, two registered primary endpoints, and formal statistical analyses based on Cox proportional-hazards models and the log-rank test.
| Feature | S-TRAC |
|---|---|
| Trial name | S-TRAC |
| Phase | Phase 3 |
| Condition | Kidney Neoplasms |
| Design | Randomized, parallel-group, triple-masked |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 674 |
| Interventions | Sunitinib malate and placebo |
| Primary endpoints | 2 registered endpoints |
| Primary endpoint types | Binary; time-to-event |
| Results posted | Yes |
| Statistical analyses posted | 3 |
| ClinicalTrials.gov | NCT00375674 |
2. Clinical Question
The central statistical question was whether assignment to sunitinib, compared with placebo, was associated with a different disease-free survival experience in participants at high risk of recurrent renal cell cancer. The registry contains two formal primary endpoint analyses: disease-free survival assessed by blinded independent central review and disease-free survival assessed by the investigator, with the latter explicitly described as stratified by the University of California Los Angeles Integrated Staging System (UISS) High Risk Group in the intent-to-treat population.
Population
Participants with kidney neoplasms who were enrolled in the phase 3 S-TRAC trial and randomized to one of two treatment arms.
Intervention
Sunitinib malate.
Comparator
Placebo.
Primary question
Does sunitinib produce a different disease-free survival outcome from placebo under the prespecified superiority framework?
3. Trial Design
Sunitinib
- Sunitinib malate
- Randomized treatment assignment
Placebo
- Placebo
- Randomized comparator assignment
The registry records a start date of August 1, 2007 and a primary completion date of April 7, 2016. The trial is listed as completed. The sponsor is Pfizer, classified in the registry as industry.
4. Endpoints
| Registered primary endpoint | Definition | Time frame |
|---|---|---|
| Disease-free Survival (DFS) — Assessed by Blinded Independent Central Review | DFS was defined as the time interval, in years, from the date of randomization to the first date of recurrence or occurrence of a secondary malignancy or death. Recurrence refers to relapse of the primary tumor in-situ or at metastatic sites. | Every 12 weeks during the first 3 years and every 6 months after that unless the participant had withdrawn consent. |
| DFS — Assessed by the Investigator, stratified by UISS High Risk Group — Intent to Treat Population | DFS was defined as the time interval, in years, from the date of randomization to the first date of recurrence or occurrence of a secondary malignancy or death. Recurrence refers to relapse of the primary tumor in-situ or at metastatic sites. | Every 12 weeks during the first 3 years and every 6 months after that unless the participant had withdrawn consent. |
The registry therefore defines DFS as a time-to-event endpoint. Participants who have not experienced the defined event at the time of their last informative assessment can contribute follow-up without necessarily contributing an observed event. That distinction is fundamental to why survival-analysis methods are used rather than ordinary comparisons of proportions.
5. Analysis Populations and Statistical Structure
The formal primary analyses use the intent-to-treat population. The registry defines this population as including all participants who were randomized, with study-drug assignment designated according to initial randomization, regardless of whether participants received study drug or received treatment thereafter.
| Analysis feature | Registry-supported description |
|---|---|
| Primary efficacy population | Intent-to-treat population |
| Group comparison | Sunitinib vs placebo |
| Primary model | Cox proportional-hazards model |
| Stratification | UISS High Risk Group |
| Primary effect measure | Hazard ratio |
| Hypothesis type | Superiority |
This structure is important because the treatment comparison is not simply a comparison of the percentage of participants who eventually experienced recurrence. The timing of recurrence, secondary malignancy, or death is incorporated into the analysis.
6. Statistical Methodology
Kaplan-Meier estimation
A time-to-event endpoint such as DFS is commonly summarized using the Kaplan-Meier estimator. It estimates the probability of remaining event-free over time while allowing participants to be censored when their follow-up ends without an observed event.
Here, di represents the number of events at time ti, while ni represents the number at risk immediately before that time.
The registry's formal analyses do not list Kaplan-Meier estimation as a method; they identify Cox proportional-hazards modeling for the primary DFS analyses and a log-rank test for the secondary overall-survival analysis. Kaplan-Meier estimation is nevertheless the natural descriptive framework for displaying a time-to-event endpoint such as DFS; it should not be confused with the inferential model used to obtain the reported hazard ratios.
Cox proportional-hazards model
The two primary DFS analyses use a Cox proportional-hazards model. The reported model is stratified by the UISS High Risk Group. The model estimates the relative hazard associated with the randomized treatment comparison while accounting for the specified stratification factor.
An HR below 1 indicates a lower estimated instantaneous event rate in the sunitinib group under the fitted model. It does not directly equal a percentage reduction in the probability of an event by a fixed date.
Log-rank test
The registry reports a log-rank test for the secondary overall-survival analysis. The log-rank test is designed for comparing survival distributions between randomized groups while accounting for the timing of observed events and censoring.
Stratified analysis
Both primary DFS analyses are described as stratified analyses. The BICR DFS analysis specifically states that the Cox proportional-hazards model was stratified by the UISS High Risk Group. Stratification allows the comparison to account for a prespecified categorical factor without requiring that factor to have one common regression coefficient in the ordinary unstratified sense.
Intention-to-treat analysis
The ITT population includes all randomized participants and preserves their initial treatment assignment. This matters because post-randomization treatment changes do not redefine the original randomized comparison. For efficacy, the resulting estimate therefore answers a question about the effect associated with treatment assignment rather than a narrower question about adherence among participants who remained on assigned treatment.
7. Primary Results: Disease-Free Survival by Blinded Independent Central Review
The first formal primary analysis evaluated disease-free survival as assessed by blinded independent central review. The registry reports a superiority analysis using a Cox proportional-hazards model stratified by UISS High Risk Group in the ITT population.
DFS hazard ratio
95% CI: 0.594–0.975 · P = 0.030
Two-sided 95% confidence interval · Sunitinib vs placebo
| Primary endpoint | Sunitinib vs placebo | 95% CI | P-value | Model |
|---|---|---|---|---|
| Disease-free Survival — Blinded Independent Central Review | HR 0.761 | 0.594–0.975 | 0.030 | Cox proportional-hazards model, stratified by UISS High Risk Group |
The reported HR of 0.761 means that, under the fitted Cox model, the estimated instantaneous rate of the DFS event in the sunitinib group was about 76.1% of that in the placebo group. Expressed as a simple relative-hazard interpretation, this corresponds to an estimated 23.9% lower hazard for the sunitinib group relative to placebo.
The HR does not mean that 23.9% fewer participants experienced recurrence, secondary malignancy, or death. It also does not mean that every participant had exactly a 23.9% reduction in individual risk. A hazard ratio is a model-based relative measure of event rates over follow-up.
The two-sided 95% confidence interval of 0.594–0.975 describes the statistical uncertainty around the estimated hazard ratio under the model and sampling framework. The interval is relatively broad compared with the point estimate, so the point estimate should not be treated as exact.
The P-value of 0.030 measures the compatibility of the observed data with the specified null hypothesis under the analysis framework. It is not a measure of effect size, clinical importance, or the probability that the treatment works. The magnitude of the effect is better conveyed by the hazard ratio and its confidence interval.
Because the analysis uses a Cox proportional-hazards model, interpretation also depends on the proportional-hazards framework. A single HR summarizes the relative hazard under that model; it is not itself a complete description of how treatment effects may vary over time.
Reading the confidence interval
The lower confidence-limit value of 0.594 and upper confidence-limit value of 0.975 provide a range of model-compatible hazard-ratio values under the stated 95% confidence framework. The interval remains below 1, while approaching 1 at its upper boundary. That combination is reflected in the reported two-sided P-value of 0.030.
8. Primary Results: Investigator-Assessed Disease-Free Survival
The second formal primary analysis evaluated DFS according to investigator assessment and explicitly identifies the UISS High Risk Group as the stratification factor. The analysis was performed in the ITT population using the Cox proportional-hazards model.
Investigator-assessed DFS hazard ratio
95% CI: 0.643–1.023 · P = 0.077
Two-sided 95% confidence interval · Sunitinib vs placebo
| Primary endpoint | Sunitinib vs placebo | 95% CI | P-value | Model |
|---|---|---|---|---|
| DFS — Investigator assessment, stratified by UISS High Risk Group | HR 0.811 | 0.643–1.023 | 0.077 | Cox proportional-hazards model, stratified by UISS High Risk Group |
The HR of 0.811 corresponds to an estimated hazard approximately 81.1% as large in the sunitinib group as in the placebo group under the fitted model, or approximately a 18.9% lower estimated hazard.
Again, this is not the same as saying that 18.9% fewer participants had an event. The hazard ratio describes a relative event-rate measure under the Cox model, not an absolute probability difference.
The 95% confidence interval is 0.643–1.023. Unlike the BICR analysis, this interval extends slightly above 1. That means the confidence interval includes values corresponding to a higher estimated hazard as well as values corresponding to a lower estimated hazard for sunitinib.
The reported P-value of 0.077 is a hypothesis-testing quantity, not an effect-size measure. It should not be interpreted as a 7.7% probability that the treatment effect is due to chance, nor as a direct probability that the null hypothesis is true.
The difference between the two primary DFS estimates is statistically informative: the BICR analysis gives HR 0.761, whereas the investigator-assessed analysis gives HR 0.811. They are related analyses of the same randomized comparison but use different assessment sources. The ClinicalTrials.gov record does not provide enough information to attribute the numerical difference to any particular source of disagreement.
9. Comparing the Two Primary DFS Analyses
The two primary analyses address the same broad endpoint but differ in how recurrence and other DFS events were assessed. The first uses blinded independent central review; the second uses investigator assessment and is explicitly identified as stratified by the UISS High Risk Group.
| Feature | BICR DFS | Investigator DFS |
|---|---|---|
| Endpoint role | Primary | Primary |
| Assessment | Blinded Independent Central Review | Investigator |
| Population | ITT | ITT |
| Comparison | Sunitinib vs placebo | Sunitinib vs placebo |
| Model | Cox proportional-hazards model | Cox proportional-hazards model |
| Stratification | UISS High Risk Group | UISS High Risk Group |
| HR | 0.761 | 0.811 |
| 95% CI | 0.594–0.975 | 0.643–1.023 |
| P-value | 0.030 | 0.077 |
Both estimates are below 1, so both point estimates favor a lower estimated DFS-event hazard in the sunitinib group. The estimates should nevertheless be interpreted separately rather than averaged or combined. They arise from distinct endpoint-assessment approaches, and the registry does not provide the information required to construct a pooled estimate.
10. Secondary Result: Overall Survival
The registry also reports a formal statistical analysis for overall survival, stratified by UISS High Risk Group in the ITT population. The analysis method is recorded as the log-rank test, with the reported hazard ratio based on a Cox proportional-hazards model stratified by UISS High Risk Group.
Overall survival hazard ratio
95% CI: 0.670–1.289 · P = 0.661
Two-sided 95% confidence interval · Sunitinib vs placebo
| Endpoint | Sunitinib vs placebo | 95% CI | P-value | Reported method |
|---|---|---|---|---|
| Overall Survival — stratified by UISS High Risk Group, ITT | HR 0.929 | 0.670–1.289 | 0.661 | Log-rank test; HR based on stratified Cox proportional-hazards model |
An OS HR of 0.929 means that the estimated instantaneous rate of death in the sunitinib group was 92.9% of the corresponding rate in the placebo group under the reported Cox model. As a simple relative-hazard transformation, this corresponds to an estimated 7.1% lower hazard for sunitinib.
The confidence interval of 0.670–1.289 is substantially wider than the point estimate alone suggests. It spans values below and above 1, so the data are compatible with a range of relative hazard values under the stated confidence framework.
The P-value of 0.661 is not an estimate of the probability that the treatment has no effect. It is a measure of the data's compatibility with the specified null hypothesis under the test framework. It also does not establish that the two treatments have identical survival outcomes.
The result should therefore be described using the estimate and confidence interval rather than reducing it to a statement based solely on whether the P-value crosses a conventional threshold.
11. Primary and Secondary Results in One View
| Endpoint | Role | HR | 95% CI | P-value | Analysis |
|---|---|---|---|---|---|
| DFS — BICR | Primary | 0.761 | 0.594–0.975 | 0.030 | Stratified Cox model |
| DFS — Investigator | Primary | 0.811 | 0.643–1.023 | 0.077 | Stratified Cox model |
| Overall Survival | Secondary | 0.929 | 0.670–1.289 | 0.661 | Log-rank; HR from stratified Cox model |
This table illustrates why clinical-trial interpretation should preserve the full statistical record. The three hazard ratios are all different, their confidence intervals have different widths, and their P-values answer different inferential questions. A single summary statement would discard information about endpoint definition, assessment method, and statistical precision.
12. Serious Adverse Events
The registry reports serious adverse events by treatment arm as affected participants divided by participants at risk.
| Arm | Affected / at risk | Simple observed fraction |
|---|---|---|
| Sunitinib | 67 / 306 | 67/306 |
| Placebo | 52 / 304 | 52/304 |
The ClinicalTrials.gov record provides the serious-adverse-event counts and denominators but do not provide a formal statistical comparison for this safety measure. Accordingly, the figures should be read descriptively rather than converted into a treatment-effect estimate or hypothesis-test conclusion.
13. Statistical Methods Explained
Why was a Cox proportional-hazards model used?
DFS and OS are time-to-event outcomes. A Cox model can use both the occurrence of an event and the amount of follow-up contributed by participants who have not yet experienced the event. It also produces a hazard ratio that summarizes the relative event rate between randomized groups under the model.
What does an HR of 0.761 mean?
An HR of 0.761 means that the estimated instantaneous event rate in the sunitinib group was 0.761 times that of the placebo group under the fitted Cox model. Equivalently, the estimated hazard was 23.9% lower because 1 − 0.761 = 0.239. That calculation describes the relative hazard; it does not provide an absolute probability difference.
Why is the confidence interval important?
The point estimate is only one estimate from the observed data. The 95% confidence interval communicates statistical precision. For the BICR DFS analysis, the interval is 0.594–0.975; for investigator-assessed DFS, it is 0.643–1.023. The widths and locations of these intervals provide information that the P-values alone cannot provide.
Why does the P-value not measure effect size?
A P-value is tied to a hypothesis-testing framework. It reflects how compatible the observed data are with the specified null hypothesis under that framework. It is influenced by both the estimated effect and the amount of information in the analysis. Effect magnitude is better described using the hazard ratio together with its confidence interval and, when available, absolute event probabilities.
Why was the analysis stratified by UISS High Risk Group?
The registry explicitly states that the primary Cox analyses were stratified by the UISS High Risk Group. Stratification allows the model to account for this factor when comparing treatment groups. It does not mean that the UISS factor itself is assigned a single ordinary regression coefficient in the same way as an unstratified covariate.
What does intention-to-treat add to the analysis?
ITT keeps participants in the treatment group to which they were originally randomized. This maintains the randomized comparison and avoids redefining the efficacy population based on subsequent treatment exposure. The registry specifically states that the ITT population includes all randomized participants regardless of whether they received study drug or received treatment thereafter.
Why should DFS and OS not be interpreted as the same endpoint?
DFS and OS capture different clinical events. The registered DFS definition includes recurrence, occurrence of a secondary malignancy, or death, beginning at randomization. OS instead concerns survival. A treatment effect on DFS therefore cannot automatically be translated into an equivalent treatment effect on OS.
14. Interpreting Hazard Ratios Without Overstating Them
The HR of 0.761 is a relative measure of the event hazard under the Cox model. The corresponding 23.9% reduction is a transformation of that hazard ratio, not an absolute 23.9-percentage-point improvement in disease-free survival.
The HR of 0.811 corresponds to an estimated 18.9% lower hazard under the same type of model. Its confidence interval, 0.643–1.023, indicates more uncertainty than the point estimate alone reveals.
The OS HR of 0.929 corresponds to an estimated 7.1% lower hazard under the reported Cox model. The 95% CI of 0.670–1.289 is broad enough to include hazard ratios above 1, so the point estimate should not be treated as a precise estimate of the underlying treatment effect.
These examples illustrate why hazard ratios should be reported with their confidence intervals. The HR supplies a compact measure of relative effect, while the confidence interval shows how much statistical uncertainty surrounds that estimate.
15. What the Results Do — and Do Not — Establish
What is directly reported
The registry reports two primary DFS analyses, both using stratified Cox proportional-hazards models in the ITT population, with HRs of 0.761 and 0.811 and their associated confidence intervals and P-values.
What the point estimates mean
Each HR below 1 corresponds to a lower estimated instantaneous event rate for sunitinib under the fitted model.
What the confidence intervals add
They quantify uncertainty around each HR. The investigator-assessed DFS interval extends above 1, while the BICR DFS interval remains below 1.
What cannot be inferred
The ClinicalTrials.gov record does not support an absolute DFS difference, median DFS, subgroup-specific treatment effects, or an individual patient's probability of benefit.
16. Assessment of the Time-to-Event Framework
The S-TRAC primary endpoints are well suited to time-to-event analysis because the registered DFS definition explicitly incorporates the interval from randomization until the first occurrence of a qualifying event. This structure preserves information about when events occur rather than reducing every participant to a simple yes/no status at one arbitrary time.
There are three statistical layers to keep distinct:
- Endpoint definition: DFS specifies which events count and establishes randomization as time zero.
- Analysis population: the ITT population preserves the randomized treatment assignment.
- Effect model: the stratified Cox model estimates the relative hazard associated with treatment assignment.
The registry's reported analysis therefore represents more than a simple comparison of recurrence percentages. It is a model-based analysis of event timing using all available follow-up consistent with the endpoint and censoring framework.
17. Limitations
- Registry-level detail: the ClinicalTrials.gov record provides the formal methods and selected estimates but do not provide every element that might appear in a full statistical analysis plan.
- No absolute DFS estimates: the ClinicalTrials.gov record does not include time-specific DFS probabilities or median DFS, so absolute differences cannot be calculated from the ClinicalTrials.gov record.
- Assessment source differs: the two primary DFS analyses use different assessment approaches, so their HRs should be interpreted separately rather than treated as interchangeable estimates.
- Hazard-ratio assumptions: Cox-model interpretation depends on the proportional-hazards framework. The ClinicalTrials.gov record does not provide diagnostics establishing how well that assumption holds over time.
- Confidence-interval width: the investigator-assessed DFS estimate has a 95% CI of 0.643–1.023, illustrating that the point estimate should not be interpreted without considering its uncertainty.
- Multiplicity information: the ClinicalTrials.gov record identifies two primary endpoints and three posted statistical analyses but do not provide a complete multiplicity-adjustment strategy. No additional multiplicity procedure should therefore be inferred.
- Missing-data and censoring details: the ClinicalTrials.gov record does not specify an imputation strategy or detailed censoring rules beyond the registered DFS assessment schedule.
- Safety comparison: serious adverse-event counts are provided descriptively, but the ClinicalTrials.gov record does not include a formal statistical comparison.
- Generalizability: the ClinicalTrials.gov record identifies the condition and high-risk trial population but does not provide the detailed baseline characteristics needed to assess applicability across specific patient subgroups.
18. Why This Trial Matters Statistically
S-TRAC is a useful teaching example because the registry data bring together several central concepts in clinical-trial biostatistics without requiring a complicated endpoint structure. The primary question is time-to-event based, the efficacy population is ITT, the primary model is a stratified Cox proportional-hazards model, and the effect measure is a hazard ratio.
| Concept | How it appears in S-TRAC |
|---|---|
| Randomization | Participants were randomized to sunitinib or placebo. |
| Parallel-group design | The trial used a parallel design with 2 arms. |
| Triple masking | The registry identifies the trial as triple-masked. |
| Intention-to-treat analysis | Primary analyses included all randomized participants according to initial assignment. |
| Time-to-event endpoint | DFS was defined from randomization to recurrence, secondary malignancy, or death. |
| Hazard ratio | DFS and OS treatment effects were reported as hazard ratios. |
| Cox proportional-hazards model | Used for both formal primary DFS analyses and to derive the OS HR. |
| Stratified analysis | Primary Cox analyses were stratified by UISS High Risk Group. |
| Log-rank test | Reported method for the secondary OS analysis. |
| Confidence intervals | 95% two-sided intervals accompany all three reported HR estimates. |
| Superiority testing | The registered hypothesis type is superiority. |
19. Statistical Interpretation of the Three Reported Hazard Ratios
Putting the three formal estimates side by side shows an important lesson in clinical-trial interpretation: an effect estimate is never complete without its endpoint definition, analysis population, statistical model, confidence interval, and assessment context.
The visual comparison is descriptive only. It does not imply that one endpoint should be prioritized over another or that the numerical differences between HRs constitute a formal test of heterogeneity. Different endpoints measure different events and can naturally produce different effect estimates.
20. Clinical Biostats Interpretation Framework
Start with the endpoint
Before reading the HR, identify exactly what constitutes an event. For S-TRAC DFS, the registry includes recurrence, secondary malignancy, or death.
Then identify the population
The primary analyses use the ITT population, maintaining the original randomized assignment.
Then identify the model
The primary DFS analyses use stratified Cox proportional-hazards models with UISS High Risk Group as the stratification factor.
Finally read HR, CI, and P-value together
The HR describes relative magnitude, the CI describes uncertainty, and the P-value addresses compatibility with the specified null hypothesis.
This sequence prevents a common statistical error: beginning with the P-value and then working backward to construct a narrative. A more informative approach begins with the clinical endpoint, identifies the population and analysis method, and then interprets the effect estimate together with its uncertainty.
21. Related Tutorials
Learn more about the methods used in this trial:
22. Related Calculators
23. Sources
- ClinicalTrials.gov: S-TRAC, NCT00375674.
- PubMed record: PMID 36216827 — PubMed.
- PubMed record: PMID 33028084 — PubMed.
- PubMed record: PMID 32582758 — PubMed.
- PubMed record: PMID 30412222 — PubMed.
- PubMed record: PMID 29773662 — PubMed.
Continue through Clinical Biostats
Build from this trial's time-to-event methods into tutorials and statistical calculators covering survival analysis, hazard ratios, confidence intervals, and clinical-trial methodology.
24. Record Summary
S-TRAC provides a clear example of a randomized phase 3 time-to-event analysis. The trial enrolled 674 participants and compared sunitinib with placebo using a randomized, parallel-group, triple-masked design. The two registered primary endpoints were disease-free survival assessed by blinded independent central review and investigator-assessed disease-free survival, with both primary analyses conducted in the ITT population using Cox proportional-hazards models stratified by UISS High Risk Group.
The BICR DFS analysis reported an HR of 0.761 with a two-sided 95% CI of 0.594–0.975 and P = 0.030. The investigator-assessed DFS analysis reported an HR of 0.811 with a two-sided 95% CI of 0.643–1.023 and P = 0.077. The secondary OS analysis reported an HR of 0.929 with a two-sided 95% CI of 0.670–1.289 and P = 0.661.
The statistical lesson is broader than any individual P-value. The treatment effect must be interpreted in the context of the endpoint definition, randomized analysis population, assessment method, stratification factor, Cox model, confidence interval, and hypothesis-testing framework. In particular, a hazard ratio is a relative model-based measure rather than an absolute probability difference, and its interpretation is strongest when accompanied by its confidence interval and the full definition of the event being analyzed.