This page separates reported trial results from statistical interpretation. Numerical results and trial characteristics on this page are restricted to the ClinicalTrials.gov record for NCT00326898.
1. Trial at a Glance
ASSURE was a randomized, double-blind, phase 3 parallel trial in oncology with 1,943 participants. The registered primary endpoint was disease-free survival (DFS), a time-to-event endpoint comparing each active-treatment arm with the placebo arm.
| Feature | ASSURE |
|---|---|
| Trial name | ASSURE |
| ClinicalTrials.gov identifier | NCT00326898 |
| Phase | Phase 3 |
| Status | Completed |
| Therapeutic area | Oncology |
| Conditions | Clear Cell Renal Cell Carcinoma; Stage I Renal Cell Cancer AJCC v6 and v7; Stage II Renal Cell Cancer AJCC v7; Stage III Renal Cell Cancer AJCC v7 |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | Double |
| Primary purpose | Treatment |
| Enrollment | 1,943 |
| Lead sponsor | National Cancer Institute (NCI) |
| Sponsor type | NIH |
2. Clinical Question
The registered clinical question was whether treatment with sunitinib malate or sorafenib tosylate, compared separately with the placebo arm, was associated with a difference in disease-free survival among patients with kidney cancer that was removed by surgery.
Population
Patients with clear cell renal cell carcinoma and specified stage I, II, or III renal cell cancer classifications whose kidney cancer was removed by surgery.
Intervention 1
Sunitinib malate, with sunitinib placebo used in the other treatment comparisons.
Intervention 2
Sorafenib tosylate, with sorafenib placebo used in the other treatment comparisons.
Primary question
Does either active-treatment arm differ from the placebo arm in disease-free survival under the registered superiority framework?
3. Trial Design
Sunitinib arm
- Sunitinib malate
- Sorafenib placebo
Sorafenib arm
- Sorafenib tosylate
- Sunitinib placebo
Placebo arm
- Sunitinib placebo
- Sorafenib placebo
The placebo structure is statistically important because the primary analyses are defined as two separate randomized comparisons against the same reference arm. The double-blind design also means that treatment assignment was masked within the registered trial design.
4. Endpoints
| Endpoint | Registry definition / assessment | Statistical role |
|---|---|---|
| Disease-free Survival (DFS) | Disease-free survival (DFS) is defined as time from randomization to recurrence, development of second primary cancer (except localized breast or prostate cancer or nonmelanoma skin cancer), or death from any cause. Patients who were alive without recurrence or qualifying second primary cancer were censored at the date of last disease evaluation. | Primary endpoint |
| 5-year Overall Survival Rate | 5-year overall survival rate; assessed every 3 months if patient is < 2 years from study entry; every 6 months if patient is 2 - 5 years from study entry. | Secondary endpoint |
For the registered DFS endpoint, disease assessment was scheduled every 3 months if the patient was less than 2 years from study entry and every 6 months if the patient was 2 - 5 years from study entry. This schedule is relevant to the interpretation of censoring because DFS depends on determining when recurrence or another qualifying event is observed.
5. Statistical Methodology
Stratified log-rank test
The formal primary analyses used a stratified log-rank test. Rather than treating all patients as belonging to one homogeneous population, the analysis incorporated prespecified stratification factors into the comparison.
The same statistical method was used for both primary comparisons: Arm A versus Arm C and Arm B versus Arm C.
Stratification factors
The registry analysis notes that the stratified log-rank test included the following stratification factors:
| Stratification factor | Role in analysis |
|---|---|
| Basis of risk group | Used to stratify the primary time-to-event comparison. |
| Histologic subtype | Used to stratify the primary time-to-event comparison. |
| Performance status | Used to stratify the primary time-to-event comparison. |
| Type of surgery | Used to stratify the primary time-to-event comparison. |
Hazard ratio
The effect measure reported for the primary DFS analyses was the hazard ratio (HR). The HR summarizes the relative instantaneous event rate between the compared groups under a time-to-event model.
The hazard ratio is a relative time-to-event measure. It is not an absolute difference in the probability of remaining disease-free, and it does not mean that the same percentage of individual patients experienced a reduction in risk.
Cox proportional-hazards model
The registry's normalized statistical-method profile identifies the Cox proportional-hazards model alongside the stratified log-rank test. This model is used to estimate a hazard ratio for time-to-event outcomes while accounting for the trial's stratification structure.
For a Cox model, the proportional-hazards assumption concerns whether the relative hazard between groups can reasonably be represented by a stable ratio over the analyzed time scale. A reported HR should therefore be understood as a model-based summary rather than as a direct transformation of an absolute survival probability.
Analysis population
Both posted primary DFS analyses used all randomized patients. This is important because treatment comparisons based on randomized assignment preserve the trial's allocation framework rather than restricting efficacy analysis to patients who remained on treatment.
6. Primary Results: Disease-Free Survival
ClinicalTrials.gov reports two formal primary DFS comparisons. Both use all randomized patients, the stratified log-rank test, and a hazard ratio as the effect measure. The hypothesis type is superiority, and the reported confidence intervals are 97.5% two-sided intervals.
Arm A: Sunitinib + Sorafenib Placebo vs Placebo
Disease-free survival hazard ratio
97.5% two-sided CI: 0.85–1.23 · P = 0.80
Analysis population: all randomized patients
| Feature | Reported analysis |
|---|---|
| Outcome | Disease-free Survival (DFS) |
| Comparison | Arm A (Sunitinib + Sorafenib Placebo) vs Arm C (Sunitinib Placebo + Sorafenib Placebo) |
| Method | Stratified logrank test |
| Effect measure | Hazard Ratio (HR) |
| Estimate | 1.02 |
| 97.5% two-sided CI | 0.85–1.23 |
| P-value | 0.80 |
| Stratification | Basis of risk group, histologic subtype, performance status, and type of surgery |
The estimated hazard ratio of 1.02 means that the estimated instantaneous rate of a DFS event in Arm A relative to Arm C was approximately 1.02 times the corresponding rate under the reported analysis. Because the estimate is very close to 1, the estimated relative difference is small in magnitude.
The HR does not mean that 2% of patients experienced an event, nor does it mean that every patient had a 2% higher risk. It is a model-based relative measure of event rates over time.
The 97.5% two-sided confidence interval of 0.85–1.23 describes uncertainty around the estimated hazard ratio. It spans values below and above 1, so the interval includes both a possible lower event hazard and a possible higher event hazard for Arm A relative to Arm C.
The P-value of 0.80 is evidence about the compatibility of the observed test statistic with the null hypothesis under the specified testing framework. It is not a measure of the size of the treatment effect, and it should not be interpreted as an 80% probability that the null hypothesis is true.
Because DFS is time-to-event data, censoring and the timing of events matter. Interpretation of the HR also depends on the appropriateness of the proportional-hazards framework used to summarize the treatment comparison.
Arm B: Sorafenib + Sunitinib Placebo vs Placebo
Disease-free survival hazard ratio
97.5% two-sided CI: 0.80–1.17 · P = 0.72
Analysis population: all randomized patients
| Feature | Reported analysis |
|---|---|
| Outcome | Disease-free Survival (DFS) |
| Comparison | Arm B (Sorafenib + Sunitinib Placebo) vs Arm C (Sunitinib Placebo + Sorafenib Placebo) |
| Method | Stratified logrank test |
| Effect measure | Hazard Ratio (HR) |
| Estimate | 0.97 |
| 97.5% two-sided CI | 0.80–1.17 |
| P-value | 0.72 |
| Stratification | Basis of risk group, histologic subtype, performance status, and type of surgery |
The estimated hazard ratio of 0.97 means that the estimated instantaneous rate of a DFS event in Arm B relative to Arm C was approximately 0.97 times the corresponding rate under the reported analysis. Expressed descriptively, this is an estimate close to the null value of 1.
The HR does not represent a 3% absolute improvement in disease-free survival, and it does not imply that every patient experienced a 3% reduction in risk. It is a relative time-to-event measure.
The 97.5% two-sided confidence interval of 0.80–1.17 indicates statistical uncertainty around the estimate and spans 1. The interval therefore includes values consistent with either a lower or higher event hazard for Arm B relative to Arm C.
The P-value of 0.72 describes the result of the reported hypothesis test; it does not quantify the magnitude or clinical importance of the estimated HR.
As with the other DFS comparison, the interpretation depends on the censoring process, the stratified analysis framework, and the assumptions behind summarizing the time-to-event comparison with a hazard ratio.
7. Secondary Results: 5-year Overall Survival Rate
ClinicalTrials.gov also reports analyses for the secondary endpoint 5-year Overall Survival Rate. These analyses use all randomized patients and report hazard ratios estimated using a stratified proportional-hazards model with Arm C, the placebo arm, as the reference group.
| Comparison | HR | 97.5% two-sided CI | Hypothesis |
|---|---|---|---|
| Arm A (Sunitinib + Sorafenib Placebo) vs Arm C | 1.17 | 0.90–1.52 | Superiority |
| Arm B (Sorafenib + Sunitinib Placebo) vs Arm C | 0.98 | 0.75–1.28 | Superiority |
Arm A vs Arm C
The reported HR was 1.17, with a 97.5% two-sided CI of 0.90–1.52. The interval includes 1, so the estimate should be interpreted together with its uncertainty rather than as a precise difference in overall survival.
Arm B vs Arm C
The reported HR was 0.98, with a 97.5% two-sided CI of 0.75–1.28. This estimate is close to 1, while the confidence interval spans both sides of 1.
The registry labels this endpoint as a 5-year overall survival rate but reports the comparative effect as a hazard ratio from a stratified proportional-hazards model. These are related but not identical statistical concepts: a five-year survival proportion is an absolute time-specific quantity, whereas the HR is a relative time-to-event measure derived from a model.
8. Serious Adverse Events by Arm
The registry provides serious adverse-event counts by randomized arm. These figures describe the number affected and the corresponding number at risk in each arm.
| Arm | Serious adverse events affected / at risk |
|---|---|
| Arm A — Sunitinib + Sorafenib Placebo | 368/625 |
| Arm B — Sorafenib + Sunitinib Placebo | 424/628 |
| Arm C — Sunitinib Placebo + Sorafenib Placebo | 81/626 |
The serious-adverse-event data should be kept separate from the efficacy hazard ratios. The DFS HRs describe a time-to-event efficacy comparison, whereas the serious-adverse-event figures are counts of affected participants relative to those at risk.
The reported safety counts also show why randomized treatment assignment does not make every endpoint statistically interchangeable. Safety is influenced by exposure and treatment-related events, while DFS incorporates recurrence, qualifying second primary cancer, death, and censoring according to the registered endpoint definition.
9. Statistical Methods Explained
Why use a stratified log-rank test for DFS?
DFS is a time-to-event endpoint, so the timing of an event matters rather than simply whether an event eventually occurred. The log-rank test compares the observed and expected event experience between randomized groups over follow-up. In ASSURE, the test was stratified by basis of risk group, histologic subtype, performance status, and type of surgery.
What does a hazard ratio of 1.02 mean?
A hazard ratio of 1.02 indicates an estimated instantaneous event rate approximately 1.02 times that of the reference group under the fitted analysis. It is therefore close to the null value of 1. It is not a statement that the probability of an event was exactly 2% higher for every participant.
Why is 1 the null value for a hazard ratio?
A hazard ratio of 1 represents equal estimated hazards between the compared groups. Values below 1 indicate a lower estimated event hazard in the numerator group, while values above 1 indicate a higher estimated event hazard. The direction must always be read from the stated comparison.
What does the 97.5% confidence interval add?
The point estimate alone gives only one estimate of the relative event hazard. The confidence interval provides a range expressing uncertainty around that estimate under the specified statistical framework. For the two primary DFS analyses, the intervals are 0.85–1.23 and 0.80–1.17, respectively, and both include the null value of 1.
Why doesn't the P-value measure effect size?
A P-value summarizes the evidence against a specified null hypothesis under the test procedure. It does not tell us how large an effect is. The estimated HR and its confidence interval are the appropriate reported quantities for understanding the magnitude and precision of the relative time-to-event effect.
Why does censoring matter for DFS?
The registered DFS definition explicitly states that patients alive without recurrence or a qualifying second primary cancer were censored at the date of last disease evaluation. Censoring means that the patient's exact future event time is not observed within the available follow-up. Time-to-event methods incorporate the information available up to that censoring point rather than treating the patient as if a later event had or had not occurred.
Why is the analysis stratified?
Stratification allows the comparison to account for specified baseline factors without treating the different strata as statistically identical. In ASSURE, basis of risk group, histologic subtype, performance status, and type of surgery were incorporated into the stratified log-rank analysis. This aligns the analysis with the trial's prespecified risk structure.
10. Understanding the Primary DFS Results Together
| Primary comparison | HR | 97.5% CI | P-value | Statistical reading |
|---|---|---|---|---|
| Sunitinib arm vs placebo arm | 1.02 | 0.85–1.23 | 0.80 | Estimate near 1; confidence interval spans 1. |
| Sorafenib arm vs placebo arm | 0.97 | 0.80–1.17 | 0.72 | Estimate near 1; confidence interval spans 1. |
Looking at the two estimates together illustrates an important principle in clinical-trial interpretation: a treatment effect should not be judged from the point estimate alone. The HR, confidence interval, P-value, analysis population, endpoint definition, and comparison group all contribute to the statistical interpretation.
The two HR estimates are also not interchangeable. The first is a comparison of Arm A with Arm C; the second compares Arm B with Arm C. Because the placebo arm is the reference for both, the appropriate interpretation is two separate treatment-versus-reference analyses.
An HR close to 1 indicates that the estimated relative event rates are close to one another under the model. It does not by itself tell us the absolute probability of remaining disease-free at a particular time. Absolute survival or disease-free probabilities require time-specific estimates, while the HR summarizes relative event rates across the time-to-event analysis.
For Arm A versus Arm C, the 97.5% two-sided interval is 0.85–1.23. For Arm B versus Arm C, it is 0.80–1.17. In each case, the interval spans 1, demonstrating why the point estimate should not be interpreted without its uncertainty interval.
The reported P-values of 0.80 and 0.72 are outputs of the corresponding statistical tests. They do not represent the probability that either treatment is effective or ineffective, nor do they measure the clinical magnitude of an effect.
11. Multiplicity and the Two Primary Comparisons
ASSURE has one registered primary endpoint, disease-free survival, but the registry reports two primary-endpoint analyses: Arm A versus Arm C and Arm B versus Arm C. This creates an important statistical distinction between the endpoint and the number of formal comparisons.
One endpoint
DFS is the single registered primary endpoint and has a specific event definition and censoring rule.
Two comparisons
The primary endpoint is evaluated separately for the sunitinib arm and the sorafenib arm against the placebo arm.
Same reference group
Both comparisons use Arm C as the placebo reference, making the direction of each HR dependent on the stated comparison.
Interpretation
Two formal comparisons should not automatically be collapsed into one overall treatment estimate or interpreted as though only one hypothesis had been tested.
The ClinicalTrials.gov record identifies both comparisons as superiority analyses but do not provide an additional multiplicity-adjustment specification. Accordingly, this page does not infer an unreported alpha-allocation procedure or adjusted significance threshold.
12. Time-to-Event Analysis: What the Endpoint Actually Measures
DFS is more informative than a simple binary endpoint because it incorporates when a qualifying event occurs. The registered definition starts the clock at randomization and ends it at recurrence, a qualifying second primary cancer, or death from any cause.
Patients alive without recurrence or qualifying second primary cancer are censored at their last disease evaluation.
This structure explains why standard time-to-event methods are appropriate. Two patients who both eventually experience recurrence can contribute different amounts of information if their recurrence times differ substantially. Similarly, a patient who remains event-free at the last evaluation contributes follow-up information even though the ultimate event time is unknown.
Why assessment frequency matters
The registry specifies disease assessment every 3 months for patients less than 2 years from study entry and every 6 months for patients 2 - 5 years from study entry. The observation schedule therefore forms part of the context in which recurrence is detected and censoring is determined.
13. Confidence Intervals and Statistical Precision
| Endpoint / comparison | Estimate | Confidence interval | What the interval shows |
|---|---|---|---|
| DFS: Arm A vs Arm C | 1.02 | 0.85–1.23 | Uncertainty extends below and above the null value of 1. |
| DFS: Arm B vs Arm C | 0.97 | 0.80–1.17 | Uncertainty extends below and above the null value of 1. |
| 5-year OS: Arm A vs Arm C | 1.17 | 0.90–1.52 | Uncertainty extends below and above the null value of 1. |
| 5-year OS: Arm B vs Arm C | 0.98 | 0.75–1.28 | Uncertainty extends below and above the null value of 1. |
All four reported comparative intervals span 1. This does not mean that the true effect is known to be exactly 1. Instead, it indicates that the reported uncertainty intervals contain values on both sides of the no-difference reference value.
The confidence interval also prevents overinterpretation of a point estimate. For example, the Arm B DFS estimate is 0.97, but its interval extends from 0.80 to 1.17. The interval communicates substantially more statistical information than the single HR alone.
14. Interpreting Hazard Ratios Carefully
An HR of 0.97 means that the estimated instantaneous event rate for the numerator group was approximately 0.97 times that of the reference group under the reported model. It does not mean that 3% of patients benefited, nor does it represent a 3-percentage-point change in disease-free survival.
An HR of 1.02 means that the estimated instantaneous event rate for the numerator group was approximately 1.02 times that of the reference group under the reported analysis. It does not mean that every participant had exactly a 2% higher event probability.
The same numerical HR can have different meanings if the order of the comparison is reversed. In ASSURE, both primary HRs are explicitly reported as active-treatment arm versus Arm C. The reference group must therefore be identified before interpreting whether an HR below or above 1 represents a lower or higher estimated event hazard.
15. Safety and Efficacy Are Different Statistical Questions
The ASSURE registry data provide both time-to-event efficacy analyses and serious-adverse-event counts. These should not be reduced to a single numerical summary because they answer different questions.
| Domain | Endpoint / measure | Statistical structure |
|---|---|---|
| Efficacy | Disease-free Survival | Time-to-event; stratified log-rank; hazard ratio |
| Overall survival | 5-year Overall Survival Rate | Time-to-event; stratified proportional-hazards model; hazard ratio |
| Safety | Serious adverse events | Affected participants / participants at risk |
This distinction matters because a treatment can have different statistical behavior across efficacy and safety outcomes. A DFS hazard ratio cannot be used to infer the frequency of serious adverse events, and serious-adverse-event counts cannot be converted into a DFS effect estimate.
16. Design Features That Matter Statistically
Why randomization matters
Randomization creates the framework for comparing treatment assignments without deliberately selecting participants into treatment groups according to their expected outcomes. The primary DFS analyses preserve this framework by using all randomized patients.
Why double-blinding matters
Blinding is particularly relevant in clinical trials because knowledge of assignment can affect participant behavior, assessment, treatment decisions, or reporting. ASSURE is registered as double-blind, so masking is a formal design feature rather than an assumption added during statistical interpretation.
Why stratification matters
Stratification is a bridge between trial design and statistical analysis. The same risk-related structure that informs randomization can be incorporated into the analysis, helping the time-to-event comparison respect important prespecified differences among strata.
17. Limitations and Interpretation Issues
- Hazard-ratio interpretation: an HR is a relative time-to-event summary and should not be interpreted as an absolute risk difference or individual-level treatment effect.
- Proportional-hazards assumption: interpretation of a Cox-model HR depends on the suitability of the proportional-hazards framework for the observed event process.
- Censoring: DFS includes censoring at the date of last disease evaluation for participants alive without recurrence or a qualifying second primary cancer.
- Two primary comparisons: the registry reports two primary analyses of the same primary endpoint, so the comparisons should be interpreted separately rather than averaged or combined.
- Confidence intervals: all four reported comparative HR intervals reported in the ClinicalTrials.gov record span the null value of 1, emphasizing uncertainty around the point estimates.
- Assessment schedule: DFS assessment occurred every 3 months for patients less than 2 years from study entry and every 6 months for patients 2 - 5 years from study entry.
- Analysis population: the posted primary analyses use all randomized patients, which is appropriate for preserving the randomized comparison but does not describe adherence or treatment exposure by itself.
- Registry scope: the ClinicalTrials.gov record does not provide additional baseline tables, subgroup estimates, median DFS, median overall survival, or a more detailed missing-data/imputation strategy. Those quantities are therefore not reproduced here.
18. What the Registry Does and Does Not Establish
What is reported
The registry reports formal primary DFS comparisons, hazard ratios, 97.5% two-sided confidence intervals, P-values, the stratified log-rank method, stratification factors, and the analysis population.
What is not reported here
The ClinicalTrials.gov record does not provide median DFS, median overall survival, time-specific DFS estimates, subgroup hazard ratios, or individual patient event-time data.
What the HR tells us
The HR summarizes the relative event hazard under the reported time-to-event analysis and provides a compact measure of the treatment comparison.
What the HR cannot tell us alone
The HR alone cannot provide an absolute survival probability, the number of patients cured, or the treatment effect experienced by each individual.
This distinction is central to statistical reporting. A complete trial analysis should distinguish between quantities that are actually reported and conclusions that require additional information. The absence of a median survival estimate in the ClinicalTrials.gov record is not a reason to infer one from the hazard ratio.
19. Why This Trial Matters Statistically
ASSURE provides a useful teaching example because several core clinical-trial concepts occur in the same analysis: randomized allocation, double-blinding, a three-arm parallel design, a time-to-event primary endpoint, stratified log-rank testing, hazard ratios, confidence intervals, P-values, and separate comparisons against a common placebo reference.
| Concept | How it appears in ASSURE |
|---|---|
| Randomization | Participants were randomized across three parallel treatment arms. |
| Blinding | The registered masking is double. |
| Time-to-event endpoint | DFS is defined from randomization to recurrence, qualifying second primary cancer, or death. |
| Stratified analysis | Primary DFS comparisons were stratified by basis of risk group, histologic subtype, performance status, and type of surgery. |
| Log-rank test | The primary DFS comparisons used a stratified log-rank test. |
| Hazard ratio | HR was the reported effect measure for the primary DFS analyses and the 5-year OS analyses. |
| Cox model | The registry's statistical-method profile identifies the Cox proportional-hazards model. |
| Confidence interval | Primary HRs are accompanied by 97.5% two-sided confidence intervals. |
| Multiple comparisons | Two primary analyses compare the two active-treatment arms separately with the placebo arm. |
| Safety analysis | Serious adverse events are reported as affected participants relative to participants at risk in each arm. |
The trial is particularly useful for understanding that a statistical analysis is not simply a collection of P-values. The design determines the comparison, the endpoint determines the data structure, the analysis method determines how the comparison is quantified, and the confidence interval determines how precisely the treatment effect has been estimated.
20. A Practical Reading of the ASSURE Results
Start with DFS
The primary endpoint is disease-free survival, measured from randomization until recurrence, a qualifying second primary cancer, or death.
Read each arm separately
Arm A is compared with Arm C, and Arm B is compared with Arm C. These are two distinct treatment comparisons.
Read the HR
The primary reported effect measure is the hazard ratio. Values close to 1 indicate estimated event hazards close to the reference group under the model.
Read the confidence interval
The 97.5% two-sided intervals show how much uncertainty surrounds each point estimate. Both primary DFS intervals include 1.
Read the P-value
The P-values of 0.80 and 0.72 summarize the corresponding hypothesis tests. They should not be mistaken for effect-size measures.
21. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
22. Related Statistical Calculators
23. Sources
- ClinicalTrials.gov: NCT00326898 — ASSURE, the official registry record and source for the trial data summarized on this page.
- Linked publication record: PubMed PMID 37433436.
- Linked publication record: PubMed PMID 34480522.
- Linked publication record: PubMed PMID 34405965.
- Linked publication record: PubMed PMID 31703971.
- Linked publication record: PubMed PMID 28945839.
Continue through the Clinical Biostats statistical tutorials
Use the related methods to deepen your understanding of randomized clinical-trial design, survival analysis, hazard ratios, confidence intervals, and stratified testing.
24. Record Summary
ASSURE is a phase 3, randomized, double-blind, parallel clinical trial with 1,943 participants and three treatment arms. Its registered primary endpoint, disease-free survival, is a time-to-event outcome defined from randomization to recurrence, qualifying second primary cancer, or death, with specified censoring for patients who remained alive without a qualifying event.
The primary statistical analyses compare each active-treatment arm with the placebo arm using stratified log-rank testing and hazard ratios. For sunitinib versus placebo, the reported DFS HR is 1.02 with a 97.5% two-sided CI of 0.85–1.23 and P = 0.80. For sorafenib versus placebo, the reported DFS HR is 0.97 with a 97.5% two-sided CI of 0.80–1.17 and P = 0.72.
The secondary 5-year overall survival analyses report HRs of 1.17 for Arm A versus Arm C and 0.98 for Arm B versus Arm C, with 97.5% two-sided confidence intervals of 0.90–1.52 and 0.75–1.28, respectively. The registry also reports serious adverse events affecting 368/625 participants in Arm A, 424/628 in Arm B, and 81/626 in Arm C.
The central statistical lesson is that these results should be read as complete comparisons rather than isolated numbers: identify the randomized comparison, define the time-to-event endpoint, understand the stratification, interpret the hazard ratio in its correct direction, examine the confidence interval, and then consider the P-value without treating it as a measure of effect size.