This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
SWOG 9346 was a randomized, parallel phase 3 trial in men with stage IV prostate cancer. The registry describes a comparison involving continuous and intermittent hormonal therapy, with overall survival evaluated as a non-inferiority endpoint and several 3-month functional and sexual outcomes evaluated using superiority analyses.
| Feature | SWOG 9346 |
|---|---|
| Trial name | SWOG 9346 |
| Phase | Phase 3 |
| Condition | Prostate Cancer |
| Brief title | SWOG-9346, Hormone Therapy in Treating Men With Stage IV Prostate Cancer |
| Design | Randomized, parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 3040 |
| Lead sponsor | SWOG Cancer Research Network |
| Status | Completed |
| Start | 1995-05 |
| Primary completion | 2013-06 |
| ClinicalTrials.gov | NCT00002651 |
2. Clinical Question
The central statistical question was whether intermittent combined androgen deprivation (CAD) could provide overall survival that was not substantially worse than continuous CAD. The registry also evaluated whether intermittent and continuous hormonal therapy differed in physical functioning, emotional functioning, erectile dysfunction, high libido, and vitality at 3 months.
Population
Men with stage IV prostate cancer, as stated in the trial's brief title.
Interventions
Bicalutamide, goserelin acetate, and clinical observation are the interventions listed in the registry.
Comparison
Continuous Hormonal Therapy versus Intermittent Hormonal Therapy for the 3-month functional analyses; the overall-survival analysis compares the registered consolidation arms.
Primary question
Is intermittent CAD not substantially worse than continuous CAD with respect to overall survival, while also evaluating differences in selected 3-month functional and sexual outcomes?
3. Trial Design
Continuous treatment strategy
- Registry safety denominator: 732
- Serious adverse events affected 19 participants
- Used as one of the treatment groups for the 3-month functional and sexual comparisons
Intermittent treatment strategy
- Registry safety denominator: 702
- Serious adverse events affected 10 participants
- Used as one of the treatment groups for the 3-month functional and sexual comparisons
The registry lists bicalutamide and goserelin acetate as drug interventions and clinical observation as an other intervention. The ClinicalTrials.gov record does not provide a more detailed treatment schedule, dose sequence, or treatment-duration description, so those details are not reproduced here.
4. Endpoints
The registry contains six primary endpoints. One is a long-term time-to-event endpoint; the other five evaluate change from baseline to 3 months in physical, emotional, sexual, or vitality measures.
| Primary endpoint | Time frame | Registry definition | Analysis |
|---|---|---|---|
| Overall Survival | Up to 15 years | Non-inferiority test to determine if intermittent combined androgen deprivation (CAD) overall survival is not substantially worse than continuous CAD overall survival. Specifically, the trial is designed for a one-sided test of the hypothesis that the hazard ratio of intermittent CAD to continuous CAD is 1.2. | Cox proportional-hazards model |
| Physical Functioning as Measured by the SF-36 | 3 months | Outcome scored on a scale of 0 to 100, with higher scores indicating better functioning. Change from Baseline in SF-36 Score at 3 Months. | 2-sided t-test |
| Emotional Functioning as Measured by the SF-36 Mental Health Inventory | 3 months | Outcome scored on a scale of 0 to 100, with higher scores indicating better functioning. Change from Baseline in SF-36 Score at 3 Months. | 2-sided t-test |
| Erectile Dysfunction | 3 months | Patients reported whether they had erectile dysfunction (a score of 1) or no erectile dysfunction (a score of 0). Analysis looks at change from Baseline to 3 Months. | 2-sided t-test |
| High Libido | 3 months | Very high, high, or moderate interest in sexual activities was scored as 1; low or very low interest was scored as 0. The outcome reports change from baseline in the percentage of participants with High Libido at 3 months. | 2-sided t-test |
| Vitality | 3 months | Outcome scored on a scale of 0 to 100, with higher scores indicating better functioning. Analysis looks at mean change from Baseline score to 3 Months. | 2-sided t-test |
5. Statistical Methodology
Overall survival and the Cox proportional-hazards model
The registry reports a Cox proportional-hazards model for overall survival. The reported effect measure is a hazard ratio comparing the intermittent and continuous CAD strategies. The analysis is appropriate for a time-to-event endpoint because it uses the timing of events as well as information from participants whose event time is not observed during follow-up.
The hazard ratio associated with a treatment indicator is exp(β). A value below 1 indicates a lower estimated instantaneous event rate for the numerator treatment group; a value above 1 indicates a higher estimated instantaneous event rate.
Non-inferiority framework
The registered overall-survival endpoint is explicitly a non-inferiority test. The trial was designed around a one-sided hypothesis involving a hazard ratio of 1.2. The registry wording states that the overall type I error rate was 0.05, the type II error rate was 0.10, and power was 0.9.
For this type of design, the key question is not whether the observed hazard ratio is statistically different from 1.0. Instead, the question is whether the data provide sufficient evidence that the intermittent strategy's hazard is below the prespecified non-inferiority boundary of 1.2.
t-tests for 3-month outcomes
The five functional and sexual endpoints were analyzed using 2-sided t-tests. The effect measure reported for each was the Mean Difference (Net). These analyses compare the change from baseline to 3 months between the continuous and intermittent hormonal therapy groups.
The registry labels the comparison as Continuous Hormonal Therapy versus Intermittent Hormonal Therapy. The direction and magnitude of the posted mean difference should therefore be read together with the endpoint's coding and scale rather than treated as a generic treatment-effect percentage.
Analysis populations
The registry specifies restricted analysis populations for the five 3-month outcomes. Only eligible participants with usable responses or form sets at both baseline and 3 months were included in the respective analyses. This is important because the denominators for these analyses are therefore not necessarily the full randomized enrollment of 3040.
| Endpoint | Analysis population specified by the registry |
|---|---|
| Physical Functioning | Only eligible patients with a usable form set for the Physical Functioning portion of the SF-36 both at baseline and 3 months. |
| Emotional Functioning | Only eligible patients with a usable form set for the SF-36 Mental Health Inventory both at baseline and 3 months. |
| Erectile Dysfunction | Only eligible patients with usable answers regarding erectile dysfunction both at baseline and 3 months. |
| High Libido | Only eligible patients with usable answers regarding libido both at baseline and 3 months. |
| Vitality | Only eligible patients with a usable form set for Vitality both at baseline and 3 months. |
6. Results: Overall Survival
Overall survival was the principal time-to-event endpoint. The registry reports a Cox proportional-hazards analysis comparing the consolidation arms, with a hazard ratio as the effect measure.
Overall survival hazard ratio
90% CI: 0.99–1.23 · P = 0.15
Non-inferiority analysis; overall-survival time frame up to 15 years
| Endpoint | Comparison | Method | Effect estimate | P-value |
|---|---|---|---|---|
| Overall Survival | Consolidation Arm I vs Consolidation Arm II | Cox proportional-hazards model | HR 1.10 90% CI 0.99–1.23 |
0.15 |
The reported hazard ratio of 1.10 means that the estimated instantaneous rate of death in the numerator group was 1.10 times that in the comparison group under the fitted Cox model. In relative terms, an HR of 1.10 corresponds to an estimated hazard 10% higher for the numerator group, but this should not be interpreted as a 10% higher probability of dying for an individual patient.
The 90% confidence interval of 0.99–1.23 describes statistical uncertainty around the estimated hazard ratio. Importantly, the upper limit extends beyond the prespecified non-inferiority boundary of 1.2. Thus, the interval includes values compatible with a hazard as high as 1.23, which is beyond the trial's stated non-inferiority threshold.
The P-value of 0.15 is not a measure of the size or clinical importance of the hazard ratio. Nor should it be read as the probability that the null hypothesis is true. In a non-inferiority design, interpretation is anchored to the prespecified margin and the direction of the hypothesis rather than simply asking whether a conventional two-sided superiority test crosses 0.05.
As with any Cox-model hazard ratio, the interpretation also depends on the model's proportional-hazards assumption and on the handling of censoring over the up-to-15-year follow-up period. The ClinicalTrials.gov record does not provide enough information to evaluate those assumptions directly.
7. Results: Physical Functioning
Physical Functioning was measured with the SF-36. The outcome was scored from 0 to 100, with higher scores indicating better functioning, and the registered analysis examined change from baseline to 3 months.
Mean difference in physical functioning
95% CI: -0.31–3.97 · P = 0.09
Continuous Hormonal Therapy vs Intermittent Hormonal Therapy
The reported mean difference of 1.83 units is the registry's net comparison of change from baseline to 3 months between the continuous and intermittent hormonal therapy groups. Because the endpoint is measured on a 0-to-100 scale, the estimate is expressed in scale units rather than as a percentage reduction in risk.
The 95% CI of -0.31–3.97 includes zero. That means the data are compatible with a small difference in either direction as well as with a positive difference of greater magnitude. The confidence interval therefore communicates more than the point estimate alone: the observed estimate is not precise enough to exclude zero at the stated confidence level.
The P-value of 0.09 is evidence against the null hypothesis only in the framework of the specified two-sided t-test; it is not a measure of the size of the difference. It also does not establish that the two treatment strategies are clinically equivalent.
The analysis population was restricted to eligible participants with usable SF-36 Physical Functioning form sets at both baseline and 3 months. Consequently, the result describes that analysis population rather than automatically representing every participant enrolled in the trial.
8. Results: Emotional Functioning
Emotional Functioning was measured with the SF-36 Mental Health Inventory. The registry defines the scale as 0 to 100, with higher scores indicating better functioning, and evaluates change from baseline at 3 months.
Mean difference in emotional functioning
95% CI: 1.00–4.76 · P = 0.003
Continuous Hormonal Therapy vs Intermittent Hormonal Therapy
The reported mean difference of 2.88 units represents the net difference in change from baseline to 3 months between the two treatment groups as defined in the registry analysis.
The 95% CI of 1.00–4.76 lies above zero. Within the framework of this two-sided comparison, the interval is therefore consistent with a positive difference for the reported comparison. The interval also gives a sense of precision: the estimated difference is not presented as a single exact quantity, but as an interval reflecting sampling uncertainty.
The P-value of 0.003 indicates strong statistical evidence against a zero mean difference under the specified t-test assumptions. It does not say that the treatment difference is 0.003 units, nor does it measure clinical importance.
Because the endpoint was analyzed only among eligible participants with usable Mental Health Inventory measurements at both baseline and 3 months, the analysis is conditioned on availability of those measurements. The ClinicalTrials.gov record does not provide information sufficient to determine how missing assessments affected the comparison.
9. Results: Erectile Dysfunction
Erectile Dysfunction was coded as a binary outcome: a score of 1 for erectile dysfunction and 0 for no erectile dysfunction. The analysis examined change from baseline to 3 months and reported a net mean difference in percentage of participants.
Mean difference in erectile dysfunction
95% CI: -14–-5 · P < 0.001
Continuous Hormonal Therapy vs Intermittent Hormonal Therapy
The reported mean difference of -10 is expressed in percentage-point units because the outcome unit is the percentage of participants. It is therefore not a hazard ratio and should not be described as a 10% relative risk reduction.
The 95% CI of -14–-5 does not include zero. The interval is consistent with a negative net difference for the reported Continuous-versus-Intermittent contrast across the range shown by the registry.
The P-value of <0.001 indicates strong statistical evidence against a zero mean difference under the posted two-sided t-test. It does not quantify the probability that one treatment strategy is superior, and it does not itself establish whether a 10-percentage-point difference is clinically important.
A methodological caution is particularly relevant here: the endpoint is binary, yet the registry reports a t-test and a mean difference in percentage of participants. The analysis should therefore be interpreted as the registry-posted statistical comparison rather than silently replacing it with a different binary-outcome model.
10. Results: High Libido
High Libido was defined by the registry as very high, high, or moderate interest in sexual activities, coded as 1, versus low or very low interest, coded as 0. The analysis reports change from baseline in the percentage of participants with High Libido at 3 months.
Mean difference in high libido
95% CI: 1–36 · P = 0.04
Continuous Hormonal Therapy vs Intermittent Hormonal Therapy
The posted mean difference of 18 is reported in percentage-point units. For a binary endpoint summarized as a percentage, a difference of 18 represents an 18-percentage-point net difference, not an 18% relative increase.
The 95% CI of 1–36 lies above zero, so the interval is consistent with a positive difference for the registry's stated treatment contrast. Its width also indicates that the exact size of the difference is uncertain: the data support a range from a relatively small positive difference to a substantially larger one.
The P-value of 0.04 is the result of the specified two-sided t-test. It should not be interpreted as the probability that the observed effect is clinically meaningful, nor as a direct measure of the magnitude of the treatment difference.
Because the trial contains multiple primary endpoints, this individual P-value should also be interpreted in the context of the overall endpoint structure. The ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure for these five 3-month superiority analyses.
11. Results: Vitality
Vitality was scored on a 0-to-100 scale, with higher scores indicating better functioning. The registered analysis examines mean change from baseline to 3 months.
Mean difference in vitality
95% CI: -0.83–3.46 · P = 0.23
Continuous Hormonal Therapy vs Intermittent Hormonal Therapy
The reported mean difference of 1.32 units is the net comparison of mean change from baseline to 3 months between the continuous and intermittent hormonal therapy groups.
The 95% CI of -0.83–3.46 includes zero. The observed estimate is therefore compatible with no difference as well as with positive or negative differences within the interval.
The P-value of 0.23 does not provide strong evidence against a zero mean difference under the specified two-sided test. It also does not demonstrate that the two strategies are equivalent: absence of statistical evidence for a difference is not the same as evidence of equivalence.
As with the other 3-month outcomes, the analysis population was limited to eligible participants with usable Vitality form sets at both baseline and 3 months. The result should therefore be understood in relation to that analysis population.
12. Results Summary
The six posted primary analyses use two distinct statistical frameworks: a Cox proportional-hazards model for overall survival and 2-sided t-tests for the five 3-month functional and sexual outcomes.
| Primary endpoint | Effect measure | Estimate | Confidence interval | P-value | Hypothesis |
|---|---|---|---|---|---|
| Overall Survival | Hazard ratio | 1.10 | 90% CI 0.99–1.23 | 0.15 | Non-inferiority or equivalence |
| Physical Functioning | Mean Difference (Net) | 1.83 | 95% CI -0.31–3.97 | 0.09 | Superiority |
| Emotional Functioning | Mean Difference (Net) | 2.88 | 95% CI 1.00–4.76 | 0.003 | Superiority |
| Erectile Dysfunction | Mean Difference (Net) | -10 | 95% CI -14–-5 | <0.001 | Superiority |
| High Libido | Mean Difference (Net) | 18 | 95% CI 1–36 | 0.04 | Superiority |
| Vitality | Mean Difference (Net) | 1.32 | 95% CI -0.83–3.46 | 0.23 | Superiority |
13. Safety Results
The ClinicalTrials.gov record reports serious adverse events by arm using an affected-participant count and an at-risk denominator.
| Safety measure | Continuous Hormonal Therapy | Intermittent Hormonal Therapy |
|---|---|---|
| Serious adverse events | 19/732 | 10/702 |
The ClinicalTrials.gov record does not provide a formal hypothesis test, confidence interval, or comparative effect estimate for serious adverse events. Accordingly, these counts are presented descriptively rather than converted into an unreported statistical comparison.
14. Statistical Methods Explained
Why was a Cox proportional-hazards model used for overall survival?
Overall survival is a time-to-event endpoint: both whether a participant experiences the event and when that event occurs are relevant. The Cox model is designed to compare event rates over follow-up while allowing participants with censored observations to contribute information up to their censoring time.
What does an overall-survival hazard ratio of 1.10 mean?
An HR of 1.10 means the estimated instantaneous event rate in the numerator group is 1.10 times that in the comparison group under the fitted model. It does not mean that 10% more participants died, nor does it represent an absolute difference in survival probability.
Why is 1.2 important for the overall-survival analysis?
The trial's registered non-inferiority hypothesis uses 1.2 as the hazard-ratio boundary. A non-inferiority analysis asks whether the data rule out a treatment effect worse than that prespecified margin. This is fundamentally different from a superiority analysis asking whether a hazard ratio differs from 1.0.
Why does the 90% confidence interval matter in the non-inferiority analysis?
The posted interval is 0.99–1.23. Its upper limit extends above 1.2. In a non-inferiority framework, the upper confidence bound is directly relevant because it represents the less favorable end of the estimated treatment-effect range. The ClinicalTrials.gov record therefore does not place the entire reported confidence interval below the stated 1.2 boundary.
Why were t-tests used for the 3-month outcomes?
The registry reports 2-sided t-tests with Mean Difference (Net) as the effect measure for Physical Functioning, Emotional Functioning, Erectile Dysfunction, High Libido, and Vitality. For the scale-based endpoints, the comparison is expressed in scale units; for the binary outcomes, the registry reports the result in percentage-of-participants units.
Does P = 0.003 mean the emotional-functioning effect is large?
No. A P-value measures how unusual the observed data would be under a specified null hypothesis and statistical model; it does not measure effect magnitude. The magnitude is described by the mean difference of 2.88, while the 95% CI of 1.00–4.76 describes uncertainty around that estimate.
Why should the five 3-month P-values not be read independently?
They belong to a set of five registered primary endpoints evaluated with superiority tests, alongside the overall-survival primary endpoint. Testing multiple endpoints creates a broader multiplicity question. The ClinicalTrials.gov record does not specify an adjustment procedure for these five superiority analyses, so the individual nominal P-values should not automatically be treated as if they represented a single isolated hypothesis test.
15. Confidence Intervals and P-values
The results illustrate why a complete statistical interpretation should report the point estimate, confidence interval, and P-value together.
| Endpoint | What the point estimate says | What the CI adds | What the P-value adds |
|---|---|---|---|
| Overall Survival | HR 1.10 indicates an estimated hazard ratio above 1 for the reported comparison. | 90% CI 0.99–1.23 shows uncertainty and extends beyond the 1.2 NI boundary. | P = 0.15 is the posted result for the non-inferiority analysis and is not an effect-size measure. |
| Physical Functioning | Mean difference 1.83 scale units. | 95% CI -0.31–3.97 includes zero. | P = 0.09 under the 2-sided t-test. |
| Emotional Functioning | Mean difference 2.88 scale units. | 95% CI 1.00–4.76 is above zero. | P = 0.003 under the 2-sided t-test. |
| Erectile Dysfunction | Mean difference -10 percentage points. | 95% CI -14–-5 is below zero. | P < 0.001 under the 2-sided t-test. |
| High Libido | Mean difference 18 percentage points. | 95% CI 1–36 is above zero. | P = 0.04 under the 2-sided t-test. |
| Vitality | Mean difference 1.32 scale units. | 95% CI -0.83–3.46 includes zero. | P = 0.23 under the 2-sided t-test. |
A confidence interval is particularly valuable because it preserves information about the plausible magnitude and direction of the effect. A P-value reduces the comparison to a measure of evidence against a specified null hypothesis; it does not tell the reader how large or clinically important the observed difference is.
16. Non-Inferiority Logic in SWOG 9346
The overall-survival endpoint is statistically different from the five 3-month superiority endpoints because the trial was asking a different question. The non-inferiority question was whether intermittent CAD was not substantially worse than continuous CAD, with a prespecified hazard-ratio threshold of 1.2.
Superiority question
Is the treatment effect different from the null value, typically represented by HR 1.0 for a hazard ratio?
Non-inferiority question
Can the data exclude a treatment effect worse than a prespecified clinically relevant margin?
SWOG 9346 margin
The registered overall-survival design specifies a hazard-ratio value of 1.2 for the one-sided non-inferiority hypothesis.
Posted result
The reported HR is 1.10 with a 90% CI of 0.99–1.23, so the upper confidence bound extends beyond the stated 1.2 boundary.
This distinction prevents a common statistical error: concluding non-inferiority merely because the observed hazard ratio is near 1.0 or because a conventional superiority test is not statistically significant. Non-inferiority must be assessed against its prespecified margin.
17. Missing Data and Analysis Populations
The registry explicitly restricts each of the five 3-month analyses to eligible participants who had usable measurements at both baseline and 3 months. That design feature is statistically important because a change-from-baseline analysis requires information at both time points.
| Issue | What the ClinicalTrials.gov record establishes |
|---|---|
| Baseline and 3-month availability | Required for each of the five functional or sexual analyses. |
| Usable measurement | Each analysis requires a usable form set or usable answer for its specific endpoint. |
| Imputation | The ClinicalTrials.gov record does not specify a missing-data imputation method. |
| Full enrollment | Enrollment was 3040, but the 3-month analyses were restricted to participants meeting their endpoint-specific analysis-population criteria. |
The absence of a reported imputation method in the ClinicalTrials.gov record should not be filled in with an assumed approach. Different missing-data assumptions can produce different estimates, particularly when missingness is related to treatment, outcome, or participant characteristics.
18. Multiplicity and Multiple Primary Endpoints
SWOG 9346 has six registered primary endpoints: overall survival plus five 3-month measures. This creates an important interpretive distinction between the endpoint-specific P-values and the trial as a whole.
| Endpoint family | Number | Hypothesis type | Posted method |
|---|---|---|---|
| Overall Survival | 1 | Non-inferiority or equivalence | Cox proportional-hazards model |
| Physical Functioning | 1 | Superiority | 2-sided t-test |
| Emotional Functioning | 1 | Superiority | 2-sided t-test |
| Erectile Dysfunction | 1 | Superiority | 2-sided t-test |
| High Libido | 1 | Superiority | 2-sided t-test |
| Vitality | 1 | Superiority | 2-sided t-test |
The ClinicalTrials.gov record does not report a multiplicity-adjustment method for the five superiority analyses. Therefore, the posted P-values should be reported as posted rather than reinterpreted as adjusted familywise-error probabilities.
19. Overall Statistical Story
The statistical structure of SWOG 9346 is unusually useful for understanding the difference between non-inferiority and superiority. Its overall-survival analysis asks whether intermittent CAD can stay within a prespecified hazard-ratio margin, while the 3-month outcomes ask whether the two treatment strategies differ in specific functional and sexual measures.
Time-to-event evidence
Overall survival was analyzed with a Cox proportional-hazards model and reported as HR 1.10 with a 90% CI of 0.99–1.23 and P = 0.15.
Functional outcomes
Physical Functioning and Vitality had confidence intervals that included zero, while Emotional Functioning had a 95% CI of 1.00–4.76.
Sexual outcomes
Erectile Dysfunction had a posted mean difference of -10, while High Libido had a posted mean difference of 18.
Safety
Serious adverse events were reported as 19/732 for continuous hormonal therapy and 10/702 for intermittent hormonal therapy.
These findings should not be collapsed into a single overall "better" or "worse" statistic. Each endpoint has a different scale, hypothesis, follow-up period, analysis population, and interpretation. The overall-survival result also requires non-inferiority logic rather than a simple superiority framework.
20. Limitations
- Non-inferiority interpretation: the overall-survival analysis uses a prespecified hazard-ratio boundary of 1.2. The posted 90% confidence interval extends to 1.23, so the interval is not entirely below the stated margin.
- Different hypothesis types: overall survival was analyzed for non-inferiority, whereas the five 3-month outcomes used superiority hypotheses. Their P-values therefore answer different statistical questions.
- Endpoint multiplicity: six primary endpoints were registered, but the ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure for the five 3-month superiority analyses.
- Restricted 3-month populations: the functional and sexual analyses required eligible participants with usable measurements at both baseline and 3 months. These populations are not necessarily identical to the full randomized cohort.
- Missing-data methodology: the ClinicalTrials.gov record does not specify an imputation strategy for missing baseline or 3-month assessments.
- Binary endpoints analyzed with t-tests: Erectile Dysfunction and High Libido were binary measures, yet the registry-posted analyses use t-tests and report net mean differences. This should be preserved when describing the registry analysis rather than silently substituting another method.
- Cox-model assumptions: the hazard ratio is model-based and depends on the proportional-hazards framework. The ClinicalTrials.gov record does not provide information sufficient to assess that assumption.
- Incomplete safety comparison: the ClinicalTrials.gov record contains serious-adverse-event counts and denominators but no formal comparative statistical analysis.
- Limited registry detail: the ClinicalTrials.gov record does not provide treatment schedules, baseline characteristics, median survival, subgroup estimates, interim-analysis details, or a detailed missing-data plan. Those items are therefore not inferred here.
21. Why This Trial Matters Statistically
SWOG 9346 is a useful teaching case because the registry combines a long-term survival endpoint with short-term patient-reported or functional outcomes and, crucially, uses different hypothesis frameworks for those endpoints.
| Concept | How it appears in SWOG 9346 |
|---|---|
| Randomization | Randomized, 2-arm, parallel phase 3 design. |
| Time-to-event analysis | Overall survival followed for up to 15 years. |
| Cox proportional-hazards model | Used for the posted overall-survival analysis. |
| Hazard ratio | Overall survival reported as HR 1.10 with a 90% CI of 0.99–1.23. |
| Non-inferiority | Overall survival evaluated against a registered hazard-ratio value of 1.2. |
| Confidence intervals | 90% for the overall-survival hazard ratio and 95% for the five 3-month mean differences. |
| t-test | 2-sided t-tests used for five 3-month primary outcomes. |
| Change from baseline | Used for Physical Functioning, Emotional Functioning, Erectile Dysfunction, High Libido, and Vitality. |
| Analysis populations | 3-month analyses require eligible participants with usable measurements at baseline and 3 months. |
| Multiplicity | Six registered primary endpoints create a multiple-endpoint interpretation issue. |
| Safety analysis | Serious adverse events reported by treatment arm with affected/at-risk counts. |
22. A Practical Reading Guide to the Six Primary Results
A useful way to read the registry results is to ask four questions for every endpoint: What is being measured? What is the effect measure? What uncertainty surrounds the estimate? and What hypothesis was actually tested?
| Endpoint | Measure | Key interpretation question |
|---|---|---|
| Overall Survival | Hazard ratio | Does the confidence interval establish that the intermittent strategy stays below the 1.2 non-inferiority margin? |
| Physical Functioning | Mean difference | How large is the change difference, and does its CI include zero? |
| Emotional Functioning | Mean difference | How precisely is the positive mean difference estimated? |
| Erectile Dysfunction | Mean difference in percentage | What does the negative percentage-point contrast mean under the registry's coding? |
| High Libido | Mean difference in percentage | How should an 18-percentage-point estimate be separated from a relative percentage change? |
| Vitality | Mean difference | Does the uncertainty interval exclude zero? |
This framework helps avoid a common mistake in trial interpretation: treating every P-value as if it answered the same question. In SWOG 9346, the overall-survival analysis is explicitly non-inferiority, while the five 3-month analyses are superiority tests.
23. Related Tutorials
Learn more about the methods used in this trial:
24. Related Calculators
25. Sources
- ClinicalTrials.gov: SWOG 9346, NCT00002651.
- Linked publication: PubMed PMID 16921051.
- Linked publication: PubMed PMID 26720308.
- Linked publication: PubMed PMID 25087673.
- Linked publication: PubMed PMID 23550669.
Continue through the Clinical Biostats statistical library
Explore the statistical methods behind randomized trials, survival analysis, confidence intervals, non-inferiority designs, and hypothesis testing.
26. Record Summary
SWOG 9346 provides a compact but technically rich example of clinical-trial statistics. The trial enrolled 3040 participants in a randomized phase 3 parallel design and registered six primary endpoints. Overall survival was evaluated with a Cox proportional-hazards model under a non-inferiority framework using a hazard-ratio value of 1.2, producing a posted HR of 1.10 with a 90% CI of 0.99–1.23 and P = 0.15. Five additional primary endpoints evaluated 3-month changes in physical functioning, emotional functioning, erectile dysfunction, high libido, and vitality using 2-sided t-tests.
The 3-month results demonstrate why point estimates, confidence intervals, and P-values need to be interpreted together. Emotional Functioning had a mean difference of 2.88 with a 95% CI of 1.00–4.76 and P = 0.003; Erectile Dysfunction had a mean difference of -10 with a 95% CI of -14–-5 and P < 0.001; High Libido had a mean difference of 18 with a 95% CI of 1–36 and P = 0.04. Physical Functioning and Vitality had confidence intervals that included zero. These results represent different endpoints and scales and should not be collapsed into a single treatment-effect statistic.