← Clinical Trials
Prostate Cancer Phase 3 Randomized NCT00002651

SWOG 9346: Complete Statistical Analysis of Intermittent vs Continuous Hormonal Therapy in Stage IV Prostate Cancer

An independent statistical review of the randomized phase 3 SWOG 9346 trial evaluating intermittent versus continuous combined androgen deprivation in men with stage IV prostate cancer, with overall survival and 3-month functional outcomes as registered primary endpoints.

SWOG Cancer Research Network  ·  Enrollment 3040  ·  Completed
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

SWOG 9346 was a randomized, parallel phase 3 trial in men with stage IV prostate cancer. The registry describes a comparison involving continuous and intermittent hormonal therapy, with overall survival evaluated as a non-inferiority endpoint and several 3-month functional and sexual outcomes evaluated using superiority analyses.

3040
Enrollment
Randomized trial
2
Arms
Parallel design
1.10
Overall Survival HR
90% CI 0.99–1.23
0.15
OS P-value
Cox model
FeatureSWOG 9346
Trial nameSWOG 9346
PhasePhase 3
ConditionProstate Cancer
Brief titleSWOG-9346, Hormone Therapy in Treating Men With Stage IV Prostate Cancer
DesignRandomized, parallel
MaskingNone
Primary purposeTreatment
Enrollment3040
Lead sponsorSWOG Cancer Research Network
StatusCompleted
Start1995-05
Primary completion2013-06
ClinicalTrials.govNCT00002651

2. Clinical Question

The central statistical question was whether intermittent combined androgen deprivation (CAD) could provide overall survival that was not substantially worse than continuous CAD. The registry also evaluated whether intermittent and continuous hormonal therapy differed in physical functioning, emotional functioning, erectile dysfunction, high libido, and vitality at 3 months.

Population

Men with stage IV prostate cancer, as stated in the trial's brief title.

Interventions

Bicalutamide, goserelin acetate, and clinical observation are the interventions listed in the registry.

Comparison

Continuous Hormonal Therapy versus Intermittent Hormonal Therapy for the 3-month functional analyses; the overall-survival analysis compares the registered consolidation arms.

Primary question

Is intermittent CAD not substantially worse than continuous CAD with respect to overall survival, while also evaluating differences in selected 3-month functional and sexual outcomes?

3. Trial Design

01
Enroll3040 participants
02
Randomize2-arm parallel design
03
Hormonal therapyContinuous or intermittent
04
AssessOverall survival and 3-month outcomes
05
AnalyzeCox model and t-tests
CONTINUOUS HORMONAL THERAPY

Continuous treatment strategy

  • Registry safety denominator: 732
  • Serious adverse events affected 19 participants
  • Used as one of the treatment groups for the 3-month functional and sexual comparisons
INTERMITTENT HORMONAL THERAPY

Intermittent treatment strategy

  • Registry safety denominator: 702
  • Serious adverse events affected 10 participants
  • Used as one of the treatment groups for the 3-month functional and sexual comparisons

The registry lists bicalutamide and goserelin acetate as drug interventions and clinical observation as an other intervention. The ClinicalTrials.gov record does not provide a more detailed treatment schedule, dose sequence, or treatment-duration description, so those details are not reproduced here.

4. Endpoints

The registry contains six primary endpoints. One is a long-term time-to-event endpoint; the other five evaluate change from baseline to 3 months in physical, emotional, sexual, or vitality measures.

Primary endpointTime frameRegistry definitionAnalysis
Overall Survival Up to 15 years Non-inferiority test to determine if intermittent combined androgen deprivation (CAD) overall survival is not substantially worse than continuous CAD overall survival. Specifically, the trial is designed for a one-sided test of the hypothesis that the hazard ratio of intermittent CAD to continuous CAD is 1.2. Cox proportional-hazards model
Physical Functioning as Measured by the SF-36 3 months Outcome scored on a scale of 0 to 100, with higher scores indicating better functioning. Change from Baseline in SF-36 Score at 3 Months. 2-sided t-test
Emotional Functioning as Measured by the SF-36 Mental Health Inventory 3 months Outcome scored on a scale of 0 to 100, with higher scores indicating better functioning. Change from Baseline in SF-36 Score at 3 Months. 2-sided t-test
Erectile Dysfunction 3 months Patients reported whether they had erectile dysfunction (a score of 1) or no erectile dysfunction (a score of 0). Analysis looks at change from Baseline to 3 Months. 2-sided t-test
High Libido 3 months Very high, high, or moderate interest in sexual activities was scored as 1; low or very low interest was scored as 0. The outcome reports change from baseline in the percentage of participants with High Libido at 3 months. 2-sided t-test
Vitality 3 months Outcome scored on a scale of 0 to 100, with higher scores indicating better functioning. Analysis looks at mean change from Baseline score to 3 Months. 2-sided t-test
Endpoint distinction: Overall survival is a time-to-event endpoint, whereas the five 3-month outcomes are analyzed as changes or differences at a fixed follow-up time. These endpoint types require different statistical interpretations. The registry's posted analyses use a Cox proportional-hazards model for overall survival and t-tests for the five 3-month outcomes.

5. Statistical Methodology

Overall survival and the Cox proportional-hazards model

The registry reports a Cox proportional-hazards model for overall survival. The reported effect measure is a hazard ratio comparing the intermittent and continuous CAD strategies. The analysis is appropriate for a time-to-event endpoint because it uses the timing of events as well as information from participants whose event time is not observed during follow-up.

Cox model concept
h(t | X) = h0(t) exp(βX)

The hazard ratio associated with a treatment indicator is exp(β). A value below 1 indicates a lower estimated instantaneous event rate for the numerator treatment group; a value above 1 indicates a higher estimated instantaneous event rate.

Non-inferiority framework

The registered overall-survival endpoint is explicitly a non-inferiority test. The trial was designed around a one-sided hypothesis involving a hazard ratio of 1.2. The registry wording states that the overall type I error rate was 0.05, the type II error rate was 0.10, and power was 0.9.

For this type of design, the key question is not whether the observed hazard ratio is statistically different from 1.0. Instead, the question is whether the data provide sufficient evidence that the intermittent strategy's hazard is below the prespecified non-inferiority boundary of 1.2.

Important non-inferiority point: the posted overall-survival estimate is HR 1.10 with a two-sided 90% CI of 0.99–1.23. Because the upper confidence limit is 1.23, it extends beyond the prespecified 1.2 non-inferiority boundary. The posted P-value of 0.15 should not be interpreted as a conventional superiority P-value establishing whether the two strategies are different; the registered hypothesis is non-inferiority.

t-tests for 3-month outcomes

The five functional and sexual endpoints were analyzed using 2-sided t-tests. The effect measure reported for each was the Mean Difference (Net). These analyses compare the change from baseline to 3 months between the continuous and intermittent hormonal therapy groups.

Mean-difference framework
Mean Difference = Mean changeContinuous − Mean changeIntermittent

The registry labels the comparison as Continuous Hormonal Therapy versus Intermittent Hormonal Therapy. The direction and magnitude of the posted mean difference should therefore be read together with the endpoint's coding and scale rather than treated as a generic treatment-effect percentage.

Analysis populations

The registry specifies restricted analysis populations for the five 3-month outcomes. Only eligible participants with usable responses or form sets at both baseline and 3 months were included in the respective analyses. This is important because the denominators for these analyses are therefore not necessarily the full randomized enrollment of 3040.

EndpointAnalysis population specified by the registry
Physical FunctioningOnly eligible patients with a usable form set for the Physical Functioning portion of the SF-36 both at baseline and 3 months.
Emotional FunctioningOnly eligible patients with a usable form set for the SF-36 Mental Health Inventory both at baseline and 3 months.
Erectile DysfunctionOnly eligible patients with usable answers regarding erectile dysfunction both at baseline and 3 months.
High LibidoOnly eligible patients with usable answers regarding libido both at baseline and 3 months.
VitalityOnly eligible patients with a usable form set for Vitality both at baseline and 3 months.

6. Results: Overall Survival

Overall survival was the principal time-to-event endpoint. The registry reports a Cox proportional-hazards analysis comparing the consolidation arms, with a hazard ratio as the effect measure.

Overall survival hazard ratio

1.10

90% CI: 0.99–1.23   ·   P = 0.15

Non-inferiority analysis; overall-survival time frame up to 15 years

EndpointComparisonMethodEffect estimateP-value
Overall Survival Consolidation Arm I vs Consolidation Arm II Cox proportional-hazards model HR 1.10
90% CI 0.99–1.23
0.15
Clinical Biostats interpretation

The reported hazard ratio of 1.10 means that the estimated instantaneous rate of death in the numerator group was 1.10 times that in the comparison group under the fitted Cox model. In relative terms, an HR of 1.10 corresponds to an estimated hazard 10% higher for the numerator group, but this should not be interpreted as a 10% higher probability of dying for an individual patient.

The 90% confidence interval of 0.99–1.23 describes statistical uncertainty around the estimated hazard ratio. Importantly, the upper limit extends beyond the prespecified non-inferiority boundary of 1.2. Thus, the interval includes values compatible with a hazard as high as 1.23, which is beyond the trial's stated non-inferiority threshold.

The P-value of 0.15 is not a measure of the size or clinical importance of the hazard ratio. Nor should it be read as the probability that the null hypothesis is true. In a non-inferiority design, interpretation is anchored to the prespecified margin and the direction of the hypothesis rather than simply asking whether a conventional two-sided superiority test crosses 0.05.

As with any Cox-model hazard ratio, the interpretation also depends on the model's proportional-hazards assumption and on the handling of censoring over the up-to-15-year follow-up period. The ClinicalTrials.gov record does not provide enough information to evaluate those assumptions directly.

7. Results: Physical Functioning

Physical Functioning was measured with the SF-36. The outcome was scored from 0 to 100, with higher scores indicating better functioning, and the registered analysis examined change from baseline to 3 months.

Mean difference in physical functioning

1.83

95% CI: -0.31–3.97   ·   P = 0.09

Continuous Hormonal Therapy vs Intermittent Hormonal Therapy

Clinical Biostats interpretation

The reported mean difference of 1.83 units is the registry's net comparison of change from baseline to 3 months between the continuous and intermittent hormonal therapy groups. Because the endpoint is measured on a 0-to-100 scale, the estimate is expressed in scale units rather than as a percentage reduction in risk.

The 95% CI of -0.31–3.97 includes zero. That means the data are compatible with a small difference in either direction as well as with a positive difference of greater magnitude. The confidence interval therefore communicates more than the point estimate alone: the observed estimate is not precise enough to exclude zero at the stated confidence level.

The P-value of 0.09 is evidence against the null hypothesis only in the framework of the specified two-sided t-test; it is not a measure of the size of the difference. It also does not establish that the two treatment strategies are clinically equivalent.

The analysis population was restricted to eligible participants with usable SF-36 Physical Functioning form sets at both baseline and 3 months. Consequently, the result describes that analysis population rather than automatically representing every participant enrolled in the trial.

8. Results: Emotional Functioning

Emotional Functioning was measured with the SF-36 Mental Health Inventory. The registry defines the scale as 0 to 100, with higher scores indicating better functioning, and evaluates change from baseline at 3 months.

Mean difference in emotional functioning

2.88

95% CI: 1.00–4.76   ·   P = 0.003

Continuous Hormonal Therapy vs Intermittent Hormonal Therapy

Clinical Biostats interpretation

The reported mean difference of 2.88 units represents the net difference in change from baseline to 3 months between the two treatment groups as defined in the registry analysis.

The 95% CI of 1.00–4.76 lies above zero. Within the framework of this two-sided comparison, the interval is therefore consistent with a positive difference for the reported comparison. The interval also gives a sense of precision: the estimated difference is not presented as a single exact quantity, but as an interval reflecting sampling uncertainty.

The P-value of 0.003 indicates strong statistical evidence against a zero mean difference under the specified t-test assumptions. It does not say that the treatment difference is 0.003 units, nor does it measure clinical importance.

Because the endpoint was analyzed only among eligible participants with usable Mental Health Inventory measurements at both baseline and 3 months, the analysis is conditioned on availability of those measurements. The ClinicalTrials.gov record does not provide information sufficient to determine how missing assessments affected the comparison.

9. Results: Erectile Dysfunction

Erectile Dysfunction was coded as a binary outcome: a score of 1 for erectile dysfunction and 0 for no erectile dysfunction. The analysis examined change from baseline to 3 months and reported a net mean difference in percentage of participants.

Mean difference in erectile dysfunction

-10

95% CI: -14–-5   ·   P < 0.001

Continuous Hormonal Therapy vs Intermittent Hormonal Therapy

Clinical Biostats interpretation

The reported mean difference of -10 is expressed in percentage-point units because the outcome unit is the percentage of participants. It is therefore not a hazard ratio and should not be described as a 10% relative risk reduction.

The 95% CI of -14–-5 does not include zero. The interval is consistent with a negative net difference for the reported Continuous-versus-Intermittent contrast across the range shown by the registry.

The P-value of <0.001 indicates strong statistical evidence against a zero mean difference under the posted two-sided t-test. It does not quantify the probability that one treatment strategy is superior, and it does not itself establish whether a 10-percentage-point difference is clinically important.

A methodological caution is particularly relevant here: the endpoint is binary, yet the registry reports a t-test and a mean difference in percentage of participants. The analysis should therefore be interpreted as the registry-posted statistical comparison rather than silently replacing it with a different binary-outcome model.

10. Results: High Libido

High Libido was defined by the registry as very high, high, or moderate interest in sexual activities, coded as 1, versus low or very low interest, coded as 0. The analysis reports change from baseline in the percentage of participants with High Libido at 3 months.

Mean difference in high libido

18

95% CI: 1–36   ·   P = 0.04

Continuous Hormonal Therapy vs Intermittent Hormonal Therapy

Clinical Biostats interpretation

The posted mean difference of 18 is reported in percentage-point units. For a binary endpoint summarized as a percentage, a difference of 18 represents an 18-percentage-point net difference, not an 18% relative increase.

The 95% CI of 1–36 lies above zero, so the interval is consistent with a positive difference for the registry's stated treatment contrast. Its width also indicates that the exact size of the difference is uncertain: the data support a range from a relatively small positive difference to a substantially larger one.

The P-value of 0.04 is the result of the specified two-sided t-test. It should not be interpreted as the probability that the observed effect is clinically meaningful, nor as a direct measure of the magnitude of the treatment difference.

Because the trial contains multiple primary endpoints, this individual P-value should also be interpreted in the context of the overall endpoint structure. The ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure for these five 3-month superiority analyses.

11. Results: Vitality

Vitality was scored on a 0-to-100 scale, with higher scores indicating better functioning. The registered analysis examines mean change from baseline to 3 months.

Mean difference in vitality

1.32

95% CI: -0.83–3.46   ·   P = 0.23

Continuous Hormonal Therapy vs Intermittent Hormonal Therapy

Clinical Biostats interpretation

The reported mean difference of 1.32 units is the net comparison of mean change from baseline to 3 months between the continuous and intermittent hormonal therapy groups.

The 95% CI of -0.83–3.46 includes zero. The observed estimate is therefore compatible with no difference as well as with positive or negative differences within the interval.

The P-value of 0.23 does not provide strong evidence against a zero mean difference under the specified two-sided test. It also does not demonstrate that the two strategies are equivalent: absence of statistical evidence for a difference is not the same as evidence of equivalence.

As with the other 3-month outcomes, the analysis population was limited to eligible participants with usable Vitality form sets at both baseline and 3 months. The result should therefore be understood in relation to that analysis population.

12. Results Summary

The six posted primary analyses use two distinct statistical frameworks: a Cox proportional-hazards model for overall survival and 2-sided t-tests for the five 3-month functional and sexual outcomes.

Primary endpointEffect measureEstimateConfidence intervalP-valueHypothesis
Overall SurvivalHazard ratio1.1090% CI 0.99–1.230.15Non-inferiority or equivalence
Physical FunctioningMean Difference (Net)1.8395% CI -0.31–3.970.09Superiority
Emotional FunctioningMean Difference (Net)2.8895% CI 1.00–4.760.003Superiority
Erectile DysfunctionMean Difference (Net)-1095% CI -14–-5<0.001Superiority
High LibidoMean Difference (Net)1895% CI 1–360.04Superiority
VitalityMean Difference (Net)1.3295% CI -0.83–3.460.23Superiority
Read the estimates on their own scales. The overall-survival result is a hazard ratio, while the 3-month results are net mean differences. The latter are reported in scale units or percentage-point units according to the endpoint. These quantities cannot be compared numerically with one another.

13. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm using an affected-participant count and an at-risk denominator.

Safety measureContinuous Hormonal TherapyIntermittent Hormonal Therapy
Serious adverse events19/73210/702
Serious adverse events · affected participants
Continuous therapy
19
Intermittent therapy
10

The ClinicalTrials.gov record does not provide a formal hypothesis test, confidence interval, or comparative effect estimate for serious adverse events. Accordingly, these counts are presented descriptively rather than converted into an unreported statistical comparison.

14. Statistical Methods Explained

Why was a Cox proportional-hazards model used for overall survival?

Overall survival is a time-to-event endpoint: both whether a participant experiences the event and when that event occurs are relevant. The Cox model is designed to compare event rates over follow-up while allowing participants with censored observations to contribute information up to their censoring time.

What does an overall-survival hazard ratio of 1.10 mean?

An HR of 1.10 means the estimated instantaneous event rate in the numerator group is 1.10 times that in the comparison group under the fitted model. It does not mean that 10% more participants died, nor does it represent an absolute difference in survival probability.

Why is 1.2 important for the overall-survival analysis?

The trial's registered non-inferiority hypothesis uses 1.2 as the hazard-ratio boundary. A non-inferiority analysis asks whether the data rule out a treatment effect worse than that prespecified margin. This is fundamentally different from a superiority analysis asking whether a hazard ratio differs from 1.0.

Why does the 90% confidence interval matter in the non-inferiority analysis?

The posted interval is 0.99–1.23. Its upper limit extends above 1.2. In a non-inferiority framework, the upper confidence bound is directly relevant because it represents the less favorable end of the estimated treatment-effect range. The ClinicalTrials.gov record therefore does not place the entire reported confidence interval below the stated 1.2 boundary.

Why were t-tests used for the 3-month outcomes?

The registry reports 2-sided t-tests with Mean Difference (Net) as the effect measure for Physical Functioning, Emotional Functioning, Erectile Dysfunction, High Libido, and Vitality. For the scale-based endpoints, the comparison is expressed in scale units; for the binary outcomes, the registry reports the result in percentage-of-participants units.

Does P = 0.003 mean the emotional-functioning effect is large?

No. A P-value measures how unusual the observed data would be under a specified null hypothesis and statistical model; it does not measure effect magnitude. The magnitude is described by the mean difference of 2.88, while the 95% CI of 1.00–4.76 describes uncertainty around that estimate.

Why should the five 3-month P-values not be read independently?

They belong to a set of five registered primary endpoints evaluated with superiority tests, alongside the overall-survival primary endpoint. Testing multiple endpoints creates a broader multiplicity question. The ClinicalTrials.gov record does not specify an adjustment procedure for these five superiority analyses, so the individual nominal P-values should not automatically be treated as if they represented a single isolated hypothesis test.

15. Confidence Intervals and P-values

The results illustrate why a complete statistical interpretation should report the point estimate, confidence interval, and P-value together.

EndpointWhat the point estimate saysWhat the CI addsWhat the P-value adds
Overall Survival HR 1.10 indicates an estimated hazard ratio above 1 for the reported comparison. 90% CI 0.99–1.23 shows uncertainty and extends beyond the 1.2 NI boundary. P = 0.15 is the posted result for the non-inferiority analysis and is not an effect-size measure.
Physical Functioning Mean difference 1.83 scale units. 95% CI -0.31–3.97 includes zero. P = 0.09 under the 2-sided t-test.
Emotional Functioning Mean difference 2.88 scale units. 95% CI 1.00–4.76 is above zero. P = 0.003 under the 2-sided t-test.
Erectile Dysfunction Mean difference -10 percentage points. 95% CI -14–-5 is below zero. P < 0.001 under the 2-sided t-test.
High Libido Mean difference 18 percentage points. 95% CI 1–36 is above zero. P = 0.04 under the 2-sided t-test.
Vitality Mean difference 1.32 scale units. 95% CI -0.83–3.46 includes zero. P = 0.23 under the 2-sided t-test.

A confidence interval is particularly valuable because it preserves information about the plausible magnitude and direction of the effect. A P-value reduces the comparison to a measure of evidence against a specified null hypothesis; it does not tell the reader how large or clinically important the observed difference is.

16. Non-Inferiority Logic in SWOG 9346

The overall-survival endpoint is statistically different from the five 3-month superiority endpoints because the trial was asking a different question. The non-inferiority question was whether intermittent CAD was not substantially worse than continuous CAD, with a prespecified hazard-ratio threshold of 1.2.

Superiority question

Is the treatment effect different from the null value, typically represented by HR 1.0 for a hazard ratio?

Non-inferiority question

Can the data exclude a treatment effect worse than a prespecified clinically relevant margin?

SWOG 9346 margin

The registered overall-survival design specifies a hazard-ratio value of 1.2 for the one-sided non-inferiority hypothesis.

Posted result

The reported HR is 1.10 with a 90% CI of 0.99–1.23, so the upper confidence bound extends beyond the stated 1.2 boundary.

This distinction prevents a common statistical error: concluding non-inferiority merely because the observed hazard ratio is near 1.0 or because a conventional superiority test is not statistically significant. Non-inferiority must be assessed against its prespecified margin.

17. Missing Data and Analysis Populations

The registry explicitly restricts each of the five 3-month analyses to eligible participants who had usable measurements at both baseline and 3 months. That design feature is statistically important because a change-from-baseline analysis requires information at both time points.

IssueWhat the ClinicalTrials.gov record establishes
Baseline and 3-month availabilityRequired for each of the five functional or sexual analyses.
Usable measurementEach analysis requires a usable form set or usable answer for its specific endpoint.
ImputationThe ClinicalTrials.gov record does not specify a missing-data imputation method.
Full enrollmentEnrollment was 3040, but the 3-month analyses were restricted to participants meeting their endpoint-specific analysis-population criteria.

The absence of a reported imputation method in the ClinicalTrials.gov record should not be filled in with an assumed approach. Different missing-data assumptions can produce different estimates, particularly when missingness is related to treatment, outcome, or participant characteristics.

18. Multiplicity and Multiple Primary Endpoints

SWOG 9346 has six registered primary endpoints: overall survival plus five 3-month measures. This creates an important interpretive distinction between the endpoint-specific P-values and the trial as a whole.

Endpoint familyNumberHypothesis typePosted method
Overall Survival1Non-inferiority or equivalenceCox proportional-hazards model
Physical Functioning1Superiority2-sided t-test
Emotional Functioning1Superiority2-sided t-test
Erectile Dysfunction1Superiority2-sided t-test
High Libido1Superiority2-sided t-test
Vitality1Superiority2-sided t-test

The ClinicalTrials.gov record does not report a multiplicity-adjustment method for the five superiority analyses. Therefore, the posted P-values should be reported as posted rather than reinterpreted as adjusted familywise-error probabilities.

19. Overall Statistical Story

The statistical structure of SWOG 9346 is unusually useful for understanding the difference between non-inferiority and superiority. Its overall-survival analysis asks whether intermittent CAD can stay within a prespecified hazard-ratio margin, while the 3-month outcomes ask whether the two treatment strategies differ in specific functional and sexual measures.

Time-to-event evidence

Overall survival was analyzed with a Cox proportional-hazards model and reported as HR 1.10 with a 90% CI of 0.99–1.23 and P = 0.15.

Functional outcomes

Physical Functioning and Vitality had confidence intervals that included zero, while Emotional Functioning had a 95% CI of 1.00–4.76.

Sexual outcomes

Erectile Dysfunction had a posted mean difference of -10, while High Libido had a posted mean difference of 18.

Safety

Serious adverse events were reported as 19/732 for continuous hormonal therapy and 10/702 for intermittent hormonal therapy.

These findings should not be collapsed into a single overall "better" or "worse" statistic. Each endpoint has a different scale, hypothesis, follow-up period, analysis population, and interpretation. The overall-survival result also requires non-inferiority logic rather than a simple superiority framework.

20. Limitations

21. Why This Trial Matters Statistically

SWOG 9346 is a useful teaching case because the registry combines a long-term survival endpoint with short-term patient-reported or functional outcomes and, crucially, uses different hypothesis frameworks for those endpoints.

ConceptHow it appears in SWOG 9346
RandomizationRandomized, 2-arm, parallel phase 3 design.
Time-to-event analysisOverall survival followed for up to 15 years.
Cox proportional-hazards modelUsed for the posted overall-survival analysis.
Hazard ratioOverall survival reported as HR 1.10 with a 90% CI of 0.99–1.23.
Non-inferiorityOverall survival evaluated against a registered hazard-ratio value of 1.2.
Confidence intervals90% for the overall-survival hazard ratio and 95% for the five 3-month mean differences.
t-test2-sided t-tests used for five 3-month primary outcomes.
Change from baselineUsed for Physical Functioning, Emotional Functioning, Erectile Dysfunction, High Libido, and Vitality.
Analysis populations3-month analyses require eligible participants with usable measurements at baseline and 3 months.
MultiplicitySix registered primary endpoints create a multiple-endpoint interpretation issue.
Safety analysisSerious adverse events reported by treatment arm with affected/at-risk counts.

22. A Practical Reading Guide to the Six Primary Results

A useful way to read the registry results is to ask four questions for every endpoint: What is being measured? What is the effect measure? What uncertainty surrounds the estimate? and What hypothesis was actually tested?

EndpointMeasureKey interpretation question
Overall SurvivalHazard ratioDoes the confidence interval establish that the intermittent strategy stays below the 1.2 non-inferiority margin?
Physical FunctioningMean differenceHow large is the change difference, and does its CI include zero?
Emotional FunctioningMean differenceHow precisely is the positive mean difference estimated?
Erectile DysfunctionMean difference in percentageWhat does the negative percentage-point contrast mean under the registry's coding?
High LibidoMean difference in percentageHow should an 18-percentage-point estimate be separated from a relative percentage change?
VitalityMean differenceDoes the uncertainty interval exclude zero?

This framework helps avoid a common mistake in trial interpretation: treating every P-value as if it answered the same question. In SWOG 9346, the overall-survival analysis is explicitly non-inferiority, while the five 3-month analyses are superiority tests.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Calculators

25. Sources

Continue through the Clinical Biostats statistical library

Explore the statistical methods behind randomized trials, survival analysis, confidence intervals, non-inferiority designs, and hypothesis testing.

26. Record Summary

SWOG 9346 provides a compact but technically rich example of clinical-trial statistics. The trial enrolled 3040 participants in a randomized phase 3 parallel design and registered six primary endpoints. Overall survival was evaluated with a Cox proportional-hazards model under a non-inferiority framework using a hazard-ratio value of 1.2, producing a posted HR of 1.10 with a 90% CI of 0.99–1.23 and P = 0.15. Five additional primary endpoints evaluated 3-month changes in physical functioning, emotional functioning, erectile dysfunction, high libido, and vitality using 2-sided t-tests.

The 3-month results demonstrate why point estimates, confidence intervals, and P-values need to be interpreted together. Emotional Functioning had a mean difference of 2.88 with a 95% CI of 1.00–4.76 and P = 0.003; Erectile Dysfunction had a mean difference of -10 with a 95% CI of -14–-5 and P < 0.001; High Libido had a mean difference of 18 with a 95% CI of 1–36 and P = 0.04. Physical Functioning and Vitality had confidence intervals that included zero. These results represent different endpoints and scales and should not be collapsed into a single treatment-effect statistic.

Clinical Biostats methodology: The purpose of this page is to distinguish the registry's reported numerical evidence from statistical interpretation. In particular, non-inferiority conclusions must be anchored to the prespecified margin, while superiority P-values should be interpreted alongside effect estimates, confidence intervals, endpoint definitions, analysis populations, and the broader multiple-endpoint structure.