← Clinical Trials
Metastatic Hormone Sensitive Prostate Cancer Phase 3 Completed NCT02677896

ARCHES: Complete Statistical Analysis of Enzalutamide in Metastatic Hormone Sensitive Prostate Cancer

An independent statistical review of the randomized phase 3 ARCHES trial evaluating enzalutamide plus androgen deprivation therapy versus placebo plus androgen deprivation therapy in patients with metastatic hormone sensitive prostate cancer.

Trial start: 2016-03-09  ·  Primary completion: 2018-10-14  ·  Enrollment: 1150
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are limited to the information reported in the ClinicalTrials.gov record.

1. Trial at a Glance

ARCHES was a randomized, parallel, quadruple-masked phase 3 trial evaluating enzalutamide plus androgen deprivation therapy (ADT) versus placebo plus ADT in patients with metastatic hormone sensitive prostate cancer. The registry reports an enrollment of 1150 participants and two primary endpoints centered on radiographic progression-free survival.

1150
Enrolled
Phase 3
3
Arms
Registry record
0.39
Primary rPFS HR
95% CI 0.30–0.50
<0.0001
Primary P-value
Two-sided
FeatureARCHES
PhasePhase 3
PopulationPatients with metastatic hormone sensitive prostate cancer
DesignRandomized, parallel-group
MaskingQuadruple
AllocationRandomized
Primary purposeTreatment
Enrollment1150
Arms3
Primary endpoints2
Primary endpoint typesBinary; time-to-event
Results postedYes
Statistical analyses posted12
Lead sponsorAstellas Pharma Global Development, Inc.
Sponsor typeIndustry
ClinicalTrials.govNCT02677896

2. Clinical Question

The central statistical question was whether adding enzalutamide to androgen deprivation therapy improves radiographic progression-free survival compared with placebo plus ADT in patients with metastatic hormone sensitive prostate cancer.

Population

Patients with metastatic hormone sensitive prostate cancer enrolled in the ARCHES phase 3 trial.

Intervention

Enzalutamide plus androgen deprivation therapy.

Comparator

Placebo plus androgen deprivation therapy.

Primary question

Does enzalutamide plus ADT improve rPFS compared with placebo plus ADT under the registered superiority framework?

3. Trial Design

01
Randomize1150 enrolled
02
MaskQuadruple masking
03
TreatEnzalutamide + ADT or placebo + ADT
04
AssessrPFS and secondary endpoints
05
AnalyzeITT and prespecified comparisons
Allocation
Randomized allocation in a parallel-group phase 3 design.
Masking
The registry describes the trial as quadruple masked.
Primary purpose
Treatment.
Hypothesis type
Superiority.
ACTIVE TREATMENT

Enzalutamide + ADT

  • Enzalutamide
  • Androgen deprivation therapy
CONTROL

Placebo + ADT

  • Placebo
  • Androgen deprivation therapy

The registry reports three arms overall, while the posted statistical comparisons in the ClinicalTrials.gov record compare enzalutamide plus ADT with placebo plus ADT. The third registry arm is described in the ClinicalTrials.gov record as a placebo cross-over enzalutamide group for the serious-adverse-event summary.

4. Endpoints

EndpointRegistry definition / time frameType
Radiographic Progression-Free Survival (rPFS) Based on Independent Central Review (ICR) of Bone Scan According to Prostate Cancer Clinical Trials Working Group 2 (PCWG2) Criteria From the date of randomization to the first objective evidence of rPD at any time or death. Maximum duration was 26.6 months). rPFS was calculated from randomization to first objective evidence of radiographic progression disease or death up to 24 weeks after study drug discontinuation without documented radiographic progression, whichever occurred first. Time-to-event
rPFS Based on ICR of Bone Scan According to Protocol Assessment Criteria From the date of randomization to the first objective evidence of rPD at any time or death. Maximum duration was 26.6 months). The registry definition describes rPFS as time from randomization to radiographic progression disease or death up to 24 weeks after study drug discontinuation without documented radiographic progression, whichever occurred first. Time-to-event / registered endpoint classification: binary

The registry defines radiographic progression disease as progressive disease by Response Evaluation Criteria in Solid Tumors version 1.1 for soft tissue disease. The two primary endpoints use different assessment criteria: one is based on PCWG2 criteria and the other on protocol assessment criteria.

Secondary endpoints reported

Secondary endpointTime frameAnalysis methodEffect measure
Overall Survival (OS)From randomization to death due to any cause; maximum duration 58.6 monthsStratified log-rank test; Cox proportional-hazards modelHazard ratio
Time to PSA ProgressionFrom randomization to first observation of PSA progression; maximum duration 26.6 monthsStratified log-rank test; Cox hazard ratioHazard ratio
Time to Start of New Antineoplastic TherapyFrom randomization to first dose administration of the first antineoplastic therapyCox proportional-hazards modelHazard ratio
PSA Undetectable RateFrom baseline to detectable PSA values; maximum duration 26.6 monthsCochran-Mantel-Haenszel testRisk difference
Objective Response Rate (ORR)From randomization up to 26.6 monthsCochran-Mantel-Haenszel testRisk difference
Time to Deterioration in Urinary SymptomsFrom randomization to first deterioration in urinary symptoms at any postbaseline visit; maximum duration not fully reported in registry-reported registry analysis textStratified log-rank test; Cox hazard ratioHazard ratio
Time to First Symptomatic Skeletal Event (SSE)From randomization to occurrence of first SSE; maximum duration 26.6 monthsLog-rank testHazard ratio
Time to Castration ResistanceFrom randomization to first castration-resistant event; maximum duration 26.6 monthsLog-rank testHazard ratio
Time to Deterioration of Quality of Life (QoL) in FACT-PFrom randomization to first date with a decline from baseline of 10 points or more in FACT-P total score (maximum duration was 26.6 months)Log-rank testHazard ratio
Time to Pain Progression Based on BPI-SFFrom randomization to first pain progression event; maximum duration 26.6 monthsLog-rank testHazard ratio

5. Primary Results: Radiographic Progression-Free Survival

The registry reports formal statistical analyses for both primary rPFS endpoints. Both comparisons use the intention-to-treat population and compare enzalutamide plus ADT with placebo plus ADT.

rPFS Based on Independent Central Review and PCWG2 Criteria

Hazard ratio for radiographic progression or death

0.39

95% CI: 0.30–0.50   ·   P < 0.0001

Two-sided confidence interval · Superiority hypothesis

The primary comparison was based on a stratified log-rank test, with the hazard ratio estimated using a Cox model. The analysis population was the ITT population, defined as all participants randomized in the study.

Primary endpointAnalysis populationComparisonMethodEstimate95% CIP-value
rPFS by ICR / PCWG2 ITT Enzalutamide + ADT vs placebo + ADT Stratified log-rank; Cox HR 0.39 0.30–0.50 <0.0001
rPFS by ICR / protocol assessment criteria ITT Enzalutamide + ADT vs placebo + ADT Log-rank; Cox proportional hazards model 0.39 0.30–0.50 <0.0001
Clinical Biostats interpretation

An HR of 0.39 means that, under the Cox proportional-hazards model, the estimated instantaneous rate of the rPFS event was approximately 39% as high in the enzalutamide-plus-ADT group as in the placebo-plus-ADT group. Equivalently, the estimate corresponds to an approximately 61% lower estimated hazard of radiographic progression or death.

The HR does not mean that 61% of patients avoided progression, that every patient had a 61% reduction in risk, or that the absolute probability of progression was reduced by 61 percentage points. A hazard ratio is a relative time-to-event measure, not an absolute risk measure.

The 95% CI of 0.30–0.50 describes statistical uncertainty around the estimated hazard ratio under the model and analysis framework. It does not describe the range of treatment effects experienced by individual patients.

The P < 0.0001 value addresses evidence against the null hypothesis under the specified statistical test. It does not measure the size or clinical importance of the treatment effect. Effect size and uncertainty are better conveyed by the HR and confidence interval.

Because the analysis is based on a Cox model, interpretation of a single HR depends on the proportional-hazards framework. The ClinicalTrials.gov record does not report a test of that assumption. The analysis also uses censoring inherent to a time-to-event endpoint, so the treatment comparison depends on the observed event and censoring process.

Why the two primary rPFS analyses are informative

The registry defines two primary rPFS endpoints that use different assessment criteria. Their posted estimates are identical: 0.39, with a 95% CI of 0.30–0.50 and P-value <0.0001 for each comparison. This provides two registry-defined analytical views of the same broad clinical time-to-event construct.

The second analysis is not a separate randomized comparison. It compares the same treatment groups in the ITT population, but according to the protocol assessment criteria specified for that endpoint.

6. Secondary Results

The registry contains formal statistical analyses for ten secondary outcomes reported in the ClinicalTrials.gov record. These analyses illustrate how the same randomized treatment comparison can be expressed through different estimands: time-to-event hazard ratios for events occurring over follow-up and risk differences for binary or rate endpoints.

Overall Survival

Hazard ratio for death

0.66

95% CI: 0.53–0.81   ·   P < 0.0001

Maximum duration: 58.6 months

The OS analysis used the ITT population. The P-value came from a stratified log-rank test with a stated significance level of 0.04; the hazard ratio and 95% confidence interval were estimated using a Cox proportional-hazards model.

Clinical Biostats interpretation

An HR of 0.66 corresponds to an estimated hazard of death approximately 66% as high in the enzalutamide-plus-ADT group as in the placebo-plus-ADT group, or an approximately 34% lower estimated hazard of death under the fitted model.

The estimate is not an absolute survival probability and does not mean that 34% of participants avoided death. The 95% CI of 0.53–0.81 quantifies uncertainty around the relative hazard estimate.

The P-value of <0.0001 is evidence against the null hypothesis under the stated stratified test and significance level. It should not be interpreted as a measure of effect magnitude. The ClinicalTrials.gov record also do not provide median OS or time-specific survival probabilities, so those quantities should not be inferred from the HR.

Other Secondary Time-to-Event Outcomes

EndpointEstimate95% CIP-valueAnalysis population
Time to PSA ProgressionHR 0.190.13–0.26<0.0001ITT
Time to Start of New Antineoplastic TherapyHR 0.380.31–0.48Not reported in registry-reported analysisITT Population
Time to Deterioration in Urinary SymptomsHR 0.880.72–1.080.2162ITT
Time to First Symptomatic Skeletal EventHR 0.520.33–0.800.0026ITT
Time to Castration ResistanceHR 0.280.22–0.36<0.0001ITT
Time to Deterioration of QoL in FACT-PHR 0.960.81–1.140.6548ITT
Time to Pain Progression based on BPI-SFHR 0.920.78–1.070.2715ITT

The pattern is not uniform across endpoints. The registry reports substantial relative reductions in the estimated hazards for PSA progression, starting a new antineoplastic therapy, first symptomatic skeletal event, and castration resistance. For urinary symptoms, FACT-P quality-of-life deterioration, and pain progression, the confidence intervals include 1 and the registry-reported P-values are not below conventional two-sided thresholds.

PSA Undetectable Rate

Difference in rate

50.5

95% CI: 45.3–55.7   ·   P < 0.0001

ITT participants with detectable PSA at baseline

The PSA undetectable-rate comparison used the Cochran-Mantel-Haenszel test and reported a difference in rate, normalized here as a risk difference.

Clinical Biostats interpretation

A reported risk difference of 50.5 represents a difference of 50.5 percentage points between the treatment-group rates as defined by the registry analysis. It is therefore an absolute rather than relative effect measure.

The 95% CI of 45.3–55.7 expresses uncertainty around that estimated difference. Unlike a hazard ratio, a risk difference can be interpreted directly on an absolute percentage-point scale.

The P-value of <0.0001 tests the treatment comparison under the specified Cochran-Mantel-Haenszel framework. It does not say that the difference is "50.5% statistically significant" or quantify its magnitude.

Objective Response Rate

Difference in response rate

19.3

95% CI: 10.4–28.2   ·   P < 0.0001

ITT participants with measurable disease at baseline

Objective response rate was analyzed using the Cochran-Mantel-Haenszel test in ITT participants with measurable disease at baseline. The registry reports a difference in rate of 19.3.

Clinical Biostats interpretation

The risk difference of 19.3 indicates an absolute difference in the analyzed response rates of 19.3 percentage points. The 95% CI of 10.4–28.2 describes uncertainty around this absolute difference.

ORR is fundamentally different from rPFS and OS. It asks whether a prespecified response occurred, whereas rPFS and OS incorporate the timing of events and censoring. A response-rate difference therefore cannot be substituted for a hazard ratio.

The P-value of <0.0001 assesses the evidence for a treatment-group difference under the stated categorical-data analysis. It does not quantify how large or clinically important the difference is.

7. Secondary Endpoint Interpretation

EndpointWhat the estimate representsHow to read it
OSRelative rate of death over timeHR below 1 favors enzalutamide + ADT in the reported comparison.
PSA progressionRelative rate of PSA-progression events over timeHR 0.19 indicates a much lower estimated event hazard under the fitted model.
New antineoplastic therapyRelative rate of starting a new antineoplastic therapyHR 0.38 indicates a lower estimated hazard of this event.
Urinary symptom deteriorationRelative rate of urinary deteriorationHR 0.88 has a CI that includes 1.
SSERelative rate of first symptomatic skeletal eventHR 0.52 indicates a lower estimated event hazard.
Castration resistanceRelative rate of first castration-resistant eventHR 0.28 indicates a lower estimated event hazard.
FACT-P QoL deteriorationRelative rate of the defined QoL deterioration eventHR 0.96 has a CI that includes 1.
Pain progressionRelative rate of first pain-progression eventHR 0.92 has a CI that includes 1.
PSA undetectable rateAbsolute difference in analyzed ratesRisk difference 50.5 percentage points.
ORRAbsolute difference in analyzed response ratesRisk difference 19.3 percentage points.

This table highlights an important statistical principle: an endpoint's clinical meaning and its statistical estimand are linked. A hazard ratio summarizes relative event rates over time; a risk difference summarizes an absolute difference in rates. These quantities should not be compared as though they were interchangeable scales.

8. Statistical Methodology

Kaplan-Meier estimation and time-to-event data

Several ARCHES endpoints are time-to-event outcomes. For such outcomes, patients can have different lengths of follow-up, and some patients may not experience the event before observation ends. Kaplan-Meier estimation is the standard nonparametric framework for describing the event-time distribution in the presence of right censoring.

Conceptual Kaplan-Meier form
S(t) = ∏ti ≤ t (1 − di/ni)

where di is the number of events at time ti and ni is the number at risk immediately before that time.

The ClinicalTrials.gov record does not provide the underlying event and censoring records or Kaplan-Meier estimates, so this page does not construct a curve from the summary statistics alone.

Log-rank test

The log-rank test compares the observed and expected numbers of events between randomized groups over the observed follow-up. It is particularly suited to randomized comparisons of time-to-event endpoints.

ARCHES used the log-rank test for both primary rPFS analyses and for several secondary time-to-event endpoints. For OS, the registry specifically identifies a stratified log-rank test.

Cox proportional-hazards model

The Cox model estimates a relative hazard while allowing the baseline hazard function to remain unspecified. In conceptual form:

Cox proportional-hazards model
h(t | X) = h0(t) exp(βX)

For a two-group treatment indicator, exp(β) is the estimated hazard ratio comparing the treatment groups, subject to the proportional-hazards framework.

The registry identifies Cox hazard ratios for the primary rPFS comparisons and multiple secondary time-to-event endpoints. For OS, it explicitly states that the hazard ratio and 95% confidence interval were estimated by a Cox proportional-hazards model.

Intention-to-treat analysis

The primary rPFS analysis population is explicitly defined as the ITT population, consisting of all participants randomized in the study. Several secondary analyses likewise specify ITT populations.

Analyzing randomized participants according to their assigned treatment preserves the treatment-comparison framework created by randomization. It also means that treatment discontinuation or other post-randomization events do not automatically cause a participant to disappear from the efficacy analysis.

Cochran-Mantel-Haenszel test

The Cochran-Mantel-Haenszel method was used for the PSA undetectable rate and objective response rate. In a stratified categorical analysis, this approach can combine information across strata while accounting for the stratification structure rather than treating all observations as belonging to one unstructured table.

For these endpoints, the registry reports a difference in rate, represented here as a risk difference. The reported estimates are 50.5 for PSA undetectable rate and 19.3 for ORR.

Stratified analysis

The primary rPFS analysis was stratified by volume of disease (low vs high) and prior docetaxel use (yes vs no) during the screening period. The same factors are specified in the registry's analysis notes for several other time-to-event comparisons.

Why stratification matters: when important baseline factors are used to structure a randomized comparison, a stratified analysis can preserve that structure in the statistical comparison. The resulting hazard ratio is interpreted within the model used for the stratified analysis rather than as an unadjusted comparison that ignores the specified strata.

9. Understanding the Primary Hazard Ratio

Relative effect

The primary rPFS estimate of 0.39 is a relative time-to-event measure. It indicates a lower estimated instantaneous rate of radiographic progression or death for enzalutamide plus ADT compared with placebo plus ADT under the Cox model.

What it does not mean

It does not mean that 39% of participants progressed, that 61% of participants were protected from progression, or that every participant had an identical 61% reduction in individual risk.

Confidence interval

The 95% CI of 0.30–0.50 indicates the statistical precision of the estimated relative effect. A confidence interval is not a range containing 95% of individual patient effects.

P-value

The P-value of <0.0001 addresses the compatibility of the observed data with the null hypothesis under the specified statistical procedure. It is not a measure of treatment magnitude, clinical importance, or probability that the treatment hypothesis is true.

10. Why the Analysis Uses Both Log-Rank Tests and Cox Models

The log-rank test and Cox model answer related but distinct statistical questions. The log-rank test provides a formal comparison of survival experience between groups, while the Cox model supplies an interpretable relative effect measure—the hazard ratio—with a confidence interval.

Log-rank test

Tests whether the event-time distributions differ between randomized groups under the specified framework.

Cox model

Estimates the relative hazard and provides a confidence interval around the treatment-effect estimate.

Kaplan-Meier

Describes the estimated event-free probability over time while accounting for right censoring.

Together

The three approaches provide complementary information about a time-to-event treatment comparison.

A statistically significant log-rank result does not make the hazard ratio itself a probability. Conversely, the hazard ratio should not be interpreted without considering its confidence interval, the endpoint definition, censoring, and the proportional-hazards framework.

11. Secondary Endpoint Results in Detail

Time to PSA Progression

Hazard ratio

0.19

95% CI: 0.13–0.26   ·   P < 0.0001

The analysis was performed in the ITT population using a stratified log-rank framework, with the Cox hazard ratio as the effect measure. An HR of 0.19 corresponds to an approximately 81% lower estimated hazard of the defined PSA-progression event under the fitted model. This does not mean that 81% of participants avoided PSA progression.

Time to Start of New Antineoplastic Therapy

Hazard ratio

0.38

95% CI: 0.31–0.48

The registry reports a Cox proportional-hazards analysis in the ITT population. A formal P-value is not included in the ClinicalTrials.gov record for this endpoint, so none is reported here.

Time to First Symptomatic Skeletal Event

Hazard ratio

0.52

95% CI: 0.33–0.80   ·   P = 0.0026

The HR of 0.52 corresponds to an approximately 48% lower estimated hazard of the first SSE under the Cox model. The confidence interval reflects uncertainty around that relative estimate.

Time to Castration Resistance

Hazard ratio

0.28

95% CI: 0.22–0.36   ·   P < 0.0001

The HR of 0.28 corresponds to an approximately 72% lower estimated hazard of the first castration-resistant event under the fitted model. This is a relative event-rate interpretation, not a percentage of participants who avoided the event.

Endpoints Without a Reported Statistical Signal Below 0.05

EndpointHR95% CIP-value
Time to Deterioration in Urinary Symptoms0.880.72–1.080.2162
Time to Deterioration of QoL in FACT-P0.960.81–1.140.6548
Time to Pain Progression based on BPI-SF0.920.78–1.070.2715

For each of these endpoints, the 95% confidence interval includes 1. The registry's P-values are also above 0.05. These results do not establish a statistically detectable treatment difference under the reported analyses. They also should not be converted into proof that the treatment groups are identical, because a nonsignificant result is not equivalent to an equivalence test.

12. Statistical Methods Explained

Why was a log-rank test used?

The primary rPFS endpoint is a time-to-event outcome. A log-rank test compares the timing of events between randomized groups while accounting for different follow-up durations and right censoring. It is therefore more appropriate for this endpoint than a simple comparison of proportions at a single time point.

Why was a Cox proportional-hazards model used?

The Cox model provides a compact relative effect measure—the hazard ratio—while allowing the baseline hazard to remain unspecified. This makes it possible to summarize the randomized comparison with an estimate and confidence interval rather than only a hypothesis-test result.

What does an HR of 0.39 mean?

Under the fitted model, the estimated instantaneous rate of the rPFS event in the enzalutamide-plus-ADT group was 0.39 times the corresponding rate in the placebo-plus-ADT group. It is a relative hazard, not a probability and not a percentage of patients who benefit.

Why does the confidence interval matter?

The point estimate alone does not communicate statistical precision. The 95% CI of 0.30–0.50 shows the range of parameter values compatible with the data under the specified inferential framework. A narrower interval generally indicates greater precision than a wider interval, although precision and clinical importance are separate concepts.

Why does the P-value not measure effect size?

A P-value is influenced by the magnitude of the observed difference, the variability and information in the data, and the statistical test. A very small P-value therefore does not mean that the treatment effect is necessarily very large. The effect estimate and confidence interval are needed to describe magnitude and precision.

Why use the ITT population?

ITT analysis retains participants according to randomized assignment, preserving the central comparison generated by randomization. This is particularly important for a treatment-effect analysis because excluding participants after randomization can introduce selection mechanisms that were not present at allocation.

Why can different endpoints produce different conclusions?

Each endpoint measures a different event. rPFS, OS, PSA progression, skeletal events, symptoms, quality of life, and pain progression are not interchangeable outcomes. A treatment can have different effects on different biological or clinical processes, so their hazard ratios should be interpreted separately rather than collapsed into one summary number.

13. Safety Results

The ClinicalTrials.gov record provides serious adverse-event counts by arm. These are reported as affected participants divided by participants at risk.

Safety groupSerious adverse eventsInterpretation
Enzalutamide + ADT244 / 572244 affected participants among 572 at risk
Placebo + ADT128 / 574128 affected participants among 574 at risk
Placebo Cross-over Enzalutamide81 / 18281 affected participants among 182 at risk

These are descriptive safety counts rather than the efficacy estimands used for rPFS and the other time-to-event endpoints. They should not be interpreted as hazard ratios or as adjusted treatment effects.

The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or P-value for serious adverse events. Accordingly, this page does not calculate one.

14. Trial Timing and Registry Status

2016-03-09

Trial start

The ARCHES registry record gives 2016-03-09 as the study start date.

2018-10-14

Primary completion

The registry gives 2018-10-14 as the primary completion date.

COMPLETED

Current registry status in the ClinicalTrials.gov record

the ClinicalTrials.gov record identifies ARCHES as completed and reports results and statistical analyses.

15. Multiplicity and Interpretation of Multiple Endpoints

ARCHES has two registered primary endpoints and multiple secondary endpoints. That structure matters because each additional statistical comparison creates another opportunity to observe a low P-value by chance if tests are considered independently.

The ClinicalTrials.gov record identifies the hypothesis type as superiority, but they do not provide a complete multiplicity-adjustment strategy for all primary and secondary endpoints. Therefore, the individual secondary P-values should be interpreted as the registry-reported results of their respective analyses rather than assumed to represent independently adjusted confirmatory tests.

Interpretation caution: a collection of several small P-values does not automatically establish that every endpoint was separately powered and error-controlled. The appropriate interpretation depends on the prespecified testing hierarchy and multiplicity strategy. Those details are not reported in the ClinicalTrials.gov record.

16. Analysis Populations and Their Consequences

Endpoint / analysisPopulation reportedWhy it matters
Primary rPFS by ICR / PCWG2ITT; all randomized participantsPreserves the randomized treatment comparison.
Primary rPFS by protocol assessment criteriaITTUses the randomized population for the second primary assessment.
OSITT populationMaintains treatment assignment as the basis of comparison.
PSA progressionITTUses randomized assignment for the time-to-event comparison.
PSA undetectable rateITT with detectable PSA at baselineRestricts the categorical analysis to the population specified for this endpoint.
ORRITT participants with measurable disease at baselineResponse can only be evaluated in participants meeting the endpoint's baseline measurement requirement.

The different populations are not necessarily contradictions. Endpoint definitions can require different evaluable characteristics. What matters statistically is that the population is specified before the comparison and interpreted consistently with the endpoint definition.

17. Stratification and Confounding Control

For the primary rPFS comparison, the registry-reported analysis notes specify stratification by volume of disease (low vs high) and prior docetaxel use (yes vs no) during the screening period.

Volume of disease

The analysis distinguishes low-volume from high-volume disease as a stratification factor.

Prior docetaxel

The analysis distinguishes participants according to prior docetaxel use during the screening period.

Stratified log-rank

The time-to-event comparison can account for the specified strata rather than ignoring them.

Stratified interpretation

The reported hazard ratio summarizes the treatment comparison within the statistical framework incorporating the specified stratification.

Stratification does not turn a randomized trial into an observational study. Rather, it incorporates prespecified baseline structure into the analysis. The distinction is important: randomization provides the principal basis for causal comparison, while stratification can improve the alignment of analysis with the trial's design.

18. Censoring and Time-to-Event Endpoints

rPFS, OS, PSA progression, new antineoplastic therapy, urinary symptoms, SSE, castration resistance, FACT-P deterioration, and pain progression are all represented as time-to-event endpoints in the statistical analyses posted on ClinicalTrials.gov.

A participant who has not experienced the endpoint by the end of observed follow-up does not necessarily provide no information. In a standard survival analysis, the participant contributes information up to the point at which their outcome becomes censored. The validity of the analysis depends on the censoring mechanism and on the prespecified endpoint and follow-up rules.

General time-to-event framework
Observed time = min(event time, censoring time)

The statistical analysis uses the observed follow-up while distinguishing an observed event from an observation that ends before the event occurs.

This is why a simple comparison of the proportion of participants with progression at the end of follow-up would not reproduce the information contained in the reported Cox and log-rank analyses.

19. Proportional-Hazards Assumption

The Cox proportional-hazards model interprets the treatment effect through a hazard ratio. The standard interpretation assumes that the relative hazard is reasonably represented as proportional over the relevant follow-up period.

The ClinicalTrials.gov record does not report a proportional-hazards diagnostic or an alternative time-varying treatment-effect model. Therefore, the HRs should be read as the model-based summaries reported by the registry, rather than as proof that the proportional-hazards assumption was perfectly satisfied at every time point.

Important distinction: an HR below 1 summarizes relative event rates. It does not by itself tell the reader whether the treatment effect is constant over time, whether the survival curves have a particular shape, or what the absolute difference in event-free probability is at a particular time point.

20. What the Registry Does Not Report in the Supplied Data

The registry-reported ARCHES data contain formal estimates for the primary endpoints and ten secondary analyses, but they do not provide several quantities that would normally enrich a full clinical-trial statistical report.

These omissions matter because a hazard ratio is only one component of a complete time-to-event interpretation. Without median times, survival probabilities, or underlying event-time data, the page should not manufacture those quantities from the reported HRs.

21. Limitations

22. Why This Trial Matters Statistically

ARCHES is a useful teaching case because it brings together randomized treatment comparison, quadruple masking, multiple time-to-event endpoints, stratified survival analysis, Cox hazard ratios, categorical-data analysis, and ITT principles within one phase 3 trial.

ConceptHow it appears in ARCHES
RandomizationThe registry describes allocation as randomized.
Parallel-group designThe design model is parallel.
BlindingThe trial is described as quadruple masked.
ITT analysisThe primary rPFS analyses use randomized participants.
Time-to-event endpointsrPFS, OS, PSA progression, new therapy, symptoms, SSE, castration resistance, QoL deterioration, and pain progression are analyzed as time-to-event outcomes.
Kaplan-Meier frameworkTime-to-event interpretation is naturally connected to survival-function estimation and censoring.
Log-rank testUsed for both primary rPFS analyses and several secondary endpoints.
Cox modelUsed to estimate hazard ratios and confidence intervals.
Stratified analysisPrimary rPFS analysis stratified by disease volume and prior docetaxel use.
Cochran-Mantel-Haenszel testUsed for PSA undetectable rate and objective response rate.
Risk differenceUsed as the effect measure for PSA undetectable rate and ORR.
MultiplicityTwo primary endpoints and multiple secondary analyses require careful interpretation of the testing framework.

23. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The registry reports a primary rPFS hazard ratio of 0.39 with a 95% CI of 0.30–0.50 and P < 0.0001 for both registered primary rPFS analyses. The analyses use randomized populations and survival-analysis methods specified in the registry.

Endpoint interpretation

The secondary results show different treatment-effect estimates across disease-control, treatment-initiation, skeletal, castration-resistance, symptom, quality-of-life, pain, PSA, and response endpoints. They should be interpreted according to their individual endpoint definitions.

Absolute vs relative effects

Hazard ratios describe relative event rates, whereas the reported PSA undetectable-rate and ORR estimates are risk differences. These effect measures answer different statistical questions.

Evidence vs inference

The reported estimates and P-values are registry results. Claims about clinical importance, long-term durability, or subgroup consistency would require additional information not contained in the ClinicalTrials.gov record.

24. Statistical Methods Explained Through ARCHES

A compact statistical map
Randomization → ITT → time-to-event endpoint → log-rank comparison → Cox HR + 95% CI

For categorical endpoints, the pathway instead becomes: randomized population → endpoint-specific eligibility → Cochran-Mantel-Haenszel comparison → risk difference + 95% CI.

This distinction helps explain why a single clinical trial can legitimately contain several different statistical tests. The method follows the measurement scale and endpoint structure.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Statistical Calculators

27. Sources

Continue through the Clinical Biostats statistical library

Use the trial's statistical concepts as a pathway into deeper tutorials, calculators, and other clinical-trial analyses.

28. Record Summary

ARCHES provides a useful example of how a phase 3 randomized trial can combine several statistical frameworks around one treatment comparison. Its two registered primary rPFS endpoints use time-to-event methodology, with log-rank testing and Cox hazard ratios in the ITT population. The ClinicalTrials.gov record reports a hazard ratio of 0.39 with a 95% CI of 0.30–0.50 and P < 0.0001 for both primary analyses. Secondary analyses extend the statistical story to overall survival, PSA progression, new antineoplastic therapy, skeletal events, castration resistance, symptoms, quality of life, pain, PSA undetectable rate, and objective response rate.

The most important statistical lesson is that these estimates should be interpreted according to their estimand. Hazard ratios describe relative event rates over time; risk differences describe absolute differences in rates. Confidence intervals describe precision, while P-values address evidence against a null hypothesis under a specified testing framework. None of these quantities should be interpreted in isolation from the endpoint definition, analysis population, censoring, stratification, and multiplicity context.

Clinical Biostats methodology: A trial-results page should reconstruct the statistical story without manufacturing unreported results. For ARCHES, that means preserving the registry's endpoint definitions, analysis populations, effect measures, confidence intervals, and P-values while clearly distinguishing reported evidence from statistical education and interpretation.