← Clinical Trials
Urothelial Cancer Phase 3 Overall Survival NCT03390504

THOR: Complete Statistical Analysis of Erdafitinib in Advanced Urothelial Cancer

An independent statistical analysis of the randomized phase 3 THOR trial evaluating erdafitinib compared with chemotherapy or pembrolizumab in participants with advanced urothelial cancer and selected fibroblast growth factor receptor (FGFR) gene aberrations.

Trial start: March 23, 2018  ·  Primary completion: April 15, 2024  ·  Enrollment: 629
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are limited to the ClinicalTrials.gov record.

1. Trial at a Glance

THOR was a randomized, open-label, parallel phase 3 study in advanced urothelial cancer. The registry reports an enrollment of 629 participants across 4 arms and posts two formal primary-endpoint analyses for overall survival, comparing erdafitinib with chemotherapy in one cohort and with pembrolizumab in another.

629
Enrollment
Randomized trial
4
Arms
Parallel design
0.65
OS HR
Erdafitinib vs chemotherapy
1.16
OS HR
Erdafitinib vs pembrolizumab
FeatureTHOR
Trial nameTHOR
PhasePhase 3
ConditionUrothelial Cancer
AllocationRandomized
DesignParallel
MaskingNone
Primary purposeTreatment
Enrollment629
Arms4
Primary endpointOverall Survival (OS)
Primary endpoint typeTime-to-event
Results postedYes
Outcome measures posted21
Statistical analyses posted2
Lead sponsorJanssen Research & Development, LLC
Sponsor typeIndustry
Registry statusActive, not recruiting

This page focuses on the statistical information explicitly contained in the ClinicalTrials.gov record. The registry data identify two primary OS comparisons and provide a hazard ratio, two-sided 95% confidence interval, and P-value for each.

2. Clinical Question

The trial evaluates overall survival from randomization to death due to any cause in participants with advanced urothelial cancer and selected FGFR gene aberrations. The statistical questions differ by cohort because erdafitinib is compared with different active comparators.

Population

Participants with advanced urothelial cancer and selected fibroblast growth factor receptor (FGFR) gene aberrations.

Intervention

Erdafitinib, with the registry identifying an 8 mg/9 mg regimen in the reported primary analyses.

Comparators

Cohort 1 used chemotherapy consisting of vinflunine 320 mg/m2 or docetaxel 75 mg/m2. Cohort 2 used pembrolizumab 200 mg.

Primary question

Does randomized treatment assignment produce a difference in overall survival between erdafitinib and the specified comparator within each cohort?

3. Trial Design

01
Randomize629 enrolled
02
4 ArmsParallel design
03
TreatmentErdafitinib or comparator
04
Follow-upOverall survival
05
AnalysisLog-rank framework

The registry describes THOR as a randomized, parallel, unmasked phase 3 treatment study. Four arms are listed in the ClinicalTrials.gov record, while the posted primary statistical analyses are organized into two cohort-level comparisons.

COHORT 1 · ERDAFITINIB VS CHEMOTHERAPY

Primary OS comparison

  • Erdafitinib 8 mg/9 mg
  • Comparator: chemotherapy
  • Vinflunine 320 mg/m2 or docetaxel 75 mg/m2
  • Analysis population: ITT
  • Method: log-rank test
COHORT 2 · ERDAFITINIB VS PEMBROLIZUMAB

Primary OS comparison

  • Erdafitinib 8 mg/9 mg
  • Comparator: pembrolizumab 200 mg
  • Analysis population: ITT
  • Method: stratified log-rank test
  • Hypothesis type: superiority
Interpretation of the four-arm design: the existence of four trial arms does not mean that all four arms were collapsed into one overall treatment comparison. The statistical analyses posted on ClinicalTrials.gov specify two separate primary OS comparisons, each with its own reported hazard ratio, confidence interval, and P-value.

4. Endpoints

EndpointRegistry definitionTime frameType
Overall Survival (OS) Overall survival was measured from the date of randomization to the date of the participant's death. From randomization (3 days prior to Cycle 1 Day 1) until death due to any cause (maximum up to 51.7 months) Time-to-event

Overall survival is the single registered primary endpoint in the ClinicalTrials.gov record. Because death is the event of interest and follow-up can differ among participants, the endpoint is appropriately treated as a time-to-event outcome rather than as a simple binary proportion.

Why the time origin matters

The registry defines the endpoint from randomization, with the time frame described as beginning 3 days prior to Cycle 1 Day 1. This establishes a common origin for the randomized comparison. Participants who remain alive through their available follow-up do not necessarily have an observed death time; in a conventional survival analysis they contribute follow-up information until censoring.

5. Statistical Methodology

Intention-to-treat analysis

The registry explicitly defines the ITT analysis set as including all randomized participants. Participants in this population were analyzed according to the treatment to which they were randomized.

ITT principle in this trial
Randomized participant  →  analyze according to randomized treatment

This preserves the treatment comparison created by randomization. It also means that the primary efficacy analysis is not restricted only to participants who completed treatment exactly as planned.

Log-rank test

The cohort 1 OS comparison was analyzed using a log-rank test. The log-rank test compares the survival experience of two groups over follow-up while accounting for the timing of events and censoring.

Stratified log-rank test

The cohort 2 OS comparison used a stratified log-rank test. Stratification allows a time-to-event comparison to account for prespecified strata when comparing treatment groups. The ClinicalTrials.gov record does not specify the individual stratification variables, so none are added here.

Hazard ratio

The effect measure reported for both primary analyses is the hazard ratio. A hazard ratio compares the estimated instantaneous event rates between treatment groups within the survival-analysis framework.

Conceptual interpretation
HR < 1  →  lower estimated instantaneous event rate in the first-listed group
HR > 1  →  higher estimated instantaneous event rate in the first-listed group

The hazard ratio is a relative time-to-event measure. It is not an absolute survival probability, a median survival time, or the proportion of participants who benefit.

Two-sided confidence intervals

Both primary analyses report two-sided 95% confidence intervals. The interval describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It should not be interpreted as a range containing the treatment effect for 95% of individual patients.

Superiority hypothesis

Both posted primary analyses are identified as superiority analyses. In this framework, the analysis asks whether the randomized treatment comparison provides evidence of a difference rather than attempting to demonstrate that two treatments are sufficiently close within a prespecified non-inferiority margin.

6. Primary Results: Overall Survival

The registry posts two formal primary-endpoint analyses for overall survival. Both use the ITT analysis set, and both report a hazard ratio with a two-sided 95% confidence interval and a P-value.

6.1 Cohort 1: Erdafitinib vs Chemotherapy

Overall survival hazard ratio

0.65

95% CI: 0.48–0.86   ·   P = 0.0031

Method: log-rank test  ·  Hypothesis: superiority  ·  Analysis: ITT

FeatureCohort 1 result
EndpointOverall Survival (OS)
ComparisonErdafitinib 8 mg/9 mg vs chemotherapy
ChemotherapyVinflunine 320 mg/m2 or docetaxel 75 mg/m2
Analysis populationIntent-to-Treat analysis set
MethodLog-rank test
Effect measureHazard ratio
Estimate0.65
95% CI0.48–0.86
P-value=0.0031
Hypothesis typeSuperiority
Clinical Biostats interpretation

The reported hazard ratio of 0.65 means that, within the reported survival-analysis framework, the estimated instantaneous rate of death for participants randomized to erdafitinib was approximately 65% of the corresponding rate in the chemotherapy comparison group. Equivalently, this corresponds to an estimated 35% lower hazard of death relative to the comparator.

The hazard ratio does not mean that 35% of participants avoided death, that survival increased by 35%, or that every participant experienced a 35% reduction in risk. It is a relative time-to-event measure based on the observed follow-up and statistical model.

The 95% confidence interval of 0.48–0.86 communicates uncertainty around the estimated hazard ratio. Because the entire reported interval is below 1, the interval is consistent with a lower estimated hazard in the erdafitinib group under the stated analysis.

The P-value of 0.0031 quantifies the evidence against the null hypothesis within the specified testing framework; it does not measure the magnitude of the treatment effect. Effect magnitude is better represented by the hazard ratio and its confidence interval.

As with other hazard-ratio analyses, interpretation depends on the suitability of the survival-analysis framework, including the relationship between hazards over time. The ClinicalTrials.gov record does not report a proportional-hazards diagnostic or a separate assessment of that assumption.

6.2 Cohort 2: Erdafitinib vs Pembrolizumab

Overall survival hazard ratio

1.16

95% CI: 0.92–1.48   ·   P = 0.2121

Method: stratified log-rank test  ·  Hypothesis: superiority  ·  Analysis: ITT

FeatureCohort 2 result
EndpointOverall Survival (OS)
ComparisonErdafitinib 8 mg/9 mg vs pembrolizumab 200 mg
Analysis populationIntent-to-Treat analysis set
MethodStratified log-rank test
Effect measureHazard ratio
Estimate1.16
95% CI0.92–1.48
P-value=0.2121
Hypothesis typeSuperiority
Clinical Biostats interpretation

The reported hazard ratio of 1.16 means that, within the reported survival-analysis framework, the estimated instantaneous rate of death for participants randomized to erdafitinib was approximately 16% higher than the corresponding rate in the pembrolizumab group.

This estimate does not establish that erdafitinib causes a 16% higher probability of death, nor does it imply that every participant experienced a 16% higher risk. A hazard ratio is a relative time-to-event measure rather than an absolute probability.

The 95% confidence interval of 0.92–1.48 spans 1. This indicates substantial statistical uncertainty around the estimated relative hazard and includes values corresponding to lower, similar, and higher hazards for erdafitinib relative to pembrolizumab.

The P-value of 0.2121 does not measure the size or clinical importance of the estimated hazard ratio. It quantifies the evidence against the specified null hypothesis under the reported statistical test. A non-small P-value should not be converted into a claim that the two treatments are equivalent or identical.

The comparison was a superiority analysis, not a non-inferiority analysis. Therefore, the absence of a statistically significant superiority result should not be interpreted as formal evidence of non-inferiority or equivalence.

7. Comparing the Two Primary Analyses

Primary OS comparisonMethodHR95% CIP-valueHypothesis
Erdafitinib vs chemotherapy Log-rank test 0.65 0.48–0.86 =0.0031 Superiority
Erdafitinib vs pembrolizumab Stratified log-rank test 1.16 0.92–1.48 =0.2121 Superiority

These estimates should be read as two distinct randomized comparisons, not as one four-arm estimate. The comparator differs between cohorts, and the reported statistical method also differs: a log-rank test is listed for the first comparison and a stratified log-rank test for the second.

Do not compare P-values as if they were effect sizes. The P-values of 0.0031 and 0.2121 answer questions about statistical evidence under their respective tests. They do not tell us that one treatment effect is "larger" simply because its P-value is smaller. The hazard ratios and confidence intervals describe the estimated relative effects.

8. Serious Adverse Events by Arm

The ClinicalTrials.gov record reports serious adverse events by arm as affected participants divided by participants at risk. These figures are presented exactly as reported in the registry.

CohortArmSerious adverse eventsAffected / at risk
Cohort 1Arm 1A: Erdafitinib 8 mg/9 mgSerious adverse events63 / 142
Cohort 1Arm 1B: ChemotherapySerious adverse events50 / 117
Cohort 2Arm 2A: Erdafitinib 8 mg/9 mgSerious adverse events69 / 173
Cohort 2Arm 2B: Pembrolizumab 200 mgSerious adverse events81 / 173

These are safety counts and denominators, not efficacy results. The ClinicalTrials.gov record does not provide confidence intervals, hypothesis tests, or adjusted comparisons for these serious-adverse-event counts, so no inferential comparison is added here.

Safety denominator matters: the affected/at-risk format makes clear that the denominator differs between the two arms in cohort 1 and is the same for the two registry-reported cohort 2 arms. The numbers should therefore be interpreted with their stated denominators rather than as raw counts alone.

9. Statistical Methods Explained

Why was a log-rank test used for overall survival?

Overall survival is a time-to-event endpoint. A participant may die during follow-up or may remain alive when follow-up ends, creating censoring. The log-rank test is designed for comparing survival distributions while incorporating the timing of observed events rather than reducing every participant to a simple alive/dead classification at one arbitrary date.

What does a hazard ratio of 0.65 mean?

A hazard ratio of 0.65 indicates an estimated instantaneous death rate that is 65% as large in the first-listed treatment group as in the comparator under the survival-analysis framework. The complementary description is a 35% lower estimated hazard. It is not a 35-percentage-point increase in survival and does not mean that 35% of patients benefited.

What does a hazard ratio of 1.16 mean?

A hazard ratio of 1.16 indicates an estimated instantaneous death rate 16% higher for the first-listed group relative to the comparator. The confidence interval of 0.92–1.48 shows why the point estimate should not be treated as a precise estimate of a higher hazard: the interval includes 1 as well as values below and above 1.

Why is the confidence interval important?

A point estimate is only one estimate of the treatment effect. The 95% confidence interval communicates how uncertain that estimate is under the statistical framework. In cohort 1, the interval is 0.48–0.86; in cohort 2, it is 0.92–1.48. The widths of these intervals also show that the reported estimates are not exact measurements.

Why does the P-value not measure effect size?

The P-value is a measure of statistical evidence against a null hypothesis under the specified testing procedure. It depends not only on the estimated effect but also on information in the data. Effect size is communicated more directly by the hazard ratio, while the confidence interval communicates uncertainty around that estimate.

Why does the ITT analysis matter?

The registry defines the ITT population as all randomized participants, analyzed according to randomized treatment. This maintains the treatment comparison created by randomization. It also avoids redefining the primary efficacy population based on later treatment exposure or adherence.

Why is this not a non-inferiority analysis?

The registry identifies both primary analyses as superiority hypotheses. Non-inferiority requires a prespecified margin and a different inferential question: whether the treatment is not worse than the comparator by more than that margin. No non-inferiority margin is reported in the ClinicalTrials.gov record, so a non-inferiority conclusion cannot be drawn from the posted superiority analysis.

10. What the Confidence Intervals Tell Us

Cohort 1

The HR estimate is 0.65 with a two-sided 95% CI of 0.48–0.86. The entire reported interval lies below 1, so the interval is consistent with a lower estimated hazard for erdafitinib under this comparison.

Cohort 2

The HR estimate is 1.16 with a two-sided 95% CI of 0.92–1.48. The interval includes 1, so the estimate is compatible with a range of relative hazards that includes no difference.

The two intervals also illustrate why a point estimate should never be interpreted in isolation. An HR of 0.65 may appear more definitive when written alone than it does when accompanied by its uncertainty interval, while an HR of 1.16 does not by itself establish a clinically meaningful increase.

11. Interpreting the P-values Correctly

ComparisonHR95% CIP-valueStatistical reading
Erdafitinib vs chemotherapy 0.65 0.48–0.86 =0.0031 The reported superiority test provides evidence against the null under the stated analysis.
Erdafitinib vs pembrolizumab 1.16 0.92–1.48 =0.2121 The reported superiority test does not provide the same level of evidence against the null under the stated analysis.

A P-value should always be read together with the effect estimate, confidence interval, endpoint definition, analysis population, and design. A statistically persuasive result does not automatically quantify clinical importance, while a non-small P-value does not prove that treatments are equivalent.

12. Time-to-Event Analysis and Censoring

The registry defines OS from randomization until death due to any cause, with a maximum time frame of up to 51.7 months. This structure naturally creates censored observations: participants who have not died by the end of their observed follow-up have not necessarily completed the event process.

Conceptual survival function
S(t) = P(T > t)

The survival function describes the probability of remaining event-free beyond time t. For overall survival, the event is death from any cause.

The log-rank framework uses the ordering and timing of events across follow-up. This is fundamentally different from simply comparing the proportion of deaths at the end of the study because it uses information accumulated throughout the observation period.

Kaplan-Meier context: the ClinicalTrials.gov record identifies time-to-event endpoints and log-rank methods, but do not provide Kaplan-Meier estimates or median survival values. Those quantities are therefore not added to this page.

13. Analysis Population and Randomization

The registry explicitly states that the ITT analysis set included all randomized participants and that participants were analyzed according to their randomized treatment. This is one of the most important design features for interpreting the primary efficacy comparisons.

ConceptTHOR registry information
AllocationRandomized
Design modelParallel
MaskingNone
Primary analysis populationIntent-to-Treat
ITT definitionAll randomized participants
Analysis assignmentAccording to randomized treatment

Randomization creates the basis for a causal treatment comparison by determining treatment assignment rather than allowing treatment choice to be determined by participant or investigator characteristics. The ITT analysis then retains that original assignment when estimating the primary efficacy comparison.

14. Stratified vs Unstratified Log-Rank Testing

The two posted primary analyses illustrate two related but distinct survival-analysis approaches. Cohort 1 uses a log-rank test, while cohort 2 uses a stratified log-rank test.

Log-rank test

The standard log-rank test compares survival experience between groups across follow-up, accounting for event timing and censoring.

Stratified log-rank test

The stratified version performs the comparison while accounting for predefined strata. The ClinicalTrials.gov record does not identify the specific strata used in cohort 2.

The fact that the methods differ is important. A statistical analysis should be interpreted from the actual method specified for that comparison rather than assuming that every arm in a multi-arm trial was evaluated using one identical test.

15. Multiplicity and the Two Primary Comparisons

THOR has two posted primary-endpoint analyses for the same registered primary endpoint, overall survival, but they address different cohort-level comparisons. One compares erdafitinib with chemotherapy; the other compares erdafitinib with pembrolizumab.

FeatureCohort 1Cohort 2
EndpointOverall SurvivalOverall Survival
First groupErdafitinib 8 mg/9 mgErdafitinib 8 mg/9 mg
ComparatorVinflunine 320 mg/m2 or docetaxel 75 mg/m2Pembrolizumab 200 mg
TestLog-rankStratified log-rank
HypothesisSuperioritySuperiority
HR0.651.16
P-value=0.0031=0.2121
Multiplicity caution: two formal statistical comparisons create a broader inferential structure than a single two-arm hypothesis test. The ClinicalTrials.gov record does not specify an alpha-allocation or multiplicity-adjustment procedure for these two posted analyses. Therefore, this page does not infer an unreported multiplicity strategy or reinterpret the P-values under one.

16. Interim Analysis, Missing Data, Crossover, and Bayesian Methods

The ClinicalTrials.gov record does not report an interim-analysis procedure, an alpha-spending rule, a non-inferiority margin, a crossover policy, a missing-data or imputation method, or a Bayesian analysis method.

Design topicWhat can be established from the ClinicalTrials.gov record
Non-inferiority marginNot reported; the posted primary analyses are superiority analyses.
CrossoverNot reported in the ClinicalTrials.gov record.
Factorial designNot reported; the design model is parallel.
Interim analysisNot reported in the ClinicalTrials.gov record.
Missing-data / imputation methodNot reported for the primary OS analysis.
Stratification variablesNot reported in the ClinicalTrials.gov record, although cohort 2 used a stratified log-rank test.
Bayesian methodsNot reported.

These omissions are analytically important because each can affect how a trial's evidence should be interpreted. However, adding a conventional method that is not reported would blur the distinction between the registry evidence and general statistical practice.

17. Limitations

18. Why This Trial Matters Statistically

THOR is a useful teaching example because it combines randomized treatment assignment, multiple active comparators, a time-to-event primary endpoint, ITT analysis, hazard ratios, confidence intervals, log-rank testing, and a stratified log-rank comparison within one trial.

ConceptHow it appears in THOR
RandomizationThe trial uses randomized allocation.
Parallel designThe registry identifies a parallel design model with 4 arms.
ITT analysisBoth posted primary analyses use all randomized participants according to randomized treatment.
Time-to-event endpointOverall survival is measured from randomization to death due to any cause.
Hazard ratioBoth primary analyses report HR as the effect measure.
Confidence intervalBoth HRs include two-sided 95% confidence intervals.
Log-rank testingCohort 1 uses a log-rank test.
Stratified log-rank testingCohort 2 uses a stratified log-rank test.
Superiority testingBoth posted primary analyses are identified as superiority hypotheses.
Multiple comparisonsTwo formal primary OS analyses are posted for different comparator cohorts.
Safety analysisSerious adverse events are reported by arm.

19. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The cohort 1 analysis reports an HR of 0.65 with a two-sided 95% CI of 0.48–0.86 and P=0.0031. The cohort 2 analysis reports an HR of 1.16 with a two-sided 95% CI of 0.92–1.48 and P=0.2121.

Clinical interpretation

Clinical interpretation requires more than the statistical test. The ClinicalTrials.gov record does not provide median survival, absolute survival probabilities, subgroup results, or other efficacy outcomes needed to characterize the full clinical magnitude of benefit across the trial.

This distinction is particularly important for time-to-event outcomes. A statistically significant hazard ratio does not, by itself, specify how long participants lived, what proportion were alive at a particular time, or how the treatment affected individual patients.

20. Understanding the Two Hazard Ratios

Cohort 1 · HR 0.65

The estimate is below 1, corresponding to an estimated 35% lower hazard of death for erdafitinib relative to the chemotherapy comparator. The 95% CI of 0.48–0.86 quantifies uncertainty around that estimate, and the reported P-value is =0.0031.

Cohort 2 · HR 1.16

The estimate is above 1, corresponding to an estimated 16% higher hazard of death for erdafitinib relative to pembrolizumab. The 95% CI of 0.92–1.48 includes 1, and the reported P-value is =0.2121.

Do not rank these hazard ratios without their comparators. A hazard ratio only has meaning relative to the treatment group and comparator that define it. The 0.65 estimate and the 1.16 estimate are not two measurements of the same treatment contrast.

21. Trial Timeline

2018-03-23

Trial start

The registry lists March 23, 2018 as the trial start date.

2024-04-15

Primary completion

The registry lists April 15, 2024 as the primary completion date.

Current registry status

Active, not recruiting

the ClinicalTrials.gov record identifies THOR as active, not recruiting.

22. What Is and Is Not Reported

The ClinicalTrials.gov record is sufficiently detailed to support a formal discussion of the primary OS analyses, but it does not support reconstruction of every statistical feature that might appear in a full clinical-trial publication.

ItemReported in the ClinicalTrials.gov record?
Overall survival endpoint definitionYes
Primary OS analysesYes
Hazard ratiosYes
95% confidence intervalsYes
P-valuesYes
ITT population definitionYes
Log-rank methodYes
Stratified log-rank methodYes
Median overall survivalNo
Kaplan-Meier estimatesNo
Primary OS event countsNo
Subgroup efficacy estimatesNo
Non-inferiority marginNo
Interim-analysis procedureNo
Missing-data / imputation procedureNo
Bayesian analysisNo

This distinction prevents an important statistical error: replacing missing registry information with assumptions based on how similar trials are often analyzed.

23. Related Tutorials

Learn more about the methods used in this trial:

24. Related Statistical Calculators

25. Sources

Continue with Clinical Biostats

Build from the statistical methods used in THOR to explore survival analysis, clinical-trial methodology, and statistical calculators.

26. Record Summary

THOR provides a clear example of how a multi-arm randomized phase 3 trial can generate distinct time-to-event comparisons. The ClinicalTrials.gov record identifies overall survival as the registered primary endpoint and provide two formal ITT analyses: erdafitinib versus chemotherapy with a log-rank test, yielding an HR of 0.65 (95% CI 0.48–0.86; P=0.0031), and erdafitinib versus pembrolizumab with a stratified log-rank test, yielding an HR of 1.16 (95% CI 0.92–1.48; P=0.2121).

The central statistical lesson is that these results must be interpreted in their proper comparison-specific context. Hazard ratios describe relative time-to-event effects rather than absolute survival, confidence intervals communicate uncertainty around those estimates, and P-values quantify evidence under a specified testing framework rather than effect magnitude. The ITT population preserves the randomized treatment comparison, while the absence of reported design details such as multiplicity procedures, interim monitoring, missing-data methods, and stratification variables limits how far the statistical interpretation can be extended beyond the information explicitly reported by the registry.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. When the registry supplies a formal estimate, confidence interval, and P-value, those quantities can be explained in context; when a methodological detail is not reported, it should not be reconstructed from assumptions about how similar trials are commonly analyzed.