← Clinical Trials
Urothelial Cancer Phase 3 Time-to-Event Analysis NCT02603432

JAVELIN Bladder 100: Complete Statistical Analysis of Avelumab in Urothelial Cancer

An independent statistical analysis of the randomized phase 3 JAVELIN Bladder 100 trial evaluating avelumab plus best supportive care versus best supportive care in patients with urothelial cancer, with emphasis on overall survival, time-to-event methodology, and interpretation of the reported statistical estimates.

Phase 3  ·  Randomized  ·  Parallel design  ·  Enrollment 700
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the information provided in the ClinicalTrials.gov record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

JAVELIN Bladder 100 was a randomized, parallel, open-label phase 3 trial evaluating avelumab in patients with urothelial cancer. The registry reports an enrollment of 700 participants and a primary time-to-event endpoint of overall survival.

700
Enrollment
Randomized trial
2
Arms
Parallel design
0.69
OS HR
95% CI 0.556–0.863
0.0005
OS P-value
Superiority analysis
FeatureJAVELIN Bladder 100
Trial nameJAVELIN Bladder 100
Brief titleA Study Of Avelumab In Patients With Locally Advanced Or Metastatic Urothelial Cancer (JAVELIN Bladder 100)
PhasePhase 3
ConditionUrothelial Cancer
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment700
Primary endpoint typeTime-to-event
Primary endpointOverall Survival (OS)
Primary analysisLog-rank test; Cox proportional-hazards model for the reported hazard ratio
SponsorPfizer
StatusCompleted
ClinicalTrials.govNCT02603432

2. Clinical Question

The statistical question was whether avelumab plus best supportive care produced a different overall-survival experience from best supportive care alone in the randomized population, under a superiority framework.

Population

Patients enrolled in the phase 3 study of avelumab in patients with locally advanced or metastatic urothelial cancer.

Intervention

Avelumab plus best supportive care (BSC).

Comparator

Best supportive care.

Primary question

Does avelumab plus BSC improve overall survival relative to BSC under the prespecified superiority analysis?

3. Trial Design

01
Randomize700 enrolled
02
Two armsAvelumab + BSC vs BSC
03
Open labelNo masking
04
Time-to-eventOverall survival
05
AnalysisLog-rank + Cox model
ARM A · Avelumab + BSC

Avelumab plus Best Supportive Care

  • Avelumab, classified in the registry as a biological intervention.
  • Best supportive care.
  • Serious adverse events: 111 affected participants among 344 at risk.
ARM B · Best Supportive Care

Best Supportive Care

  • Best supportive care.
  • Serious adverse events: 73 affected participants among 345 at risk.

The registry also records an avelumab intervention following the planned interim analysis. The ClinicalTrials.gov record does not provide additional details about that intervention sequence, crossover rules, or the statistical consequences of the interim analysis, so those features are not inferred here.

Allocation
Randomized allocation was used to compare the two parallel treatment groups.
Masking
The registry describes the study as having no masking.
Primary purpose
Treatment.
Hypothesis
Superiority of avelumab plus BSC versus BSC for the primary time-to-event analysis.

4. Trial Timeline

April 25, 2016

Trial start

The registry lists 2016-04-25 as the study start date.

October 21, 2019

Primary completion

The registry lists 2019-10-21 as the primary completion date.

Completed

Registry status

The ClinicalTrials.gov record classifies the trial as completed and report results.

5. Endpoints

EndpointRegistry definition / time frameType
Overall Survival (OS) From randomization to discontinuation from the study, death or date of censoring, whichever occurred first (for a maximum duration of 41 months) Time-to-event
Time to Deterioration (TTD) Based on National Comprehensive Cancer Network- Functional Assessment of Cancer Therapy (NCCN-FACT) Bladder Symptom Index- 18 (FBlSI-18) Disease Related Symptoms-Physical Subscale (DRS-P) Scores From randomization up to the 90-Day Follow-up Visit (maximum duration of up to 41 months) Time-to-event
Endpoint definition note: The primary endpoint definition and time frame are reproduced from the ClinicalTrials.gov record. The registry text provided for the primary time frame ends with “for a maximu”, so this page does not infer the missing continuation.

6. Primary Endpoint: Overall Survival

Overall survival was defined as the time in months from the date of randomization to the date of death due to any cause. Participants last known to be alive were censored at the date of last contact. The registry states that the analysis was performed using the Kaplan-Meier method.

FeaturePrimary OS analysis
Analysis populationThe full analysis set included all randomized participants.
Groups comparedAvelumab + Best Supportive Care (BSC) vs Best Supportive Care
Statistical testLog Rank
Effect measureHazard Ratio (HR)
Hypothesis typeSuperiority
ModelCox's Proportional Hazard model
Confidence interval95%, two-sided

Hazard ratio for overall survival

0.69

95% CI: 0.556–0.863   ·   P = 0.0005

Analysis population: full analysis set containing all randomized participants.

Clinical Biostats interpretation

The reported OS hazard ratio of 0.69 means that, under the fitted Cox proportional-hazards model, the estimated instantaneous rate of death in the avelumab + BSC group was approximately 31% lower than the corresponding estimated rate in the BSC group.

That statement is a relative, model-based interpretation. An HR of 0.69 does not mean that 31% of patients were prevented from dying, that each individual patient experienced exactly a 31% reduction in risk, or that survival time for every patient was increased by 31%.

The 95% two-sided confidence interval of 0.556–0.863 describes statistical uncertainty around the estimated hazard ratio. It is not an interval containing the survival experience of individual patients. Its interpretation also depends on the model and censoring assumptions underlying the analysis.

The P-value of 0.0005 addresses the evidence against the null hypothesis within the reported superiority analysis. It does not measure the magnitude of the treatment effect, clinical importance, or the probability that the treatment is effective.

Because the hazard ratio came from a Cox proportional-hazards model, interpretation also depends on the proportional-hazards framework. A single HR summarizes the relative event rate under that model rather than providing a complete description of how treatment effects may evolve over time.

Why the primary analysis combines Kaplan-Meier, log-rank, and Cox methods

These methods answer related but distinct questions. The Kaplan-Meier method estimates the survival experience over time while accommodating right censoring. The log-rank test provides a formal comparison of the time-to-event distributions between randomized groups. The Cox model provides the reported hazard ratio, translating the treatment comparison into a relative hazard measure.

Conceptual survival function
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at time ti, while ni is the number of participants at risk immediately before that time.

7. Secondary Endpoint: Time to Deterioration

The registry reports a secondary time-to-event analysis for time to deterioration based on the NCCN-FACT Bladder Symptom Index-18 (FBlSI-18) Disease Related Symptoms-Physical Subscale (DRS-P) Scores.

FeatureSecondary TTD analysis
Analysis populationThe full analysis set included all randomized participants.
Groups comparedAvelumab + Best Supportive Care (BSC) vs Best Supportive Care
Statistical testLog Rank
Effect measureCox Proportional Hazard
Confidence interval95%, one-sided
Hypothesis typeSuperiority
Time frameFrom randomization up to the 90-Day Follow-up Visit (maximum duration of up to 41 months)

Hazard ratio for time to deterioration

1.26

95% CI lower bound: 0.901   ·   P = 0.9130

The registry reports a one-sided 95% confidence interval and a Cox proportional-hazard effect measure.

Clinical Biostats interpretation

The reported hazard ratio of 1.26 is above 1. Under the Cox model, this corresponds to a higher estimated instantaneous hazard of the deterioration event in the avelumab + BSC group relative to the BSC group for this endpoint.

The estimate should not be converted into a statement that patients were “26% worse” or that every patient had a 26% higher probability of deterioration. A hazard ratio is a model-based relative measure of event rates over time, not an individual-level probability.

The registry reports a one-sided 95% confidence interval with a lower bound of 0.901. Because only the lower bound is reported in the ClinicalTrials.gov record, no upper confidence limit is inferred here.

The P-value of 0.9130 is a measure used in the reported hypothesis-testing framework; it is not an effect-size measure. It should not be read as the probability that the null hypothesis is true.

As with the OS analysis, interpretation of the Cox hazard ratio relies on the proportional-hazards model and on appropriate handling of censoring and event times.

8. Statistical Methodology

Kaplan-Meier estimation

The primary OS endpoint is a time-to-event variable. The registry states that overall survival was analyzed using the Kaplan-Meier method. This method is designed for situations in which participants may have different follow-up times and some observations are censored before an event occurs.

For OS, a participant last known to be alive contributes information through the date of last contact and is then censored. This is fundamentally different from treating every participant without an observed death as if the participant had survived for the entire study period.

Log-rank testing

The registry reports a Log Rank method for both the primary OS analysis and the secondary time-to-deterioration analysis. The log-rank test compares the observed and expected numbers of events between treatment groups across event times, providing a formal hypothesis test for differences in survival experience.

Conceptual log-rank comparison
Observed events   vs   Expected events under the null hypothesis

The test uses the ordering and timing of events while accounting for the number of participants remaining at risk at each event time.

Cox proportional-hazards model

The registry states that the primary OS analysis was performed using a Cox's Proportional Hazard model. The secondary time-to-deterioration analysis likewise reports a Cox proportional-hazard effect measure. The Cox model estimates the relative hazard between groups while leaving the baseline hazard unspecified.

Hazard-ratio interpretation
HR < 1  →  lower estimated instantaneous event rate in the treatment group
HR = 1  →  equal estimated instantaneous event rates
HR > 1  →  higher estimated instantaneous event rate in the treatment group

These interpretations describe relative event rates under the Cox model. They are not equivalent to absolute risk differences or ratios of median survival times.

Intention-to-treat principle

The registry identifies the full analysis set as including all randomized participants for both posted analyses. This is consistent with an intention-to-treat approach in which the primary treatment comparison remains anchored to randomized assignment rather than being restricted to participants who completed treatment.

Right censoring

Time-to-event analyses require a rule for participants whose event has not occurred by the time their available follow-up ends. The OS definition explicitly states that participants last known to be alive were censored at the date of last contact. Proper censoring allows their observed follow-up to contribute information without pretending that an event occurred after follow-up ended.

Superiority testing

The primary OS analysis is identified as a superiority analysis. The inferential question is therefore whether the randomized groups differ in the direction specified by the superiority framework, rather than whether one treatment is merely no worse than another by a prespecified non-inferiority margin.

9. Statistical Methods Explained

Why was a log-rank test used?

Overall survival is a time-to-event endpoint, and not every participant necessarily experiences the event during the period of observation. The log-rank test is designed to compare survival distributions while using the available follow-up and accounting for censoring. It therefore fits the structure of the registered OS endpoint more directly than a simple comparison of proportions.

What does an OS hazard ratio of 0.69 mean?

Under the reported Cox model, an HR of 0.69 corresponds to an estimated instantaneous death rate that is approximately 31% lower in the avelumab + BSC group than in the BSC group. It does not mean that 31% of participants survived because of treatment or that each participant's individual risk was reduced by exactly 31%.

Why is the confidence interval important?

The point estimate is only one summary of the randomized comparison. The 95% two-sided confidence interval of 0.556–0.863 shows the statistical uncertainty surrounding the estimated OS hazard ratio under the analysis framework. A narrower interval would generally indicate greater precision than a wider interval, although precision is not the same thing as clinical importance.

Why does the P-value not measure treatment effect size?

The P-value is tied to a hypothesis-testing procedure. It describes how compatible the observed data are with the null hypothesis under that procedure. The effect size is represented here by the hazard ratio, while the confidence interval provides information about its uncertainty. A very small P-value can occur with a modest effect in a sufficiently informative study, and a large effect estimate can have substantial uncertainty in a small or event-limited study.

Why is censoring important for overall survival?

Some participants may still be alive when their available follow-up ends. Rather than treating them as if their eventual survival time were known, the analysis censors them at the last known date alive. This allows the participant's observed follow-up to contribute information without assigning an unobserved death time.

What is the difference between the OS and TTD hazard ratios?

The OS analysis reports an HR of 0.69, while the secondary TTD analysis reports an HR of 1.26. These are estimates for different endpoints. The first concerns death from any cause; the second concerns deterioration defined using the specified FBlSI-18 DRS-P score framework. The two estimates therefore should not be interpreted as contradictory measurements of the same outcome.

10. Analysis Population and Interpretation

Both posted statistical analyses use the full analysis set, which the registry defines as including all randomized participants. This is important because randomized treatment assignment is the foundation of the comparative inference.

PrincipleApplication in the reported analyses
RandomizationParticipants were randomly allocated to the trial arms.
Analysis populationThe full analysis set included all randomized participants.
Primary endpointOverall survival, a time-to-event endpoint.
Secondary endpointTime to deterioration based on the specified FBlSI-18 DRS-P score.
Survival estimationKaplan-Meier method for the primary OS endpoint.
Hypothesis testingLog-rank testing.
Effect estimationHazard ratio from a Cox proportional-hazards model.
Primary hypothesisSuperiority.

The use of all randomized participants strengthens the connection between the reported efficacy comparison and the original randomized treatment assignment. It also means that treatment discontinuation, treatment exposure, and subsequent events do not automatically remove participants from the efficacy analysis.

11. Safety Results

The ClinicalTrials.gov record reports serious adverse events by treatment arm. These figures are presented as affected participants divided by participants at risk.

Safety measureAvelumab + BSCBest Supportive Care
Serious adverse events111/34473/345
Clinical Biostats interpretation

The registry reports 111/344 participants affected by serious adverse events in the avelumab + BSC group and 73/345 in the BSC group. These are arm-specific affected/at-risk counts as reported in the ClinicalTrials.gov record.

These safety counts should be interpreted separately from the OS hazard ratio. Efficacy and safety answer different statistical questions, and a serious-adverse-event count does not provide a direct measure of the treatment effect on overall survival.

The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or P-value for these serious adverse-event counts, so none is inferred.

12. Planned Analysis and Interim-Analysis Considerations

The trial data state that an avelumab intervention was recorded following the planned interim analysis. The ClinicalTrials.gov record does not provide the interim-analysis boundary, alpha-spending function, information fraction, stopping rule, or numerical interim-analysis results.

For a time-to-event superiority trial, a planned interim analysis can be used to evaluate accumulating evidence before the final analysis. If formal repeated-look monitoring is used, the statistical design generally needs to account for the possibility that the same accumulating data are examined more than once. The exact error-control method, however, is not specified in the ClinicalTrials.gov record and therefore is not attributed to this study here.

Interpretation boundary: The existence of a planned interim analysis does not by itself establish a particular alpha-spending method, stopping boundary, or multiplicity adjustment. Those design details should be taken directly from the protocol or statistical analysis plan when available.

13. Multiplicity, Stratification, and Other Design Features

The ClinicalTrials.gov record supports a primary superiority analysis and a separate secondary time-to-event analysis. They do not provide enough information to establish a formal multiplicity hierarchy, alpha allocation across endpoints, stratification factors, Bayesian methods, non-inferiority margins, or a prespecified imputation strategy.

Design topicWhat the ClinicalTrials.gov record supports
SuperiorityYes. The primary OS analysis is identified as a superiority analysis.
Non-inferiority marginNot reported in the ClinicalTrials.gov record.
RandomizationYes. Allocation is reported as randomized.
Factorial designNot reported; the design model is parallel.
Interim analysisA planned interim analysis is referenced by the intervention record.
Alpha spendingNot reported in the ClinicalTrials.gov record.
Multiplicity adjustmentNot reported in the ClinicalTrials.gov record.
Missing-data imputationNot reported in the ClinicalTrials.gov record.
Bayesian methodsNot reported in the ClinicalTrials.gov record.
Stratification factorsNot reported in the ClinicalTrials.gov record.
CrossoverNot reported in the ClinicalTrials.gov record.

This distinction is important because statistical methods should be attributed to the evidence available rather than inferred from what is common in similar trials. The reported Cox model and log-rank test are supported directly by the registry data; more specialized design claims are not.

14. Understanding the Primary Result

Relative effect

The OS hazard ratio of 0.69 indicates an estimated relative reduction in the instantaneous hazard of death under the fitted model. Expressed as a simple derived interpretation, 1 − 0.69 = 0.31, so the estimated hazard is approximately 31% lower in the avelumab + BSC group relative to BSC.

Precision

The reported 95% two-sided confidence interval extends from 0.556 to 0.863. This interval quantifies uncertainty around the estimated hazard ratio; it does not provide a range of outcomes for individual participants.

Statistical evidence

The reported P = 0.0005 is evidence from the specified superiority hypothesis test. It should not be interpreted as a probability that the null hypothesis is true, nor as a numerical measure of how large or clinically meaningful the effect is.

Model dependence

The hazard ratio is generated by a Cox proportional-hazards model. Therefore, the interpretation is conditional on the model framework, including the proportional-hazards assumption. A hazard ratio is not a replacement for a full description of survival probabilities over time.

15. Interpreting the Secondary Time-to-Deterioration Result

The secondary endpoint provides a useful example of why every hazard ratio must be interpreted in the context of its endpoint.

EndpointHazard ratio95% CI informationP-value
Overall Survival0.690.556–0.863, two-sided0.0005
Time to Deterioration1.26Lower bound 0.901, one-sided0.9130

The OS HR below 1 and the TTD HR above 1 are not directly comparable as if they represented the same event. Overall survival measures death from any cause, whereas the secondary endpoint measures time to deterioration under a specific symptom-score definition.

The contrast also illustrates why a trial should be interpreted endpoint by endpoint. A treatment can have different estimated effects on distinct clinical outcomes, and the direction of a hazard ratio depends on how the event itself is defined.

16. Important Limitations

17. Why This Trial Matters Statistically

JAVELIN Bladder 100 provides a compact example of how randomized clinical-trial evidence is translated into a time-to-event analysis. The primary outcome is overall survival, which requires methods that can incorporate different follow-up durations and censoring. The registry then combines a log-rank test for the treatment comparison with a Cox model for the hazard-ratio estimate.

The trial is also useful for understanding why statistical interpretation should not stop at the P-value. The reported OS result contains three distinct pieces of information: the hazard ratio describes the estimated relative event rate, the confidence interval describes uncertainty around that estimate, and the P-value addresses the specified hypothesis test.

The secondary time-to-deterioration analysis provides an additional teaching point. Its HR of 1.26 is not an alternative estimate of the OS effect; it is an estimate for a different event definition. Understanding the endpoint before interpreting the direction of a hazard ratio is therefore essential.

ConceptHow it appears in JAVELIN Bladder 100
RandomizationRandomized allocation in a parallel phase 3 design.
Intention-to-treat analysisThe full analysis set included all randomized participants.
Time-to-event endpointOverall survival was the registered primary endpoint.
Kaplan-Meier estimationThe registry states that OS analysis was performed using the Kaplan-Meier method.
Log-rank testReported for both the primary OS and secondary TTD analyses.
Hazard ratioReported for both time-to-event analyses.
Cox modelThe primary OS analysis used a Cox's Proportional Hazard model.
Confidence intervalOS has a two-sided 95% CI; TTD has a one-sided 95% CI with the registry-reported lower bound.
SuperiorityThe primary OS analysis is identified as a superiority hypothesis.
CensoringParticipants last known to be alive were censored at last contact for OS.
Interim analysisA planned interim analysis is referenced in the intervention data.
Safety analysisSerious adverse-event counts are reported by arm.

18. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The primary randomized comparison reports an OS HR of 0.69 with a 95% two-sided CI of 0.556–0.863 and P = 0.0005 under a log-rank/Cox time-to-event framework.

Clinical interpretation

The statistical result describes the relative time-to-death experience observed under the trial's analysis framework. The ClinicalTrials.gov record does not provide enough additional clinical outcome information to quantify absolute survival differences or median survival times.

This distinction prevents a common statistical error: treating a statistically significant hazard ratio as though it automatically provides every clinically relevant measure of benefit. The hazard ratio is one component of the evidence. Absolute survival probabilities, median survival, event counts, patient-reported outcomes, and other clinical measures would provide additional perspectives when available.

19. Sources

The numerical trial results and methodological statements on this page are restricted to the ClinicalTrials.gov record. The linked PubMed records are provided as the publications associated with the registry record; no additional numerical results from those publications are incorporated into this analysis.

20. Related Tutorials

Learn more about the methods used in this trial:

21. Related Statistical Calculators

Continue through the Clinical Biostats statistical library

Explore tutorials and calculators covering the survival-analysis methods used to interpret randomized clinical trials.

22. Record Summary

JAVELIN Bladder 100 provides a clear example of randomized time-to-event analysis. The phase 3 trial used randomized parallel allocation, a primary overall-survival endpoint, Kaplan-Meier estimation, a log-rank test, and a Cox proportional-hazards model. The reported primary analysis compared avelumab plus best supportive care with best supportive care in the full analysis set of randomized participants and produced an OS hazard ratio of 0.69, with a 95% two-sided confidence interval of 0.556–0.863 and P = 0.0005.

The secondary time-to-deterioration analysis demonstrates the importance of endpoint-specific interpretation. Its reported hazard ratio was 1.26, with a one-sided 95% confidence interval lower bound of 0.901 and P = 0.9130. Because this endpoint concerns deterioration defined by a specific symptom-score framework rather than death, its hazard ratio should not be treated as another estimate of the OS effect.

The trial is therefore particularly useful for teaching the relationship among randomization, time-to-event endpoints, Kaplan-Meier estimation, log-rank testing, Cox hazard ratios, confidence intervals, and P-values. It also illustrates why reported statistical results must be kept separate from unsupported assumptions about interim monitoring, multiplicity, stratification, missing-data methods, or crossover.

Clinical Biostats methodology: A trial-results page should distinguish the reported statistical evidence from educational interpretation. Hazard ratios, confidence intervals, P-values, censoring rules, and analysis populations should be interpreted in the context of the endpoint and model that generated them.