This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the information provided in the ClinicalTrials.gov record.
1. Trial at a Glance
JAVELIN Bladder 100 was a randomized, parallel, open-label phase 3 trial evaluating avelumab in patients with urothelial cancer. The registry reports an enrollment of 700 participants and a primary time-to-event endpoint of overall survival.
| Feature | JAVELIN Bladder 100 |
|---|---|
| Trial name | JAVELIN Bladder 100 |
| Brief title | A Study Of Avelumab In Patients With Locally Advanced Or Metastatic Urothelial Cancer (JAVELIN Bladder 100) |
| Phase | Phase 3 |
| Condition | Urothelial Cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 700 |
| Primary endpoint type | Time-to-event |
| Primary endpoint | Overall Survival (OS) |
| Primary analysis | Log-rank test; Cox proportional-hazards model for the reported hazard ratio |
| Sponsor | Pfizer |
| Status | Completed |
| ClinicalTrials.gov | NCT02603432 |
2. Clinical Question
The statistical question was whether avelumab plus best supportive care produced a different overall-survival experience from best supportive care alone in the randomized population, under a superiority framework.
Population
Patients enrolled in the phase 3 study of avelumab in patients with locally advanced or metastatic urothelial cancer.
Intervention
Avelumab plus best supportive care (BSC).
Comparator
Best supportive care.
Primary question
Does avelumab plus BSC improve overall survival relative to BSC under the prespecified superiority analysis?
3. Trial Design
Avelumab plus Best Supportive Care
- Avelumab, classified in the registry as a biological intervention.
- Best supportive care.
- Serious adverse events: 111 affected participants among 344 at risk.
Best Supportive Care
- Best supportive care.
- Serious adverse events: 73 affected participants among 345 at risk.
The registry also records an avelumab intervention following the planned interim analysis. The ClinicalTrials.gov record does not provide additional details about that intervention sequence, crossover rules, or the statistical consequences of the interim analysis, so those features are not inferred here.
4. Trial Timeline
Trial start
The registry lists 2016-04-25 as the study start date.
Primary completion
The registry lists 2019-10-21 as the primary completion date.
Registry status
The ClinicalTrials.gov record classifies the trial as completed and report results.
5. Endpoints
| Endpoint | Registry definition / time frame | Type |
|---|---|---|
| Overall Survival (OS) | From randomization to discontinuation from the study, death or date of censoring, whichever occurred first (for a maximum duration of 41 months) | Time-to-event |
| Time to Deterioration (TTD) Based on National Comprehensive Cancer Network- Functional Assessment of Cancer Therapy (NCCN-FACT) Bladder Symptom Index- 18 (FBlSI-18) Disease Related Symptoms-Physical Subscale (DRS-P) Scores | From randomization up to the 90-Day Follow-up Visit (maximum duration of up to 41 months) | Time-to-event |
6. Primary Endpoint: Overall Survival
Overall survival was defined as the time in months from the date of randomization to the date of death due to any cause. Participants last known to be alive were censored at the date of last contact. The registry states that the analysis was performed using the Kaplan-Meier method.
| Feature | Primary OS analysis |
|---|---|
| Analysis population | The full analysis set included all randomized participants. |
| Groups compared | Avelumab + Best Supportive Care (BSC) vs Best Supportive Care |
| Statistical test | Log Rank |
| Effect measure | Hazard Ratio (HR) |
| Hypothesis type | Superiority |
| Model | Cox's Proportional Hazard model |
| Confidence interval | 95%, two-sided |
Hazard ratio for overall survival
95% CI: 0.556–0.863 · P = 0.0005
Analysis population: full analysis set containing all randomized participants.
The reported OS hazard ratio of 0.69 means that, under the fitted Cox proportional-hazards model, the estimated instantaneous rate of death in the avelumab + BSC group was approximately 31% lower than the corresponding estimated rate in the BSC group.
That statement is a relative, model-based interpretation. An HR of 0.69 does not mean that 31% of patients were prevented from dying, that each individual patient experienced exactly a 31% reduction in risk, or that survival time for every patient was increased by 31%.
The 95% two-sided confidence interval of 0.556–0.863 describes statistical uncertainty around the estimated hazard ratio. It is not an interval containing the survival experience of individual patients. Its interpretation also depends on the model and censoring assumptions underlying the analysis.
The P-value of 0.0005 addresses the evidence against the null hypothesis within the reported superiority analysis. It does not measure the magnitude of the treatment effect, clinical importance, or the probability that the treatment is effective.
Because the hazard ratio came from a Cox proportional-hazards model, interpretation also depends on the proportional-hazards framework. A single HR summarizes the relative event rate under that model rather than providing a complete description of how treatment effects may evolve over time.
Why the primary analysis combines Kaplan-Meier, log-rank, and Cox methods
These methods answer related but distinct questions. The Kaplan-Meier method estimates the survival experience over time while accommodating right censoring. The log-rank test provides a formal comparison of the time-to-event distributions between randomized groups. The Cox model provides the reported hazard ratio, translating the treatment comparison into a relative hazard measure.
Here, di represents the number of events at time ti, while ni is the number of participants at risk immediately before that time.
7. Secondary Endpoint: Time to Deterioration
The registry reports a secondary time-to-event analysis for time to deterioration based on the NCCN-FACT Bladder Symptom Index-18 (FBlSI-18) Disease Related Symptoms-Physical Subscale (DRS-P) Scores.
| Feature | Secondary TTD analysis |
|---|---|
| Analysis population | The full analysis set included all randomized participants. |
| Groups compared | Avelumab + Best Supportive Care (BSC) vs Best Supportive Care |
| Statistical test | Log Rank |
| Effect measure | Cox Proportional Hazard |
| Confidence interval | 95%, one-sided |
| Hypothesis type | Superiority |
| Time frame | From randomization up to the 90-Day Follow-up Visit (maximum duration of up to 41 months) |
Hazard ratio for time to deterioration
95% CI lower bound: 0.901 · P = 0.9130
The registry reports a one-sided 95% confidence interval and a Cox proportional-hazard effect measure.
The reported hazard ratio of 1.26 is above 1. Under the Cox model, this corresponds to a higher estimated instantaneous hazard of the deterioration event in the avelumab + BSC group relative to the BSC group for this endpoint.
The estimate should not be converted into a statement that patients were “26% worse” or that every patient had a 26% higher probability of deterioration. A hazard ratio is a model-based relative measure of event rates over time, not an individual-level probability.
The registry reports a one-sided 95% confidence interval with a lower bound of 0.901. Because only the lower bound is reported in the ClinicalTrials.gov record, no upper confidence limit is inferred here.
The P-value of 0.9130 is a measure used in the reported hypothesis-testing framework; it is not an effect-size measure. It should not be read as the probability that the null hypothesis is true.
As with the OS analysis, interpretation of the Cox hazard ratio relies on the proportional-hazards model and on appropriate handling of censoring and event times.
8. Statistical Methodology
Kaplan-Meier estimation
The primary OS endpoint is a time-to-event variable. The registry states that overall survival was analyzed using the Kaplan-Meier method. This method is designed for situations in which participants may have different follow-up times and some observations are censored before an event occurs.
For OS, a participant last known to be alive contributes information through the date of last contact and is then censored. This is fundamentally different from treating every participant without an observed death as if the participant had survived for the entire study period.
Log-rank testing
The registry reports a Log Rank method for both the primary OS analysis and the secondary time-to-deterioration analysis. The log-rank test compares the observed and expected numbers of events between treatment groups across event times, providing a formal hypothesis test for differences in survival experience.
The test uses the ordering and timing of events while accounting for the number of participants remaining at risk at each event time.
Cox proportional-hazards model
The registry states that the primary OS analysis was performed using a Cox's Proportional Hazard model. The secondary time-to-deterioration analysis likewise reports a Cox proportional-hazard effect measure. The Cox model estimates the relative hazard between groups while leaving the baseline hazard unspecified.
HR = 1 → equal estimated instantaneous event rates
HR > 1 → higher estimated instantaneous event rate in the treatment group
These interpretations describe relative event rates under the Cox model. They are not equivalent to absolute risk differences or ratios of median survival times.
Intention-to-treat principle
The registry identifies the full analysis set as including all randomized participants for both posted analyses. This is consistent with an intention-to-treat approach in which the primary treatment comparison remains anchored to randomized assignment rather than being restricted to participants who completed treatment.
Right censoring
Time-to-event analyses require a rule for participants whose event has not occurred by the time their available follow-up ends. The OS definition explicitly states that participants last known to be alive were censored at the date of last contact. Proper censoring allows their observed follow-up to contribute information without pretending that an event occurred after follow-up ended.
Superiority testing
The primary OS analysis is identified as a superiority analysis. The inferential question is therefore whether the randomized groups differ in the direction specified by the superiority framework, rather than whether one treatment is merely no worse than another by a prespecified non-inferiority margin.
9. Statistical Methods Explained
Why was a log-rank test used?
Overall survival is a time-to-event endpoint, and not every participant necessarily experiences the event during the period of observation. The log-rank test is designed to compare survival distributions while using the available follow-up and accounting for censoring. It therefore fits the structure of the registered OS endpoint more directly than a simple comparison of proportions.
What does an OS hazard ratio of 0.69 mean?
Under the reported Cox model, an HR of 0.69 corresponds to an estimated instantaneous death rate that is approximately 31% lower in the avelumab + BSC group than in the BSC group. It does not mean that 31% of participants survived because of treatment or that each participant's individual risk was reduced by exactly 31%.
Why is the confidence interval important?
The point estimate is only one summary of the randomized comparison. The 95% two-sided confidence interval of 0.556–0.863 shows the statistical uncertainty surrounding the estimated OS hazard ratio under the analysis framework. A narrower interval would generally indicate greater precision than a wider interval, although precision is not the same thing as clinical importance.
Why does the P-value not measure treatment effect size?
The P-value is tied to a hypothesis-testing procedure. It describes how compatible the observed data are with the null hypothesis under that procedure. The effect size is represented here by the hazard ratio, while the confidence interval provides information about its uncertainty. A very small P-value can occur with a modest effect in a sufficiently informative study, and a large effect estimate can have substantial uncertainty in a small or event-limited study.
Why is censoring important for overall survival?
Some participants may still be alive when their available follow-up ends. Rather than treating them as if their eventual survival time were known, the analysis censors them at the last known date alive. This allows the participant's observed follow-up to contribute information without assigning an unobserved death time.
What is the difference between the OS and TTD hazard ratios?
The OS analysis reports an HR of 0.69, while the secondary TTD analysis reports an HR of 1.26. These are estimates for different endpoints. The first concerns death from any cause; the second concerns deterioration defined using the specified FBlSI-18 DRS-P score framework. The two estimates therefore should not be interpreted as contradictory measurements of the same outcome.
10. Analysis Population and Interpretation
Both posted statistical analyses use the full analysis set, which the registry defines as including all randomized participants. This is important because randomized treatment assignment is the foundation of the comparative inference.
| Principle | Application in the reported analyses |
|---|---|
| Randomization | Participants were randomly allocated to the trial arms. |
| Analysis population | The full analysis set included all randomized participants. |
| Primary endpoint | Overall survival, a time-to-event endpoint. |
| Secondary endpoint | Time to deterioration based on the specified FBlSI-18 DRS-P score. |
| Survival estimation | Kaplan-Meier method for the primary OS endpoint. |
| Hypothesis testing | Log-rank testing. |
| Effect estimation | Hazard ratio from a Cox proportional-hazards model. |
| Primary hypothesis | Superiority. |
The use of all randomized participants strengthens the connection between the reported efficacy comparison and the original randomized treatment assignment. It also means that treatment discontinuation, treatment exposure, and subsequent events do not automatically remove participants from the efficacy analysis.
11. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm. These figures are presented as affected participants divided by participants at risk.
| Safety measure | Avelumab + BSC | Best Supportive Care |
|---|---|---|
| Serious adverse events | 111/344 | 73/345 |
The registry reports 111/344 participants affected by serious adverse events in the avelumab + BSC group and 73/345 in the BSC group. These are arm-specific affected/at-risk counts as reported in the ClinicalTrials.gov record.
These safety counts should be interpreted separately from the OS hazard ratio. Efficacy and safety answer different statistical questions, and a serious-adverse-event count does not provide a direct measure of the treatment effect on overall survival.
The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or P-value for these serious adverse-event counts, so none is inferred.
12. Planned Analysis and Interim-Analysis Considerations
The trial data state that an avelumab intervention was recorded following the planned interim analysis. The ClinicalTrials.gov record does not provide the interim-analysis boundary, alpha-spending function, information fraction, stopping rule, or numerical interim-analysis results.
For a time-to-event superiority trial, a planned interim analysis can be used to evaluate accumulating evidence before the final analysis. If formal repeated-look monitoring is used, the statistical design generally needs to account for the possibility that the same accumulating data are examined more than once. The exact error-control method, however, is not specified in the ClinicalTrials.gov record and therefore is not attributed to this study here.
13. Multiplicity, Stratification, and Other Design Features
The ClinicalTrials.gov record supports a primary superiority analysis and a separate secondary time-to-event analysis. They do not provide enough information to establish a formal multiplicity hierarchy, alpha allocation across endpoints, stratification factors, Bayesian methods, non-inferiority margins, or a prespecified imputation strategy.
| Design topic | What the ClinicalTrials.gov record supports |
|---|---|
| Superiority | Yes. The primary OS analysis is identified as a superiority analysis. |
| Non-inferiority margin | Not reported in the ClinicalTrials.gov record. |
| Randomization | Yes. Allocation is reported as randomized. |
| Factorial design | Not reported; the design model is parallel. |
| Interim analysis | A planned interim analysis is referenced by the intervention record. |
| Alpha spending | Not reported in the ClinicalTrials.gov record. |
| Multiplicity adjustment | Not reported in the ClinicalTrials.gov record. |
| Missing-data imputation | Not reported in the ClinicalTrials.gov record. |
| Bayesian methods | Not reported in the ClinicalTrials.gov record. |
| Stratification factors | Not reported in the ClinicalTrials.gov record. |
| Crossover | Not reported in the ClinicalTrials.gov record. |
This distinction is important because statistical methods should be attributed to the evidence available rather than inferred from what is common in similar trials. The reported Cox model and log-rank test are supported directly by the registry data; more specialized design claims are not.
14. Understanding the Primary Result
The OS hazard ratio of 0.69 indicates an estimated relative reduction in the instantaneous hazard of death under the fitted model. Expressed as a simple derived interpretation, 1 − 0.69 = 0.31, so the estimated hazard is approximately 31% lower in the avelumab + BSC group relative to BSC.
The reported 95% two-sided confidence interval extends from 0.556 to 0.863. This interval quantifies uncertainty around the estimated hazard ratio; it does not provide a range of outcomes for individual participants.
The reported P = 0.0005 is evidence from the specified superiority hypothesis test. It should not be interpreted as a probability that the null hypothesis is true, nor as a numerical measure of how large or clinically meaningful the effect is.
The hazard ratio is generated by a Cox proportional-hazards model. Therefore, the interpretation is conditional on the model framework, including the proportional-hazards assumption. A hazard ratio is not a replacement for a full description of survival probabilities over time.
15. Interpreting the Secondary Time-to-Deterioration Result
The secondary endpoint provides a useful example of why every hazard ratio must be interpreted in the context of its endpoint.
| Endpoint | Hazard ratio | 95% CI information | P-value |
|---|---|---|---|
| Overall Survival | 0.69 | 0.556–0.863, two-sided | 0.0005 |
| Time to Deterioration | 1.26 | Lower bound 0.901, one-sided | 0.9130 |
The OS HR below 1 and the TTD HR above 1 are not directly comparable as if they represented the same event. Overall survival measures death from any cause, whereas the secondary endpoint measures time to deterioration under a specific symptom-score definition.
The contrast also illustrates why a trial should be interpreted endpoint by endpoint. A treatment can have different estimated effects on distinct clinical outcomes, and the direction of a hazard ratio depends on how the event itself is defined.
16. Important Limitations
- Registry-level information: the ClinicalTrials.gov record contains the primary OS result and one secondary time-to-deterioration analysis but do not provide a complete statistical analysis plan.
- Incomplete endpoint time-frame text: the registry-reported primary OS time frame ends with “for a maximu”, so the missing registry continuation is not reconstructed.
- Hazard-ratio assumptions: both reported hazard-ratio interpretations depend on the Cox proportional-hazards framework.
- Censoring: time-to-event analyses depend on appropriate handling and interpretation of censored observations. The registry explicitly describes last-contact censoring for participants last known to be alive for OS.
- Secondary endpoint interpretation: the TTD analysis has a one-sided confidence interval in the ClinicalTrials.gov record, with only its lower bound reported. No upper bound is inferred.
- Multiplicity: the ClinicalTrials.gov record does not specify how type I error was allocated across the primary and secondary analyses or across any interim looks.
- Interim analysis: a planned interim analysis is referenced, but its formal statistical boundaries and error-control method are not reported.
- Safety comparison: serious adverse-event counts are available by arm, but no formal comparative safety analysis is reported.
- Missing-data methods: the ClinicalTrials.gov record does not specify an imputation strategy for missing assessments.
- Generalizability: the ClinicalTrials.gov recordset identifies the condition and trial design but does not provide a detailed baseline-characteristic table from which broader population comparisons could be made.
17. Why This Trial Matters Statistically
JAVELIN Bladder 100 provides a compact example of how randomized clinical-trial evidence is translated into a time-to-event analysis. The primary outcome is overall survival, which requires methods that can incorporate different follow-up durations and censoring. The registry then combines a log-rank test for the treatment comparison with a Cox model for the hazard-ratio estimate.
The trial is also useful for understanding why statistical interpretation should not stop at the P-value. The reported OS result contains three distinct pieces of information: the hazard ratio describes the estimated relative event rate, the confidence interval describes uncertainty around that estimate, and the P-value addresses the specified hypothesis test.
The secondary time-to-deterioration analysis provides an additional teaching point. Its HR of 1.26 is not an alternative estimate of the OS effect; it is an estimate for a different event definition. Understanding the endpoint before interpreting the direction of a hazard ratio is therefore essential.
| Concept | How it appears in JAVELIN Bladder 100 |
|---|---|
| Randomization | Randomized allocation in a parallel phase 3 design. |
| Intention-to-treat analysis | The full analysis set included all randomized participants. |
| Time-to-event endpoint | Overall survival was the registered primary endpoint. |
| Kaplan-Meier estimation | The registry states that OS analysis was performed using the Kaplan-Meier method. |
| Log-rank test | Reported for both the primary OS and secondary TTD analyses. |
| Hazard ratio | Reported for both time-to-event analyses. |
| Cox model | The primary OS analysis used a Cox's Proportional Hazard model. |
| Confidence interval | OS has a two-sided 95% CI; TTD has a one-sided 95% CI with the registry-reported lower bound. |
| Superiority | The primary OS analysis is identified as a superiority hypothesis. |
| Censoring | Participants last known to be alive were censored at last contact for OS. |
| Interim analysis | A planned interim analysis is referenced in the intervention data. |
| Safety analysis | Serious adverse-event counts are reported by arm. |
18. Statistical Interpretation vs Clinical Interpretation
Statistical interpretation
The primary randomized comparison reports an OS HR of 0.69 with a 95% two-sided CI of 0.556–0.863 and P = 0.0005 under a log-rank/Cox time-to-event framework.
Clinical interpretation
The statistical result describes the relative time-to-death experience observed under the trial's analysis framework. The ClinicalTrials.gov record does not provide enough additional clinical outcome information to quantify absolute survival differences or median survival times.
This distinction prevents a common statistical error: treating a statistically significant hazard ratio as though it automatically provides every clinically relevant measure of benefit. The hazard ratio is one component of the evidence. Absolute survival probabilities, median survival, event counts, patient-reported outcomes, and other clinical measures would provide additional perspectives when available.
19. Sources
- ClinicalTrials.gov: JAVELIN Bladder 100, NCT02603432.
- Linked publication: PubMed record, PMID 38728337.
- Linked publication: PubMed record, PMID 37884606.
- Linked publication: PubMed record, PMID 37149458.
- Linked publication: PubMed record, PMID 37121850.
- Linked publication: PubMed record, PMID 37071838.
The numerical trial results and methodological statements on this page are restricted to the ClinicalTrials.gov record. The linked PubMed records are provided as the publications associated with the registry record; no additional numerical results from those publications are incorporated into this analysis.
20. Related Tutorials
Learn more about the methods used in this trial:
21. Related Statistical Calculators
Continue through the Clinical Biostats statistical library
Explore tutorials and calculators covering the survival-analysis methods used to interpret randomized clinical trials.
22. Record Summary
JAVELIN Bladder 100 provides a clear example of randomized time-to-event analysis. The phase 3 trial used randomized parallel allocation, a primary overall-survival endpoint, Kaplan-Meier estimation, a log-rank test, and a Cox proportional-hazards model. The reported primary analysis compared avelumab plus best supportive care with best supportive care in the full analysis set of randomized participants and produced an OS hazard ratio of 0.69, with a 95% two-sided confidence interval of 0.556–0.863 and P = 0.0005.
The secondary time-to-deterioration analysis demonstrates the importance of endpoint-specific interpretation. Its reported hazard ratio was 1.26, with a one-sided 95% confidence interval lower bound of 0.901 and P = 0.9130. Because this endpoint concerns deterioration defined by a specific symptom-score framework rather than death, its hazard ratio should not be treated as another estimate of the OS effect.
The trial is therefore particularly useful for teaching the relationship among randomization, time-to-event endpoints, Kaplan-Meier estimation, log-rank testing, Cox hazard ratios, confidence intervals, and P-values. It also illustrates why reported statistical results must be kept separate from unsupported assumptions about interim monitoring, multiplicity, stratification, missing-data methods, or crossover.