← Clinical Trials
Renal Cell Cancer Phase 3 Time-to-Event Analysis NCT02684006

JAVELIN Renal 101: Complete Statistical Analysis of Avelumab + Axitinib in Renal Cell Cancer

An independent statistical analysis of the randomized phase 3 JAVELIN Renal 101 trial comparing avelumab plus axitinib with sunitinib in advanced renal cell cancer, with emphasis on progression-free survival, overall survival, hazard ratios, confidence intervals, and log-rank testing.

Randomized  ·  Parallel design  ·  886 enrolled  ·  Completed
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the trial data posted on ClinicalTrials.gov for JAVELIN Renal 101. The official registry record is available at ClinicalTrials.gov.

Independent analysis: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. View the JAVELIN Renal 101 registry record.

1. Trial at a Glance

JAVELIN Renal 101 was a randomized, parallel, open-label phase 3 trial evaluating avelumab with axitinib versus sunitinib in advanced renal cell cancer. The registry reports 886 enrolled participants, two treatment arms, two primary time-to-event endpoints, and six posted statistical analyses.

886
Enrolled
Phase 3
2
Treatment arms
Randomized, parallel
0.61
Primary PFS HR
95% CI 0.475–0.790
0.86
Primary OS HR
95% CI 0.701–1.057
FeatureJAVELIN Renal 101
Trial nameJAVELIN Renal 101
Brief titleA Study of Avelumab With Axitinib Versus Sunitinib In Advanced Renal Cell Cancer (JAVELIN Renal 101)
PhasePhase 3
StatusCompleted
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment886
Primary endpoints2
Primary endpoint typeTime-to-event
Effect measureHazard ratio
Primary hypothesis typeSuperiority
Statistical analyses posted6
Lead sponsorPfizer
Start2016-03-23
Primary completion2023-08-31

2. Clinical Question

The central statistical question was whether treatment with avelumab plus axitinib differed from sunitinib with respect to the registered primary time-to-event endpoints in participants with renal cell cancer.

Population

Participants in the phase 3 JAVELIN Renal 101 trial with renal cell cancer. The primary endpoint analyses were performed in the subset of randomized participants who were PD-L1 positive.

Intervention

Avelumab (MSB0010718C) plus axitinib (AG-013736).

Comparator

Sunitinib.

Primary question

Does avelumab plus axitinib produce a different time-to-event outcome from sunitinib for the registered primary endpoints, under a superiority framework?

3. Trial Design

01
Randomize886 enrolled
02
Two armsAvelumab + axitinib or sunitinib
03
FollowTime-to-event outcomes
04
AssessPFS / OS
05
AnalyzeLog-rank / hazard ratio
ARM A

Avelumab + Axitinib

  • Avelumab (MSB0010718C)
  • Axitinib (AG-013736)
ARM B

Sunitinib

  • Sunitinib
Allocation
Randomized
Design model
Parallel
Masking
None
Primary purpose
Treatment

The trial's statistical profile is dominated by time-to-event analysis. Both registered primary endpoints are defined from randomization to an event or censoring, and the posted formal analyses use a log-rank test with a hazard ratio as the effect measure.

4. Endpoints

The registry identifies two primary endpoints and provides additional secondary time-to-event analyses. The wording and time frames below follow the ClinicalTrials.gov record.

EndpointRegistry time frameTypePopulation
Progression Free Survival (PFS) as Assessed by Blinded Independent Central Review (BICR) in Programmed Death-Ligand 1 (PD-L1) Positive Participants From date of randomization to the first documentation of PD or death due to any cause or censoring date, whichever occurred first (maximum up to approximately 26 months) Time-to-event Subset of randomized participants who were PD-L1 positive
Overall Survival (OS) in PD-L1 Positive Participants From the date of randomization to the date of death due to any cause or censoring date, whichever occurred first (maximum up to approximately 89 months) Time-to-event Subset of randomized participants who were PD-L1 positive
PFS as Assessed by BICR in Participants Irrespective of PD-L1 Expression From date of randomization to the first documentation of PD or death due to any cause or censoring date, whichever occurred first (maximum up to approximately 26 months) Time-to-event All randomized participants
OS in Participants Irrespective of PD-L1 Expression From date of randomization to the date of death due to any cause or censoring date, whichever occurred first (maximum up to approximately 89 months) Time-to-event All randomized participants
PFS as Assessed by Investigator in Participants Irrespective of PD-L1 Expression From date of randomization until PD, whichever occurred first (maximum up to approximately 89 months) Time-to-event All randomized participants
Progression-Free Survival on Next-line Therapy (PFS2) in Participants Irrespective of PD-L1 Expression From date of randomization until PD or death, whichever occurred first (maximum up to approximately 89 months) Time-to-event All randomized participants
Registry endpoint definition: The registry definition for PFS states that PFS is the time from randomization to the first documentation of progressive disease according to RECIST v1.1 or death due to any cause, whichever occurred first, as assessed by BICR. The registry-reported OS definition states that OS is the time from randomization to death due to any cause, with participants last known to be alive censored at the date of last contact; the analysis was performed using the Kaplan-Meier method.

5. Analysis Populations

The registry distinguishes the full analysis set from the subset used for the PD-L1-positive primary analyses.

Endpoint groupAnalysis population
Primary PFS and OS FAS included all participants who were randomized. Analysis was performed on the subset of randomized participants who were PD-L1 positive.
Secondary PFS and OS irrespective of PD-L1 expression FAS included all randomized participants.
Investigator-assessed PFS irrespective of PD-L1 expression FAS included all randomized participants.
PFS2 irrespective of PD-L1 expression FAS included all randomized participants.

This distinction is statistically important. The primary PFS and OS estimates do not describe exactly the same analysis population as the secondary analyses that include participants irrespective of PD-L1 expression. Consequently, the corresponding hazard ratios should be read as estimates from different populations rather than as interchangeable measures of one treatment effect.

6. Primary Results: Progression-Free Survival

The primary PFS endpoint was assessed by blinded independent central review in PD-L1-positive participants. The registry reports a formal log-rank analysis with a hazard ratio comparing avelumab plus axitinib with sunitinib.

Primary PFS hazard ratio

0.61

95% CI: 0.475–0.790   ·   P = 0.0001

Analysis: log-rank test  ·  Hypothesis: superiority

FeatureReported result
EndpointProgression Free Survival (PFS) as Assessed by Blinded Independent Central Review (BICR) in Programmed Death-Ligand 1 (PD-L1) Positive Participants
Groups comparedAvelumab + Axitinib vs Sunitinib
Effect measureHazard Ratio (HR)
Estimate0.61
95% CI0.475–0.790
P-value0.0001
Hypothesis typeSuperiority
Clinical Biostats interpretation

An HR of 0.61 means that the estimated instantaneous rate of the PFS event was about 61% as high in the avelumab-plus-axitinib group as in the sunitinib group under the time-to-event model used for the comparison. Expressed as a relative difference, this corresponds to an estimated 39% lower hazard of the event.

The HR does not mean that 39% of patients avoided progression, nor does it mean that every patient experienced exactly a 39% reduction in risk. It is a relative, model-based time-to-event measure.

The 95% CI of 0.475–0.790 describes uncertainty around the estimated hazard ratio under the statistical model and sampling framework. It does not describe the range of individual patient outcomes.

The P-value of 0.0001 addresses the evidence against the null hypothesis in the reported testing framework. It is not a measure of the magnitude of the treatment effect. The magnitude is conveyed by the HR and its confidence interval.

Because PFS is a time-to-event endpoint, censoring is part of the analysis. Interpretation of a Cox-type hazard ratio also ordinarily depends on the proportional-hazards assumption if a single HR is used as a summary of the treatment effect over follow-up.

7. Primary Results: Overall Survival

The second primary endpoint was overall survival in PD-L1-positive participants. OS was defined in the ClinicalTrials.gov record as time from randomization to death from any cause, with participants last known to be alive censored at their last contact.

Primary OS hazard ratio

0.86

95% CI: 0.701–1.057   ·   P = 0.1509

Analysis: log-rank test  ·  Hypothesis: superiority

FeatureReported result
EndpointOverall Survival (OS) in PD-L1 Positive Participants
Groups comparedAvelumab + Axitinib vs Sunitinib
Effect measureHazard Ratio (HR)
Estimate0.86
95% CI0.701–1.057
P-value0.1509
Hypothesis typeSuperiority
Clinical Biostats interpretation

An HR of 0.86 means that the estimated instantaneous rate of death in the avelumab-plus-axitinib group was 86% of that in the sunitinib group under the reported time-to-event comparison. As a relative hazard measure, this corresponds to an estimated 14% lower hazard of death.

The estimate should not be converted into a statement that 14% of patients survived longer or that individual patients experienced a 14% reduction in mortality risk. A hazard ratio summarizes the relative event rate over the analyzed follow-up.

The 95% CI of 0.701–1.057 crosses 1.00. Thus, the interval includes hazard-ratio values representing both a lower and a higher estimated hazard for the avelumab-plus-axitinib group relative to sunitinib.

The P-value of 0.1509 is a measure of statistical evidence against the null hypothesis in the reported superiority test. It does not quantify the size or clinical importance of the observed HR.

OS also differs conceptually from PFS because death is the event and participants who remain alive are censored at their last known contact. Interpretation therefore depends on the completeness and timing of follow-up as well as the treatment comparison itself.

8. Primary Endpoints Side by Side

Primary endpointHR95% CIP-valueMethod
PFS in PD-L1-positive participants 0.61 0.475–0.790 0.0001 Log-rank
OS in PD-L1-positive participants 0.86 0.701–1.057 0.1509 Log-rank

These two endpoints illustrate why a clinical trial should not be reduced to a single P-value. PFS and OS answer different time-to-event questions, and the ClinicalTrials.gov record produce different effect estimates and different levels of statistical evidence.

PFS

The reported HR of 0.61 is below 1, with a 95% CI of 0.475–0.790 and P = 0.0001.

OS

The reported HR of 0.86 is below 1, but its 95% CI of 0.701–1.057 includes 1 and the reported P-value is 0.1509.

9. Secondary Results

The registry also reports four additional secondary analyses in the ClinicalTrials.gov record. These extend the time-to-event assessment beyond the PD-L1-positive primary population.

Secondary endpointHR95% CIP-valueMethod
PFS by BICR, irrespective of PD-L1 expression 0.69 0.563–0.840 0.0002 Log-rank
OS, irrespective of PD-L1 expression 0.88 0.749–1.039 0.1338 Log-rank
PFS by investigator, irrespective of PD-L1 expression 0.66 0.565–0.768 <.0001 Log-rank
PFS2, irrespective of PD-L1 expression 0.64 0.551–0.754 Not reported Not reported

Secondary PFS by BICR

Hazard ratio

0.69

95% CI: 0.563–0.840   ·   P = 0.0002

In the all-randomized population irrespective of PD-L1 expression, the BICR-assessed PFS HR was 0.69. This is a relative hazard estimate corresponding to an approximately 31% lower estimated hazard of the PFS event, not a statement about the percentage of patients who benefit.

Secondary OS irrespective of PD-L1 expression

Hazard ratio

0.88

95% CI: 0.749–1.039   ·   P = 0.1338

The OS HR of 0.88 corresponds to an estimated 12% lower hazard of death under the reported model. The 95% CI of 0.749–1.039 includes 1.00, so the interval is compatible with both a lower and a higher hazard for the avelumab-plus-axitinib group relative to sunitinib.

Secondary investigator-assessed PFS

Hazard ratio

0.66

95% CI: 0.565–0.768   ·   P < 0.0001

The investigator-assessed PFS result produced an HR of 0.66, corresponding to an approximately 34% lower estimated hazard of the PFS event. This is a secondary endpoint and uses investigator assessment rather than BICR, so it should be distinguished from the primary BICR-based PFS analysis.

Progression-Free Survival on Next-line Therapy

Reported hazard ratio

0.64

95% CI: 0.551–0.754   ·   P-value not reported

The registry reports an HR of 0.64 for PFS2, with a 95% CI of 0.551–0.754. The ClinicalTrials.gov record does not report a formal analysis method or P-value for this result. Accordingly, the HR and confidence interval can be described, but no additional hypothesis-testing interpretation should be attributed to the registry.

Do not treat the secondary results as interchangeable with the primary results. The analyses differ in endpoint definition, population, assessment method, or reported statistical information. In particular, the PFS2 entry has no reported analysis method or P-value in the ClinicalTrials.gov record.

10. Safety: Serious Adverse Events

The ClinicalTrials.gov record reports serious adverse events by treatment arm as affected participants divided by participants at risk.

Treatment groupAffected / at risk
Avelumab + Axitinib 231 / 434
Sunitinib 166 / 439

These figures should be kept in their reported form. They describe the number of participants affected among those at risk for the registry-reported serious-adverse-event measure; they are not a hazard ratio and are not a time-to-event efficacy endpoint.

Statistical distinction: efficacy and safety answer different questions. The primary PFS and OS comparisons are anchored to randomized treatment assignment, whereas an adverse-event count describes observed safety experience in the reported treatment groups. A serious-adverse-event frequency should therefore not be interpreted as another estimate of treatment efficacy.

11. Statistical Methodology

Time-to-event analysis

All two primary endpoints and the registry-reported secondary efficacy endpoints are classified as time-to-event outcomes. Rather than recording only whether an event occurred, these analyses incorporate the time from randomization to the event or to censoring.

Generic time-to-event structure
Time-to-event = date of event − date of randomization

When an event has not occurred during the available observation period, the participant can contribute follow-up until the prespecified censoring point.

This structure is particularly important for PFS because participants can have different lengths of follow-up. A simple proportion progressing would discard information about when progression occurred and would not handle censoring in the same way as a survival analysis.

Kaplan-Meier estimation

The registry-reported OS definition explicitly states that analysis was performed using the Kaplan-Meier method. Kaplan-Meier estimation constructs an estimated survival function while accounting for right-censored observations.

Kaplan-Meier estimator
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at time ti, while ni represents the number at risk immediately before that event time.

Log-rank test

The registry reports a log-rank test for the formal primary PFS and OS analyses and for the reported secondary PFS and OS analyses. The log-rank test compares the observed and expected numbers of events between treatment groups across event times.

Conceptually, the test asks whether the observed pattern of events over follow-up is compatible with the null hypothesis that the groups have the same underlying survival experience.

Hazard ratio

The effect measure reported for the statistical analyses is the hazard ratio. For these comparisons, an HR below 1 indicates a lower estimated instantaneous event rate in the avelumab-plus-axitinib group relative to sunitinib, while an HR above 1 indicates a higher estimated event rate.

Interpretation
HR < 1  →  lower estimated instantaneous event rate in the treatment group

The hazard ratio is not a relative risk, not a risk difference, and not the percentage of participants who benefit.

Confidence intervals

The primary analyses use two-sided 95% confidence intervals. The interval gives a range of values reflecting statistical uncertainty around the estimated hazard ratio under the analysis framework. It should not be interpreted as a range containing 95% of individual patient effects.

Superiority testing

The statistical analyses posted on ClinicalTrials.gov identify the hypothesis type as superiority. Thus, the inferential question is whether the treatment groups differ in the direction specified by the superiority framework rather than whether one treatment is merely no worse than another within a non-inferiority margin.

12. Statistical Methods Explained

Why is a time-to-event analysis appropriate for PFS?

PFS incorporates both whether a participant experiences progression or death and when that event occurs. Participants without an event by their last adequate assessment can be censored, allowing the analysis to use unequal follow-up times rather than treating all participants as if they had identical observation periods.

What does an HR of 0.61 mean?

An HR of 0.61 means the estimated instantaneous rate of the event is 61% as high in the avelumab-plus-axitinib group as in the sunitinib group under the fitted time-to-event comparison. It does not mean that 61% of participants had an event or that every participant experienced the same relative effect.

Why use a log-rank test?

The log-rank test is designed for comparing time-to-event distributions between groups while accounting for the timing of events and censoring. It is therefore naturally aligned with PFS and OS rather than with simple binary outcomes.

Why is the confidence interval important when the HR is below 1?

The point estimate is only one estimate from the observed data. The 95% CI shows how precisely the hazard ratio has been estimated. For primary PFS, the interval is 0.475–0.790; for primary OS, it is 0.701–1.057. The latter includes 1.00, which is the no-difference value for a hazard ratio.

Why does the P-value not tell us the size of the treatment effect?

A P-value quantifies statistical evidence against a null hypothesis under the specified testing framework. It is affected by the amount of information in the analysis and does not directly describe how large or clinically important the treatment effect is. The HR and its confidence interval provide the effect estimate and its precision.

Why should the primary and secondary populations be kept separate?

The primary analyses were performed in PD-L1-positive randomized participants, whereas several secondary analyses included all randomized participants irrespective of PD-L1 expression. Combining these results without identifying their populations would make the statistical interpretation less precise.

What caution applies to interpreting a single hazard ratio?

A single HR summarizes a time-to-event comparison under a model-based framework. Its usual interpretation as a relative hazard over follow-up is most straightforward when the proportional-hazards assumption is reasonable. The HR should therefore be considered alongside the underlying survival distributions and censoring pattern when those are available.

13. Understanding the Primary PFS Result

Effect size

The primary PFS HR of 0.61 is below 1. On the relative hazard scale, the estimated event rate was approximately 39% lower in the avelumab-plus-axitinib group than in the sunitinib group under the reported comparison.

Precision

The 95% CI of 0.475–0.790 provides a measure of uncertainty around the point estimate. It remains below 1.00 throughout the reported interval, so the estimated treatment effect is consistently on the lower-hazard side of the no-difference value within that interval.

Statistical evidence

The reported P-value is 0.0001. This is evidence against the null hypothesis under the reported superiority testing framework. It should not be read as the probability that the null hypothesis is true, nor as the probability that the treatment effect is clinically important.

What the result cannot establish by itself

The HR does not provide an absolute increase in months of PFS, the proportion of participants who ultimately benefit, or the experience of any particular individual. Those questions require absolute survival estimates, event-time summaries, or individual-level information that is not contained in the ClinicalTrials.gov record.

14. Understanding the Primary OS Result

Effect size

The primary OS HR of 0.86 is below 1, corresponding to an estimated 14% lower hazard of death in the avelumab-plus-axitinib group relative to sunitinib under the reported comparison.

Precision

The 95% CI extends from 0.701 to 1.057. Because this interval crosses 1.00, the reported uncertainty includes values on both sides of the no-difference hazard ratio.

Statistical evidence

The reported P-value is 0.1509. It does not meet the conventional 0.05 threshold often used as a descriptive reference point, although the ClinicalTrials.gov record identifies the hypothesis as superiority and do not provide additional alpha-allocation details.

Do not equate a non-small P-value with proof of no effect. The OS estimate is 0.86, but the confidence interval includes 1.00. The statistically appropriate description is that the registry-reported analysis estimates a lower hazard, while the reported interval and P-value do not provide strong evidence against the null under the stated superiority framework.

15. Comparing BICR and Investigator-Assessed PFS

The ClinicalTrials.gov record contains two secondary PFS analyses in the population irrespective of PD-L1 expression: one assessed by BICR and one assessed by the investigator.

AssessmentHR95% CIP-value
BICR 0.69 0.563–0.840 0.0002
Investigator 0.66 0.565–0.768 <.0001

The estimates are numerically different but both are below 1. The difference between the estimates should not automatically be interpreted as evidence that one assessment method is biased or that the treatment effect truly differs by assessor. The two analyses involve different assessment processes, and the ClinicalTrials.gov record does not provide an agreement analysis or a formal comparison between their hazard ratios.

BICR role

The primary PFS endpoint was assessed by blinded independent central review, which separates the primary radiologic assessment from knowledge of treatment assignment.

Investigator role

The secondary investigator-assessed PFS provides a separate assessment of progression in the all-randomized population.

16. Censoring and What It Means for Interpretation

Censoring is fundamental to the interpretation of both PFS and OS. For PFS, the registry-reported definition identifies censoring as part of the endpoint time frame. For OS, participants last known to be alive are censored at their date of last contact.

Conceptual example
Observed follow-up = min(event time, censoring time)

A censored participant contributes information up to the censoring point, but the analysis does not treat that participant as having experienced the event after censoring.

The validity of a survival analysis therefore depends partly on whether the censoring mechanism is appropriately handled. A participant who leaves observation early is not equivalent to a participant who remains under observation without experiencing the event. Kaplan-Meier and related methods are designed to preserve the available information while accounting for this distinction.

17. Why the Analysis Population Matters

The primary endpoint results are based on the subset of randomized participants who were PD-L1 positive. Several secondary results are based on all randomized participants irrespective of PD-L1 expression.

QuestionPrimary analysisSeveral secondary analyses
Population Randomized participants who were PD-L1 positive All randomized participants
PFS assessment BICR BICR or investigator, depending on endpoint
Effect measure Hazard ratio Hazard ratio

This distinction is more than a reporting detail. A hazard ratio is always conditional on the population being analyzed. An HR estimated among PD-L1-positive participants and an HR estimated among all randomized participants answer related but different questions.

18. Multiplicity and Multiple Endpoints

The trial has two registered primary endpoints, both under a superiority hypothesis, plus multiple secondary analyses. Multiple endpoints create an important statistical-design issue because testing several hypotheses can increase the chance of observing at least one apparently positive result by chance if no appropriate error-control strategy is specified.

Two primary endpoints

PFS and OS are both designated primary endpoints. Their individual P-values should therefore be understood within the trial's prespecified multiplicity framework rather than interpreted in isolation.

Secondary endpoints

The ClinicalTrials.gov record identifies additional PFS, OS, investigator-assessed PFS, and PFS2 analyses. Their role is distinct from the two registered primary endpoints.

What the ClinicalTrials.gov record does not establish: the ClinicalTrials.gov record does not specify an alpha-allocation procedure, gatekeeping sequence, hierarchical testing strategy, or other multiplicity-adjustment details. Those details should not be invented from the observed P-values.

19. Interim Analysis, Crossover, and Other Design Topics

The ClinicalTrials.gov record identifies randomization, parallel design, no masking, superiority hypotheses, and log-rank testing. They do not provide a non-inferiority margin, crossover rule, factorial structure, interim-analysis boundary, missing-data imputation procedure, stratification factors, or Bayesian analysis method.

Design topicWhat is supported by the ClinicalTrials.gov record
RandomizationYes — allocation is randomized.
Parallel designYes — design model is parallel.
MaskingNone.
Superiority testingYes — reported for the statistical analyses.
Non-inferiority marginNot reported in the ClinicalTrials.gov record.
CrossoverNot reported in the ClinicalTrials.gov record.
Factorial designNot reported in the ClinicalTrials.gov record.
Interim-analysis procedureNot reported in the ClinicalTrials.gov record.
Missing-data / imputation procedureNot reported in the ClinicalTrials.gov record.
Stratification factorsNot reported in the ClinicalTrials.gov record.
Bayesian methodsNot reported in the ClinicalTrials.gov record.

This is an important methodological boundary. Absence of a detail in the ClinicalTrials.gov record does not justify reconstructing it from another source or from general expectations about phase 3 oncology trials.

20. What the Hazard Ratio Does — and Does Not — Mean

Relative effect

An HR below 1 indicates a lower estimated instantaneous event rate for avelumab plus axitinib relative to sunitinib. For example, the primary PFS HR of 0.61 corresponds to an estimated 39% lower hazard of the PFS event.

Not an absolute effect

A hazard ratio cannot tell us from the ClinicalTrials.gov record how many additional months an individual participant will remain progression-free or alive. Absolute effects require survival probabilities, median times, or other absolute measures.

Not a probability of benefit

An HR of 0.61 does not mean that a participant has a 61% probability of benefiting, nor does an HR of 0.86 mean that an individual has an 86% probability of survival.

Confidence interval

The 95% CI describes uncertainty around the estimated HR. For PFS, the interval is 0.475–0.790. For OS, the interval is 0.701–1.057. These intervals should be considered alongside the corresponding point estimates rather than treated as ranges of individual outcomes.

21. Results Summary

EndpointPopulationHR95% CIP-value
Primary PFS by BICR PD-L1 positive 0.61 0.475–0.790 0.0001
Primary OS PD-L1 positive 0.86 0.701–1.057 0.1509
Secondary PFS by BICR Irrespective of PD-L1 expression 0.69 0.563–0.840 0.0002
Secondary OS Irrespective of PD-L1 expression 0.88 0.749–1.039 0.1338
Secondary investigator-assessed PFS Irrespective of PD-L1 expression 0.66 0.565–0.768 <.0001
Secondary PFS2 Irrespective of PD-L1 expression 0.64 0.551–0.754 Not reported

The ClinicalTrials.gov record shows a consistent pattern of hazard ratios below 1 for the reported PFS endpoints and PFS2, while the reported OS hazard ratios are closer to 1. That pattern is descriptive of the registry-reported estimates; it does not by itself establish why the endpoints differ or whether the treatment effects are heterogeneous across endpoints.

22. Important Limitations and Interpretation Issues

23. Why This Trial Matters Statistically

JAVELIN Renal 101 is a useful teaching example because the ClinicalTrials.gov record bring several core clinical-trial concepts together without requiring the analysis to be reduced to a single number.

ConceptHow it appears in JAVELIN Renal 101
RandomizationThe trial uses randomized allocation to two treatment arms.
Parallel designThe two treatment groups are evaluated in a parallel design.
Time-to-event endpointsBoth registered primary endpoints are time-to-event outcomes.
Kaplan-Meier estimationThe registry-reported OS definition states that analysis was performed using the Kaplan-Meier method.
Hazard ratioAll six registry-reported statistical analyses use hazard ratio as the effect measure.
Log-rank testThe formal reported analyses for PFS and OS use a log-rank test where a method is reported.
Confidence intervalsBoth primary analyses report two-sided 95% confidence intervals.
SuperiorityThe analyses posted on ClinicalTrials.gov identify superiority as the hypothesis type.
Analysis populationsPrimary analyses use PD-L1-positive randomized participants; several secondary analyses use all randomized participants.
Assessment methodPrimary PFS is assessed by BICR, while a secondary PFS analysis is investigator-assessed.
MultiplicityTwo registered primary endpoints and multiple secondary analyses require careful interpretation of individual P-values.
SafetySerious adverse events are reported separately by treatment arm.

The statistical lesson is that a randomized trial produces a collection of related but distinct estimands. PFS and OS are not interchangeable. PD-L1-positive and all-randomized analyses are not interchangeable. BICR and investigator assessments are not interchangeable. And an HR, confidence interval, and P-value each answer different parts of the statistical question.

24. A Practical Reading Sequence for This Trial

Step 1 · Identify the estimand

Start with the endpoint definition and population. Determine whether the analysis concerns PFS, OS, or PFS2 and whether it is restricted to PD-L1-positive participants.

Step 2 · Identify the effect measure

The analyses posted on ClinicalTrials.gov use hazard ratios. Establish whether the reported HR is below, equal to, or above 1 before considering its precision.

Step 3 · Read the confidence interval

Assess the width of the 95% CI and whether it includes 1.00. This provides information about uncertainty that the point estimate alone cannot provide.

Step 4 · Read the P-value

Use the P-value as evidence against the null hypothesis under the specified testing framework, not as a measure of effect size.

Step 5 · Check the analysis population

Do not compare estimates from PD-L1-positive participants and all randomized participants as though they were estimates from exactly the same population.

Step 6 · Check what is missing

The ClinicalTrials.gov record does not provide median survival, survival probabilities, stratification factors, interim rules, or missing-data procedures. Those quantities should not be reconstructed from the reported HRs.

25. Sources

Source boundary: The numerical trial results and methodological facts on this page are restricted to the registry-reported JAVELIN Renal 101 trial data. The PubMed links are provided as the associated publication records identified in that data; publication-specific numbers or claims not contained in the ClinicalTrials.gov record is intentionally not added.

26. Related Tutorials

Learn more about the methods used in this trial:

27. Related Calculators

Continue through the Clinical Biostats statistical pathway

Use the related tutorials and calculators to explore the survival-analysis concepts illustrated by JAVELIN Renal 101.

28. Record Summary

JAVELIN Renal 101 provides a clear example of how a randomized phase 3 trial can be analyzed through several related time-to-event questions. The two registered primary endpoints are PFS assessed by BICR and OS in PD-L1-positive participants. The registry-reported primary PFS analysis reports an HR of 0.61 with a 95% CI of 0.475–0.790 and P = 0.0001, while the primary OS analysis reports an HR of 0.86 with a 95% CI of 0.701–1.057 and P = 0.1509.

The secondary analyses extend the picture to participants irrespective of PD-L1 expression and include both BICR- and investigator-assessed PFS, OS, and PFS2. The reported HRs are 0.69, 0.88, 0.66, and 0.64, respectively, with the PFS2 analysis lacking a reported method and P-value in the ClinicalTrials.gov record.

The central statistical lesson is to interpret each estimate in its proper context: endpoint definition, analysis population, assessment method, effect measure, confidence interval, and hypothesis-testing framework. A hazard ratio provides a relative time-to-event measure, while the confidence interval describes its precision and the P-value addresses statistical evidence under the specified test. None of these quantities, by itself, describes the complete clinical experience of every participant.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. For JAVELIN Renal 101, that means preserving the registry-defined populations and endpoints, reporting the registry-reported hazard ratios and confidence intervals exactly, identifying where a formal method was not reported, and avoiding unsupported reconstruction of additional trial characteristics.