← Clinical Trials
Ovarian Cancer Phase 3 Time-to-Event NCT02718417

JAVELIN Ovarian 100: Complete Statistical Analysis of Avelumab in Ovarian Cancer

An independent statistical analysis of the randomized phase 3 JAVELIN Ovarian 100 trial evaluating avelumab-containing treatment strategies in previously untreated patients with epithelial ovarian cancer.

JAVELIN Ovarian 100  ·  NCT02718417  ·  Phase 3  ·  998 enrolled
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

JAVELIN Ovarian 100 was a randomized, parallel, open-label phase 3 trial in ovarian cancer. The registry reports 998 enrolled participants, 3 arms, a time-to-event primary endpoint, and formal statistical analyses based on the full analysis set of all randomized participants.

998
Enrolled
ClinicalTrials.gov
3
Arms
Parallel design
1.43
Primary PFS HR
95% CI 1.051–1.946
1.14
Primary PFS HR
95% CI 0.832–1.565
FeatureJAVELIN Ovarian 100
Trial nameJAVELIN Ovarian 100
Brief titleAvelumab in Previously Untreated Patients With Epithelial Ovarian Cancer (JAVELIN OVARIAN 100)
PhasePhase 3
ConditionOvarian Cancer
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment998
Primary endpoint typeTime-to-event
Primary endpoint count1
Results postedYes
Outcome measures posted36
Statistical analyses posted6
Primary-endpoint analyses2
Primary analyses with estimate + CI2
Registry statusTerminated
Lead sponsorPfizer

2. Clinical Question

The trial addressed whether avelumab-containing treatment strategies could be evaluated against chemotherapy followed by observation in previously untreated epithelial ovarian cancer, using progression-free survival as the registered primary endpoint.

Population

Previously untreated patients with epithelial ovarian cancer, according to the trial's brief title.

Intervention strategies

The registry data identify avelumab in two treatment strategies: chemotherapy followed by avelumab, and chemotherapy plus avelumab followed by avelumab.

Comparator

Chemotherapy followed by observation.

Primary question

How does progression-free survival compare between each avelumab-containing strategy and chemotherapy followed by observation?

3. Trial Design

01
Randomize998 enrolled
02
3 armsParallel design
03
TreatmentChemotherapy ± avelumab
04
Follow-upPFS / OS assessment
05
AnalysisLog-rank + Cox model
Allocation
Randomized allocation was used in a parallel-group phase 3 design.
Masking
The registry identifies the trial as having no masking.
Primary purpose
Treatment.
Primary endpoint
Progression-Free Survival (PFS) as Assessed by Blinded Independent Central Review (BICR).
ARM 1

Chemotherapy Followed by Avelumab

  • Chemotherapy
  • Followed by avelumab
  • Serious AEs: 92/328
ARM 2

Chemotherapy + Avelumab Followed by Avel

  • Chemotherapy plus avelumab
  • Followed by the registry-described avelumab strategy
  • Serious AEs: 118/329
ARM 3

Chemotherapy Followed by Observation

  • Chemotherapy
  • Followed by observation
  • Serious AEs: 64/334
Three-arm structure matters. The posted primary analyses compare each avelumab-containing strategy with chemotherapy followed by observation. This is different from a single two-arm comparison because the statistical interpretation must keep the randomized treatment strategies distinct.

4. Trial Timeline

MilestoneDate
Trial start2016-05-19
Primary completion2018-09-07
Registry statusTerminated

The ClinicalTrials.gov record identifies the trial as terminated and gives the start and primary-completion dates above. Those dates should not be interpreted as additional efficacy results.

5. Endpoints

EndpointRegistry definition / time frameType
Progression-Free Survival (PFS) as Assessed by Blinded Independent Central Review (BICR) Baseline to progression of disease or discontinuation from the study or death, whichever occurred first (maximum duration of 27 months) Time-to-event

Primary endpoint definition

The registry describes BICR-assessed PFS as the duration from randomization until disease progression or death. PFS data were censored on the date of the last adequate tumor assessment for participants who did not have an event, who started a new anti-cancer therapy prior to an event, or who had an event after 2 or more missing tumor assessments. The registry definition continues with progression as per Response Evaluation Criteria in Solid Tumors (RECIST) version 1.1: as at least a 20 percent (%) increase in the sum of diameters of target lesions, taking as reference the smallest sum on study (this includes the baseline sum if that is the smallest on study). In addition to the relative increase of 20%, the sum must have also demonstrated an absolute increase of at least 5 millimeters (mm). The appearance of one or more new lesions was also considered progression. Analysis was performed using Kaplan-Meier method..

Secondary endpoints with posted formal analyses

EndpointTime frameAnalysis type
Overall Survival Baseline to discontinuation from the study or death, whichever occurred first (maximum duration of 27 months) Time-to-event; log-rank; stratified Cox model
Progression-Free Survival (PFS) as Assessed by Investigator Baseline to progression of disease or discontinuation from the study or death, whichever occurred first (maximum duration of 27 months) Time-to-event; log-rank

6. Analysis Populations and Stratification

The registry states that the full analysis set included all randomized participants for the posted primary and secondary time-to-event analyses. This is consistent with an intention-to-treat analysis principle: randomized participants remain associated with their randomized treatment group for the efficacy comparison.

Full analysis set

All randomized participants were included in the full analysis set used for the posted efficacy analyses.

Intention-to-treat concept

The analysis text explicitly identifies intention-to-treat analysis as a concept associated with the primary and secondary analyses.

Stratified analysis

The primary analysis notes specify that the Cox proportional-hazards model was stratified by the randomization strata and that a stratified log-rank test was used.

What is not specified here

The ClinicalTrials.gov record does not identify the individual randomization-stratum variables, so they are not reproduced or inferred on this page.

7. Primary Results: BICR-Assessed Progression-Free Survival

The registered primary endpoint was BICR-assessed progression-free survival. Two formal primary analyses were posted, each comparing an avelumab-containing strategy with chemotherapy followed by observation.

Chemotherapy Followed by Avelumab vs Chemotherapy Followed by Observation

Hazard ratio for progression or death

1.43

95% CI: 1.051–1.946   ·   P = 0.9890

Full analysis set: all randomized participants

FeatureReported result
EndpointProgression-Free Survival (PFS) as Assessed by Blinded Independent Central Review (BICR)
ComparisonChemotherapy Followed by Avelumab vs Chemotherapy Followed by Observation
MethodStratified log-rank test and Cox proportional-hazards model stratified by randomization strata
Effect measureHazard Ratio (HR)
Estimate1.43
95% CI1.051–1.946
P-value0.9890
Hypothesis typeSuperiority
Clinical Biostats interpretation

An HR of 1.43 means that, under the fitted Cox model and for this comparison, the estimated instantaneous rate of the PFS event was 1.43 times the corresponding rate in the chemotherapy-followed-by-observation group. Equivalently, 1.43 represents an estimated 43% higher hazard relative to the comparator; it is not a statement that 43% more patients progressed.

The 95% CI of 1.051–1.946 quantifies uncertainty around the estimated hazard ratio under the model and sampling framework. It does not describe the range of effects that individual patients experienced.

The p-value of 0.9890 is a measure associated with the statistical testing procedure; it is not a measure of effect size, clinical importance, or the probability that the treatment is effective. The registry reports this p-value together with a two-sided 95% confidence interval, and the registry values should be reported as given rather than mathematically reconciled or replaced.

Because the analysis uses a Cox proportional-hazards model, the interpretation of a single HR also depends on the model's proportional-hazards framework. Censoring rules are part of the PFS definition and therefore affect the information contributing to the analysis.

Chemotherapy + Avelumab Followed by Avelumab vs Chemotherapy Followed by Observation

Hazard ratio for progression or death

1.14

95% CI: 0.832–1.565   ·   P = 0.7935

Full analysis set: all randomized participants

FeatureReported result
EndpointProgression-Free Survival (PFS) as Assessed by Blinded Independent Central Review (BICR)
ComparisonChemotherapy + Avelumab Followed by Avelumab vs Chemotherapy Followed by Observation
MethodStratified log-rank test and Cox proportional-hazards model stratified by randomization strata
Effect measureHazard Ratio (HR)
Estimate1.14
95% CI0.832–1.565
P-value0.7935
Hypothesis typeSuperiority
Clinical Biostats interpretation

An HR of 1.14 means that the fitted model estimated the instantaneous PFS event rate at 1.14 times that of the chemotherapy-followed-by-observation group for this randomized comparison. This is an estimated relative hazard, not a 14% difference in the proportion of patients who experienced progression or death.

The 95% CI of 0.832–1.565 spans 1.00, illustrating substantial uncertainty about the direction and magnitude of the relative hazard on the scale represented by the confidence interval.

The reported p-value of 0.7935 is a result of the statistical testing procedure and should not be interpreted as an effect-size measure or as the probability that the null hypothesis is true. The registry reports a two-sided 95% confidence interval and the p-value separately.

As with the first primary analysis, the Cox-model interpretation relies on its proportional-hazards framework, while the log-rank comparison incorporates the trial's stratified time-to-event analysis. The analysis population was the full analysis set of all randomized participants.

Important reporting point: The ClinicalTrials.gov record contains an apparent numerical tension between the reported two-sided confidence intervals and the corresponding p-values for the two primary analyses. This page does not recompute, reverse-engineer, or substitute values. The reported estimates, confidence intervals, and p-values are presented exactly as reported in the ClinicalTrials.gov record.

8. Secondary Endpoint Results: Overall Survival

Overall survival was analyzed as a secondary time-to-event endpoint in the full analysis set. The registry reports two formal comparisons using the same stratified log-rank and stratified Cox-model framework.

ComparisonHR95% CIP-value
Chemotherapy Followed by Avelumab vs Chemotherapy Followed by Observation 1.53 0.760–3.080 0.8848
Chemotherapy + Avelumab Followed by Avelumab vs Chemotherapy Followed by Observation 1.55 0.776–3.111 0.8953

How to interpret these OS estimates

First comparison

The reported HR of 1.53 represents a model-based estimate of the relative instantaneous death hazard for chemotherapy followed by avelumab versus chemotherapy followed by observation.

Second comparison

The reported HR of 1.55 represents the corresponding model-based estimate for chemotherapy plus avelumab followed by avelumab versus chemotherapy followed by observation.

Both confidence intervals are wide and include 1.00. The ClinicalTrials.gov record does not provide median overall survival or other absolute survival estimates, so those quantities are not reported here.

9. Secondary Endpoint Results: Investigator-Assessed PFS

Investigator-assessed PFS was also analyzed as a secondary time-to-event endpoint. The registry reports two comparisons against chemotherapy followed by observation.

ComparisonHR95% CIP-valueHypothesis type
Chemotherapy Followed by Avelumab vs Chemotherapy Followed by Observation 1.21 0.935–1.578 0.9278 Other / not stated
Chemotherapy + Avelumab Followed by Avelumab vs Chemotherapy Followed by Observation 0.90 0.688–1.189 0.2367 Other / not stated

The investigator-assessed analyses are useful as a separate assessment of PFS, but they should not be silently substituted for the registered primary BICR endpoint. Different assessment sources can produce different event times and censoring patterns, which is one reason trials distinguish central-review and investigator-assessed endpoints.

10. Statistical Methodology

Kaplan-Meier estimation

PFS and overall survival are time-to-event endpoints. Kaplan-Meier estimation is the standard descriptive framework for representing the event-time distribution while accommodating right censoring. A participant who has not experienced the event by the last relevant follow-up can contribute information up to the censoring time.

Kaplan-Meier concept
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at an event time and ni represents the number at risk immediately before that time.

The ClinicalTrials.gov record does not provide Kaplan-Meier estimates, median event times, or underlying individual event/censoring records. Therefore, this page does not construct a Kaplan-Meier curve or infer one from the hazard ratios.

Stratified log-rank test

The registry identifies the log-rank test as the reported primary and secondary comparison method. For the primary analyses, the analysis notes specifically state that a stratified log-rank test was used.

A log-rank test compares the observed and expected numbers of events between randomized groups over follow-up. Stratification allows the comparison to account for the randomization strata rather than treating all participants as belonging to one unstratified risk set.

Stratified Cox proportional-hazards model

The primary analysis notes state that the analysis was performed using a Cox proportional-hazards model stratified by the randomization strata, together with a stratified log-rank test.

Hazard-ratio interpretation
HR = estimated hazard in comparison group ÷ estimated hazard in reference group

An HR below 1 indicates a lower estimated instantaneous event hazard in the numerator group; an HR above 1 indicates a higher estimated instantaneous event hazard. The HR is not an absolute risk difference.

Intention-to-treat analysis

The analysis text identifies intention-to-treat analysis among the concepts associated with the efficacy analyses. In practical terms, analyzing randomized participants according to their assigned treatment preserves the treatment comparison created by randomization and avoids defining efficacy groups solely by treatment exposure.

Stratification

The registry states that the Cox model was stratified by the randomization strata and that a stratified log-rank test was used. The individual strata themselves are not specified in the ClinicalTrials.gov record, so no additional stratification variables are asserted here.

11. Censoring and Missing Tumor Assessments

The registered BICR PFS definition contains explicit censoring rules. Participants without an event were censored at the date of the last adequate tumor assessment. The definition also specifies censoring for participants who started a new anti-cancer therapy before an event and for participants with an event after 2 or more missing tumor assessments.

Why censoring matters

Time-to-event methods use information from participants up to their event or censoring time. The censoring rule therefore determines which portion of follow-up contributes to the PFS estimate.

Missing assessments

Because the registry definition explicitly addresses missing tumor assessments, missingness is not simply ignored. Its effect depends on the prespecified event and censoring rules.

New anti-cancer therapy

The registry-reported PFS definition specifies a censoring rule for participants who started a new anti-cancer therapy before an event.

What is not reported here

The ClinicalTrials.gov record does not specify a separate statistical imputation model. No additional imputation method is therefore attributed to the trial.

12. Multiplicity and the Three-Arm Structure

The trial contains three randomized arms and two posted primary-endpoint comparisons against chemotherapy followed by observation. That structure creates an important statistical distinction between the existence of a treatment effect estimate and the interpretation of multiple hypothesis tests.

ComparisonPrimary endpointFormal methodHypothesis type
Chemotherapy Followed by Avelumab vs Chemotherapy Followed by Observation BICR-assessed PFS Stratified log-rank + stratified Cox model Superiority
Chemotherapy + Avelumab Followed by Avelumab vs Chemotherapy Followed by Observation BICR-assessed PFS Stratified log-rank + stratified Cox model Superiority

The ClinicalTrials.gov record does not state a multiplicity-adjustment procedure, alpha-allocation scheme, or hierarchical testing sequence for these two primary comparisons. Consequently, no such procedure is inferred on this page.

Why this matters: when multiple treatment comparisons are tested, the probability of at least one false-positive finding can depend on how the hypotheses are handled jointly. A complete statistical analysis normally distinguishes the prespecified multiplicity strategy from nominal p-values. Here, only the information explicitly reported in the ClinicalTrials.gov record is reported.

13. Safety Results

The registry data provide serious adverse-event counts by arm as affected participants divided by participants at risk.

ArmSerious adverse eventsAffected / at risk
Chemotherapy Followed by Avelumab 92 participants 92/328
Chemotherapy + Avelumab Followed by Avel 118 participants 118/329
Chemotherapy Followed by Observation 64 participants 64/334

These are serious adverse-event counts, not overall adverse-event rates. They should therefore not be substituted for other safety endpoints such as treatment-emergent adverse events, grade-specific adverse events, discontinuations, or deaths unless those measures are separately reported.

Safety denominator principle
Observed safety proportion = affected participants ÷ participants at risk

The ClinicalTrials.gov record already provide the affected and at-risk counts. This page retains the reported fractions rather than calculating new percentages.

14. Statistical Methods Explained

Why was a log-rank test used?

PFS and overall survival are time-to-event outcomes, so a simple comparison of means would not appropriately use the available follow-up information. The log-rank test compares the timing of events between randomized groups while accommodating censoring. In this trial, the registry specifically reports a stratified log-rank test for the primary analyses.

What does an HR of 1.43 mean?

For the chemotherapy-followed-by-avelumab versus chemotherapy-followed-by-observation primary comparison, an HR of 1.43 is a model-based estimate that the instantaneous rate of progression or death was 1.43 times that of the reference group. It does not mean that 43% of patients experienced progression, nor does it describe an absolute difference in PFS probability.

Why is the confidence interval important?

A point estimate alone gives only one estimate of the relative treatment effect. The 95% confidence interval supplies information about statistical precision under the model and sampling framework. A wide interval indicates greater uncertainty than a narrow interval; the interval does not represent the range of individual patient effects.

Why does the full analysis set matter?

The registry defines the full analysis set as all randomized participants. Keeping randomized participants in the efficacy analysis maintains the comparison established by randomization. It also prevents post-randomization treatment exposure from becoming the sole basis for defining the efficacy population.

Why distinguish BICR PFS from investigator-assessed PFS?

The trial reports both BICR-assessed PFS and investigator-assessed PFS. These are not interchangeable measurements. BICR provides an independent central assessment, while investigator assessment represents the study-site evaluation. Their results can differ because progression classification and timing can differ between assessment processes.

What does the p-value tell us?

A p-value is tied to a specified statistical testing framework. It measures how compatible the observed data are with the null hypothesis under that framework; it does not measure the magnitude of the treatment effect, the clinical importance of the result, or the probability that the null hypothesis is true. For this trial, the reported p-values should be considered together with the hazard ratios, confidence intervals, analysis population, and prespecified testing structure.

Why should the HR not be treated as a risk ratio?

A hazard ratio compares instantaneous event rates within a time-to-event model. A risk ratio compares probabilities over a specified time period. Because they use different quantities, an HR of 1.14 does not mean that the probability of progression or death is 14% higher at every time point.

15. Interpreting the Primary Results Together

The two primary comparisons are most clearly understood as separate randomized contrasts rather than as a single pooled estimate.

Primary comparisonHR95% CIP-valueStatistical reading
Chemotherapy Followed by Avelumab vs Chemotherapy Followed by Observation 1.43 1.051–1.946 0.9890 Estimated HR above 1; CI does not contain 1
Chemotherapy + Avelumab Followed by Avelumab vs Chemotherapy Followed by Observation 1.14 0.832–1.565 0.7935 Estimated HR above 1; CI contains 1

The first comparison has a point estimate above 1 and a 95% confidence interval entirely above 1, while the second has a point estimate above 1 with a confidence interval spanning 1. The registry nevertheless reports p-values of 0.9890 and 0.7935, respectively. Because the ClinicalTrials.gov record does not provide enough information to establish why those numerical elements differ in this way, the appropriate educational approach is to preserve the reported quantities and avoid reverse-engineering an alternative test.

Clinical Biostats interpretation

The most important lesson is that a clinical-trial result is not represented adequately by a single p-value. The treatment contrast, endpoint definition, analysis population, hazard ratio, confidence interval, statistical test, and multiplicity framework all contribute to the interpretation.

For a time-to-event endpoint, the HR is a relative model-based measure, while the confidence interval describes uncertainty around that measure. Neither one provides an absolute probability of progression at a particular time without additional survival estimates.

16. What the Hazard Ratio Does — and Does Not — Mean

Statistical interpretation

A hazard ratio of 1.21, for example, would mean that the fitted model estimates an instantaneous event hazard 1.21 times that of the reference group for the corresponding comparison. The analogous interpretation applies to the reported HRs of 1.43, 1.14, 1.53, 1.55, and 0.90.

It does not mean that the corresponding percentage of patients experienced an event, nor does it mean that every participant had the same proportional change in risk.

Why the confidence interval matters

The confidence interval places the point estimate in a range reflecting statistical uncertainty under the model and sampling framework. For the primary analyses, the intervals are 1.051–1.946 and 0.832–1.565. These intervals should be read alongside the corresponding HR estimates rather than treated as estimates of individual patient outcomes.

Why absolute measures would add information

Hazard ratios summarize relative time-to-event effects, but they do not directly communicate absolute event probabilities or median survival. The ClinicalTrials.gov record does not report those additional measures for this analysis, so they are not inferred.

17. Limitations

18. Why This Trial Matters Statistically

JAVELIN Ovarian 100 is a useful teaching case because it combines randomized three-arm treatment allocation with a primary time-to-event endpoint, independent central assessment, stratified survival analysis, multiple treatment comparisons, and both central-review and investigator-assessed PFS.

ConceptHow it appears in JAVELIN Ovarian 100
RandomizationRandomized allocation in a phase 3 parallel design.
Three-arm designThree treatment strategies are represented in the registry data.
Time-to-event analysisPFS is the registered primary endpoint; overall survival and investigator-assessed PFS are secondary analyzed endpoints.
BICR assessmentThe primary endpoint is PFS as assessed by blinded independent central review.
Kaplan-Meier estimationA standard descriptive framework for the registered time-to-event endpoints, although no KM estimates are reported in the ClinicalTrials.gov record.
Hazard ratioThe primary and secondary formal analyses use hazard ratios as the effect measure.
Log-rank testingThe registry identifies log-rank as the formal comparison method.
Stratified analysisThe primary analysis uses a Cox model stratified by randomization strata and a stratified log-rank test.
Intention-to-treatThe full analysis set includes all randomized participants.
CensoringThe registered PFS definition specifies several censoring rules.
MultiplicityTwo primary comparisons are posted for the same primary endpoint.
Safety denominatorsSerious adverse events are reported as affected participants divided by participants at risk for each arm.

19. Related Tutorials

Learn more about the methods used in this trial:

20. Related Calculators

21. Sources

Continue with Clinical Biostats statistical methods

Use the related tutorials and calculators to examine the survival-analysis concepts that appear throughout randomized clinical trials.

22. Record Summary

JAVELIN Ovarian 100 provides a compact example of how a randomized three-arm phase 3 trial can be analyzed through time-to-event methods. The registered primary endpoint was BICR-assessed progression-free survival, with two posted comparisons using a stratified log-rank test and a Cox proportional-hazards model stratified by randomization strata. The full analysis set included all randomized participants.

The posted primary analyses report HRs of 1.43 and 1.14, with corresponding 95% confidence intervals of 1.051–1.946 and 0.832–1.565. Secondary analyses provide additional HR estimates for overall survival and investigator-assessed PFS. The registry also reports serious adverse-event counts of 92/328, 118/329, and 64/334 across the three treatment strategies.

The statistical lesson is broader than any individual number: a rigorous interpretation requires attention to the randomized comparison, endpoint definition, analysis population, censoring rules, effect measure, confidence interval, hypothesis test, stratification, and the relationship among multiple comparisons. The ClinicalTrials.gov record supports those methodological conclusions without requiring assumptions about unreported median survival, subgroup effects, or additional statistical procedures.

Clinical Biostats methodology: A trial-results page should distinguish reported evidence from statistical interpretation. The objective is to explain what the reported analysis estimates, how the method works, and what conclusions the ClinicalTrials.gov record can and cannot support.