This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
JAVELIN Ovarian 100 was a randomized, parallel, open-label phase 3 trial in ovarian cancer. The registry reports 998 enrolled participants, 3 arms, a time-to-event primary endpoint, and formal statistical analyses based on the full analysis set of all randomized participants.
| Feature | JAVELIN Ovarian 100 |
|---|---|
| Trial name | JAVELIN Ovarian 100 |
| Brief title | Avelumab in Previously Untreated Patients With Epithelial Ovarian Cancer (JAVELIN OVARIAN 100) |
| Phase | Phase 3 |
| Condition | Ovarian Cancer |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 998 |
| Primary endpoint type | Time-to-event |
| Primary endpoint count | 1 |
| Results posted | Yes |
| Outcome measures posted | 36 |
| Statistical analyses posted | 6 |
| Primary-endpoint analyses | 2 |
| Primary analyses with estimate + CI | 2 |
| Registry status | Terminated |
| Lead sponsor | Pfizer |
2. Clinical Question
The trial addressed whether avelumab-containing treatment strategies could be evaluated against chemotherapy followed by observation in previously untreated epithelial ovarian cancer, using progression-free survival as the registered primary endpoint.
Population
Previously untreated patients with epithelial ovarian cancer, according to the trial's brief title.
Intervention strategies
The registry data identify avelumab in two treatment strategies: chemotherapy followed by avelumab, and chemotherapy plus avelumab followed by avelumab.
Comparator
Chemotherapy followed by observation.
Primary question
How does progression-free survival compare between each avelumab-containing strategy and chemotherapy followed by observation?
3. Trial Design
Chemotherapy Followed by Avelumab
- Chemotherapy
- Followed by avelumab
- Serious AEs: 92/328
Chemotherapy + Avelumab Followed by Avel
- Chemotherapy plus avelumab
- Followed by the registry-described avelumab strategy
- Serious AEs: 118/329
Chemotherapy Followed by Observation
- Chemotherapy
- Followed by observation
- Serious AEs: 64/334
4. Trial Timeline
| Milestone | Date |
|---|---|
| Trial start | 2016-05-19 |
| Primary completion | 2018-09-07 |
| Registry status | Terminated |
The ClinicalTrials.gov record identifies the trial as terminated and gives the start and primary-completion dates above. Those dates should not be interpreted as additional efficacy results.
5. Endpoints
| Endpoint | Registry definition / time frame | Type |
|---|---|---|
| Progression-Free Survival (PFS) as Assessed by Blinded Independent Central Review (BICR) | Baseline to progression of disease or discontinuation from the study or death, whichever occurred first (maximum duration of 27 months) | Time-to-event |
Primary endpoint definition
The registry describes BICR-assessed PFS as the duration from randomization until disease progression or death. PFS data were censored on the date of the last adequate tumor assessment for participants who did not have an event, who started a new anti-cancer therapy prior to an event, or who had an event after 2 or more missing tumor assessments. The registry definition continues with progression as per Response Evaluation Criteria in Solid Tumors (RECIST) version 1.1: as at least a 20 percent (%) increase in the sum of diameters of target lesions, taking as reference the smallest sum on study (this includes the baseline sum if that is the smallest on study). In addition to the relative increase of 20%, the sum must have also demonstrated an absolute increase of at least 5 millimeters (mm). The appearance of one or more new lesions was also considered progression. Analysis was performed using Kaplan-Meier method..
Secondary endpoints with posted formal analyses
| Endpoint | Time frame | Analysis type |
|---|---|---|
| Overall Survival | Baseline to discontinuation from the study or death, whichever occurred first (maximum duration of 27 months) | Time-to-event; log-rank; stratified Cox model |
| Progression-Free Survival (PFS) as Assessed by Investigator | Baseline to progression of disease or discontinuation from the study or death, whichever occurred first (maximum duration of 27 months) | Time-to-event; log-rank |
6. Analysis Populations and Stratification
The registry states that the full analysis set included all randomized participants for the posted primary and secondary time-to-event analyses. This is consistent with an intention-to-treat analysis principle: randomized participants remain associated with their randomized treatment group for the efficacy comparison.
Full analysis set
All randomized participants were included in the full analysis set used for the posted efficacy analyses.
Intention-to-treat concept
The analysis text explicitly identifies intention-to-treat analysis as a concept associated with the primary and secondary analyses.
Stratified analysis
The primary analysis notes specify that the Cox proportional-hazards model was stratified by the randomization strata and that a stratified log-rank test was used.
What is not specified here
The ClinicalTrials.gov record does not identify the individual randomization-stratum variables, so they are not reproduced or inferred on this page.
7. Primary Results: BICR-Assessed Progression-Free Survival
The registered primary endpoint was BICR-assessed progression-free survival. Two formal primary analyses were posted, each comparing an avelumab-containing strategy with chemotherapy followed by observation.
Chemotherapy Followed by Avelumab vs Chemotherapy Followed by Observation
Hazard ratio for progression or death
95% CI: 1.051–1.946 · P = 0.9890
Full analysis set: all randomized participants
| Feature | Reported result |
|---|---|
| Endpoint | Progression-Free Survival (PFS) as Assessed by Blinded Independent Central Review (BICR) |
| Comparison | Chemotherapy Followed by Avelumab vs Chemotherapy Followed by Observation |
| Method | Stratified log-rank test and Cox proportional-hazards model stratified by randomization strata |
| Effect measure | Hazard Ratio (HR) |
| Estimate | 1.43 |
| 95% CI | 1.051–1.946 |
| P-value | 0.9890 |
| Hypothesis type | Superiority |
An HR of 1.43 means that, under the fitted Cox model and for this comparison, the estimated instantaneous rate of the PFS event was 1.43 times the corresponding rate in the chemotherapy-followed-by-observation group. Equivalently, 1.43 represents an estimated 43% higher hazard relative to the comparator; it is not a statement that 43% more patients progressed.
The 95% CI of 1.051–1.946 quantifies uncertainty around the estimated hazard ratio under the model and sampling framework. It does not describe the range of effects that individual patients experienced.
The p-value of 0.9890 is a measure associated with the statistical testing procedure; it is not a measure of effect size, clinical importance, or the probability that the treatment is effective. The registry reports this p-value together with a two-sided 95% confidence interval, and the registry values should be reported as given rather than mathematically reconciled or replaced.
Because the analysis uses a Cox proportional-hazards model, the interpretation of a single HR also depends on the model's proportional-hazards framework. Censoring rules are part of the PFS definition and therefore affect the information contributing to the analysis.
Chemotherapy + Avelumab Followed by Avelumab vs Chemotherapy Followed by Observation
Hazard ratio for progression or death
95% CI: 0.832–1.565 · P = 0.7935
Full analysis set: all randomized participants
| Feature | Reported result |
|---|---|
| Endpoint | Progression-Free Survival (PFS) as Assessed by Blinded Independent Central Review (BICR) |
| Comparison | Chemotherapy + Avelumab Followed by Avelumab vs Chemotherapy Followed by Observation |
| Method | Stratified log-rank test and Cox proportional-hazards model stratified by randomization strata |
| Effect measure | Hazard Ratio (HR) |
| Estimate | 1.14 |
| 95% CI | 0.832–1.565 |
| P-value | 0.7935 |
| Hypothesis type | Superiority |
An HR of 1.14 means that the fitted model estimated the instantaneous PFS event rate at 1.14 times that of the chemotherapy-followed-by-observation group for this randomized comparison. This is an estimated relative hazard, not a 14% difference in the proportion of patients who experienced progression or death.
The 95% CI of 0.832–1.565 spans 1.00, illustrating substantial uncertainty about the direction and magnitude of the relative hazard on the scale represented by the confidence interval.
The reported p-value of 0.7935 is a result of the statistical testing procedure and should not be interpreted as an effect-size measure or as the probability that the null hypothesis is true. The registry reports a two-sided 95% confidence interval and the p-value separately.
As with the first primary analysis, the Cox-model interpretation relies on its proportional-hazards framework, while the log-rank comparison incorporates the trial's stratified time-to-event analysis. The analysis population was the full analysis set of all randomized participants.
8. Secondary Endpoint Results: Overall Survival
Overall survival was analyzed as a secondary time-to-event endpoint in the full analysis set. The registry reports two formal comparisons using the same stratified log-rank and stratified Cox-model framework.
| Comparison | HR | 95% CI | P-value |
|---|---|---|---|
| Chemotherapy Followed by Avelumab vs Chemotherapy Followed by Observation | 1.53 | 0.760–3.080 | 0.8848 |
| Chemotherapy + Avelumab Followed by Avelumab vs Chemotherapy Followed by Observation | 1.55 | 0.776–3.111 | 0.8953 |
How to interpret these OS estimates
First comparison
The reported HR of 1.53 represents a model-based estimate of the relative instantaneous death hazard for chemotherapy followed by avelumab versus chemotherapy followed by observation.
Second comparison
The reported HR of 1.55 represents the corresponding model-based estimate for chemotherapy plus avelumab followed by avelumab versus chemotherapy followed by observation.
Both confidence intervals are wide and include 1.00. The ClinicalTrials.gov record does not provide median overall survival or other absolute survival estimates, so those quantities are not reported here.
9. Secondary Endpoint Results: Investigator-Assessed PFS
Investigator-assessed PFS was also analyzed as a secondary time-to-event endpoint. The registry reports two comparisons against chemotherapy followed by observation.
| Comparison | HR | 95% CI | P-value | Hypothesis type |
|---|---|---|---|---|
| Chemotherapy Followed by Avelumab vs Chemotherapy Followed by Observation | 1.21 | 0.935–1.578 | 0.9278 | Other / not stated |
| Chemotherapy + Avelumab Followed by Avelumab vs Chemotherapy Followed by Observation | 0.90 | 0.688–1.189 | 0.2367 | Other / not stated |
The investigator-assessed analyses are useful as a separate assessment of PFS, but they should not be silently substituted for the registered primary BICR endpoint. Different assessment sources can produce different event times and censoring patterns, which is one reason trials distinguish central-review and investigator-assessed endpoints.
10. Statistical Methodology
Kaplan-Meier estimation
PFS and overall survival are time-to-event endpoints. Kaplan-Meier estimation is the standard descriptive framework for representing the event-time distribution while accommodating right censoring. A participant who has not experienced the event by the last relevant follow-up can contribute information up to the censoring time.
Here, di represents the number of events at an event time and ni represents the number at risk immediately before that time.
The ClinicalTrials.gov record does not provide Kaplan-Meier estimates, median event times, or underlying individual event/censoring records. Therefore, this page does not construct a Kaplan-Meier curve or infer one from the hazard ratios.
Stratified log-rank test
The registry identifies the log-rank test as the reported primary and secondary comparison method. For the primary analyses, the analysis notes specifically state that a stratified log-rank test was used.
A log-rank test compares the observed and expected numbers of events between randomized groups over follow-up. Stratification allows the comparison to account for the randomization strata rather than treating all participants as belonging to one unstratified risk set.
Stratified Cox proportional-hazards model
The primary analysis notes state that the analysis was performed using a Cox proportional-hazards model stratified by the randomization strata, together with a stratified log-rank test.
An HR below 1 indicates a lower estimated instantaneous event hazard in the numerator group; an HR above 1 indicates a higher estimated instantaneous event hazard. The HR is not an absolute risk difference.
Intention-to-treat analysis
The analysis text identifies intention-to-treat analysis among the concepts associated with the efficacy analyses. In practical terms, analyzing randomized participants according to their assigned treatment preserves the treatment comparison created by randomization and avoids defining efficacy groups solely by treatment exposure.
Stratification
The registry states that the Cox model was stratified by the randomization strata and that a stratified log-rank test was used. The individual strata themselves are not specified in the ClinicalTrials.gov record, so no additional stratification variables are asserted here.
11. Censoring and Missing Tumor Assessments
The registered BICR PFS definition contains explicit censoring rules. Participants without an event were censored at the date of the last adequate tumor assessment. The definition also specifies censoring for participants who started a new anti-cancer therapy before an event and for participants with an event after 2 or more missing tumor assessments.
Why censoring matters
Time-to-event methods use information from participants up to their event or censoring time. The censoring rule therefore determines which portion of follow-up contributes to the PFS estimate.
Missing assessments
Because the registry definition explicitly addresses missing tumor assessments, missingness is not simply ignored. Its effect depends on the prespecified event and censoring rules.
New anti-cancer therapy
The registry-reported PFS definition specifies a censoring rule for participants who started a new anti-cancer therapy before an event.
What is not reported here
The ClinicalTrials.gov record does not specify a separate statistical imputation model. No additional imputation method is therefore attributed to the trial.
12. Multiplicity and the Three-Arm Structure
The trial contains three randomized arms and two posted primary-endpoint comparisons against chemotherapy followed by observation. That structure creates an important statistical distinction between the existence of a treatment effect estimate and the interpretation of multiple hypothesis tests.
| Comparison | Primary endpoint | Formal method | Hypothesis type |
|---|---|---|---|
| Chemotherapy Followed by Avelumab vs Chemotherapy Followed by Observation | BICR-assessed PFS | Stratified log-rank + stratified Cox model | Superiority |
| Chemotherapy + Avelumab Followed by Avelumab vs Chemotherapy Followed by Observation | BICR-assessed PFS | Stratified log-rank + stratified Cox model | Superiority |
The ClinicalTrials.gov record does not state a multiplicity-adjustment procedure, alpha-allocation scheme, or hierarchical testing sequence for these two primary comparisons. Consequently, no such procedure is inferred on this page.
13. Safety Results
The registry data provide serious adverse-event counts by arm as affected participants divided by participants at risk.
| Arm | Serious adverse events | Affected / at risk |
|---|---|---|
| Chemotherapy Followed by Avelumab | 92 participants | 92/328 |
| Chemotherapy + Avelumab Followed by Avel | 118 participants | 118/329 |
| Chemotherapy Followed by Observation | 64 participants | 64/334 |
These are serious adverse-event counts, not overall adverse-event rates. They should therefore not be substituted for other safety endpoints such as treatment-emergent adverse events, grade-specific adverse events, discontinuations, or deaths unless those measures are separately reported.
The ClinicalTrials.gov record already provide the affected and at-risk counts. This page retains the reported fractions rather than calculating new percentages.
14. Statistical Methods Explained
Why was a log-rank test used?
PFS and overall survival are time-to-event outcomes, so a simple comparison of means would not appropriately use the available follow-up information. The log-rank test compares the timing of events between randomized groups while accommodating censoring. In this trial, the registry specifically reports a stratified log-rank test for the primary analyses.
What does an HR of 1.43 mean?
For the chemotherapy-followed-by-avelumab versus chemotherapy-followed-by-observation primary comparison, an HR of 1.43 is a model-based estimate that the instantaneous rate of progression or death was 1.43 times that of the reference group. It does not mean that 43% of patients experienced progression, nor does it describe an absolute difference in PFS probability.
Why is the confidence interval important?
A point estimate alone gives only one estimate of the relative treatment effect. The 95% confidence interval supplies information about statistical precision under the model and sampling framework. A wide interval indicates greater uncertainty than a narrow interval; the interval does not represent the range of individual patient effects.
Why does the full analysis set matter?
The registry defines the full analysis set as all randomized participants. Keeping randomized participants in the efficacy analysis maintains the comparison established by randomization. It also prevents post-randomization treatment exposure from becoming the sole basis for defining the efficacy population.
Why distinguish BICR PFS from investigator-assessed PFS?
The trial reports both BICR-assessed PFS and investigator-assessed PFS. These are not interchangeable measurements. BICR provides an independent central assessment, while investigator assessment represents the study-site evaluation. Their results can differ because progression classification and timing can differ between assessment processes.
What does the p-value tell us?
A p-value is tied to a specified statistical testing framework. It measures how compatible the observed data are with the null hypothesis under that framework; it does not measure the magnitude of the treatment effect, the clinical importance of the result, or the probability that the null hypothesis is true. For this trial, the reported p-values should be considered together with the hazard ratios, confidence intervals, analysis population, and prespecified testing structure.
Why should the HR not be treated as a risk ratio?
A hazard ratio compares instantaneous event rates within a time-to-event model. A risk ratio compares probabilities over a specified time period. Because they use different quantities, an HR of 1.14 does not mean that the probability of progression or death is 14% higher at every time point.
15. Interpreting the Primary Results Together
The two primary comparisons are most clearly understood as separate randomized contrasts rather than as a single pooled estimate.
| Primary comparison | HR | 95% CI | P-value | Statistical reading |
|---|---|---|---|---|
| Chemotherapy Followed by Avelumab vs Chemotherapy Followed by Observation | 1.43 | 1.051–1.946 | 0.9890 | Estimated HR above 1; CI does not contain 1 |
| Chemotherapy + Avelumab Followed by Avelumab vs Chemotherapy Followed by Observation | 1.14 | 0.832–1.565 | 0.7935 | Estimated HR above 1; CI contains 1 |
The first comparison has a point estimate above 1 and a 95% confidence interval entirely above 1, while the second has a point estimate above 1 with a confidence interval spanning 1. The registry nevertheless reports p-values of 0.9890 and 0.7935, respectively. Because the ClinicalTrials.gov record does not provide enough information to establish why those numerical elements differ in this way, the appropriate educational approach is to preserve the reported quantities and avoid reverse-engineering an alternative test.
The most important lesson is that a clinical-trial result is not represented adequately by a single p-value. The treatment contrast, endpoint definition, analysis population, hazard ratio, confidence interval, statistical test, and multiplicity framework all contribute to the interpretation.
For a time-to-event endpoint, the HR is a relative model-based measure, while the confidence interval describes uncertainty around that measure. Neither one provides an absolute probability of progression at a particular time without additional survival estimates.
16. What the Hazard Ratio Does — and Does Not — Mean
A hazard ratio of 1.21, for example, would mean that the fitted model estimates an instantaneous event hazard 1.21 times that of the reference group for the corresponding comparison. The analogous interpretation applies to the reported HRs of 1.43, 1.14, 1.53, 1.55, and 0.90.
It does not mean that the corresponding percentage of patients experienced an event, nor does it mean that every participant had the same proportional change in risk.
The confidence interval places the point estimate in a range reflecting statistical uncertainty under the model and sampling framework. For the primary analyses, the intervals are 1.051–1.946 and 0.832–1.565. These intervals should be read alongside the corresponding HR estimates rather than treated as estimates of individual patient outcomes.
Hazard ratios summarize relative time-to-event effects, but they do not directly communicate absolute event probabilities or median survival. The ClinicalTrials.gov record does not report those additional measures for this analysis, so they are not inferred.
17. Limitations
- Registry-level detail: the ClinicalTrials.gov record provides the registered endpoint, analysis methods, estimates, confidence intervals, p-values, and selected safety counts, but not a complete statistical analysis plan.
- Three-arm design: the primary endpoint has two reported treatment-versus-observation comparisons, making the interpretation of multiple testing important.
- Multiplicity information: the ClinicalTrials.gov record does not state an alpha-allocation or multiplicity-adjustment procedure for the two primary comparisons.
- Hazard-ratio assumptions: Cox-model HRs are model-based and rely on the proportional-hazards framework. The ClinicalTrials.gov record does not provide a formal assessment of that assumption.
- Censoring: PFS uses prespecified censoring rules involving last adequate tumor assessment, new anti-cancer therapy, and missing tumor assessments.
- No median survival estimates in the ClinicalTrials.gov record: median PFS and overall survival are not included in the ClinicalTrials.gov record and therefore are not reported.
- No subgroup estimates: the ClinicalTrials.gov record does not provide subgroup-specific efficacy estimates, so no subgroup conclusions are drawn.
- Safety scope: the ClinicalTrials.gov record is limited to serious adverse events by arm and does not establish the complete safety profile.
- Internal numerical tension: the registry-reported primary p-values and confidence intervals should be preserved as reported rather than reconciled by an independent calculation.
18. Why This Trial Matters Statistically
JAVELIN Ovarian 100 is a useful teaching case because it combines randomized three-arm treatment allocation with a primary time-to-event endpoint, independent central assessment, stratified survival analysis, multiple treatment comparisons, and both central-review and investigator-assessed PFS.
| Concept | How it appears in JAVELIN Ovarian 100 |
|---|---|
| Randomization | Randomized allocation in a phase 3 parallel design. |
| Three-arm design | Three treatment strategies are represented in the registry data. |
| Time-to-event analysis | PFS is the registered primary endpoint; overall survival and investigator-assessed PFS are secondary analyzed endpoints. |
| BICR assessment | The primary endpoint is PFS as assessed by blinded independent central review. |
| Kaplan-Meier estimation | A standard descriptive framework for the registered time-to-event endpoints, although no KM estimates are reported in the ClinicalTrials.gov record. |
| Hazard ratio | The primary and secondary formal analyses use hazard ratios as the effect measure. |
| Log-rank testing | The registry identifies log-rank as the formal comparison method. |
| Stratified analysis | The primary analysis uses a Cox model stratified by randomization strata and a stratified log-rank test. |
| Intention-to-treat | The full analysis set includes all randomized participants. |
| Censoring | The registered PFS definition specifies several censoring rules. |
| Multiplicity | Two primary comparisons are posted for the same primary endpoint. |
| Safety denominators | Serious adverse events are reported as affected participants divided by participants at risk for each arm. |
19. Related Tutorials
Learn more about the methods used in this trial:
20. Related Calculators
21. Sources
- ClinicalTrials.gov: NCT02718417 — JAVELIN Ovarian 100.
- Linked publication: PubMed record for PMID 34363762.
Continue with Clinical Biostats statistical methods
Use the related tutorials and calculators to examine the survival-analysis concepts that appear throughout randomized clinical trials.
22. Record Summary
JAVELIN Ovarian 100 provides a compact example of how a randomized three-arm phase 3 trial can be analyzed through time-to-event methods. The registered primary endpoint was BICR-assessed progression-free survival, with two posted comparisons using a stratified log-rank test and a Cox proportional-hazards model stratified by randomization strata. The full analysis set included all randomized participants.
The posted primary analyses report HRs of 1.43 and 1.14, with corresponding 95% confidence intervals of 1.051–1.946 and 0.832–1.565. Secondary analyses provide additional HR estimates for overall survival and investigator-assessed PFS. The registry also reports serious adverse-event counts of 92/328, 118/329, and 64/334 across the three treatment strategies.
The statistical lesson is broader than any individual number: a rigorous interpretation requires attention to the randomized comparison, endpoint definition, analysis population, censoring rules, effect measure, confidence interval, hypothesis test, stratification, and the relationship among multiple comparisons. The ClinicalTrials.gov record supports those methodological conclusions without requiring assumptions about unreported median survival, subgroup effects, or additional statistical procedures.