This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
EVOKE-01 was a randomized, open-label, parallel-group phase 3 trial comparing sacituzumab govitecan-hziy (SG) with docetaxel in participants with advanced or metastatic non-small cell lung cancer (NSCLC). The registry reports 603 participants, one primary time-to-event endpoint, and posted analyses for overall survival plus four secondary endpoints.
| Feature | EVOKE-01 |
|---|---|
| Phase | Phase 3 |
| Condition | Non-Small Cell Lung Cancer |
| Brief title | Study of Sacituzumab Govitecan (SG) Versus Docetaxel in Participants With Advanced or Metastatic Non-Small Cell Lung Cancer (NSCLC) |
| Allocation | Randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 603 |
| Primary endpoint | Overall Survival (OS) |
| Primary endpoint type | Time-to-event |
| Results posted | Yes |
| Outcome measures posted | 9 |
| Statistical analyses posted | 5 |
| Hypothesis type | Superiority |
| Lead sponsor | Gilead Sciences |
2. Clinical Question
The central statistical question was whether sacituzumab govitecan-hziy improves overall survival compared with docetaxel in participants with advanced or metastatic non-small cell lung cancer.
Population
Participants with advanced or metastatic non-small cell lung cancer (NSCLC), as defined by the EVOKE-01 registry record.
Intervention
Sacituzumab govitecan-hziy (SG), classified in the registry as a biological intervention.
Comparator
Docetaxel, classified in the registry as a drug intervention.
Primary question
Does randomized assignment to SG produce a difference in overall survival compared with randomized assignment to docetaxel?
3. Trial Design
Sacituzumab Govitecan
- Sacituzumab govitecan-hziy (SG)
- Registry intervention type: biological
Docetaxel
- Docetaxel
- Registry intervention type: drug
The registry identifies EVOKE-01 as a randomized, parallel-group, unmasked phase 3 treatment study. The available trial data do not provide treatment-specific randomized sample sizes, dosing schedules, stratification factors, crossover rules, or treatment-duration rules, so those features are not inferred here.
Trial timeline
Trial start
The registered study start date was November 17, 2021.
Primary completion
The registered primary completion date was November 29, 2023.
Active, not recruiting
the ClinicalTrials.gov record identifies the study status as ACTIVE_NOT_RECRUITING.
4. Endpoints
| Endpoint | Role | Time frame | Endpoint type | Registry analysis |
|---|---|---|---|---|
| Overall Survival (OS) | Primary | Up to 24.4 months | Time-to-event | Kaplan-Meier estimation; log-rank test; hazard ratio |
| Progression-free Survival (PFS) Assessed by Investigator Per RECIST Version 1.1 | Secondary | Up to 24.4 months | Time-to-event | Log-rank test; hazard ratio |
| Objective Response Rate (ORR) Assessed by Investigator Per RECIST Version 1.1 | Secondary | Up to 24.4 months | Binary | Cochran-Mantel-Haenszel test; risk difference |
| Time to First Deterioration in Shortness of Breath Domain as Measured by NSCLC-SAQ | Secondary | Up to 24.4 months | Time-to-event | Log-rank test; hazard ratio |
| Time to First Deterioration in NSCLC-SAQ Total Score | Secondary | Up to 24.4 months | Time-to-event | Log-rank test; hazard ratio |
Primary endpoint definition: Overall Survival
The registry defines OS as the time from the date of randomization until the date of death from any cause. OS was estimated using the Kaplan-Meier estimate. Participants without documentation of death were censored on the date they were last known to be alive.
5. Statistical Methodology
Intention-to-treat analysis
The primary OS analysis used the Intent to Treat (ITT) Analysis Set, defined in the registry as including all randomized participants according to the treatment arm to which each participant was randomized. This preserves the randomized comparison rather than redefining treatment groups according to treatment actually received.
Kaplan-Meier estimation
The registry states that OS was estimated using the Kaplan-Meier method. Kaplan-Meier estimation is designed for time-to-event data in which some participants may be censored before the event occurs. In EVOKE-01, participants without documentation of death were censored on the date they were last known to be alive.
Here, di represents the number of events at an observed event time and ni represents the number of participants at risk immediately before that time.
Log-rank testing
The formal time-to-event analyses posted in the registry use the log-rank test. The log-rank procedure compares the observed pattern of events between treatment groups across follow-up while accounting for the timing of events and censoring.
Hazard ratio
For OS, the reported effect measure is a hazard ratio. A hazard ratio compares the estimated instantaneous event rates between the two randomized groups over the analyzed follow-up under the time-to-event model and analysis framework.
The HR is a relative time-to-event measure. It is not a risk ratio, an absolute difference in survival probability, or the percentage of participants who benefit.
Cochran-Mantel-Haenszel analysis
The registry uses a Cochran-Mantel-Haenszel test for investigator-assessed ORR. The corresponding effect measure is the risk difference, reported as the difference in response proportions between SG and docetaxel.
Confidence intervals
The posted statistical analyses report two-sided 95% confidence intervals for the treatment-effect estimates. A confidence interval communicates the statistical precision of an estimated effect under the model and sampling framework; it does not describe the range of individual patient outcomes.
6. Results: Primary Endpoint — Overall Survival
The primary endpoint was overall survival through the registry time frame of up to 24.4 months. The ITT analysis compared sacituzumab govitecan with docetaxel using a log-rank test and reported a hazard ratio as the effect measure.
Overall survival hazard ratio
95% CI: 0.68–1.04 · P = 0.0534
Two-sided confidence interval · Superiority hypothesis
| Primary endpoint | SG | Docetaxel | Effect estimate | P-value |
|---|---|---|---|---|
| Overall Survival (OS) | ITT | ITT | HR 0.84 (95% CI 0.68–1.04) | 0.0534 |
An OS hazard ratio of 0.84 means that the estimated hazard of death in the SG group was 0.84 times the estimated hazard in the docetaxel group under the reported time-to-event analysis. Expressed as a relative hazard interpretation, this corresponds to a 16% lower estimated hazard for SG relative to docetaxel.
The HR does not mean that 16% fewer participants died, that individual patients had exactly a 16% reduction in their probability of death, or that survival was extended by a fixed percentage. Those interpretations require absolute survival estimates or other measures that are not provided in the ClinicalTrials.gov record.
The two-sided 95% confidence interval of 0.68–1.04 describes uncertainty around the estimated hazard ratio. Because the interval extends above 1, the data are compatible with a range of relative effects that includes an estimated hazard higher than that in the comparator group as well as effects below 1.
The P-value of 0.0534 measures the statistical evidence against the null hypothesis under the specified testing framework; it does not measure the size, clinical importance, or probability of the treatment effect. A P-value should therefore be interpreted together with the hazard ratio and confidence interval rather than as a measure of effect magnitude.
The ClinicalTrials.gov record does not report a proportional-hazards diagnostic, so the extent to which a single HR adequately summarizes the relative event patterns over the entire follow-up cannot be assessed from these data alone.
7. Secondary Endpoint Results
Progression-free Survival
Investigator-assessed PFS per RECIST Version 1.1 was analyzed in participants in the ITT Analysis Set. The registry reports a log-rank test and hazard ratio for the SG versus docetaxel comparison.
PFS hazard ratio
95% CI: 0.77–1.11 · P = 0.1938
Two-sided confidence interval · Superiority hypothesis
The reported PFS hazard ratio of 0.92 corresponds to an estimated hazard that was 0.92 times that of docetaxel in the reported analysis. As a relative hazard interpretation, this is an estimated 8% lower hazard for SG.
The 95% confidence interval of 0.77–1.11 indicates meaningful statistical uncertainty around that estimate and includes 1.00. The P-value of 0.1938 is evidence from the specified test against its null hypothesis; it is not an estimate of effect size and should not be interpreted as the probability that SG has no effect.
Because PFS is a time-to-event endpoint, interpretation also depends on censoring and on how progression events are defined and assessed. The ClinicalTrials.gov record identifies investigator assessment per RECIST Version 1.1 but do not provide enough underlying event-time information to reconstruct the PFS distribution.
Objective Response Rate
Investigator-assessed ORR per RECIST Version 1.1 was analyzed as a binary endpoint in the ITT Analysis Set. The reported statistical method was the Cochran-Mantel-Haenszel test, with difference in proportions as the effect measure.
Risk difference in objective response
95% CI: −10.1 to 1.5 · P = 0.9255
Percentage points · Two-sided confidence interval · Superiority hypothesis
The reported risk difference of −4.3 percentage points represents the reported difference in response proportions between the SG and docetaxel groups, using the registry's effect-measure convention. A negative value indicates that the estimated response proportion in the SG group was lower than in the comparator group by 4.3 percentage points under this analysis.
The 95% confidence interval, −10.1 to 1.5 percentage points, spans zero. Thus, the uncertainty interval includes both a larger negative difference and a positive difference in response proportions.
The P-value of 0.9255 is a test statistic for the specified comparison. It is not the probability that the two treatments are equivalent, nor does it quantify the clinical importance of the observed difference.
Because the ClinicalTrials.gov record does not provide the response counts or response percentages for each arm, the risk difference should be interpreted directly as reported rather than reconstructed into separate arm-specific response rates.
Time to First Deterioration in Shortness of Breath
The registry reports time to first deterioration in the shortness of breath domain as measured by the Non-small Cell Lung Cancer Symptom Assessment Questionnaire (NSCLC-SAQ). Participants in the ITT Analysis Set were analyzed using a log-rank test and hazard ratio.
Hazard ratio for first deterioration in shortness of breath
95% CI: 0.61–0.91 · P = 0.0020
Two-sided confidence interval · Superiority hypothesis
The hazard ratio of 0.75 means that the estimated instantaneous hazard of the first recorded deterioration in the specified shortness-of-breath domain was 0.75 times that in the docetaxel group under the reported analysis. This corresponds to a 25% lower estimated hazard for SG.
The 95% confidence interval of 0.61–0.91 gives the reported range of statistical uncertainty around the estimate and remains below 1.00. The P-value of 0.0020 measures evidence against the null hypothesis under the specified testing framework; it does not indicate that there is a 0.2% probability that the observed treatment effect arose by chance, nor does it measure the magnitude of the effect.
The endpoint is a time-to-event measure. A hazard ratio therefore does not tell us the absolute percentage of participants who eventually experienced deterioration or the median time to deterioration. Those quantities are not included in the ClinicalTrials.gov record.
Time to First Deterioration in NSCLC-SAQ Total Score
The registry also reports time to first deterioration in the NSCLC-SAQ Total Score. Participants in the ITT Analysis Set were analyzed using a log-rank test and hazard ratio.
Hazard ratio for first deterioration in total score
95% CI: 0.66–0.97 · P = 0.0125
Two-sided confidence interval · Superiority hypothesis
The reported hazard ratio of 0.80 corresponds to an estimated instantaneous hazard of first deterioration in the NSCLC-SAQ Total Score that was 0.80 times that in the docetaxel group. This is a 20% lower estimated hazard for SG under the reported analysis.
The 95% confidence interval of 0.66–0.97 quantifies uncertainty around the estimate and remains below 1.00. The P-value of 0.0125 provides statistical evidence under the specified test, but it does not measure effect size or establish the magnitude of benefit for an individual participant.
As with the other time-to-event endpoints, interpretation depends on censoring and on the event definition. The ClinicalTrials.gov record does not provide the underlying deterioration times or censoring records, so no absolute deterioration probability or median time should be inferred.
8. Results Summary
| Endpoint | Role | Method | Effect | 95% CI | P-value |
|---|---|---|---|---|---|
| Overall Survival | Primary | Log-rank | HR 0.84 | 0.68–1.04 | 0.0534 |
| Progression-free Survival | Secondary | Log-rank | HR 0.92 | 0.77–1.11 | 0.1938 |
| Objective Response Rate | Secondary | Cochran-Mantel-Haenszel | Risk difference −4.3 | −10.1 to 1.5 | 0.9255 |
| Time to first deterioration in shortness of breath | Secondary | Log-rank | HR 0.75 | 0.61–0.91 | 0.0020 |
| Time to first deterioration in NSCLC-SAQ Total Score | Secondary | Log-rank | HR 0.80 | 0.66–0.97 | 0.0125 |
The five posted statistical analyses show that the direction and magnitude of the estimated effects differ across endpoints. The OS estimate was below 1, as were all four secondary effect estimates, but the corresponding confidence intervals and P-values vary substantially. This is one reason a trial should be interpreted endpoint by endpoint rather than reduced to a single numerical result.
9. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment arm as the number affected among those at risk. No other safety-event categories or treatment-emergent adverse-event summaries are included in the ClinicalTrials.gov record.
| Safety measure | Sacituzumab Govitecan (SG) | Docetaxel |
|---|---|---|
| Serious adverse events | 137 / 296 | 124 / 288 |
The affected/at-risk figures correspond to approximately 46.3% and 43.1% when expressed as simple descriptive proportions of the registry-reported counts. These percentages are derived only for visual orientation from the reported numerators and denominators; the registry does not provide a formal between-arm statistical analysis for serious adverse events in the ClinicalTrials.gov record.
10. Statistical Methods Explained
Why was a log-rank test used for overall survival?
Overall survival is a time-to-event endpoint because the analysis uses both the timing of death and the follow-up information available for participants who have not yet had the event. A log-rank test is designed to compare survival experience between randomized groups while incorporating the timing of events and censoring. EVOKE-01's registry specifically reports the log-rank method for OS, PFS, and both deterioration endpoints.
What does an OS hazard ratio of 0.84 mean?
A hazard ratio of 0.84 means that the estimated instantaneous hazard of death for SG was 0.84 times the estimated hazard for docetaxel under the reported analysis. It can be expressed as a 16% lower estimated hazard. It does not mean that 16% of participants benefited, that survival time increased by 16%, or that an individual participant's probability of death fell by exactly 16%.
Why is the confidence interval important?
The point estimate alone does not show how precisely the treatment effect was estimated. The OS 95% confidence interval runs from 0.68 to 1.04. This interval demonstrates that the estimate has uncertainty and includes 1.00, the value corresponding to equal hazards under a conventional hazard-ratio interpretation. The confidence interval is therefore essential for understanding the precision and range of effects compatible with the statistical analysis.
Why does the P-value not measure effect size?
A P-value summarizes evidence against a null hypothesis under the specified statistical model and testing framework. It does not tell us how large the treatment effect is. For example, the deterioration endpoints have P-values of 0.0020 and 0.0125, but their effect sizes must be understood from their hazard ratios of 0.75 and 0.80 and their corresponding confidence intervals.
How is a risk difference different from a hazard ratio?
A risk difference compares two proportions directly. For ORR, EVOKE-01 reports a risk difference of −4.3 percentage points. A hazard ratio instead compares event rates over time. The two measures cannot be interpreted interchangeably: a risk difference describes a difference in proportions, while a hazard ratio describes a relative time-to-event measure.
Why does the ITT analysis matter?
The primary OS analysis included all randomized participants according to their assigned treatment arm. This intention-to-treat principle preserves the comparison created by randomization. It avoids redefining the treatment groups based on post-randomization events such as adherence or treatment exposure, although the exact handling of such events and missing information is not specified in the ClinicalTrials.gov record.
What should be considered when interpreting a hazard ratio?
A hazard ratio compresses a time-to-event comparison into a single relative measure. That summary is most straightforward when the relative hazards are reasonably represented by a common proportional-hazards relationship over follow-up. The registry-reported EVOKE-01 data do not report a proportional-hazards assessment, so the adequacy of that assumption cannot be evaluated from the ClinicalTrials.gov record.
11. Understanding the Primary Result
The primary OS estimate was HR 0.84. Under the reported analysis, the estimated instantaneous hazard of death was 16% lower for SG relative to docetaxel.
The 95% confidence interval was 0.68–1.04. This interval is substantially more informative than the point estimate alone because it communicates the uncertainty surrounding the estimated treatment effect.
The reported two-sided P-value was 0.0534. A P-value is not an effect-size measure and should not be used as a substitute for the hazard ratio or its confidence interval.
The ClinicalTrials.gov record does not provide median OS, Kaplan-Meier survival probabilities at specific time points, numbers of deaths by treatment arm, or a reconstructed survival curve. Those quantities are therefore not presented here.
12. Endpoint-by-Endpoint Statistical Perspective
| Question | OS | PFS | ORR | Shortness of breath deterioration | NSCLC-SAQ total-score deterioration |
|---|---|---|---|---|---|
| Endpoint type | Time-to-event | Time-to-event | Binary | Time-to-event | Time-to-event |
| Analysis set | ITT | ITT | ITT | ITT | ITT |
| Primary statistical method | Log-rank | Log-rank | Cochran-Mantel-Haenszel | Log-rank | Log-rank |
| Effect measure | Hazard ratio | Hazard ratio | Risk difference | Hazard ratio | Hazard ratio |
| Estimate | 0.84 | 0.92 | −4.3 | 0.75 | 0.80 |
This endpoint structure illustrates why clinical-trial interpretation requires more than identifying which P-values are below a conventional threshold. OS and PFS are time-to-event endpoints, while ORR is binary. The symptom endpoints are also time-to-event measures, but their events represent deterioration rather than death or progression. Each estimate therefore answers a different statistical question.
13. Missing Design Information and What It Means for Interpretation
The registry-reported EVOKE-01 registry data provide the principal design classification and the posted statistical analyses, but they do not provide several design details that would normally be important in a full statistical analysis plan.
| Design topic | Information in the ClinicalTrials.gov record | Interpretation |
|---|---|---|
| Stratification | Not reported | No stratification factors are inferred. |
| Non-inferiority margin | Not applicable to the reported hypothesis type | The registered hypothesis type is superiority. |
| Crossover | Not reported | No crossover effect is inferred. |
| Factorial design | Not reported | The registered design model is parallel. |
| Multiplicity strategy | Not reported | The ClinicalTrials.gov record does not establish a familywise error-control strategy across endpoints. |
| Interim analysis | Not reported | No interim-analysis rule is inferred. |
| Missing-data / imputation strategy | Not reported | No imputation method is inferred. |
| Bayesian methods | Not reported | No Bayesian analysis is inferred. |
This distinction is important. Absence of a design feature from the ClinicalTrials.gov record does not establish that the underlying protocol contained no such provision; it means only that the feature is not available in the ClinicalTrials.gov record.
14. Statistical Interpretation of the Secondary Endpoints
PFS
The estimated hazard ratio was 0.92, with a 95% CI of 0.77–1.11 and P = 0.1938. The confidence interval includes 1.00.
Objective response
The reported risk difference was −4.3 percentage points, with a 95% CI of −10.1 to 1.5 and P = 0.9255.
Shortness of breath
The reported hazard ratio was 0.75, with a 95% CI of 0.61–0.91 and P = 0.0020.
NSCLC-SAQ total score
The reported hazard ratio was 0.80, with a 95% CI of 0.66–0.97 and P = 0.0125.
These results should not be collapsed into a single composite effect. Each endpoint has its own estimand, event definition, statistical scale, and uncertainty. In particular, a response-rate difference cannot be directly compared numerically with a hazard ratio.
15. Limitations
- Summary-data limitation: the ClinicalTrials.gov record contains selected statistical analyses but not the underlying participant-level or event-time data.
- Limited absolute-effect information: the ClinicalTrials.gov record does not report median OS, median PFS, fixed-time survival probabilities, or arm-specific response percentages.
- Hazard-ratio assumptions: the registry reports hazard ratios but does not provide a proportional-hazards assessment in the ClinicalTrials.gov record.
- Endpoint multiplicity: the ClinicalTrials.gov record identifies one primary endpoint and four posted secondary analyses but do not provide a multiplicity-adjustment strategy. The individual P-values therefore should not automatically be interpreted as independent confirmatory tests within a familywise error framework.
- Analysis-plan detail: stratification factors, interim-monitoring rules, missing-data methods, and other protocol-level statistical details are not provided in the ClinicalTrials.gov record.
- Open-label design: the registry identifies masking as none. This is particularly relevant when interpreting endpoints involving investigator assessment or patient-reported symptoms because knowledge of treatment assignment can potentially influence assessment or reporting. The ClinicalTrials.gov record does not quantify such an effect.
- Safety denominators: serious adverse-event counts are reported as 137/296 for SG and 124/288 for docetaxel, which differ from the overall enrollment of 603. The ClinicalTrials.gov record does not explain the denominator construction.
- Generalizability: the registry record describes a specific phase 3 population with advanced or metastatic NSCLC; applicability to populations outside the enrolled study population requires information not provided here.
16. Why This Trial Matters Statistically
EVOKE-01 is useful as a statistical teaching case because the same randomized comparison generates several distinct estimands: overall survival, progression-free survival, objective response, and patient-reported deterioration endpoints. The trial therefore demonstrates how different endpoint types require different statistical tools and how apparently similar measures can answer substantially different questions.
| Concept | How it appears in EVOKE-01 |
|---|---|
| Randomization | Participants were randomized to SG or docetaxel. |
| Parallel-group design | The registered design model is parallel. |
| Intention-to-treat analysis | The primary OS analysis included all randomized participants according to assigned treatment arm. |
| Kaplan-Meier estimation | Used to estimate overall survival. |
| Log-rank testing | Used for OS, PFS, and both time-to-deterioration endpoints. |
| Hazard ratio | Used as the effect measure for OS, PFS, and the two deterioration endpoints. |
| Confidence interval | Two-sided 95% intervals were reported for the posted treatment-effect estimates. |
| Cochran-Mantel-Haenszel test | Used for the investigator-assessed binary ORR endpoint. |
| Risk difference | Used as the reported effect measure for ORR. |
| Patient-reported endpoint analysis | Time to first deterioration was evaluated for the shortness-of-breath domain and NSCLC-SAQ Total Score. |
17. Statistical Methods Explained in More Depth
Kaplan-Meier estimation and censoring
For a time-to-event endpoint, participants may enter the analysis without having experienced the event by the end of their available observation. Kaplan-Meier estimation allows those participants to contribute information up to the point at which they are censored. For EVOKE-01 OS, the registry explicitly states that participants without documentation of death were censored on the date they were last known to be alive.
Why an ITT population is natural for a randomized superiority trial
The ITT principle analyzes participants according to randomized assignment. This maintains the treatment comparison established at randomization and avoids allowing post-randomization treatment decisions to redefine the principal efficacy comparison. In EVOKE-01, the primary OS analysis is explicitly based on the ITT Analysis Set.
Why ORR uses a different statistical scale
ORR is recorded as a binary outcome rather than as a time-to-event variable. The registry therefore uses a Cochran-Mantel-Haenszel test and reports a difference in proportions. This produces an effect on the percentage-point scale rather than the hazard-ratio scale used for OS and the other time-to-event outcomes.
Why a P-value cannot establish clinical importance
A P-value and an effect estimate answer different questions. The P-value addresses statistical evidence under a null hypothesis, while the effect estimate describes the observed magnitude on its particular scale. A complete interpretation therefore considers the estimate, confidence interval, endpoint definition, analysis population, and clinical context together.
Why confidence intervals are more informative than point estimates alone
A point estimate such as HR 0.84 is only one estimate from the data. The corresponding 95% CI of 0.68–1.04 shows the uncertainty surrounding it. Confidence intervals are particularly useful when the estimate is near the null value because they show how much statistical uncertainty remains around the treatment effect.
Why open-label status matters for some endpoints
The registry identifies masking as none. For OS, death is a comparatively concrete event, although endpoint ascertainment and censoring remain important. For investigator-assessed response and patient-reported deterioration, awareness of treatment assignment can potentially affect assessment or reporting. The ClinicalTrials.gov record does not quantify the magnitude of any such influence, so the open-label design is a methodological consideration rather than a quantified bias estimate.
18. Primary Endpoint vs Secondary Endpoints
| Endpoint | Registry role | Interpretive emphasis |
|---|---|---|
| Overall Survival | Primary | Time from randomization to death from any cause; Kaplan-Meier estimate and log-rank comparison. |
| Progression-free Survival | Secondary | Investigator-assessed RECIST Version 1.1 time-to-event endpoint. |
| Objective Response Rate | Secondary | Investigator-assessed RECIST Version 1.1 binary endpoint. |
| Time to First Deterioration in Shortness of Breath | Secondary | Patient-reported time-to-deterioration endpoint measured by NSCLC-SAQ. |
| Time to First Deterioration in NSCLC-SAQ Total Score | Secondary | Patient-reported time-to-deterioration endpoint. |
The distinction between primary and secondary endpoints is part of the statistical architecture of a trial. The ClinicalTrials.gov record identifies OS as the single registered primary endpoint and the other four posted analyses as secondary endpoints. They do not provide the formal multiplicity procedure, if any, used to control error across the full endpoint family.
19. What the Hazard Ratio Does — and Does Not — Mean
The EVOKE-01 OS hazard ratio was 0.84. Under the reported analysis, the estimated instantaneous hazard of death in the SG group was approximately 84% of the estimated hazard in the docetaxel group.
Equivalently, this is a 16% lower estimated hazard. It does not mean a 16% lower absolute mortality rate, a 16% improvement in median survival, or that every participant experienced the same proportional reduction.
The 95% CI of 0.68–1.04 describes uncertainty around the estimated hazard ratio. Because the interval includes 1.00, the ClinicalTrials.gov record does not establish a precise treatment effect separated from the null on the conventional hazard-ratio scale.
The P-value of 0.0534 describes statistical evidence under the specified two-sided testing framework. It should not be read as a probability that the treatment is effective or ineffective, nor as a measure of the magnitude of the treatment effect.
20. A Closer Look at the Time-to-Event Endpoints
Four of the five posted statistical analyses use time-to-event methods: OS, PFS, time to first deterioration in the shortness-of-breath domain, and time to first deterioration in NSCLC-SAQ Total Score. This common framework does not make the endpoints interchangeable.
Overall survival
The event is death from any cause. The registry explicitly defines the starting point as randomization.
Progression-free survival
The endpoint is investigator-assessed PFS per RECIST Version 1.1. The ClinicalTrials.gov record does not provide its full event definition beyond the registry endpoint label.
Shortness of breath
The event is first deterioration in the shortness-of-breath domain measured by the NSCLC-SAQ.
NSCLC-SAQ total score
The event is first deterioration in the NSCLC-SAQ Total Score.
Because these endpoints have different event definitions, their hazard ratios describe different quantities even though the statistical scale is the same. A hazard ratio of 0.75 for one endpoint cannot be interpreted as if it were the same clinical outcome as an OS hazard ratio of 0.84.
21. Serious Adverse Events: Statistical Perspective
The ClinicalTrials.gov record provides counts affected and at risk rather than a formal treatment-effect estimate. For SG, 137 of 296 participants were affected; for docetaxel, 124 of 288 were affected.
These figures are descriptive safety summaries. The ClinicalTrials.gov record does not provide a confidence interval, P-value, risk difference, risk ratio, or other formal statistical comparison for serious adverse events.
This distinction illustrates an important statistical principle: a table of adverse-event counts is not automatically a hypothesis test. A formal safety comparison would require a prespecified analysis population, an estimand, an appropriate effect measure, and an associated measure of uncertainty. Those details are not included in the ClinicalTrials.gov record.
22. What the Supplied Data Do Not Establish
- Median survival: no median OS or median PFS is provided in the ClinicalTrials.gov record.
- Absolute survival rates: no OS or PFS survival probability at a specific time point is provided.
- Event counts by arm: the ClinicalTrials.gov record does not report the number of OS or PFS events by treatment group.
- Subgroups: no subgroup estimates are included in the statistical analyses posted on ClinicalTrials.gov.
- Stratification: no randomization or analysis stratification factors are reported in the ClinicalTrials.gov record.
- Interim monitoring: no interim-analysis schedule or alpha-spending rule is reported.
- Multiplicity: no formal multiplicity-adjustment procedure is reported.
- Missing-data methods: no imputation or sensitivity-analysis strategy is reported.
- Bayesian analysis: no Bayesian method is reported.
- Crossover: no crossover procedure is reported.
These omissions are not evidence that such methods were absent from the underlying protocol. They simply define the boundary of what can be responsibly concluded from the ClinicalTrials.gov record.
23. Why This Trial Is a Useful Statistical Teaching Example
EVOKE-01 connects several core concepts in clinical-trial statistics. The primary endpoint is a time-to-event outcome analyzed using Kaplan-Meier estimation and a log-rank test, with a hazard ratio as the effect measure. Secondary endpoints demonstrate how the same randomized comparison can be evaluated through additional time-to-event outcomes and through a binary response endpoint using the Cochran-Mantel-Haenszel framework.
| Learning concept | EVOKE-01 example |
|---|---|
| Randomization | Randomized assignment to SG or docetaxel. |
| ITT analysis | Primary OS analysis included all randomized participants according to assigned treatment. |
| Kaplan-Meier estimation | Used for OS estimation. |
| Log-rank test | Used for OS, PFS, and time-to-deterioration analyses. |
| Hazard ratio | Reported for four time-to-event endpoints. |
| Confidence intervals | Two-sided 95% CIs accompany the posted treatment-effect estimates. |
| Cochran-Mantel-Haenszel test | Used for investigator-assessed ORR. |
| Risk difference | Reported for ORR as a difference in proportions. |
| Patient-reported outcomes | Time to deterioration was assessed using NSCLC-SAQ domains and total score. |
24. Related Tutorials
Learn more about the methods used in this trial:
25. Related Statistical Calculators
26. Sources
- ClinicalTrials.gov: NCT05089734 — EVOKE-01.
- Linked publication: PubMed record for PMID 38843511.
Continue through the Clinical Biostats statistical pathway
Explore the statistical methods behind randomized clinical trials, time-to-event endpoints, treatment-effect measures, confidence intervals, and categorical-data analysis.
27. Record Summary
EVOKE-01 provides a useful example of how a randomized phase 3 clinical trial can generate evidence across several endpoint types. The primary endpoint, overall survival, was analyzed in the ITT population using Kaplan-Meier estimation and a log-rank test, with a reported hazard ratio of 0.84, 95% CI 0.68–1.04, and P = 0.0534. Secondary analyses included PFS, ORR, time to first deterioration in the NSCLC-SAQ shortness-of-breath domain, and time to first deterioration in the NSCLC-SAQ Total Score. Their reported effect measures were hazard ratios for the time-to-event endpoints and a risk difference for ORR.
The statistical lesson is broader than any individual P-value. Time-to-event endpoints require attention to censoring, event definitions, Kaplan-Meier estimation, and hazard-ratio interpretation. Binary endpoints require a different effect measure and testing framework. Patient-reported deterioration endpoints introduce another clinically distinct class of time-to-event outcome. Across all of these analyses, estimates should be interpreted together with their confidence intervals, analysis populations, endpoint definitions, and the limitations of the available statistical information.