This page separates reported trial results from statistical interpretation. Trial-specific numerical results and endpoint definitions are taken from the ClinicalTrials.gov record. The official registry record is ClinicalTrials.gov NCT03653507.
1. Trial at a Glance
GLOW was a randomized, double-blind, parallel phase 3 trial comparing zolbetuximab plus CAPOX with placebo plus CAPOX in subjects with CLDN 18.2-positive, HER2-negative, locally advanced unresectable or metastatic gastric or gastroesophageal junction adenocarcinoma.
| Feature | GLOW |
|---|---|
| Phase | Phase 3 |
| Population | CLDN 18.2-positive, HER2-negative, locally advanced unresectable or metastatic gastric or gastroesophageal junction adenocarcinoma |
| Design | Randomized, double-blind, parallel |
| Allocation | Randomized |
| Primary endpoint | Progression Free Survival (PFS) |
| Primary endpoint type | Time-to-event |
| Hypothesis type | Superiority |
| Enrollment | 507 |
| Trial status | Active, not recruiting |
| Start | November 28, 2018 |
| Primary completion | October 25, 2022 |
| Lead sponsor | Astellas Pharma Global Development, Inc. |
| ClinicalTrials.gov | NCT03653507 |
2. Clinical Question
The central question was whether first-line zolbetuximab added to CAPOX could improve progression-free survival compared with placebo plus CAPOX in subjects with CLDN 18.2-positive, HER2-negative, locally advanced unresectable or metastatic gastric or gastroesophageal junction adenocarcinoma.
Population
Subjects with CLDN 18.2-positive, HER2-negative, locally advanced unresectable or metastatic gastric or gastroesophageal junction adenocarcinoma.
Intervention
Zolbetuximab plus CAPOX, with zolbetuximab, oxaliplatin, and capecitabine as the registered interventions.
Comparator
Placebo plus CAPOX, with placebo, oxaliplatin, and capecitabine as the registered interventions.
Primary question
Does zolbetuximab plus CAPOX improve PFS relative to placebo plus CAPOX under a superiority framework?
3. Trial Design
Zolbetuximab + CAPOX
- Zolbetuximab
- Oxaliplatin
- Capecitabine
Placebo + CAPOX
- Placebo
- Oxaliplatin
- Capecitabine
The registry describes the trial as randomized, double-blind, and parallel, with a primary purpose of treatment. The design therefore creates a direct randomized comparison between the two treatment strategies while masking participants and relevant trial personnel according to the registered double-blind design.
4. Randomization, Stratification, and Analysis Populations
The primary PFS analysis was conducted in the full analysis set (FAS). The registry reports that the log-rank analyses were stratified by three factors: region, number of organs with metastatic sites, and prior gastrectomy.
| Analysis population | Role reported in the registry analyses |
|---|---|
| Full analysis set (FAS) | Used for the primary PFS analysis and the reported secondary endpoint analyses except duration of response. |
| FAS — All Objective Responders | Used for the Duration Of Response (DOR) analysis. |
Stratification factors
| Factor | Registered strata |
|---|---|
| Region | Asia vs Non-Asia |
| Number of organs with metastatic sites | 0 to 2 vs ≥ 3 |
| Prior gastrectomy | Yes or No |
Stratification is important because it allows the treatment comparison to account for prespecified factors that may be related to prognosis or treatment allocation. In this trial, the same three factors are explicitly identified in the statistical-analysis notes for the reported log-rank analyses.
5. Endpoints
| Endpoint | Registered definition / time frame | Analysis |
|---|---|---|
| Progression Free Survival (PFS) | PFS was defined as the time from the date of randomization until the date of radiological progressive disease (PD) per RECIST 1.1 by independent review committee (IRC), or death from any cause, whichever was earliest. Time frame: from randomization until 61 months and 12 days. | Stratified log-rank test; hazard ratio |
| Overall Survival (OS) | From the date of randomization until 61 months and 12 days. | Stratified log-rank test; hazard ratio |
| Time to Confirmed Deterioration (TTCD): physical functioning | Using physical functioning as measured by EORTC QLQ-C30; from randomization until 61 months and 12 days. | Stratified log-rank test; hazard ratio |
| Time to Confirmed Deterioration (TTCD): abdominal pain and discomfort | Using Oesophago-gastric Questionnaire (OG25) on abdominal pain and discomfort as measured by EORTC QLQ-OG25 plus STO22 Belching Subscale; from randomization until 61 months and 12 days. | Stratified log-rank test; hazard ratio |
| Time to Confirmed Deterioration (TTCD): global health status | Using global health status as measured by EORTC QLQ-C30; from randomization until 61 months and 12 days. | Stratified log-rank test; hazard ratio |
| Duration Of Response (DOR) | From first response (CR/PR) until 61 months and 12 days. | Stratified log-rank test; hazard ratio |
| Objective Response Rate (ORR) | From the date of randomization until 61 months and 12 days. | Cochran-Mantel-Haenszel test |
Primary endpoint definition in statistical terms
PFS is a composite time-to-event endpoint. A participant can experience the event because of radiological progressive disease or because of death from any cause, with whichever occurs first defining the event time. Participants who have not experienced the defined event contribute follow-up information until the applicable censoring point.
6. Statistical Methodology
Kaplan-Meier estimation
Progression-free survival and the other time-to-event outcomes can be represented using Kaplan-Meier estimation. The Kaplan-Meier estimator is designed for data in which some participants may not have experienced the event by the end of their observed follow-up.
where di is the number of events at time ti and ni is the number at risk immediately before that time.
The registry does not provide Kaplan-Meier event-time coordinates in the ClinicalTrials.gov record, so this page does not attempt to reconstruct a survival curve from the summary hazard ratio and confidence interval.
Stratified log-rank test
The primary PFS comparison was performed using a log-rank test with stratification. The reported analysis was stratified by region, number of organs with metastatic sites, and prior gastrectomy.
The log-rank framework compares the observed and expected numbers of events between randomized groups over follow-up. Stratification performs that comparison while maintaining separate risk-set contributions across the prespecified strata.
Hazard ratio
The reported effect measure for the time-to-event analyses is the hazard ratio (HR). An HR below 1 indicates a lower estimated instantaneous event rate in the zolbetuximab-plus-CAPOX group relative to the comparator under the time-to-event model.
A hazard ratio is a relative time-to-event measure. It is not the same as an absolute risk difference, a probability of being progression-free at a particular time, or the percentage of patients who benefit.
Cochran-Mantel-Haenszel test
Objective Response Rate was analyzed with a Cochran-Mantel-Haenszel test. This is a categorical-data method that can provide a treatment comparison while accounting for stratification variables.
Superiority testing
The registry identifies the hypothesis type for the reported analyses as superiority. The objective is therefore to evaluate whether the treatment comparison provides evidence of a difference favoring the intervention rather than whether the intervention is merely no worse than a prespecified non-inferiority margin.
7. Primary Result: Progression-Free Survival
PFS was the registered primary endpoint. The analysis population was the FAS, and the analysis used a stratified log-rank test. The reported time frame was from the date of randomization until 61 months and 12 days.
Hazard ratio for progression or death
95% CI: 0.552–0.860 · P = 0.0005
CAPOX + zolbetuximab versus the comparator, FAS, two-sided confidence interval.
| Primary endpoint | Analysis population | Method | Effect estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Progression Free Survival (PFS) | FAS | Stratified log-rank test | HR 0.689 | 0.552–0.860 | 0.0005 |
The analysis was stratified by region (Asia vs Non-Asia), number of organs with metastatic sites (0 to 2 vs ≥ 3), and prior gastrectomy (Yes or No).
The reported HR of 0.689 means that the estimated instantaneous rate of the PFS event was about 68.9% of that in the comparator group under the reported time-to-event analysis. Expressed as a relative quantity, 1 − 0.689 corresponds to approximately a 31.1% lower estimated hazard for progression or death.
The HR does not mean that 31.1% of participants avoided progression, that every participant experienced a 31.1% reduction in risk, or that the absolute improvement in PFS was 31.1 percentage points. Those are different quantities.
The 95% CI of 0.552–0.860 describes statistical uncertainty around the estimated HR under the analysis framework. Because the interval is entirely below 1, the reported estimate is compatible with a lower hazard in the zolbetuximab-plus-CAPOX group across the confidence-interval range.
The p-value of 0.0005 addresses evidence against the null hypothesis under the specified testing framework. It does not measure the size of the treatment effect. Effect size is communicated by the HR, while precision is communicated by the confidence interval.
As with other hazard-ratio analyses, interpretation also depends on the time-to-event framework and censoring assumptions. A single HR is a relative summary over the analyzed follow-up and should not be interpreted as a constant individual-level risk reduction.
8. Secondary Result: Overall Survival
Overall Survival was reported as a secondary time-to-event endpoint. The analysis used the FAS and a stratified log-rank test over the registered period from randomization until 61 months and 12 days.
Hazard ratio for overall survival
95% CI: 0.622–0.936 · P = 0.0047
CAPOX + zolbetuximab versus the comparator, FAS, two-sided confidence interval.
The reported OS HR of 0.763 corresponds to an estimated instantaneous rate of death about 76.3% of that in the comparator group under the reported time-to-event analysis. The corresponding relative expression is approximately a 23.7% lower estimated hazard of death.
This does not mean that 23.7% of patients were saved, that survival increased by 23.7 percentage points, or that an individual patient's probability of death was reduced by exactly 23.7%.
The 95% CI of 0.622–0.936 indicates uncertainty around the estimated relative treatment effect. The interval remains below 1, while its width shows that the exact magnitude of the relative effect is not known with perfect precision.
The p-value of 0.0047 indicates the strength of evidence against the null hypothesis within the reported superiority testing framework. It is not a measure of how clinically important the HR is, and it should not be read as a probability that the treatment works.
Because OS is a time-to-event endpoint, censoring and the underlying hazard structure remain important to interpretation. The ClinicalTrials.gov record does not provide median OS or time-specific survival probabilities, so those quantities are not reported on this page.
9. Secondary Time-to-Confirmed-Deterioration Results
Three Time to Confirmed Deterioration (TTCD) endpoints were reported. Each used the FAS and a stratified log-rank test with the same three stratification factors used for the primary PFS analysis.
| Endpoint | HR | 95% CI | P-value |
|---|---|---|---|
| TTCD using physical functioning measured by EORTC QLQ-C30 | 1.012 | 0.772–1.328 | 0.4654 |
| TTCD using OG25 abdominal pain and discomfort measured by EORTC QLQ-OG25 plus STO22 Belching Subscale | 1.030 | 0.673–1.577 | 0.4478 |
| TTCD using global health status measured by EORTC QLQ-C30 | 0.869 | 0.655–1.153 | 0.1670 |
These are hazard-ratio analyses of time to confirmed deterioration rather than direct comparisons of mean questionnaire scores. An HR near 1 indicates similar estimated instantaneous deterioration rates between groups under the reported model; an HR above or below 1 indicates the direction of the estimated relative hazard.
10. Secondary Result: Duration Of Response
Duration Of Response was analyzed from first response (CR/PR) until 61 months and 12 days. Unlike the other reported time-to-event analyses, the analysis population was the FAS — All Objective Responders.
Hazard ratio for duration of response
95% CI: 0.552–1.105 · P = 0.0826
CAPOX + zolbetuximab versus the comparator among objective responders.
The analysis used a stratified log-rank test with stratification by region, number of organs with metastatic sites, and prior gastrectomy.
The DOR result should be interpreted differently from the primary PFS analysis because it is restricted to participants who achieved an objective response. Conditioning on response changes the analysis population and therefore changes the question being answered.
11. Secondary Result: Objective Response Rate
Objective Response Rate was analyzed as a binary endpoint from randomization until 61 months and 12 days. The registry reports a Cochran-Mantel-Haenszel analysis in the FAS.
Formal categorical comparison
Cochran-Mantel-Haenszel test
The registry statistical-analysis data do not report an ORR effect estimate or confidence interval.
The registry supplies a p-value of 0.2219 for the ORR comparison but does not provide an ORR effect estimate or confidence interval in the ClinicalTrials.gov record.
A Cochran-Mantel-Haenszel analysis normally evaluates the association between treatment and a binary outcome while accounting for specified strata. Without the corresponding response proportions or an effect estimate, it would be inappropriate to infer the magnitude or direction of the ORR difference from the p-value alone.
This illustrates an important reporting principle: a p-value without an effect estimate does not communicate the size of a treatment difference. For a binary endpoint, response proportions together with an appropriate effect measure and confidence interval would provide a more complete description.
12. Safety Results
The ClinicalTrials.gov record reports serious adverse events by intervention as affected / at risk. These values are reported at the intervention level and should not automatically be treated as mutually exclusive randomized-arm counts because participants receiving combination therapy are exposed to more than one registered intervention.
| Registered intervention | Serious adverse events affected / at risk |
|---|---|
| Zolbetuximab | 123 / 254 |
| Capecitabine | 249 / 503 |
| Oxaliplatin | 249 / 503 |
| Placebo | 126 / 249 |
13. Statistical Methods Explained
Why was a log-rank test used for PFS?
PFS records the time from randomization until progression or death, so it is a time-to-event endpoint. The log-rank test is designed to compare event-time distributions between treatment groups while incorporating the timing of events and censored observations rather than reducing every participant to a simple yes/no outcome at one fixed date.
What does an HR of 0.689 mean?
An HR of 0.689 means that the estimated instantaneous rate of the PFS event in the zolbetuximab-plus-CAPOX group was 0.689 times that of the comparator under the reported analysis. It can be expressed as approximately a 31.1% lower estimated hazard, but it is not a 31.1-percentage-point improvement in PFS and does not mean that every patient has the same relative reduction.
Why was the analysis stratified?
The registry explicitly reports stratification by region, number of organs with metastatic sites, and prior gastrectomy. Stratified analysis allows the treatment comparison to account for these prespecified factors rather than treating the entire population as if those strata were irrelevant to the comparison.
What does the 95% confidence interval tell us?
The 95% CI describes statistical uncertainty around the estimated treatment effect under the specified analysis framework. For PFS, the interval was 0.552–0.860. It does not describe the range of outcomes that individual patients will experience, nor does it mean that 95% of patients have an HR somewhere inside that interval.
Why does the p-value not measure effect size?
The p-value quantifies evidence against a null hypothesis under a specified statistical model and testing framework. It does not tell us whether the treatment effect is large or small. The HR communicates relative effect size, while the confidence interval communicates the uncertainty around that estimate.
Why is the DOR analysis population different from the PFS population?
PFS begins at randomization and includes all participants in the FAS. DOR begins at first response and was analyzed in the FAS among all objective responders. Consequently, DOR answers a conditional question about the duration of response among responders rather than the overall randomized treatment effect from the point of randomization.
What is the difference between PFS and TTCD?
PFS is defined around radiological progressive disease or death. TTCD endpoints in this registry are based on confirmed deterioration in specified patient-reported quality-of-life domains. They therefore capture different types of clinical events and should not be interpreted as interchangeable measures.
14. Interpreting the Primary Hazard Ratio
The primary PFS HR of 0.689 is a relative time-to-event estimate comparing CAPOX plus zolbetuximab with the comparator under the reported stratified analysis. The estimate is below 1, indicating a lower estimated hazard of progression or death in the zolbetuximab group.
It does not mean that 31.1% more patients were progression-free, that progression was prevented in 31.1% of patients, or that the treatment produces the same relative benefit at every individual follow-up time.
The 95% CI of 0.552–0.860 places uncertainty around the estimated HR. The entire interval lies below 1, but the interval still permits a range of plausible effect magnitudes under the statistical framework.
The p-value of 0.0005 addresses statistical evidence against the null hypothesis. It does not replace the HR or its confidence interval when describing the magnitude and precision of the treatment effect.
15. PFS and OS Answer Different Questions
Both PFS and OS are time-to-event endpoints, but their events are different. PFS captures the first occurrence of radiological progression or death, whereas OS concerns death from any cause. Consequently, the two hazard ratios should not be treated as interchangeable measurements of the same outcome.
| Feature | PFS | OS |
|---|---|---|
| Role | Primary endpoint | Secondary endpoint |
| Event framework | Radiological PD or death, whichever is earliest | Death from any cause |
| Analysis population | FAS | FAS |
| Method | Stratified log-rank test | Stratified log-rank test |
| HR | 0.689 | 0.763 |
| 95% CI | 0.552–0.860 | 0.622–0.936 |
| P-value | 0.0005 | 0.0047 |
The difference between these endpoints illustrates why a clinical trial can generate several complementary measures of treatment effect. PFS captures disease-control timing, while OS captures survival regardless of the immediate cause of death.
16. Stratified Analysis in This Trial
The reported log-rank analyses were stratified using three variables. Stratification is especially useful when the trial design identifies factors that should be accounted for in the primary comparison.
Region
Asia versus Non-Asia. Regional differences can be incorporated into the stratified comparison rather than allowing the overall event comparison to ignore this prespecified factor.
Metastatic organ burden
0 to 2 organs versus ≥ 3 organs with metastatic sites. This represents a registered stratification factor for the reported time-to-event analyses.
Prior gastrectomy
Yes or No. This was also included as a stratification factor in the reported log-rank analyses.
Why it matters
The resulting comparison is a stratified treatment comparison rather than an unadjusted log-rank comparison that ignores the registered strata.
17. Time-to-Event Endpoints and Censoring
Time-to-event analysis is necessary when participants can have different lengths of observed follow-up. In GLOW, the registered PFS definition begins at randomization and ends when radiological PD or death occurs, whichever is earliest.
A participant who has not experienced the event during observed follow-up may contribute censored information. Censoring allows that participant to contribute information about the event-free period that was actually observed rather than being discarded simply because the event did not occur during follow-up.
18. Confidence Intervals and Statistical Precision
The confidence intervals reported for the time-to-event endpoints provide an important complement to their point estimates.
| Endpoint | HR | 95% CI | Interpretive role |
|---|---|---|---|
| PFS | 0.689 | 0.552–0.860 | Primary treatment-effect estimate with uncertainty interval |
| OS | 0.763 | 0.622–0.936 | Secondary treatment-effect estimate with uncertainty interval |
| TTCD — physical functioning | 1.012 | 0.772–1.328 | Relative deterioration-hazard estimate |
| TTCD — abdominal pain/discomfort | 1.030 | 0.673–1.577 | Relative deterioration-hazard estimate |
| TTCD — global health status | 0.869 | 0.655–1.153 | Relative deterioration-hazard estimate |
| DOR | 0.781 | 0.552–1.105 | Relative response-duration hazard among responders |
The widths of these intervals also illustrate why a point estimate should never be interpreted in isolation. An estimate such as 0.781 for DOR is accompanied by a range extending above 1, while the primary PFS interval remains below 1.
19. Limitations
- Summary-data limitation: the ClinicalTrials.gov record does not provide median PFS, median OS, time-specific survival probabilities, event counts by randomized arm, or Kaplan-Meier coordinates. Those quantities are therefore not reported here.
- Effect-estimate limitation for ORR: the registry statistical-analysis record provides a Cochran-Mantel-Haenszel p-value of 0.2219 but does not provide the ORR effect estimate or confidence interval in the ClinicalTrials.gov record.
- Safety reporting structure: serious adverse events are reported by registered intervention rather than as a conventional mutually exclusive randomized-arm table. Combination therapy means the same participant can be represented under multiple intervention records.
- Hazard-ratio interpretation: a single HR is a relative time-to-event summary and should not be interpreted as an individual-level probability or an absolute risk difference.
- Censoring: time-to-event estimates depend on the handling and assumptions associated with censored observations. The ClinicalTrials.gov record does not provide the detailed censoring rules needed for an independent reconstruction.
- Subgroup information: the statistical analyses posted on ClinicalTrials.gov identify stratification factors but do not provide subgroup-specific efficacy estimates. No subgroup conclusions should therefore be inferred from the stratification variables alone.
- Multiplicity: the ClinicalTrials.gov record does not describe an alpha-allocation or multiplicity-adjustment strategy. No additional multiplicity claims are made on this page.
- Non-inferiority: this was a superiority framework, and the ClinicalTrials.gov record does not identify a non-inferiority margin. Non-inferiority logic therefore does not apply to the reported primary analysis.
20. Why This Trial Matters Statistically
GLOW is a useful teaching case because the registry results bring together several core clinical-trial methods without requiring the analysis to be reduced to a single p-value.
| Concept | How it appears in GLOW |
|---|---|
| Randomization | The phase 3 trial used randomized allocation between two parallel treatment strategies. |
| Blinding | The registry identifies the study as double-blind. |
| Time-to-event analysis | PFS was the primary endpoint, with OS, three TTCD endpoints, and DOR also analyzed as time-to-event outcomes. |
| Kaplan-Meier framework | The PFS endpoint is a time-to-event measure for which Kaplan-Meier estimation is the standard descriptive framework. |
| Log-rank testing | The reported primary PFS analysis and several secondary endpoints used the log-rank test. |
| Hazard ratio | HR was the reported effect measure for PFS, OS, TTCD, and DOR. |
| Confidence intervals | Two-sided 95% CIs were reported for the time-to-event HR estimates. |
| Stratified analysis | Region, metastatic-organ count, and prior gastrectomy were used as stratification factors. |
| Categorical analysis | ORR was analyzed with the Cochran-Mantel-Haenszel test. |
| Analysis populations | Most reported analyses used the FAS, while DOR used the FAS among all objective responders. |
The statistical story is therefore broader than the primary p-value. The trial combines randomized treatment allocation, blinded assessment, a prespecified time-to-event primary endpoint, stratified hypothesis testing, relative effect estimation, uncertainty intervals, and categorical secondary analysis.
21. Related Tutorials
Learn more about the methods used in this trial:
22. Related Statistical Calculators
23. Sources
- ClinicalTrials.gov: GLOW — NCT03653507.
- PubMed record: PMID 42095986.
- PubMed record: PMID 40680855.
- PubMed record: PMID 38861294.
- PubMed record: PMID 37524953.
Continue through the Clinical Biostats statistical tutorials
Explore the underlying survival-analysis, clinical-trial, categorical-data, and inference methods used to understand randomized trial results.
24. Record Summary
GLOW provides a clear example of a randomized phase 3 time-to-event analysis. The trial used a double-blind parallel design with 507 enrolled subjects and evaluated zolbetuximab plus CAPOX against placebo plus CAPOX. PFS was the registered primary endpoint and was analyzed in the FAS using a stratified log-rank test, with an HR of 0.689, 95% CI 0.552–0.860, and P-value 0.0005. Secondary analyses included OS, three measures of time to confirmed deterioration, DOR, and ORR, illustrating how different statistical methods answer different clinical questions.
The most important statistical distinction is between effect size, precision, and evidence against a null hypothesis. The hazard ratio describes the relative time-to-event effect, the confidence interval describes uncertainty around that estimate, and the p-value describes evidence under the specified testing framework. Keeping those quantities conceptually separate makes the trial results easier to interpret without overstating what any individual statistic can establish.