This page separates reported trial results from statistical interpretation. Numerical results and trial characteristics are restricted to the ClinicalTrials.gov record and the associated posted analysis fields.
1. Trial at a Glance
GADOLIN was a randomized, open-label, parallel-group phase 3 trial evaluating bendamustine alone versus obinutuzumab plus bendamustine in participants with rituximab-refractory, indolent non-Hodgkin's lymphoma.
| Feature | GADOLIN |
|---|---|
| Trial name | GADOLIN |
| Phase | Phase 3 |
| Condition | Non-Hodgkin's Lymphoma |
| Population | Participants with rituximab-refractory, indolent non-Hodgkin's lymphoma |
| Design | Randomized, parallel-group, open-label |
| Allocation | Randomized |
| Primary purpose | Treatment |
| Enrollment | 413 |
| Interventions | Obinutuzumab; Bendamustine |
| Primary endpoints | Number of participants with progressive disease as assessed by IRC or death; Progression-Free Survival as assessed by IRC |
| Hypothesis type | Superiority |
| Status | Completed |
| ClinicalTrials.gov | NCT01059630 |
| Lead sponsor | Genentech, Inc. |
2. Clinical Question
The central statistical question was whether the treatment strategy containing obinutuzumab plus bendamustine differed from bendamustine alone with respect to disease progression or death and other reported efficacy endpoints in participants with rituximab-refractory, indolent non-Hodgkin's lymphoma.
Population
Participants with rituximab-refractory, indolent non-Hodgkin's lymphoma.
Intervention
Obinutuzumab plus bendamustine.
Comparator
Bendamustine alone.
Primary question
Does the obinutuzumab-plus-bendamustine strategy improve the registered efficacy outcomes relative to bendamustine alone?
3. Trial Design
Bendamustine
- Bendamustine alone
- Comparator treatment strategy in the randomized parallel-group design
Obinutuzumab plus bendamustine
- Obinutuzumab
- Bendamustine
- Intervention treatment strategy in the randomized parallel-group design
4. Trial Timing and Registry Structure
Trial start
The registry lists April 30, 2010 as the trial start date.
Primary completion
The registry lists September 30, 2014 as the primary completion date.
Registry status
The trial is listed as completed, with results posted and 33 outcome measures and 12 statistical analyses posted.
5. Primary Endpoints
| Endpoint | Registered definition / time frame | Statistical classification |
|---|---|---|
| Number of Participants With Progressive Disease (PD) as Assessed by Independent Review Committee (IRC) or Death | Baseline until PD or death, whichever occurred first. Assessment was registered at baseline and at specified study visits including 14 days prior to Cycle 4 Day 1. | Binary |
| Progression-Free Survival (PFS) as Assessed by IRC | Baseline until PD or death, whichever occurred first. Assessment was registered at baseline and at specified study visits including 14 days prior to Cycle 4 Day 1. | Time-to-event |
How the registered definitions differ
The first primary endpoint is expressed as a participant-level occurrence of progressive disease or death. The second expresses the same broad clinical event process as a time-to-event endpoint, preserving the timing of progression or death rather than reducing follow-up to a simple yes/no classification.
The registry defines PFS as the time from randomization to the first occurrence of PD or death as assessed by an IRC according to the modified response criteria for indolent Non-Hodgkin's lymphoma. The registry describes PD using criteria including the appearance of a new lesion more than 1.5 cm in any axis during or at the end of therapy and at least a 50% increase from nadir in the sum of product diameter of a relevant lesion measurement.
6. Analysis Populations and Statistical Framework
| Population / feature | Registry information |
|---|---|
| Primary PFS analysis population | ITT population |
| Response analysis population | ITT population; for objective response, the number analyzed signified participants who had at least one post-baseline assessment |
| Induction response analysis | ITT population; the number analyzed signified participants who had reached the end-of-induction response assessment |
| Duration of response analysis | ITT population; the number analyzed signified participants who had objective response at any time during the study |
| DFS in CR analysis | ITT population; the number analyzed signified participants who had an objective response of CR |
| Primary time-to-event method | Log-rank test |
| Categorical response method | Cochran-Mantel-Haenszel test |
| Hypothesis type | Superiority |
The ITT designation is important because randomized participants remain associated with their assigned treatment group for the efficacy comparison. This preserves the treatment comparison created by randomization and avoids redefining the primary efficacy population based solely on treatment exposure or subsequent events.
7. Statistical Methodology
Log-rank testing for time-to-event endpoints
The registry reports the log-rank test for the primary PFS analysis and for several secondary time-to-event analyses. The log-rank framework compares the observed pattern of events between randomized groups over follow-up while accounting for censoring.
The log-rank statistic is based on comparing observed and expected events across event times. It is therefore different from a simple comparison of proportions at one fixed time point.
Hazard ratios
The reported effect measure for the primary PFS analysis is the hazard ratio. A hazard ratio compares the estimated instantaneous event rates between groups within a time-to-event framework.
The hazard ratio is not a probability, not a percentage of participants who benefit, and not the same quantity as a relative risk.
Cochran-Mantel-Haenszel testing
The registry reports the Cochran-Mantel-Haenszel test for objective-response endpoints. This is a categorical-data method that can compare treatment groups while accounting for stratification variables when those strata are defined for the analysis.
Confidence intervals
The posted analyses use two-sided 95% confidence intervals for their reported effect measures. A confidence interval communicates statistical precision around an estimate. It should not be interpreted as a range containing a fixed probability that the true treatment effect lies inside the interval.
Intention-to-treat analysis
The primary PFS analysis was conducted in the ITT population. In a randomized trial, this approach keeps participants associated with their randomized assignment for efficacy analysis, helping preserve the comparability established at randomization.
8. Primary Results
Progression-Free Survival as Assessed by IRC
Primary PFS hazard ratio
95% CI: 0.40–0.70 · P < 0.0001
Analysis population: ITT · Method: log-rank test · Hypothesis: superiority
| Primary endpoint | Comparison | Method | Effect estimate | 95% CI | P-value |
|---|---|---|---|---|---|
| Progression-Free Survival as Assessed by IRC | Bendamustine alone vs obinutuzumab + bendamustine | Log-rank test | HR 0.53 | 0.40–0.70 | <0.0001 |
The reported hazard ratio of 0.53 means that the estimated instantaneous rate of progression or death was approximately 53% as high in the obinutuzumab-plus-bendamustine group as in the bendamustine-alone group under the time-to-event analysis. Equivalently, 1 − 0.53 = 0.47, so the estimate corresponds to an approximately 47% lower estimated hazard of progression or death.
It does not mean that 47% of participants avoided progression, that 47% were cured, or that each individual participant experienced exactly a 47% reduction in risk.
The 95% confidence interval of 0.40–0.70 describes uncertainty around the estimated hazard ratio. Its interpretation concerns the statistical estimate, not the range of outcomes that individual participants might experience.
The p-value of <0.0001 addresses the compatibility of the observed data with the null hypothesis under the specified testing framework. It does not measure the magnitude or clinical importance of the effect. The hazard ratio and its confidence interval provide the effect-size information.
Because this is a time-to-event analysis, interpretation also depends on censoring and on the suitability of summarizing the treatment contrast with a hazard ratio. A hazard ratio is a relative rate measure over follow-up, not an absolute difference in the probability of progression or death at a particular time point.
Number of Participants With Progressive Disease or Death
For a binary endpoint of this type, a categorical comparison could ordinarily be conducted using an appropriate contingency-table method, and the trial's posted statistical methods include the Cochran-Mantel-Haenszel test for categorical efficacy endpoints. However, the ClinicalTrials.gov record does not explicitly attach a formal statistical-analysis result to this particular primary binary endpoint.
9. Secondary Efficacy Results
PFS as Assessed by Investigator
Investigator-assessed PFS
95% CI: 0.45–0.73 · P < 0.0001
Time frame: baseline until PD or death, whichever occurred first, up to 8.5 years overall.
The investigator-assessed analysis produced a hazard ratio below 1, with the confidence interval entirely below 1. This provides a second time-to-event assessment using investigator evaluation rather than the IRC assessment used for the primary PFS analysis.
An HR of 0.57 corresponds to an estimated instantaneous rate of progression or death approximately 57% as high in the obinutuzumab-plus-bendamustine group, or an approximately 43% lower estimated hazard based on 1 − 0.57.
The 95% CI of 0.45–0.73 communicates the precision of this estimate. It does not say that the treatment reduces every patient's individual risk by the same percentage.
The p-value of <0.0001 concerns evidence against the null hypothesis under the reported test; it is not an effect-size measure. This secondary result should also be interpreted in the context of the overall endpoint hierarchy and the fact that multiple outcomes were reported.
Objective Response as Assessed by IRC
| Endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Objective response as assessed by IRC | Difference in response rate | 0.98 | 0.58–1.65 | 0.9298 |
The registry reports a Cochran-Mantel-Haenszel analysis in the ITT population, with the analyzed population defined as participants who had at least one post-baseline assessment. The reported effect measure is a difference in response rate, with an estimate of 0.98 and a two-sided 95% CI of 0.58–1.65.
The estimate is a reported difference in response rate; it should not be reinterpreted as a hazard ratio or as a relative risk. The ClinicalTrials.gov record does not provide the underlying response percentages in this statistical-analysis entry, so they are not reconstructed here.
The confidence interval gives the statistical uncertainty around the reported difference measure. The p-value of 0.9298 is evidence from the specified test, not a measure of the magnitude of the observed response difference.
Objective Response as Assessed by Investigator
| Endpoint | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| Objective response as assessed by investigator | Difference in response rate | -0.90 | -8.44 to 6.64 | 0.7857 |
This investigator-assessed response analysis also used the Cochran-Mantel-Haenszel test. The reported estimate is -0.90, with a two-sided 95% confidence interval from -8.44 to 6.64.
Objective Response at End of Induction Treatment
| Assessment | Effect measure | Estimate | 95% CI | P-value |
|---|---|---|---|---|
| IRC assessment | Difference in response rate | 2.24 | -7.20 to 11.69 | 0.8347 |
| Investigator assessment | Difference in response rate | 8.55 | -0.22 to 17.32 | 0.0466 |
The registry defines these induction-treatment analyses around the end of induction treatment and indicates that the analyzed population consisted of participants who had reached the end-of-induction response assessment. The IRC and investigator assessments should be treated as distinct measurements rather than averaged or combined.
Duration of Response
| Endpoint | Effect measure | Estimate | 95% CI |
|---|---|---|---|
| Duration of Response as Assessed by IRC | Hazard ratio | 0.43 | 0.31–0.61 |
| Duration of Response as Assessed by Investigator | Hazard ratio | 0.51 | 0.39–0.67 |
Both duration-of-response analyses were restricted to participants who had objective response at any time during the study. The ClinicalTrials.gov record does not report a formal analysis method or p-value for these endpoints, so the page does not assign a specific inferential test beyond the reported hazard-ratio framework.
Disease-Free Survival in Participants With Complete Response
| Endpoint | Effect measure | Estimate | 95% CI |
|---|---|---|---|
| DFS in participants with CR, IRC assessment | Hazard ratio | 0.13 | 0.04–0.45 |
| DFS in participants with CR, investigator assessment | Hazard ratio | 0.48 | 0.29–0.81 |
These analyses were conducted among participants who had an objective response of CR. That restriction is important: this is a selected responder population, not the full randomized ITT population. Consequently, these estimates answer a different question from the primary ITT PFS analysis.
Event-Free Survival
IRC-assessed event-free survival
95% CI: 0.44–0.74 · P = 0.0001
Analysis population: ITT · Method: log-rank test
The registry defines the endpoint as event-free survival as assessed by IRC, with a time frame from baseline until PD or death, whichever occurred first, up to approximately 5 years.
Overall Survival
Overall survival
95% CI: 0.57–1.03 · P = 0.0810
Time frame: baseline until death, up to 8.5 years overall.
The overall-survival analysis used the ITT population and a log-rank test. The estimated hazard ratio was below 1, but the two-sided 95% confidence interval includes 1.00 and the reported p-value was 0.0810.
An HR of 0.77 corresponds to an estimated instantaneous death rate approximately 77% as high in the obinutuzumab-plus-bendamustine group, or an approximately 23% lower estimated hazard based on 1 − 0.77.
The 95% CI of 0.57–1.03 indicates substantial uncertainty around the estimate and includes the null value of 1.00. The p-value of 0.0810 does not measure the size of the observed effect; it summarizes evidence against the null under the specified testing framework.
OS is also particularly sensitive to subsequent treatment and the length of follow-up. The registry's reported analysis should therefore be distinguished from the PFS analyses rather than treated as an interchangeable endpoint.
10. Results Summary
| Endpoint | Analysis | Effect | 95% CI | P-value |
|---|---|---|---|---|
| Primary PFS, IRC | ITT; log-rank | HR 0.53 | 0.40–0.70 | <0.0001 |
| PFS, investigator | ITT; log-rank | HR 0.57 | 0.45–0.73 | <0.0001 |
| Objective response, IRC | ITT; Cochran-Mantel-Haenszel | Difference 0.98 | 0.58–1.65 | 0.9298 |
| Objective response, investigator | ITT; Cochran-Mantel-Haenszel | Difference -0.90 | -8.44 to 6.64 | 0.7857 |
| End-of-induction response, IRC | ITT; Cochran-Mantel-Haenszel | Difference 2.24 | -7.20 to 11.69 | 0.8347 |
| End-of-induction response, investigator | ITT; Cochran-Mantel-Haenszel | Difference 8.55 | -0.22 to 17.32 | 0.0466 |
| Duration of response, IRC | Responder population | HR 0.43 | 0.31–0.61 | Not reported |
| Duration of response, investigator | Responder population | HR 0.51 | 0.39–0.67 | Not reported |
| DFS in CR, IRC | CR population | HR 0.13 | 0.04–0.45 | Not reported |
| DFS in CR, investigator | CR population | HR 0.48 | 0.29–0.81 | Not reported |
| EFS, IRC | ITT; log-rank | HR 0.57 | 0.44–0.74 | 0.0001 |
| Overall survival | ITT; log-rank | HR 0.77 | 0.57–1.03 | 0.0810 |
11. Statistical Methods Explained
Why was a log-rank test used for PFS?
PFS is a time-to-event endpoint because the analysis records not only whether progression or death occurred, but also when the event occurred. The log-rank test is designed to compare survival-type event-time distributions while incorporating right-censored observations. A simple comparison of proportions would discard much of that timing information.
What does an HR of 0.53 mean?
An HR of 0.53 means that the estimated instantaneous rate of the event was approximately 53% as high in the obinutuzumab-plus-bendamustine group as in the bendamustine-alone group under the reported analysis. It can be expressed as an approximately 47% lower estimated hazard, but it should not be interpreted as a 47% absolute reduction in the probability of progression or death.
Why is the confidence interval important?
The point estimate alone does not show how precisely the treatment effect was estimated. For the primary PFS analysis, the 95% CI is 0.40–0.70. The interval describes uncertainty around the estimated hazard ratio under the statistical framework; it is not a prediction interval for individual patient outcomes.
Why doesn't a p-value measure treatment effect size?
A p-value measures the strength of evidence against a specified null hypothesis under the assumptions of the statistical test. It is influenced by both the size of an effect and the amount of information available. Two studies can therefore have different p-values despite similar effect estimates, or similar p-values with materially different effect sizes.
Why does ITT matter?
The primary PFS analysis used the ITT population. Keeping participants in their randomized groups protects the original treatment comparison from being redefined after randomization. This is especially important when treatment discontinuation, subsequent therapy, or other post-randomization events occur.
Why distinguish IRC and investigator assessments?
The trial reports both independent-review-committee and investigator assessments. These are different sources of endpoint classification. Agreement between them can provide useful contextual information, but the estimates should not be averaged or treated as independent randomized experiments.
Why are the DFS-in-CR analyses different from primary PFS?
The DFS analyses in participants with CR are restricted to a subset defined by response. The primary PFS analysis is based on the ITT population. Conditioning on an outcome that occurs after randomization changes the population being analyzed and therefore changes the statistical question.
12. Reading the Primary Hazard Ratio Correctly
The primary IRC-assessed PFS HR of 0.53 is a relative time-to-event measure. Under the reported analysis, it indicates a lower estimated instantaneous rate of progression or death for the obinutuzumab-plus-bendamustine group relative to bendamustine alone.
It does not mean that 53% of participants were event-free, that exactly 47% of participants benefited, or that the absolute probability of progression or death was reduced by 47% at every time point.
The 95% CI of 0.40–0.70 provides the precision context for the estimate. The interval is entirely below 1.00, but the width of the interval still matters: an estimate of 0.53 is not equivalent to knowing the treatment effect exactly.
The p-value of <0.0001 describes statistical evidence under the reported test. It should not be used as a substitute for the hazard ratio or its confidence interval when communicating the magnitude and precision of the treatment effect.
13. Safety Results
The ClinicalTrials.gov record provides serious adverse-event counts by randomized treatment arm. They do not provide a broader complete adverse-event table in the ClinicalTrials.gov record, so the safety analysis is limited to this reported measure.
| Safety measure | Bendamustine alone | Obinutuzumab + bendamustine |
|---|---|---|
| Serious adverse events | 76 / 203 | 91 / 204 |
The denominators show that the reported serious-adverse-event measure is based on 203 participants in the bendamustine-alone arm and 204 in the obinutuzumab-plus-bendamustine arm. These counts should not be converted into an unreported inferential comparison or combined with efficacy results to create a single benefit-risk statistic.
14. Assessment of the Statistical Evidence
Primary PFS signal
The primary IRC-assessed PFS analysis reports HR 0.53 with a two-sided 95% CI of 0.40–0.70 and P < 0.0001.
Investigator PFS
The investigator-assessed PFS analysis reports HR 0.57 with a two-sided 95% CI of 0.45–0.73 and P < 0.0001.
Response endpoints
Response analyses use the Cochran-Mantel-Haenszel test, with different estimates reported for IRC and investigator assessments.
Overall survival
The reported OS HR is 0.77 with a two-sided 95% CI of 0.57–1.03 and P = 0.0810.
These results illustrate why clinical-trial interpretation should not be reduced to one p-value. The primary endpoint is a time-to-event comparison; several secondary endpoints use different estimands and populations; and overall survival has its own estimate and uncertainty interval. Each result should be interpreted according to the endpoint and population that generated it.
15. Multiplicity and Endpoint Interpretation
The registry identifies 2 primary endpoints and reports a total of 12 statistical analyses. The ClinicalTrials.gov record does not specify an alpha-allocation strategy, hierarchical testing procedure, multiplicity adjustment, or interim-analysis alpha-spending plan.
| Feature | What the ClinicalTrials.gov record establishes | Interpretive implication |
|---|---|---|
| Primary endpoints | 2 | The primary inferential framework involves more than one endpoint. |
| Statistical analyses posted | 12 | Several efficacy outcomes were formally analyzed. |
| Hypothesis type | Superiority | The reported framework evaluates whether outcomes differ in the superiority direction rather than testing non-inferiority. |
| Multiplicity adjustment | Not specified in the ClinicalTrials.gov record | Individual p-values should not automatically be interpreted as independent confirmatory tests with a known familywise error allocation. |
| Interim analysis | Not specified in the ClinicalTrials.gov record | No interim boundary or alpha-spending method is inferred. |
In particular, the reported p-value of 0.0466 for investigator-assessed objective response at the end of induction treatment should not be isolated from the broader set of endpoints and analyses. The ClinicalTrials.gov record does not establish a multiplicity-control procedure for interpreting that secondary result as a standalone confirmatory claim.
16. What Is Not Established by the Supplied Record
This distinction matters statistically. A detailed trial page should not manufacture a complete statistical protocol from common clinical-trial conventions. When a specific design or analysis feature is not documented in the ClinicalTrials.gov record, the appropriate approach is to identify the evidence boundary rather than assume that a familiar method was used.
17. Limitations
- Incomplete primary binary analysis detail: results were posted for the binary PD-or-death primary endpoint, but the ClinicalTrials.gov record does not contain an estimate, confidence interval, or p-value for that endpoint.
- Multiple endpoints: the registry contains two primary endpoints and multiple secondary efficacy analyses. Individual p-values should be interpreted within the overall testing structure rather than as automatically independent confirmatory findings.
- Different analysis populations: primary PFS uses the ITT population, whereas duration-of-response and DFS analyses are restricted to responders or participants with CR.
- Assessment source: IRC and investigator assessments are distinct endpoint assessments and can produce different estimates.
- Hazard-ratio interpretation: a hazard ratio is a relative time-to-event measure and should not be presented as an absolute risk reduction or a probability of benefit.
- Censoring: time-to-event analyses depend on how censoring is handled. The ClinicalTrials.gov record does not provide the detailed censoring rules or missing-data procedures.
- Proportional hazards: the ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption. A single HR can be less informative when relative hazards change substantially over time.
- Safety scope: the ClinicalTrials.gov record provides serious adverse-event counts by arm but not a complete safety table, so broader safety conclusions are not drawn.
- Missing design details: the ClinicalTrials.gov record does not establish stratification factors, interim monitoring, alpha spending, Bayesian methods, or other specialized design features.
18. Why This Trial Matters Statistically
GADOLIN is a useful teaching case because the registry contains several different statistical questions within one randomized trial. The primary PFS endpoint is a time-to-event comparison; objective response is categorical; duration of response and DFS are analyzed in selected responder populations; and overall survival provides a separate time-to-event outcome.
| Concept | How it appears in GADOLIN |
|---|---|
| Randomization | Randomized, parallel-group phase 3 design |
| ITT analysis | Primary PFS and multiple secondary efficacy analyses use the ITT population |
| Time-to-event endpoints | PFS, duration of response, DFS, EFS and OS |
| Hazard ratio | Primary and secondary time-to-event effects are reported as HRs |
| Confidence intervals | Two-sided 95% CIs accompany the principal reported effect estimates |
| Log-rank test | Used for the primary IRC-assessed PFS analysis and several secondary survival analyses |
| Cochran-Mantel-Haenszel test | Used for reported objective-response analyses |
| Independent assessment | IRC-assessed endpoints are reported separately from investigator assessments |
| Responder populations | Duration-of-response and DFS analyses use populations defined by response |
| Multiplicity | Multiple primary and secondary analyses require careful interpretation of individual p-values |
| Safety comparison | Serious adverse events are reported by randomized arm |
19. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
20. Related Statistical Calculators
21. Sources
- ClinicalTrials.gov: NCT01059630 — GADOLIN.
- PubMed: PubMed record 31462735.
- PubMed: PubMed record 31050355.
- PubMed: PubMed record 29584548.
- PubMed: PubMed record 27345636.
Continue through the Clinical Biostats statistical pathway
Explore the statistical methods behind randomized trials, time-to-event endpoints, categorical comparisons, confidence intervals, and treatment-effect measures.
22. Record Summary
GADOLIN provides a useful statistical example of how one randomized phase 3 trial can generate several distinct estimands and analysis populations. Its primary IRC-assessed PFS analysis used an ITT population and a log-rank test, reporting an HR of 0.53 with a two-sided 95% CI of 0.40–0.70 and P < 0.0001. Secondary analyses extend the statistical picture through investigator-assessed PFS, categorical response endpoints, duration of response, disease-free survival among participants with CR, event-free survival, and overall survival.
The most important statistical lesson is that these results should not be collapsed into a single number. The hazard ratio describes a relative time-to-event effect; the confidence interval describes its statistical precision; the p-value describes evidence against a null hypothesis; and the analysis population determines which clinical question is actually being answered. Response and survival endpoints also represent different outcomes and should remain distinct.