← Clinical Trials
Indolent NHL Phase 3 Superiority NCT01059630

GADOLIN: Complete Statistical Analysis of Obinutuzumab Plus Bendamustine in Indolent Non-Hodgkin's Lymphoma

An independent statistical analysis of the randomized phase 3 GADOLIN trial comparing bendamustine alone with obinutuzumab plus bendamustine in participants with rituximab-refractory, indolent non-Hodgkin's lymphoma.

Completed  ·  Enrollment 413  ·  Primary completion September 2014
Scope of this record

This page separates reported trial results from statistical interpretation. Numerical results and trial characteristics are restricted to the ClinicalTrials.gov record and the associated posted analysis fields.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

GADOLIN was a randomized, open-label, parallel-group phase 3 trial evaluating bendamustine alone versus obinutuzumab plus bendamustine in participants with rituximab-refractory, indolent non-Hodgkin's lymphoma.

413
Enrollment
Participants
2
Arms
Parallel groups
0.53
Primary PFS HR
95% CI 0.40–0.70
<0.0001
Primary PFS P-value
Two-sided
FeatureGADOLIN
Trial nameGADOLIN
PhasePhase 3
ConditionNon-Hodgkin's Lymphoma
PopulationParticipants with rituximab-refractory, indolent non-Hodgkin's lymphoma
DesignRandomized, parallel-group, open-label
AllocationRandomized
Primary purposeTreatment
Enrollment413
InterventionsObinutuzumab; Bendamustine
Primary endpointsNumber of participants with progressive disease as assessed by IRC or death; Progression-Free Survival as assessed by IRC
Hypothesis typeSuperiority
StatusCompleted
ClinicalTrials.govNCT01059630
Lead sponsorGenentech, Inc.

2. Clinical Question

The central statistical question was whether the treatment strategy containing obinutuzumab plus bendamustine differed from bendamustine alone with respect to disease progression or death and other reported efficacy endpoints in participants with rituximab-refractory, indolent non-Hodgkin's lymphoma.

Population

Participants with rituximab-refractory, indolent non-Hodgkin's lymphoma.

Intervention

Obinutuzumab plus bendamustine.

Comparator

Bendamustine alone.

Primary question

Does the obinutuzumab-plus-bendamustine strategy improve the registered efficacy outcomes relative to bendamustine alone?

3. Trial Design

01
Randomize413 participants
02
Two armsParallel groups
03
TreatmentBendamustine alone or obinutuzumab + bendamustine
04
AssessmentPD, death, response and survival outcomes
05
Follow-upSeveral registered time frames
ARM · BENDAMUSTINE ALONE

Bendamustine

  • Bendamustine alone
  • Comparator treatment strategy in the randomized parallel-group design
ARM · OBINUTUZUMAB + BENDAMUSTINE

Obinutuzumab plus bendamustine

  • Obinutuzumab
  • Bendamustine
  • Intervention treatment strategy in the randomized parallel-group design
Open-label design: the registry identifies the masking status as none. This means the trial was not described as blinded. For endpoints involving independent review, the distinction between treatment assignment and the independent assessment of disease status is statistically important.

4. Trial Timing and Registry Structure

April 30, 2010

Trial start

The registry lists April 30, 2010 as the trial start date.

September 30, 2014

Primary completion

The registry lists September 30, 2014 as the primary completion date.

Completed

Registry status

The trial is listed as completed, with results posted and 33 outcome measures and 12 statistical analyses posted.

5. Primary Endpoints

EndpointRegistered definition / time frameStatistical classification
Number of Participants With Progressive Disease (PD) as Assessed by Independent Review Committee (IRC) or Death Baseline until PD or death, whichever occurred first. Assessment was registered at baseline and at specified study visits including 14 days prior to Cycle 4 Day 1. Binary
Progression-Free Survival (PFS) as Assessed by IRC Baseline until PD or death, whichever occurred first. Assessment was registered at baseline and at specified study visits including 14 days prior to Cycle 4 Day 1. Time-to-event

How the registered definitions differ

The first primary endpoint is expressed as a participant-level occurrence of progressive disease or death. The second expresses the same broad clinical event process as a time-to-event endpoint, preserving the timing of progression or death rather than reducing follow-up to a simple yes/no classification.

The registry defines PFS as the time from randomization to the first occurrence of PD or death as assessed by an IRC according to the modified response criteria for indolent Non-Hodgkin's lymphoma. The registry describes PD using criteria including the appearance of a new lesion more than 1.5 cm in any axis during or at the end of therapy and at least a 50% increase from nadir in the sum of product diameter of a relevant lesion measurement.

6. Analysis Populations and Statistical Framework

Population / featureRegistry information
Primary PFS analysis populationITT population
Response analysis populationITT population; for objective response, the number analyzed signified participants who had at least one post-baseline assessment
Induction response analysisITT population; the number analyzed signified participants who had reached the end-of-induction response assessment
Duration of response analysisITT population; the number analyzed signified participants who had objective response at any time during the study
DFS in CR analysisITT population; the number analyzed signified participants who had an objective response of CR
Primary time-to-event methodLog-rank test
Categorical response methodCochran-Mantel-Haenszel test
Hypothesis typeSuperiority

The ITT designation is important because randomized participants remain associated with their assigned treatment group for the efficacy comparison. This preserves the treatment comparison created by randomization and avoids redefining the primary efficacy population based solely on treatment exposure or subsequent events.

7. Statistical Methodology

Log-rank testing for time-to-event endpoints

The registry reports the log-rank test for the primary PFS analysis and for several secondary time-to-event analyses. The log-rank framework compares the observed pattern of events between randomized groups over follow-up while accounting for censoring.

Core time-to-event comparison
H0: survival distributions are equivalent between randomized groups

The log-rank statistic is based on comparing observed and expected events across event times. It is therefore different from a simple comparison of proportions at one fixed time point.

Hazard ratios

The reported effect measure for the primary PFS analysis is the hazard ratio. A hazard ratio compares the estimated instantaneous event rates between groups within a time-to-event framework.

Interpretation
HR < 1  →  lower estimated instantaneous event rate in the obinutuzumab + bendamustine group

The hazard ratio is not a probability, not a percentage of participants who benefit, and not the same quantity as a relative risk.

Cochran-Mantel-Haenszel testing

The registry reports the Cochran-Mantel-Haenszel test for objective-response endpoints. This is a categorical-data method that can compare treatment groups while accounting for stratification variables when those strata are defined for the analysis.

Confidence intervals

The posted analyses use two-sided 95% confidence intervals for their reported effect measures. A confidence interval communicates statistical precision around an estimate. It should not be interpreted as a range containing a fixed probability that the true treatment effect lies inside the interval.

Intention-to-treat analysis

The primary PFS analysis was conducted in the ITT population. In a randomized trial, this approach keeps participants associated with their randomized assignment for efficacy analysis, helping preserve the comparability established at randomization.

8. Primary Results

Progression-Free Survival as Assessed by IRC

Primary PFS hazard ratio

0.53

95% CI: 0.40–0.70   ·   P < 0.0001

Analysis population: ITT  ·  Method: log-rank test  ·  Hypothesis: superiority

Primary endpointComparisonMethodEffect estimate95% CIP-value
Progression-Free Survival as Assessed by IRC Bendamustine alone vs obinutuzumab + bendamustine Log-rank test HR 0.53 0.40–0.70 <0.0001
Clinical Biostats interpretation

The reported hazard ratio of 0.53 means that the estimated instantaneous rate of progression or death was approximately 53% as high in the obinutuzumab-plus-bendamustine group as in the bendamustine-alone group under the time-to-event analysis. Equivalently, 1 − 0.53 = 0.47, so the estimate corresponds to an approximately 47% lower estimated hazard of progression or death.

It does not mean that 47% of participants avoided progression, that 47% were cured, or that each individual participant experienced exactly a 47% reduction in risk.

The 95% confidence interval of 0.40–0.70 describes uncertainty around the estimated hazard ratio. Its interpretation concerns the statistical estimate, not the range of outcomes that individual participants might experience.

The p-value of <0.0001 addresses the compatibility of the observed data with the null hypothesis under the specified testing framework. It does not measure the magnitude or clinical importance of the effect. The hazard ratio and its confidence interval provide the effect-size information.

Because this is a time-to-event analysis, interpretation also depends on censoring and on the suitability of summarizing the treatment contrast with a hazard ratio. A hazard ratio is a relative rate measure over follow-up, not an absolute difference in the probability of progression or death at a particular time point.

Number of Participants With Progressive Disease or Death

Results-posting distinction: ClinicalTrials.gov identifies the binary PD-or-death endpoint as a primary endpoint and indicates that results were posted for it. However, the ClinicalTrials.gov record does not provide a formal statistical-analysis record with an estimate, confidence interval, or p-value for this endpoint. The page therefore does not construct or infer a treatment comparison for this result.

For a binary endpoint of this type, a categorical comparison could ordinarily be conducted using an appropriate contingency-table method, and the trial's posted statistical methods include the Cochran-Mantel-Haenszel test for categorical efficacy endpoints. However, the ClinicalTrials.gov record does not explicitly attach a formal statistical-analysis result to this particular primary binary endpoint.

9. Secondary Efficacy Results

PFS as Assessed by Investigator

Investigator-assessed PFS

HR 0.57

95% CI: 0.45–0.73   ·   P < 0.0001

Time frame: baseline until PD or death, whichever occurred first, up to 8.5 years overall.

The investigator-assessed analysis produced a hazard ratio below 1, with the confidence interval entirely below 1. This provides a second time-to-event assessment using investigator evaluation rather than the IRC assessment used for the primary PFS analysis.

Clinical Biostats interpretation

An HR of 0.57 corresponds to an estimated instantaneous rate of progression or death approximately 57% as high in the obinutuzumab-plus-bendamustine group, or an approximately 43% lower estimated hazard based on 1 − 0.57.

The 95% CI of 0.45–0.73 communicates the precision of this estimate. It does not say that the treatment reduces every patient's individual risk by the same percentage.

The p-value of <0.0001 concerns evidence against the null hypothesis under the reported test; it is not an effect-size measure. This secondary result should also be interpreted in the context of the overall endpoint hierarchy and the fact that multiple outcomes were reported.

Objective Response as Assessed by IRC

EndpointEffect measureEstimate95% CIP-value
Objective response as assessed by IRCDifference in response rate0.980.58–1.650.9298

The registry reports a Cochran-Mantel-Haenszel analysis in the ITT population, with the analyzed population defined as participants who had at least one post-baseline assessment. The reported effect measure is a difference in response rate, with an estimate of 0.98 and a two-sided 95% CI of 0.58–1.65.

Clinical Biostats interpretation

The estimate is a reported difference in response rate; it should not be reinterpreted as a hazard ratio or as a relative risk. The ClinicalTrials.gov record does not provide the underlying response percentages in this statistical-analysis entry, so they are not reconstructed here.

The confidence interval gives the statistical uncertainty around the reported difference measure. The p-value of 0.9298 is evidence from the specified test, not a measure of the magnitude of the observed response difference.

Objective Response as Assessed by Investigator

EndpointEffect measureEstimate95% CIP-value
Objective response as assessed by investigatorDifference in response rate-0.90-8.44 to 6.640.7857

This investigator-assessed response analysis also used the Cochran-Mantel-Haenszel test. The reported estimate is -0.90, with a two-sided 95% confidence interval from -8.44 to 6.64.

Objective Response at End of Induction Treatment

AssessmentEffect measureEstimate95% CIP-value
IRC assessmentDifference in response rate2.24-7.20 to 11.690.8347
Investigator assessmentDifference in response rate8.55-0.22 to 17.320.0466

The registry defines these induction-treatment analyses around the end of induction treatment and indicates that the analyzed population consisted of participants who had reached the end-of-induction response assessment. The IRC and investigator assessments should be treated as distinct measurements rather than averaged or combined.

Duration of Response

EndpointEffect measureEstimate95% CI
Duration of Response as Assessed by IRCHazard ratio0.430.31–0.61
Duration of Response as Assessed by InvestigatorHazard ratio0.510.39–0.67

Both duration-of-response analyses were restricted to participants who had objective response at any time during the study. The ClinicalTrials.gov record does not report a formal analysis method or p-value for these endpoints, so the page does not assign a specific inferential test beyond the reported hazard-ratio framework.

Disease-Free Survival in Participants With Complete Response

EndpointEffect measureEstimate95% CI
DFS in participants with CR, IRC assessmentHazard ratio0.130.04–0.45
DFS in participants with CR, investigator assessmentHazard ratio0.480.29–0.81

These analyses were conducted among participants who had an objective response of CR. That restriction is important: this is a selected responder population, not the full randomized ITT population. Consequently, these estimates answer a different question from the primary ITT PFS analysis.

Event-Free Survival

IRC-assessed event-free survival

HR 0.57

95% CI: 0.44–0.74   ·   P = 0.0001

Analysis population: ITT  ·  Method: log-rank test

The registry defines the endpoint as event-free survival as assessed by IRC, with a time frame from baseline until PD or death, whichever occurred first, up to approximately 5 years.

Overall Survival

Overall survival

HR 0.77

95% CI: 0.57–1.03   ·   P = 0.0810

Time frame: baseline until death, up to 8.5 years overall.

The overall-survival analysis used the ITT population and a log-rank test. The estimated hazard ratio was below 1, but the two-sided 95% confidence interval includes 1.00 and the reported p-value was 0.0810.

Clinical Biostats interpretation

An HR of 0.77 corresponds to an estimated instantaneous death rate approximately 77% as high in the obinutuzumab-plus-bendamustine group, or an approximately 23% lower estimated hazard based on 1 − 0.77.

The 95% CI of 0.57–1.03 indicates substantial uncertainty around the estimate and includes the null value of 1.00. The p-value of 0.0810 does not measure the size of the observed effect; it summarizes evidence against the null under the specified testing framework.

OS is also particularly sensitive to subsequent treatment and the length of follow-up. The registry's reported analysis should therefore be distinguished from the PFS analyses rather than treated as an interchangeable endpoint.

10. Results Summary

EndpointAnalysisEffect95% CIP-value
Primary PFS, IRCITT; log-rankHR 0.530.40–0.70<0.0001
PFS, investigatorITT; log-rankHR 0.570.45–0.73<0.0001
Objective response, IRCITT; Cochran-Mantel-HaenszelDifference 0.980.58–1.650.9298
Objective response, investigatorITT; Cochran-Mantel-HaenszelDifference -0.90-8.44 to 6.640.7857
End-of-induction response, IRCITT; Cochran-Mantel-HaenszelDifference 2.24-7.20 to 11.690.8347
End-of-induction response, investigatorITT; Cochran-Mantel-HaenszelDifference 8.55-0.22 to 17.320.0466
Duration of response, IRCResponder populationHR 0.430.31–0.61Not reported
Duration of response, investigatorResponder populationHR 0.510.39–0.67Not reported
DFS in CR, IRCCR populationHR 0.130.04–0.45Not reported
DFS in CR, investigatorCR populationHR 0.480.29–0.81Not reported
EFS, IRCITT; log-rankHR 0.570.44–0.740.0001
Overall survivalITT; log-rankHR 0.770.57–1.030.0810
Do not read the table as a single hypothesis test. The registry contains multiple primary and secondary endpoints, different analysis populations, different endpoint definitions, and different assessment sources. A collection of individual p-values does not automatically constitute a single multiplicity-adjusted inferential statement.

11. Statistical Methods Explained

Why was a log-rank test used for PFS?

PFS is a time-to-event endpoint because the analysis records not only whether progression or death occurred, but also when the event occurred. The log-rank test is designed to compare survival-type event-time distributions while incorporating right-censored observations. A simple comparison of proportions would discard much of that timing information.

What does an HR of 0.53 mean?

An HR of 0.53 means that the estimated instantaneous rate of the event was approximately 53% as high in the obinutuzumab-plus-bendamustine group as in the bendamustine-alone group under the reported analysis. It can be expressed as an approximately 47% lower estimated hazard, but it should not be interpreted as a 47% absolute reduction in the probability of progression or death.

Why is the confidence interval important?

The point estimate alone does not show how precisely the treatment effect was estimated. For the primary PFS analysis, the 95% CI is 0.40–0.70. The interval describes uncertainty around the estimated hazard ratio under the statistical framework; it is not a prediction interval for individual patient outcomes.

Why doesn't a p-value measure treatment effect size?

A p-value measures the strength of evidence against a specified null hypothesis under the assumptions of the statistical test. It is influenced by both the size of an effect and the amount of information available. Two studies can therefore have different p-values despite similar effect estimates, or similar p-values with materially different effect sizes.

Why does ITT matter?

The primary PFS analysis used the ITT population. Keeping participants in their randomized groups protects the original treatment comparison from being redefined after randomization. This is especially important when treatment discontinuation, subsequent therapy, or other post-randomization events occur.

Why distinguish IRC and investigator assessments?

The trial reports both independent-review-committee and investigator assessments. These are different sources of endpoint classification. Agreement between them can provide useful contextual information, but the estimates should not be averaged or treated as independent randomized experiments.

Why are the DFS-in-CR analyses different from primary PFS?

The DFS analyses in participants with CR are restricted to a subset defined by response. The primary PFS analysis is based on the ITT population. Conditioning on an outcome that occurs after randomization changes the population being analyzed and therefore changes the statistical question.

12. Reading the Primary Hazard Ratio Correctly

What the estimate means

The primary IRC-assessed PFS HR of 0.53 is a relative time-to-event measure. Under the reported analysis, it indicates a lower estimated instantaneous rate of progression or death for the obinutuzumab-plus-bendamustine group relative to bendamustine alone.

What it does not mean

It does not mean that 53% of participants were event-free, that exactly 47% of participants benefited, or that the absolute probability of progression or death was reduced by 47% at every time point.

Why the CI matters

The 95% CI of 0.40–0.70 provides the precision context for the estimate. The interval is entirely below 1.00, but the width of the interval still matters: an estimate of 0.53 is not equivalent to knowing the treatment effect exactly.

Why the p-value matters differently

The p-value of <0.0001 describes statistical evidence under the reported test. It should not be used as a substitute for the hazard ratio or its confidence interval when communicating the magnitude and precision of the treatment effect.

13. Safety Results

The ClinicalTrials.gov record provides serious adverse-event counts by randomized treatment arm. They do not provide a broader complete adverse-event table in the ClinicalTrials.gov record, so the safety analysis is limited to this reported measure.

Safety measureBendamustine aloneObinutuzumab + bendamustine
Serious adverse events76 / 20391 / 204
Serious adverse events: affected / at risk
Bendamustine alone
76 / 203
Obinutuzumab + bendamustine
91 / 204

The denominators show that the reported serious-adverse-event measure is based on 203 participants in the bendamustine-alone arm and 204 in the obinutuzumab-plus-bendamustine arm. These counts should not be converted into an unreported inferential comparison or combined with efficacy results to create a single benefit-risk statistic.

14. Assessment of the Statistical Evidence

Primary PFS signal

The primary IRC-assessed PFS analysis reports HR 0.53 with a two-sided 95% CI of 0.40–0.70 and P < 0.0001.

Investigator PFS

The investigator-assessed PFS analysis reports HR 0.57 with a two-sided 95% CI of 0.45–0.73 and P < 0.0001.

Response endpoints

Response analyses use the Cochran-Mantel-Haenszel test, with different estimates reported for IRC and investigator assessments.

Overall survival

The reported OS HR is 0.77 with a two-sided 95% CI of 0.57–1.03 and P = 0.0810.

These results illustrate why clinical-trial interpretation should not be reduced to one p-value. The primary endpoint is a time-to-event comparison; several secondary endpoints use different estimands and populations; and overall survival has its own estimate and uncertainty interval. Each result should be interpreted according to the endpoint and population that generated it.

15. Multiplicity and Endpoint Interpretation

The registry identifies 2 primary endpoints and reports a total of 12 statistical analyses. The ClinicalTrials.gov record does not specify an alpha-allocation strategy, hierarchical testing procedure, multiplicity adjustment, or interim-analysis alpha-spending plan.

FeatureWhat the ClinicalTrials.gov record establishesInterpretive implication
Primary endpoints2The primary inferential framework involves more than one endpoint.
Statistical analyses posted12Several efficacy outcomes were formally analyzed.
Hypothesis typeSuperiorityThe reported framework evaluates whether outcomes differ in the superiority direction rather than testing non-inferiority.
Multiplicity adjustmentNot specified in the ClinicalTrials.gov recordIndividual p-values should not automatically be interpreted as independent confirmatory tests with a known familywise error allocation.
Interim analysisNot specified in the ClinicalTrials.gov recordNo interim boundary or alpha-spending method is inferred.

In particular, the reported p-value of 0.0466 for investigator-assessed objective response at the end of induction treatment should not be isolated from the broader set of endpoints and analyses. The ClinicalTrials.gov record does not establish a multiplicity-control procedure for interpreting that secondary result as a standalone confirmatory claim.

16. What Is Not Established by the Supplied Record

Evidence boundary: The ClinicalTrials.gov record does not provide baseline characteristic tables, subgroup hazard ratios, median PFS, median OS, Kaplan-Meier coordinates, an explicit proportional-hazards assessment, missing-data or imputation rules, a prespecified stratification scheme, a Bayesian method, a non-inferiority margin, a crossover analysis, a factorial structure, or an interim-analysis boundary. These features are therefore not inferred or reconstructed here.

This distinction matters statistically. A detailed trial page should not manufacture a complete statistical protocol from common clinical-trial conventions. When a specific design or analysis feature is not documented in the ClinicalTrials.gov record, the appropriate approach is to identify the evidence boundary rather than assume that a familiar method was used.

17. Limitations

18. Why This Trial Matters Statistically

GADOLIN is a useful teaching case because the registry contains several different statistical questions within one randomized trial. The primary PFS endpoint is a time-to-event comparison; objective response is categorical; duration of response and DFS are analyzed in selected responder populations; and overall survival provides a separate time-to-event outcome.

ConceptHow it appears in GADOLIN
RandomizationRandomized, parallel-group phase 3 design
ITT analysisPrimary PFS and multiple secondary efficacy analyses use the ITT population
Time-to-event endpointsPFS, duration of response, DFS, EFS and OS
Hazard ratioPrimary and secondary time-to-event effects are reported as HRs
Confidence intervalsTwo-sided 95% CIs accompany the principal reported effect estimates
Log-rank testUsed for the primary IRC-assessed PFS analysis and several secondary survival analyses
Cochran-Mantel-Haenszel testUsed for reported objective-response analyses
Independent assessmentIRC-assessed endpoints are reported separately from investigator assessments
Responder populationsDuration-of-response and DFS analyses use populations defined by response
MultiplicityMultiple primary and secondary analyses require careful interpretation of individual p-values
Safety comparisonSerious adverse events are reported by randomized arm

19. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

20. Related Statistical Calculators

21. Sources

Continue through the Clinical Biostats statistical pathway

Explore the statistical methods behind randomized trials, time-to-event endpoints, categorical comparisons, confidence intervals, and treatment-effect measures.

22. Record Summary

GADOLIN provides a useful statistical example of how one randomized phase 3 trial can generate several distinct estimands and analysis populations. Its primary IRC-assessed PFS analysis used an ITT population and a log-rank test, reporting an HR of 0.53 with a two-sided 95% CI of 0.40–0.70 and P < 0.0001. Secondary analyses extend the statistical picture through investigator-assessed PFS, categorical response endpoints, duration of response, disease-free survival among participants with CR, event-free survival, and overall survival.

The most important statistical lesson is that these results should not be collapsed into a single number. The hazard ratio describes a relative time-to-event effect; the confidence interval describes its statistical precision; the p-value describes evidence against a null hypothesis; and the analysis population determines which clinical question is actually being answered. Response and survival endpoints also represent different outcomes and should remain distinct.

Clinical Biostats methodology: A rigorous trial-results page should distinguish the reported estimate from its interpretation, preserve the stated analysis population and endpoint definition, and avoid filling undocumented statistical details with assumptions based on common trial practice.