This page separates reported trial results from statistical interpretation. Numerical results are taken from the ClinicalTrials.gov trial data posted on ClinicalTrials.gov for BELINDA. The registry provides one formal statistical analysis for the registered primary endpoint.
1. Trial at a Glance
BELINDA was a randomized, parallel, phase 3 trial in adult patients with aggressive B-cell non-Hodgkin lymphoma. The trial compared a tisagenlecleucel treatment strategy with a standard-of-care treatment strategy and used event-free survival assessed by a Blinded Independent Review Committee as its registered primary endpoint.
| Feature | BELINDA |
|---|---|
| Trial name | BELINDA |
| NCT ID | NCT03570892 |
| Phase | Phase 3 |
| Status | COMPLETED |
| Therapeutic area | Hematology |
| Condition | Non-Hodgkin Lymphoma |
| Enrollment | 330.0 |
| Allocation | RANDOMIZED |
| Design model | PARALLEL |
| Masking | SINGLE |
| Primary purpose | TREATMENT |
| Lead sponsor | Novartis Pharmaceuticals |
| Sponsor type | INDUSTRY |
| Start | 2019-05-09 |
| Primary completion | 2021-05-06 |
2. Clinical Question
The central statistical question was whether the tisagenlecleucel treatment strategy differed from the standard-of-care treatment strategy with respect to event-free survival in the randomized trial population.
Population
Adult patients with aggressive B-cell non-Hodgkin lymphoma.
Intervention
Tisagenlecleucel after optional bridging and lymphodepleting chemotherapy.
Comparator
Platinum-based immunochemotherapy followed in responding patients with high dose chemotherapy and autologous hematopoietic stem cell transplant (HSCT).
Primary question
Does the tisagenlecleucel treatment strategy improve event-free survival relative to the standard-of-care treatment strategy?
3. Trial Design
Tisagenlecleucel
- Tisagenlecleucel after optional bridging chemotherapy.
- Lymphodepleting chemotherapy.
- Evaluated as a treatment strategy in the randomized comparison.
Standard of care
- Platinum-based immunochemotherapy.
- In responding patients, high dose chemotherapy.
- Autologous hematopoietic stem cell transplant (HSCT).
The registry describes the allocation as RANDOMIZED, the design model as PARALLEL, and masking as SINGLE. These design descriptors establish the basic structure of the comparison but do not, by themselves, specify every operational feature of treatment delivery or assessment.
4. Analysis Population and Primary Endpoint
The posted primary analysis used the full analysis set (FAS). The registry defines this population as all subjects to whom study treatment was assigned by randomization.
| Analysis feature | Registry information |
|---|---|
| Analysis population | Full analysis set (FAS) |
| FAS definition | All subjects to whom study treatment was assigned by randomization |
| Endpoint | Event-free Survival (EFS) Per Blinded Independent Review Committee (BIRC) Assessment |
| Endpoint type | Time-to-event |
| Time frame | appro. 24 months |
| Groups compared | Tisagenlecleucel Treatment Strategy vs Standard of Care Treatment Strategy |
| Analysis method | Stratified log-rank test one-sided |
| Effect measure | Unadjusted stratified cox model hazard r |
| Hypothesis type | Superiority |
5. Endpoints
Primary Endpoint: Event-free Survival
The registered primary endpoint was Event-free Survival (EFS) Per Blinded Independent Review Committee (BIRC) Assessment, with a time frame of appro. 24 months.
Event-free survival (EFS) is defined as the time from the date of randomization to the date of the first documented disease progression or stable disease at or after the week 12 (+/- 1 week) assessment, as assessed by Blinded Independent Review Committee (BIRC) per Lugano criteria, or death due to any cause, at any time.
This definition is important because EFS is not simply a measure of radiographic progression. The registry definition includes the specified disease-status event at or after the week 12 assessment as well as death from any cause. Consequently, the endpoint combines disease-related and mortality events within a single time-to-event framework.
6. Statistical Methodology
Stratified log-rank test
The posted formal analysis used a one-sided stratified log-rank test. The log-rank test is designed for comparing time-to-event distributions between treatment groups while accounting for the fact that some participants may not have experienced the event by the end of observation.
In a stratified version, the comparison is conducted across predefined strata rather than treating the entire randomized population as a single homogeneous risk set. The ClinicalTrials.gov record identifies the method as a stratified log-rank test but does not provide the individual stratification factors in the ClinicalTrials.gov record.
The basic logic compares the number of events observed in each treatment group with the number expected under the null hypothesis, over the observed event times. A stratified analysis extends this comparison across strata.
Hazard ratio
The registry reports an unadjusted stratified Cox model hazard ratio, normalized in the ClinicalTrials.gov record to the effect measure "Hazard ratio." The hazard ratio summarizes the relative event rate between the two randomized treatment strategies within the time-to-event model.
An HR below 1 indicates a lower estimated instantaneous event rate in the tisagenlecleucel treatment strategy; an HR above 1 indicates a higher estimated instantaneous event rate under the fitted model.
Full analysis set
The primary analysis used the full analysis set, consisting of all subjects to whom study treatment was assigned by randomization. This is closely aligned with the intention-to-treat principle because treatment comparisons remain anchored to the randomized assignment rather than being restricted to participants who completed treatment according to protocol.
Superiority hypothesis
The statistical-analysis record identifies the hypothesis type as superiority. Thus, the formal question was whether the tisagenlecleucel treatment strategy could demonstrate an advantage in the time-to-event endpoint rather than whether it merely remained within a prespecified non-inferiority margin.
One-sided testing and two-sided confidence interval
The registry records a one-sided stratified log-rank test while reporting a 95% two-sided confidence interval for the hazard ratio. These are different components of the statistical presentation: the p-value corresponds to the reported hypothesis-testing procedure, while the confidence interval describes uncertainty around the estimated hazard ratio using a two-sided interval.
7. Results
Primary Endpoint: Event-free Survival
Hazard ratio for event-free survival
95% CI: 0.82–1.40 · One-sided P = 0.694
Analysis population: full analysis set (FAS)
| Primary endpoint | Tisagenlecleucel Treatment Strategy vs Standard of Care Treatment Strategy |
|---|---|
| Endpoint | Event-free Survival (EFS) Per Blinded Independent Review Committee (BIRC) Assessment |
| Time frame | appro. 24 months |
| Analysis | Stratified log-rank test one-sided |
| Effect measure | Unadjusted stratified cox model hazard r |
| Hazard ratio | 1.07 |
| 95% CI | 0.82–1.40 |
| P-value | = 0.694 |
| Hypothesis | Superiority |
The reported hazard ratio of 1.07 means that the estimated instantaneous rate of an EFS event was 1.07 times that of the standard-of-care treatment strategy in the reported stratified Cox model. Expressed as a simple relative comparison, this corresponds to an estimated event hazard that was about 7% higher for the tisagenlecleucel treatment strategy.
The HR does not mean that 7% more patients experienced an event, nor does it represent a difference in the percentage of patients who were event-free. A hazard ratio is a relative time-to-event measure; it summarizes the modeled relationship between event rates over follow-up.
The 95% confidence interval of 0.82–1.40 describes uncertainty around the estimated hazard ratio. Because the interval includes 1, the data are compatible with a range of relative hazard values extending from below 1 to above 1. The interval therefore provides important information about precision that is not conveyed by the point estimate alone.
The one-sided P = 0.694 is a measure used for the specified hypothesis test; it is not a measure of the size of the treatment effect. A p-value does not tell us that the effect is "7%" or "69.4%" and should not be interpreted as the probability that the treatment works or does not work.
Because EFS is a time-to-event endpoint, the interpretation also depends on censoring and the assumptions underlying the hazard-ratio model. The ClinicalTrials.gov record does not provide enough information to independently evaluate proportional hazards, the censoring pattern, or the detailed censoring rules beyond the registered endpoint definition.
8. Understanding the EFS Result
The BELINDA primary result illustrates why a time-to-event analysis should be read as a package of related quantities rather than as a single number. The hazard ratio, confidence interval, and p-value answer different questions.
What the HR says
The HR of 1.07 is the model-based relative estimate of the instantaneous EFS event rate for the tisagenlecleucel strategy compared with standard of care.
What the HR does not say
It does not give an absolute percentage-point difference in event-free survival, a probability that an individual patient benefits, or a percentage of patients with an event.
What the CI says
The 95% CI of 0.82–1.40 indicates uncertainty around the estimated relative effect and includes the null value of 1.
What the p-value says
The reported one-sided P = 0.694 quantifies the result under the specified superiority testing framework; it does not quantify clinical importance or effect size.
Why the null value is 1
For a hazard ratio, the null value is 1 because an HR of 1 represents equal modeled hazard between the two groups. Values below 1 favor a lower estimated event hazard for the numerator group, while values above 1 indicate a higher estimated event hazard.
The interval spans both sides of the null value of 1. The point estimate alone therefore should not be treated as a complete description of the uncertainty surrounding the primary result.
9. Statistical Methods Explained
Why was a time-to-event endpoint used?
EFS measures not only whether an event occurs but also when it occurs. This makes it possible to incorporate different follow-up times and right-censored observations. The registered endpoint defines the event from randomization through the first qualifying disease event or death.
Why use a stratified log-rank test?
The stratified log-rank test compares the time-to-event experience of the randomized groups while accounting for strata. This is different from a simple comparison of event proportions because it uses the timing of events and accommodates censoring. The registry specifically reports the method as a one-sided stratified log-rank test.
What does a hazard ratio of 1.07 mean?
An HR of 1.07 means that the estimated instantaneous EFS event rate for the tisagenlecleucel treatment strategy was 1.07 times the corresponding rate for the standard-of-care strategy under the reported stratified Cox model. It does not mean that 1.07 times as many patients experienced an event.
Why is the confidence interval important?
The 95% CI of 0.82–1.40 communicates the uncertainty surrounding the point estimate. A point estimate of 1.07 by itself could look close to 1, but the interval shows that the data are compatible with hazard ratios below and above the null value. Precision and direction should therefore be considered together.
Why doesn't the p-value measure effect size?
The P = 0.694 describes the evidence from the specified hypothesis test under its null model. It does not describe how large the treatment effect is. Effect size is represented here by the hazard ratio, while uncertainty around that estimate is represented by the confidence interval.
Why does the full analysis set matter?
The FAS comprised all subjects to whom study treatment was assigned by randomization. Keeping randomized participants in the efficacy analysis preserves the treatment assignment created by randomization and avoids defining the primary comparison solely by subsequent treatment exposure or completion.
10. Censoring and Time-to-Event Interpretation
Time-to-event analysis is particularly useful when participants are followed for different lengths of time. A participant who has not experienced an EFS event by the end of available follow-up does not simply disappear from the analysis; that observation can be treated as censored according to the prespecified analysis rules.
The registry-reported BELINDA statistical-analysis record does not provide the number of EFS events, censoring counts, censoring distribution, or detailed censoring rules. Those quantities therefore cannot be reconstructed from the ClinicalTrials.gov record.
11. Non-Inferiority, Superiority, and the Null Hypothesis
The ClinicalTrials.gov record identifies the hypothesis type as superiority. This is an important distinction because BELINDA's primary statistical question was not framed in the ClinicalTrials.gov record as a non-inferiority question with a prespecified margin.
| Concept | BELINDA primary analysis |
|---|---|
| Hypothesis type | Superiority |
| Primary endpoint | Event-free Survival (EFS) Per Blinded Independent Review Committee (BIRC) Assessment |
| Test | Stratified log-rank test one-sided |
| Effect measure | Hazard ratio |
| Reported estimate | 1.07 |
| 95% confidence interval | 0.82–1.40 |
| Reported p-value | = 0.694 |
In a superiority framework, the null value for the hazard ratio is 1. The analysis asks whether the evidence supports a treatment difference in the prespecified direction rather than asking whether the treatment remains within a non-inferiority margin. No non-inferiority margin is reported in the registry-reported BELINDA data, so none is introduced here.
12. Multiplicity and Interim Analysis
The ClinicalTrials.gov record does not report a multiplicity adjustment, an interim-analysis procedure, an alpha-spending strategy, or a formal endpoint hierarchy. These design features are therefore not characterized here as part of the BELINDA statistical analysis.
13. Missing Data and Imputation
The ClinicalTrials.gov record does not report a missing-data or imputation method for the primary EFS analysis. Because EFS is a time-to-event endpoint, incomplete observation is generally handled through censoring within the survival-analysis framework rather than by automatically replacing an unobserved event time with a single imputed value.
However, the specific BELINDA censoring and missing-data rules are not contained in the ClinicalTrials.gov record. No additional imputation procedure is therefore attributed to the trial.
14. Safety Results
The ClinicalTrials.gov record reports serious adverse events by treatment strategy. These figures describe the number of participants affected and the number at risk in each arm.
| Safety measure | Affected / At risk |
|---|---|
| Tisagenlecleucel Treatment Strategy | 76 / 162 |
| Standard of Care Treatment Strategy | 82 / 160 |
| All Participants | 158 / 322 |
Tisagenlecleucel strategy
Serious adverse events were reported in 76 of 162 participants in the ClinicalTrials.gov record.
Standard of care strategy
Serious adverse events were reported in 82 of 160 participants in the ClinicalTrials.gov record.
The safety denominator differs from the overall trial enrollment of 330. The ClinicalTrials.gov record reports 162 and 160 participants at risk for the two treatment strategies, respectively, for a total of 322. This difference should be preserved rather than silently replacing the safety denominator with the randomized enrollment.
15. Design Features That Shape Interpretation
These descriptors are statistically relevant because randomization establishes the basis for the comparative efficacy analysis, while the parallel structure means the primary comparison is between independently assigned treatment strategies. The ClinicalTrials.gov record does not provide enough detail to infer additional blinding procedures, crossover rules, interim monitoring, or stratification factors.
16. Primary Result in Context
The EFS hazard ratio was 1.07. On the hazard-ratio scale, the estimate is above the null value of 1, corresponding to an estimated instantaneous EFS event rate about 7% higher for the tisagenlecleucel treatment strategy than for the standard-of-care treatment strategy under the reported model.
The 95% CI of 0.82–1.40 is substantially wider than a point estimate alone. It crosses the null value of 1, so the estimated treatment effect should not be characterized solely from the direction of the point estimate.
The reported one-sided P = 0.694 is the formal test result posted on ClinicalTrials.gov for the superiority analysis. It is not an effect-size measure and should not be converted into a percentage of treatment efficacy or a probability that the null hypothesis is true.
The ClinicalTrials.gov record does not report median EFS, EFS event counts by arm, Kaplan-Meier time-point estimates, subgroup estimates, or a detailed event/censoring table. Those quantities are not reconstructed or inferred here.
17. Limitations
- Limited posted statistical detail: the ClinicalTrials.gov record contains one formal statistical analysis for the primary endpoint. They do not provide the complete statistical-analysis plan or all operational details of the survival analysis.
- Single primary result: the posted formal analysis covers EFS. The ClinicalTrials.gov record does not provide formal statistical analyses for the other posted outcome measures.
- No median EFS: median event-free survival is not contained in the ClinicalTrials.gov record and is therefore not reported here.
- No event counts by treatment arm: the ClinicalTrials.gov record does not provide the number of EFS events in each randomized group.
- No Kaplan-Meier estimates: the ClinicalTrials.gov record does not provide time-specific EFS probabilities or a complete Kaplan-Meier curve.
- Model assumptions: the hazard-ratio interpretation depends on the underlying time-to-event modeling framework. The ClinicalTrials.gov record does not provide enough detail to independently assess proportional-hazards behavior.
- Stratification details: the analysis is identified as stratified, but the ClinicalTrials.gov record does not identify the individual stratification factors.
- Missing-data details: the ClinicalTrials.gov record does not state a specific missing-data or imputation procedure.
- Multiplicity and interim monitoring: no multiplicity adjustment or interim-analysis procedure is provided in the ClinicalTrials.gov record.
- Safety comparison: serious adverse events are reported descriptively by arm, but no formal comparative safety analysis is included in the ClinicalTrials.gov record.
18. Why This Trial Matters Statistically
BELINDA is a useful teaching case because its primary analysis brings several fundamental clinical-trial concepts together in a compact time-to-event framework. The treatment comparison begins with randomization, defines a clinically structured event-free endpoint, uses the full analysis set, compares the groups with a stratified log-rank test, and expresses the relative effect through a Cox-model hazard ratio.
| Concept | How it appears in BELINDA |
|---|---|
| Randomization | Randomized allocation to two treatment strategies |
| Parallel design | Two parallel treatment strategies |
| Full analysis set | All subjects assigned study treatment by randomization |
| Time-to-event endpoint | EFS measured from randomization to the first qualifying event or death |
| BIRC assessment | EFS definition specifies assessment by a Blinded Independent Review Committee |
| Stratified log-rank test | Primary hypothesis test |
| Cox model | Unadjusted stratified Cox model used for the reported hazard ratio |
| Hazard ratio | 1.07 with 95% CI 0.82–1.40 |
| Confidence interval | Quantifies uncertainty around the HR estimate |
| Superiority testing | Primary hypothesis type |
| One-sided p-value | P = 0.694 for the reported primary analysis |
| Safety denominators | Serious adverse events reported among 162 and 160 participants at risk |
The most important statistical lesson is that a clinical-trial result cannot be reduced to its p-value. The EFS result has to be read simultaneously through the endpoint definition, analysis population, time-to-event method, hazard ratio, confidence interval, and hypothesis-testing framework.
19. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
20. Related Statistical Calculators
21. Sources
- ClinicalTrials.gov: BELINDA, NCT03570892.
- Linked publication: PubMed record, PMID 34904798.
- Linked publication: PubMed record, PMID 34515338.
- Linked publication: PubMed record, PMID 33288485.
Continue through the Clinical Biostats statistical library
Explore the underlying survival-analysis methods, hypothesis-testing concepts, and statistical calculators associated with randomized clinical trials.
22. Record Summary
BELINDA provides a clear example of a randomized phase 3 time-to-event analysis. The registered primary endpoint was event-free survival per Blinded Independent Review Committee assessment, defined from randomization to the first documented disease progression or stable disease at or after the week 12 (+/- 1 week) assessment, or death due to any cause. The primary analysis used the full analysis set, a one-sided stratified log-rank test, and an unadjusted stratified Cox model hazard ratio.
The reported EFS hazard ratio was 1.07, with a 95% CI of 0.82–1.40 and a one-sided P = 0.694. The statistical interpretation depends on considering all three quantities together: the hazard ratio describes the estimated relative event rate, the confidence interval describes uncertainty around that estimate, and the p-value describes the result of the specified superiority hypothesis test.
The ClinicalTrials.gov record also reports serious adverse events affecting 76/162 participants in the tisagenlecleucel treatment strategy and 82/160 participants in the standard-of-care treatment strategy. These safety figures are descriptive in the ClinicalTrials.gov record and are not accompanied by a formal comparative safety analysis.