This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.
1. Trial at a Glance
PROCLAIM was a randomized, open-label, parallel phase 3 treatment trial enrolling 598 participants with non-small cell lung cancer. The registered primary endpoint was overall survival, analyzed as a time-to-event endpoint using the log-rank test and a hazard ratio.
| Feature | PROCLAIM |
|---|---|
| Phase | Phase 3 |
| Condition | Non Small Cell Lung Cancer |
| Brief title | Chemotherapy and Radiation in Treating Participants With Stage 3 Non-Small Cell Lung Cancer |
| Design | Randomized, parallel, unmasked |
| Primary purpose | Treatment |
| Enrollment | 598 |
| Primary endpoint | Overall Survival |
| Primary endpoint type | Time-to-event |
| Primary hypothesis type | Superiority |
| Lead sponsor | Eli Lilly and Company |
| Status | Completed |
| Trial dates | Start: 2008-09; Primary completion: 2014-10 |
| ClinicalTrials.gov | NCT00686959 |
2. Clinical Question
The primary statistical question was whether overall survival differed between participants assigned to pemetrexed + cisplatin and thoracic radiation therapy and those assigned to etoposide + cisplatin and thoracic radiation therapy.
Population
Participants with non-small cell lung cancer, with the brief trial title specifying stage 3 disease.
Intervention
Pemetrexed + cisplatin and thoracic radiation therapy.
Comparator
Etoposide + cisplatin and thoracic radiation therapy.
Primary question
Does the intervention produce a different overall-survival experience from the comparator under a superiority framework?
3. Trial Design
Pemetrexed-based chemoradiation
- Pemetrexed
- Cisplatin
- Thoracic Radiation Therapy (TRT)
Etoposide-based chemoradiation
- Etoposide
- Cisplatin
- Thoracic Radiation Therapy (TRT)
The registry identifies the allocation as RANDOMIZED, the design model as PARALLEL, masking as NONE, and the primary purpose as TREATMENT. Randomization is important statistically because, under the trial design, treatment assignment rather than baseline prognosis determines the treatment groups in expectation. That supports a direct comparison of outcomes between the randomized groups.
4. Endpoints
| Endpoint | Registry definition / time frame | Statistical approach reported |
|---|---|---|
| Overall Survival | Baseline to Date of Death from Any Cause (Up to 71.4 Months). OS time is from baseline to the date of death from any cause. Participants not known to have died by the data cut-off were censored at the last contact date known to be alive. OS was summarized using Kaplan-Meier estimates. | Log-rank test; hazard ratio; Kaplan-Meier estimation |
| Progression-free Survival (PFS) | Baseline to Measured Progressive Disease or Death from Any Cause (Up to 66.6 Months) | Log-rank test; hazard ratio |
| Objective Response Rate | Complete Response (CR) + Partial Response (PR), baseline to measured progressive disease (up to 7 months) | Log-rank test; two-sided P-value reported |
| First Site of Disease Failure | Baseline to relapse (up to 66.6 months) | Fisher exact test for specified relapse locations |
| Swallowing Diary | Baseline through 30 days post study; percentage of participants with a post-baseline swallowing diary score ≥4 | Fisher exact test |
The registry lists one primary endpoint: overall survival. The remaining posted outcomes are secondary endpoints. This distinction matters because a primary endpoint is generally the endpoint around which the main confirmatory statistical question is constructed, whereas secondary endpoints provide additional information about efficacy, disease failure, or participant outcomes.
5. Primary Endpoint: Overall Survival
The registered primary endpoint was overall survival from baseline to death from any cause, with censoring at the last known alive contact for participants not known to have died by the data cut-off. The registry specifies Kaplan-Meier estimation and reports a formal log-rank comparison with a hazard ratio.
Overall Survival hazard ratio
95% CI: 0.79–1.20 · P = 0.831
Analysis: all randomized participants; two-sided confidence interval; superiority hypothesis
| Primary analysis feature | Reported value |
|---|---|
| Endpoint | Overall Survival |
| Time frame | Baseline to Date of Death from Any Cause (Up to 71.4 Months) |
| Analysis population | All randomized participants |
| Method | Log Rank |
| Effect measure | Hazard Ratio (HR) |
| Estimate | 0.98 |
| 95% CI | 0.79–1.20 |
| P-value | 0.831 |
| Censored participants | Arm A: 124; Arm B: 117 |
The reported HR of 0.98 is close to 1.00. In a hazard-ratio framework, an estimate of 1 would correspond to equal estimated instantaneous event rates between the two groups. An HR of 0.98 therefore represents a very small estimated relative difference in the instantaneous rate of death, with Arm A having the lower estimated hazard in this analysis.
The HR does not mean that 98% of participants survived, that mortality was reduced by 2 percentage points, or that an individual participant had exactly a 2% lower probability of death. A hazard ratio is a relative time-to-event measure, not an absolute survival probability.
The 95% CI of 0.79–1.20 shows the statistical uncertainty around the reported estimate. Importantly, the interval includes 1.00, so the data are compatible with a range of relative hazard differences in either direction under the stated model and sampling framework.
The P-value of 0.831 measures how compatible the observed comparison is with the null hypothesis used by the statistical test. It does not measure the size or clinical importance of the treatment effect. A large P-value is not itself evidence that the treatments are identical; it indicates that this analysis did not produce strong statistical evidence against the null under the specified test.
The analysis also involves right-censoring: 124 participants in Arm A and 117 in Arm B were censored. Kaplan-Meier and log-rank methods use information from participants up to their censoring times rather than treating censored observations as deaths. Interpretation of a hazard ratio also depends on the suitability of the time-to-event model; the ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption.
What the primary result does not establish
The primary result does not provide a median overall survival time, a survival probability at a specific time point, or an absolute risk difference, because those quantities are not included in the ClinicalTrials.gov record. It also does not by itself establish equivalence or non-inferiority: the registered hypothesis type is superiority, not non-inferiority.
6. Secondary Endpoint: Progression-free Survival
Progression-free survival was defined from baseline to measured progressive disease or death from any cause, with a time frame of up to 66.6 months. The analysis population was all randomized participants. The registry reports a log-rank analysis with a hazard ratio.
Progression-free survival hazard ratio
95% CI: 0.71–1.04 · P = 0.130
Analysis: all randomized participants; two-sided confidence interval; superiority hypothesis
| Feature | Reported value |
|---|---|
| Endpoint | Progression-free Survival (PFS) |
| Time frame | Baseline to Measured Progressive Disease or Death from Any Cause (Up to 66.6 Months) |
| Analysis population | All randomized participants |
| Method | Log Rank |
| Effect measure | Hazard Ratio (HR) |
| Estimate | 0.86 |
| 95% CI | 0.71–1.04 |
| P-value | 0.130 |
| Censored participants | Arm A: 99; Arm B: 87 |
An HR of 0.86 corresponds to an estimated instantaneous rate of progression or death that is 0.86 times the corresponding rate in the comparator group, under the hazard-ratio model. Expressed as a simple relative-hazard interpretation, 0.86 corresponds to a 14% lower estimated hazard in Arm A relative to Arm B.
That does not mean that 14% more participants avoided progression, nor does it mean that individual patients experienced a 14% longer progression-free survival. Absolute PFS probabilities and median PFS are not reported in the ClinicalTrials.gov record.
The 95% CI of 0.71–1.04 crosses 1.00. The interval therefore includes both a potentially lower hazard and a value slightly above 1.00 under the reported analysis. The P-value of 0.130 is a test result, not an effect-size measure, and should not be interpreted as a percentage probability that the treatment works or does not work.
Because PFS is a time-to-event endpoint, censoring and the definition of progression are central to its interpretation. The registry specifies the event as measured progressive disease or death from any cause. The ClinicalTrials.gov record does not provide a formal assessment of proportional hazards or the detailed censoring rules beyond the information summarized in the analysis record.
7. Secondary Endpoint: Objective Response Rate
The registry defines objective response rate as Complete Response (CR) + Partial Response (PR), measured from baseline to measured progressive disease, with a time frame of up to 7 months. The analysis population was all randomized participants.
| Feature | Reported value |
|---|---|
| Endpoint | Objective Response Rate (Complete Response [CR] + Partial Response [PR]) |
| Time frame | Baseline to Measured Progressive Disease (Up to 7 Months) |
| Analysis population | All randomized participants |
| Method | Log Rank |
| P-value | 0.458 |
| Confidence interval | Two-sided |
| Hypothesis | Superiority |
The registry analysis does not provide the response percentages, response counts, or a hazard ratio for this endpoint. Consequently, the reported result should not be converted into an unreported response-rate difference or ratio.
The reported P-value of 0.458 is the result of the registry's reported log-rank analysis. Because the registry does not provide an effect estimate or response percentages in the ClinicalTrials.gov record, the P-value cannot be used to reconstruct the magnitude of any difference between treatment groups.
This illustrates an important reporting principle: a P-value without an effect estimate does not tell the reader how large a treatment difference was observed. For a percentage endpoint, an informative report would ordinarily include the response proportion in each group together with an appropriate measure of uncertainty or between-group effect.
The endpoint is also labeled as an outcome with an endpoint type of time-to-event in the registry analysis, even though its outcome unit is percentage of participants. The page therefore preserves the registry's terminology rather than substituting an unreported statistical framework.
8. Secondary Endpoint: First Site of Disease Failure
The registry evaluates first site of disease failure in terms of relapse from baseline to relapse, up to 66.6 months. These analyses were restricted to all randomized participants with objective PD. Fisher exact tests were used for three specified relapse locations.
| Relapse category | Method | Two-sided P-value |
|---|---|---|
| Relapsed within the radiation treatment field | Fisher exact test | 0.132 |
| Relapsed inside thorax, outside of radiation field | Fisher exact test | 0.337 |
| Relapsed distant disease | Fisher exact test | 0.457 |
Fisher's exact test is appropriate for comparing categorical outcomes when exact inference is useful, particularly when cell counts may be small. It evaluates the allocation of categorical outcomes between groups under a specified null hypothesis; it does not itself provide an effect-size estimate.
Here, the ClinicalTrials.gov record reports three two-sided P-values: 0.132, 0.337, and 0.457. None should be converted into an unreported percentage difference or risk ratio. The ClinicalTrials.gov record also do not provide the underlying counts for each relapse category.
Because these are multiple secondary comparisons, interpretation should also distinguish the individual test results from a broader claim about the overall pattern of disease failure. The ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure for these three analyses.
9. Secondary Endpoint: Swallowing Diary
The registry also reports the percentage of participants with a post-baseline swallowing diary score ≥4, measured from baseline through 30 days post study. The analysis population consisted of all randomized participants with at least one post-baseline swallowing diary score.
Swallowing diary comparison
Fisher exact test · two-sided · superiority hypothesis
| Feature | Reported value |
|---|---|
| Endpoint | Percentage of Participants With a Post Baseline Swallowing Diary Score ≥4 |
| Time frame | Baseline through 30 Days Post Study |
| Analysis population | All randomized participants with at least one post baseline swallowing diary score |
| Method | Fisher exact test |
| P-value | 0.150 |
| Confidence interval | Two-sided |
The registry reports a two-sided Fisher exact test P-value of 0.150. The result does not provide an effect estimate or the percentage in either treatment group in the ClinicalTrials.gov record, so the magnitude of any observed difference cannot be quantified from this record alone.
The analysis population is also narrower than the overall randomized population: participants needed at least one post-baseline swallowing diary score. That distinction matters because eligibility for the analysis depends on availability of post-baseline data. The ClinicalTrials.gov record does not specify an imputation procedure for missing swallowing diary assessments.
10. Statistical Methodology
Kaplan-Meier estimation
The primary overall-survival endpoint is explicitly summarized using Kaplan-Meier estimates. Kaplan-Meier estimation is designed for time-to-event data with right censoring. Instead of requiring every participant to have an observed death, it uses each participant's observed follow-up until death or censoring.
Here, di represents the number of events at event time ti, while ni is the number at risk immediately before that time.
Log-rank test
The log-rank test compares time-to-event experience between groups across the observed follow-up. It is particularly useful when the question concerns whether the survival distributions differ rather than whether a single fixed-time proportion differs.
In PROCLAIM, the registry reports the log-rank method for overall survival and progression-free survival, and also reports it for objective response rate. The page preserves that registry-reported methodology rather than replacing it with an inferred method.
Hazard ratio
A hazard ratio summarizes the relative instantaneous event rate between two groups under a time-to-event model. An HR below 1 indicates a lower estimated hazard for the numerator group, while an HR above 1 indicates a higher estimated hazard.
A hazard ratio is not an absolute risk difference, not a probability of survival, and not necessarily a constant relative difference in cumulative event probability at every time point.
Fisher exact test
Fisher's exact test evaluates a two-group comparison for a categorical outcome using the exact distribution of the observed table under the null hypothesis. It can be especially useful when expected cell counts are small, although the ClinicalTrials.gov record does not report the cell counts for the PROCLAIM relapse or swallowing analyses.
Confidence intervals
A 95% confidence interval describes the uncertainty associated with an estimated parameter under the statistical model and sampling framework. For the PROCLAIM overall-survival HR, the interval is 0.79–1.20. For PFS, it is 0.71–1.04.
Because both intervals include 1.00, neither interval excludes the null value for a hazard ratio. This is consistent with the corresponding P-values of 0.831 for OS and 0.130 for PFS.
11. Statistical Methods Explained
Why was a log-rank test used for overall survival?
Overall survival records both whether an event occurred and when it occurred. A simple comparison of the proportion dead at a single time point would discard much of that information and would not naturally handle participants whose follow-up ends before death. The log-rank test uses the ordering of event times across the groups and accommodates right-censored observations.
What does an overall-survival HR of 0.98 mean?
Under the hazard-ratio interpretation, 0.98 means that the estimated instantaneous rate of death for Arm A relative to Arm B was 0.98. Equivalently, 1 − 0.98 = 0.02, so the point estimate corresponds to a 2% lower estimated hazard in Arm A. That arithmetic describes the point estimate only; it does not establish a 2% reduction in absolute mortality.
Why does the 95% CI matter more than the point estimate alone?
A point estimate is only one summary of the observed comparison. The 95% CI of 0.79–1.20 shows that considerable uncertainty surrounds the OS estimate. It includes values below 1 and above 1, so the data are not precise enough to isolate a narrow range of relative hazard differences.
Why doesn't P = 0.831 mean there is an 83.1% probability that the treatments are equivalent?
A P-value is calculated under a null hypothesis and describes the extremeness of the observed data, or data more extreme, under that hypothesis. It is not the posterior probability that the null hypothesis is true and it does not establish equivalence. A separate equivalence or non-inferiority design would require prespecified margins and a corresponding hypothesis-testing framework.
Why use Fisher exact testing for relapse categories?
The relapse outcomes are categorical: a participant with objective progression can be classified according to a specified first site of disease failure. Fisher's exact test provides an exact two-group test for a categorical contingency table. The ClinicalTrials.gov record does not provide the underlying counts, so the test results cannot be translated into unreported effect sizes.
Why is randomization important to the statistical analysis?
Randomization creates the treatment groups through a prespecified allocation mechanism rather than allowing participants or investigators to choose treatment. This supports the causal interpretation of between-group efficacy comparisons, subject to the trial's conduct, follow-up, endpoint definitions, and analysis assumptions.
12. Censoring and Analysis Populations
The overall-survival analysis used all randomized participants. Arm A had 124 participants censored and Arm B had 117 participants censored. The PFS analysis also used all randomized participants, with 99 censored in Arm A and 87 in Arm B.
| Endpoint | Analysis population | Arm A censored | Arm B censored |
|---|---|---|---|
| Overall Survival | All randomized participants | 124 | 117 |
| Progression-free Survival | All randomized participants | 99 | 87 |
| Objective Response Rate | All randomized participants | Not reported | Not reported |
| First Site of Disease Failure | All randomized participants with objective PD | Not reported | Not reported |
| Swallowing Diary | All randomized participants with at least one post-baseline swallowing diary score | Not reported | Not reported |
Censoring is not equivalent to an unsuccessful outcome. For example, a participant who remains alive at the data cut-off contributes survival information up to the last known time they were observed alive. The validity of standard survival analysis depends on assumptions about the relationship between censoring and the event process; the ClinicalTrials.gov record does not report a detailed missing-data or censoring sensitivity analysis.
13. Safety Results
The ClinicalTrials.gov record reports serious adverse events by randomized arm. The safety measure is presented as the number of affected participants divided by the number at risk.
| Safety measure | Arm A | Arm B |
|---|---|---|
| Serious adverse events | 134 / 283 | 145 / 272 |
| Arm definition | Pemetrexed + Cisplatin and TRT | Etoposide + Cisplatin and TRT |
The displayed percentages in the graphic are simple visual representations of the registry-reported affected/at-risk counts and are not additional reported trial estimates. The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or P-value for serious adverse events.
14. What the Primary Hazard Ratio Does — and Does Not — Mean
The OS hazard ratio of 0.98 means that the estimated instantaneous rate of death in Arm A was 0.98 times that in Arm B under the reported analysis. Expressed as a simple relative-hazard calculation, this is a 2% lower estimated hazard for Arm A at the point estimate.
It does not mean that 2% fewer participants died, that survival probability increased by 2 percentage points, or that every participant experienced the same relative difference.
The 95% CI of 0.79–1.20 communicates uncertainty around the estimated hazard ratio. The interval includes the null value of 1.00, as well as values representing lower and higher hazards for Arm A relative to Arm B.
The P-value of 0.831 is evidence about the statistical compatibility of the observed result with the null hypothesis used by the log-rank test. It is not a measure of effect size and should not be interpreted as a probability that one treatment is better, worse, or equivalent to the other.
Hazard ratios are relative measures. A complete clinical interpretation normally benefits from absolute survival estimates at meaningful time points and median survival when available. Those quantities are not included in the registry-reported PROCLAIM statistical-analysis data and therefore are not reported here.
15. Interpreting the Secondary PFS Result
The PFS estimate of 0.86 is numerically farther below 1 than the OS estimate of 0.98. A simple interpretation of the point estimate is that the estimated instantaneous rate of progression or death was 14% lower in Arm A. However, the 95% CI of 0.71–1.04 includes 1.00 and the P-value is 0.130.
This is a useful example of why effect size, precision, and hypothesis testing should be read together. The point estimate alone suggests a relative difference; the confidence interval shows uncertainty around that estimate; and the P-value provides the result of the specified hypothesis test. None of these quantities should be interpreted in isolation.
16. Endpoint-Specific Statistical Map
| Endpoint | Type | Method | Effect measure | P-value |
|---|---|---|---|---|
| Overall Survival | Time-to-event | Log-rank | HR 0.98 (95% CI 0.79–1.20) | 0.831 |
| Progression-free Survival | Time-to-event | Log-rank | HR 0.86 (95% CI 0.71–1.04) | 0.130 |
| Objective Response Rate | Time-to-event in registry-reported registry classification | Log-rank | Not reported | 0.458 |
| Relapse: radiation treatment field | Binary | Fisher exact | Not reported | 0.132 |
| Relapse: inside thorax, outside radiation field | Binary | Fisher exact | Not reported | 0.337 |
| Relapse: distant disease | Binary | Fisher exact | Not reported | 0.457 |
| Swallowing diary score ≥4 | Binary | Fisher exact | Not reported | 0.150 |
The statistical map illustrates an important feature of clinical-trial reporting: not every endpoint is summarized with the same effect measure. Time-to-event outcomes naturally support hazard ratios and survival estimates, while categorical outcomes can be compared using exact tests. A P-value alone does not replace an effect measure.
17. Multiplicity and the Superiority Framework
The ClinicalTrials.gov record identifies Overall Survival as the single registered primary endpoint and identify the hypothesis type as superiority. Seven statistical analyses are posted in total: one primary analysis and six secondary analyses in the ClinicalTrials.gov record.
| Feature | Registry information |
|---|---|
| Registered primary endpoints | 1 |
| Primary endpoint | Overall Survival |
| Primary hypothesis | Superiority |
| Statistical analyses posted | 7 |
| Primary analyses with estimate + CI | 1 |
| Non-inferiority margin | Not reported in the ClinicalTrials.gov record |
| Interim-analysis method | Not reported in the ClinicalTrials.gov record |
| Bayesian method | Not reported in the ClinicalTrials.gov record |
Because the primary endpoint is superiority, the interpretation of the OS result is based on the reported superiority analysis rather than a non-inferiority framework. A non-inferiority analysis would require a prespecified margin defining the largest clinically acceptable loss of efficacy; no such margin is included in the ClinicalTrials.gov record.
18. Missing Data and Imputation
The ClinicalTrials.gov record does not report an explicit missing-data or imputation method. This matters particularly for the swallowing diary endpoint, whose analysis population requires at least one post-baseline swallowing diary score.
For time-to-event endpoints, censoring is part of the endpoint analysis rather than a conventional imputation of a missing event time. The registry explicitly describes censoring for overall survival and provides censored-participant counts for OS and PFS. The ClinicalTrials.gov record does not describe sensitivity analyses under alternative missing-data assumptions.
OS
Participants not known to have died by the data cut-off were censored at the last contact date known to be alive.
PFS
The analysis includes all randomized participants and reports censored counts, but the ClinicalTrials.gov record does not provide a fuller missing-data strategy.
Swallowing diary
The analysis requires at least one post-baseline swallowing diary score.
Other endpoints
No additional imputation procedures are specified in the ClinicalTrials.gov record.
19. Randomization and Causal Interpretation
Randomization is the foundation of the main comparative inference in PROCLAIM. With two randomized parallel arms, the treatment groups are intended to differ systematically in treatment assignment while balancing prognostic factors in expectation.
That design does not make every observed difference automatically causal. Causal interpretation still depends on adherence to the randomized design, follow-up, outcome ascertainment, censoring, and prespecified analysis. For the primary OS endpoint, the use of all randomized participants maintains the treatment assignment framework specified in the ClinicalTrials.gov record.
The strength of the randomized comparison comes from assigning treatment before the outcome is observed, rather than from the P-value alone.
20. Why Time-to-Event Analysis Is Central Here
Both the primary OS endpoint and the secondary PFS endpoint are time-to-event outcomes. This changes the statistical problem compared with a simple binary endpoint. The analysis needs to preserve information about when an event occurs and how long each participant remains under observation.
For OS, the event is death from any cause. For PFS, the event is measured progressive disease or death from any cause. A participant without an event at the relevant data cut-off contributes follow-up information until censoring.
Event timing
The analysis uses the timing of events rather than only whether an event eventually occurred.
Censoring
Participants can contribute partial follow-up without being classified as having experienced the event.
Kaplan-Meier
Provides an estimate of the event-free survival function over time.
Log-rank
Provides a formal comparison of the time-to-event experience between groups.
21. Important Limitations and Interpretation Issues
- Limited effect reporting for secondary endpoints: the ClinicalTrials.gov record does not provide response percentages for objective response rate or effect estimates for the categorical relapse and swallowing outcomes.
- No median survival values reported: the ClinicalTrials.gov record does not report median OS or median PFS.
- No baseline table reported: the ClinicalTrials.gov record does not include detailed baseline demographic or disease characteristics by treatment arm.
- No formal proportional-hazards assessment reported: the OS and PFS hazard ratios should therefore be interpreted as model-based relative measures without claiming that the proportional-hazards assumption was formally demonstrated.
- No multiplicity strategy reported: the record reports multiple secondary analyses but does not specify an adjustment procedure in the ClinicalTrials.gov record.
- No interim-analysis details reported: the ClinicalTrials.gov record does not document an interim monitoring boundary or alpha-spending strategy.
- No non-inferiority framework: the hypothesis type is superiority, and no non-inferiority margin is reported.
- No Bayesian methods reported: the reported methodology is frequentist, using log-rank and Fisher exact tests.
- Analysis populations differ across endpoints: OS and PFS use all randomized participants, whereas the swallowing analysis requires at least one post-baseline diary score and relapse analyses require objective PD.
- Safety denominators differ: serious adverse events are reported using affected and at-risk counts of 283 and 272 for the two arms, rather than the full enrollment total.
22. Why This Trial Matters Statistically
PROCLAIM is a useful teaching case because its registry record connects randomized treatment assignment with time-to-event analysis, categorical exact testing, censoring, confidence intervals, and multiple secondary endpoints.
| Concept | How it appears in PROCLAIM |
|---|---|
| Randomization | Participants were randomized to two parallel treatment arms. |
| Superiority testing | The primary hypothesis type is superiority. |
| Kaplan-Meier estimation | Overall survival was summarized using Kaplan-Meier estimates. |
| Hazard ratio | OS HR 0.98; PFS HR 0.86. |
| Confidence interval | 95% two-sided CIs are reported for the OS and PFS hazard ratios. |
| Log-rank testing | Used for the primary OS analysis and reported secondary analyses. |
| Fisher exact testing | Used for specified relapse locations and the swallowing diary endpoint. |
| Censoring | 124 vs 117 OS participants and 99 vs 87 PFS participants were censored by arm. |
| Analysis populations | Some secondary analyses use restricted populations rather than all randomized participants. |
| Multiplicity | Seven statistical analyses are posted, while the ClinicalTrials.gov record does not specify a multiplicity procedure. |
23. Statistical Methods Explained: A Deeper View
Why isn't a hazard ratio the same as a risk ratio?
A risk ratio compares probabilities over a specified period. A hazard ratio compares instantaneous event rates under a time-to-event model. The two measures answer different questions and need not have the same numerical value.
Why can a confidence interval include 1 even when the point estimate is below 1?
The point estimate is based on the observed data, while the confidence interval incorporates sampling uncertainty. For PFS, the point estimate is 0.86, but the 95% CI extends from 0.71 to 1.04. The observed estimate is below 1, while plausible parameter values under the stated confidence framework include values above 1.
Why are censored patients still useful?
A censored participant provides information about remaining event-free up to the censoring time. Treating that participant as if an event occurred at censoring would incorrectly introduce an event that was not observed.
Why are OS and PFS not interchangeable?
OS uses death from any cause as the event. PFS uses measured progressive disease or death from any cause. A participant can therefore experience progression before death, meaning the two endpoints measure different points in the disease course.
Why can the analysis populations differ?
An endpoint can require information that is not needed for another endpoint. The swallowing analysis requires at least one post-baseline diary score, while relapse analyses require objective PD. These eligibility conditions define different analysis populations and should be stated explicitly.
Why should secondary P-values not be treated as rankings?
Each P-value corresponds to a particular statistical question. A smaller P-value does not automatically indicate a larger or more important treatment effect, particularly when the underlying effect estimates are not reported. Multiple endpoints also create a broader multiplicity question that cannot be resolved from individual P-values alone.
24. Reported Results Summary
| Endpoint | Arm A vs Arm B | 95% CI | P-value | Method |
|---|---|---|---|---|
| Overall Survival | HR 0.98 | 0.79–1.20 | 0.831 | Log-rank |
| Progression-free Survival | HR 0.86 | 0.71–1.04 | 0.130 | Log-rank |
| Objective Response Rate | Effect estimate not reported | Two-sided | 0.458 | Log-rank |
| Relapse: radiation treatment field | Effect estimate not reported | Two-sided | 0.132 | Fisher exact |
| Relapse: inside thorax, outside radiation field | Effect estimate not reported | Two-sided | 0.337 | Fisher exact |
| Relapse: distant disease | Effect estimate not reported | Two-sided | 0.457 | Fisher exact |
| Swallowing diary score ≥4 | Effect estimate not reported | Two-sided | 0.150 | Fisher exact |
The two endpoints with reported effect estimates are both time-to-event outcomes. Their point estimates are below 1, but their two-sided 95% confidence intervals include 1.00. The other five posted analyses supply P-values but not effect estimates in the ClinicalTrials.gov record, so their magnitude cannot be reconstructed without introducing information outside the permitted trial record.
25. Clinical Interpretation vs Statistical Interpretation
Statistical interpretation
The primary OS analysis reported an HR of 0.98 with a 95% CI of 0.79–1.20 and P = 0.831. The secondary PFS analysis reported an HR of 0.86 with a 95% CI of 0.71–1.04 and P = 0.130.
What remains unreported here
The ClinicalTrials.gov record does not provide median OS, median PFS, time-specific survival percentages, detailed response percentages, baseline characteristics, or formal multiplicity and interim-analysis specifications.
This distinction is important. A statistical analysis should describe what the reported estimates establish, what they leave uncertain, and which quantities are simply not available in the ClinicalTrials.gov record. Filling those gaps from memory or from an outside publication would change the evidentiary basis of the page.
26. Related Tutorials
Learn more about the methods used in this trial:
27. Related Calculators
28. Sources
- ClinicalTrials.gov: PROCLAIM, NCT00686959.
- PubMed: Publication associated with PMID 31125266.
- PubMed: Publication associated with PMID 29976505.
Continue with the statistical methods behind PROCLAIM
Explore the survival-analysis, hypothesis-testing, confidence-interval, and categorical-data methods that appear in the trial's registered statistical analyses.
29. Record Summary
PROCLAIM provides a compact example of how a randomized phase 3 oncology trial can combine several statistical frameworks. Its registered primary endpoint was overall survival, analyzed with Kaplan-Meier estimation and a log-rank comparison summarized by a hazard ratio. The reported OS estimate was 0.98 with a 95% CI of 0.79–1.20 and P = 0.831. PFS was analyzed similarly, with an HR of 0.86, 95% CI 0.71–1.04, and P = 0.130. Secondary analyses used both log-rank and Fisher exact methods.
The most important statistical lesson is that these numbers should be interpreted together rather than independently. The hazard ratio communicates relative event rates, the confidence interval communicates precision, and the P-value addresses the specified hypothesis test. For categorical endpoints where only P-values are reported, the magnitude of the treatment difference cannot be inferred. The analysis populations and censoring rules also matter: the primary OS analysis used all randomized participants, while some secondary analyses required objective progression or post-baseline diary data.