← Clinical Trials
Stage 3 NSCLC Phase 3 Overall Survival NCT00686959

PROCLAIM: Complete Statistical Analysis of Chemoradiation in Stage 3 Non-Small Cell Lung Cancer

An independent statistical analysis of the randomized phase 3 PROCLAIM trial comparing pemetrexed + cisplatin and thoracic radiation therapy with etoposide + cisplatin and thoracic radiation therapy in participants with stage 3 non-small cell lung cancer.

ClinicalTrials.gov record  ·  Completed  ·  Enrollment 598  ·  2008-09 to 2014-10
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical trial results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

PROCLAIM was a randomized, open-label, parallel phase 3 treatment trial enrolling 598 participants with non-small cell lung cancer. The registered primary endpoint was overall survival, analyzed as a time-to-event endpoint using the log-rank test and a hazard ratio.

598
Enrolled
2 randomized arms
2
Arms
Parallel design
0.98
Overall Survival HR
95% CI 0.79–1.20
0.831
OS P-value
Two-sided log-rank test
FeaturePROCLAIM
PhasePhase 3
ConditionNon Small Cell Lung Cancer
Brief titleChemotherapy and Radiation in Treating Participants With Stage 3 Non-Small Cell Lung Cancer
DesignRandomized, parallel, unmasked
Primary purposeTreatment
Enrollment598
Primary endpointOverall Survival
Primary endpoint typeTime-to-event
Primary hypothesis typeSuperiority
Lead sponsorEli Lilly and Company
StatusCompleted
Trial datesStart: 2008-09; Primary completion: 2014-10
ClinicalTrials.govNCT00686959

2. Clinical Question

The primary statistical question was whether overall survival differed between participants assigned to pemetrexed + cisplatin and thoracic radiation therapy and those assigned to etoposide + cisplatin and thoracic radiation therapy.

Population

Participants with non-small cell lung cancer, with the brief trial title specifying stage 3 disease.

Intervention

Pemetrexed + cisplatin and thoracic radiation therapy.

Comparator

Etoposide + cisplatin and thoracic radiation therapy.

Primary question

Does the intervention produce a different overall-survival experience from the comparator under a superiority framework?

3. Trial Design

01
Randomize598 participants
02
Arm APemetrexed + cisplatin + TRT
03
Arm BEtoposide + cisplatin + TRT
04
FollowSurvival and secondary outcomes
05
AnalyzeTime-to-event and categorical methods
ARM A

Pemetrexed-based chemoradiation

  • Pemetrexed
  • Cisplatin
  • Thoracic Radiation Therapy (TRT)
ARM B

Etoposide-based chemoradiation

  • Etoposide
  • Cisplatin
  • Thoracic Radiation Therapy (TRT)

The registry identifies the allocation as RANDOMIZED, the design model as PARALLEL, masking as NONE, and the primary purpose as TREATMENT. Randomization is important statistically because, under the trial design, treatment assignment rather than baseline prognosis determines the treatment groups in expectation. That supports a direct comparison of outcomes between the randomized groups.

No masking: the registry identifies the trial as unmasked. That does not prevent analysis of the primary survival endpoint, because death is an objectively defined event, but lack of masking can be more relevant to outcomes that involve assessment, reporting, or participant experience.

4. Endpoints

EndpointRegistry definition / time frameStatistical approach reported
Overall SurvivalBaseline to Date of Death from Any Cause (Up to 71.4 Months). OS time is from baseline to the date of death from any cause. Participants not known to have died by the data cut-off were censored at the last contact date known to be alive. OS was summarized using Kaplan-Meier estimates.Log-rank test; hazard ratio; Kaplan-Meier estimation
Progression-free Survival (PFS)Baseline to Measured Progressive Disease or Death from Any Cause (Up to 66.6 Months)Log-rank test; hazard ratio
Objective Response RateComplete Response (CR) + Partial Response (PR), baseline to measured progressive disease (up to 7 months)Log-rank test; two-sided P-value reported
First Site of Disease FailureBaseline to relapse (up to 66.6 months)Fisher exact test for specified relapse locations
Swallowing DiaryBaseline through 30 days post study; percentage of participants with a post-baseline swallowing diary score ≥4Fisher exact test

The registry lists one primary endpoint: overall survival. The remaining posted outcomes are secondary endpoints. This distinction matters because a primary endpoint is generally the endpoint around which the main confirmatory statistical question is constructed, whereas secondary endpoints provide additional information about efficacy, disease failure, or participant outcomes.

5. Primary Endpoint: Overall Survival

The registered primary endpoint was overall survival from baseline to death from any cause, with censoring at the last known alive contact for participants not known to have died by the data cut-off. The registry specifies Kaplan-Meier estimation and reports a formal log-rank comparison with a hazard ratio.

Overall Survival hazard ratio

0.98

95% CI: 0.79–1.20   ·   P = 0.831

Analysis: all randomized participants; two-sided confidence interval; superiority hypothesis

Primary analysis featureReported value
EndpointOverall Survival
Time frameBaseline to Date of Death from Any Cause (Up to 71.4 Months)
Analysis populationAll randomized participants
MethodLog Rank
Effect measureHazard Ratio (HR)
Estimate0.98
95% CI0.79–1.20
P-value0.831
Censored participantsArm A: 124; Arm B: 117
Clinical Biostats interpretation

The reported HR of 0.98 is close to 1.00. In a hazard-ratio framework, an estimate of 1 would correspond to equal estimated instantaneous event rates between the two groups. An HR of 0.98 therefore represents a very small estimated relative difference in the instantaneous rate of death, with Arm A having the lower estimated hazard in this analysis.

The HR does not mean that 98% of participants survived, that mortality was reduced by 2 percentage points, or that an individual participant had exactly a 2% lower probability of death. A hazard ratio is a relative time-to-event measure, not an absolute survival probability.

The 95% CI of 0.79–1.20 shows the statistical uncertainty around the reported estimate. Importantly, the interval includes 1.00, so the data are compatible with a range of relative hazard differences in either direction under the stated model and sampling framework.

The P-value of 0.831 measures how compatible the observed comparison is with the null hypothesis used by the statistical test. It does not measure the size or clinical importance of the treatment effect. A large P-value is not itself evidence that the treatments are identical; it indicates that this analysis did not produce strong statistical evidence against the null under the specified test.

The analysis also involves right-censoring: 124 participants in Arm A and 117 in Arm B were censored. Kaplan-Meier and log-rank methods use information from participants up to their censoring times rather than treating censored observations as deaths. Interpretation of a hazard ratio also depends on the suitability of the time-to-event model; the ClinicalTrials.gov record does not report a formal assessment of the proportional-hazards assumption.

What the primary result does not establish

The primary result does not provide a median overall survival time, a survival probability at a specific time point, or an absolute risk difference, because those quantities are not included in the ClinicalTrials.gov record. It also does not by itself establish equivalence or non-inferiority: the registered hypothesis type is superiority, not non-inferiority.

6. Secondary Endpoint: Progression-free Survival

Progression-free survival was defined from baseline to measured progressive disease or death from any cause, with a time frame of up to 66.6 months. The analysis population was all randomized participants. The registry reports a log-rank analysis with a hazard ratio.

Progression-free survival hazard ratio

0.86

95% CI: 0.71–1.04   ·   P = 0.130

Analysis: all randomized participants; two-sided confidence interval; superiority hypothesis

FeatureReported value
EndpointProgression-free Survival (PFS)
Time frameBaseline to Measured Progressive Disease or Death from Any Cause (Up to 66.6 Months)
Analysis populationAll randomized participants
MethodLog Rank
Effect measureHazard Ratio (HR)
Estimate0.86
95% CI0.71–1.04
P-value0.130
Censored participantsArm A: 99; Arm B: 87
Clinical Biostats interpretation

An HR of 0.86 corresponds to an estimated instantaneous rate of progression or death that is 0.86 times the corresponding rate in the comparator group, under the hazard-ratio model. Expressed as a simple relative-hazard interpretation, 0.86 corresponds to a 14% lower estimated hazard in Arm A relative to Arm B.

That does not mean that 14% more participants avoided progression, nor does it mean that individual patients experienced a 14% longer progression-free survival. Absolute PFS probabilities and median PFS are not reported in the ClinicalTrials.gov record.

The 95% CI of 0.71–1.04 crosses 1.00. The interval therefore includes both a potentially lower hazard and a value slightly above 1.00 under the reported analysis. The P-value of 0.130 is a test result, not an effect-size measure, and should not be interpreted as a percentage probability that the treatment works or does not work.

Because PFS is a time-to-event endpoint, censoring and the definition of progression are central to its interpretation. The registry specifies the event as measured progressive disease or death from any cause. The ClinicalTrials.gov record does not provide a formal assessment of proportional hazards or the detailed censoring rules beyond the information summarized in the analysis record.

7. Secondary Endpoint: Objective Response Rate

The registry defines objective response rate as Complete Response (CR) + Partial Response (PR), measured from baseline to measured progressive disease, with a time frame of up to 7 months. The analysis population was all randomized participants.

FeatureReported value
EndpointObjective Response Rate (Complete Response [CR] + Partial Response [PR])
Time frameBaseline to Measured Progressive Disease (Up to 7 Months)
Analysis populationAll randomized participants
MethodLog Rank
P-value0.458
Confidence intervalTwo-sided
HypothesisSuperiority

The registry analysis does not provide the response percentages, response counts, or a hazard ratio for this endpoint. Consequently, the reported result should not be converted into an unreported response-rate difference or ratio.

Clinical Biostats interpretation

The reported P-value of 0.458 is the result of the registry's reported log-rank analysis. Because the registry does not provide an effect estimate or response percentages in the ClinicalTrials.gov record, the P-value cannot be used to reconstruct the magnitude of any difference between treatment groups.

This illustrates an important reporting principle: a P-value without an effect estimate does not tell the reader how large a treatment difference was observed. For a percentage endpoint, an informative report would ordinarily include the response proportion in each group together with an appropriate measure of uncertainty or between-group effect.

The endpoint is also labeled as an outcome with an endpoint type of time-to-event in the registry analysis, even though its outcome unit is percentage of participants. The page therefore preserves the registry's terminology rather than substituting an unreported statistical framework.

8. Secondary Endpoint: First Site of Disease Failure

The registry evaluates first site of disease failure in terms of relapse from baseline to relapse, up to 66.6 months. These analyses were restricted to all randomized participants with objective PD. Fisher exact tests were used for three specified relapse locations.

Relapse categoryMethodTwo-sided P-value
Relapsed within the radiation treatment fieldFisher exact test0.132
Relapsed inside thorax, outside of radiation fieldFisher exact test0.337
Relapsed distant diseaseFisher exact test0.457
Clinical Biostats interpretation

Fisher's exact test is appropriate for comparing categorical outcomes when exact inference is useful, particularly when cell counts may be small. It evaluates the allocation of categorical outcomes between groups under a specified null hypothesis; it does not itself provide an effect-size estimate.

Here, the ClinicalTrials.gov record reports three two-sided P-values: 0.132, 0.337, and 0.457. None should be converted into an unreported percentage difference or risk ratio. The ClinicalTrials.gov record also do not provide the underlying counts for each relapse category.

Because these are multiple secondary comparisons, interpretation should also distinguish the individual test results from a broader claim about the overall pattern of disease failure. The ClinicalTrials.gov record does not specify a multiplicity-adjustment procedure for these three analyses.

9. Secondary Endpoint: Swallowing Diary

The registry also reports the percentage of participants with a post-baseline swallowing diary score ≥4, measured from baseline through 30 days post study. The analysis population consisted of all randomized participants with at least one post-baseline swallowing diary score.

Swallowing diary comparison

P = 0.150

Fisher exact test   ·   two-sided   ·   superiority hypothesis

FeatureReported value
EndpointPercentage of Participants With a Post Baseline Swallowing Diary Score ≥4
Time frameBaseline through 30 Days Post Study
Analysis populationAll randomized participants with at least one post baseline swallowing diary score
MethodFisher exact test
P-value0.150
Confidence intervalTwo-sided
Clinical Biostats interpretation

The registry reports a two-sided Fisher exact test P-value of 0.150. The result does not provide an effect estimate or the percentage in either treatment group in the ClinicalTrials.gov record, so the magnitude of any observed difference cannot be quantified from this record alone.

The analysis population is also narrower than the overall randomized population: participants needed at least one post-baseline swallowing diary score. That distinction matters because eligibility for the analysis depends on availability of post-baseline data. The ClinicalTrials.gov record does not specify an imputation procedure for missing swallowing diary assessments.

10. Statistical Methodology

Kaplan-Meier estimation

The primary overall-survival endpoint is explicitly summarized using Kaplan-Meier estimates. Kaplan-Meier estimation is designed for time-to-event data with right censoring. Instead of requiring every participant to have an observed death, it uses each participant's observed follow-up until death or censoring.

Conceptual form
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents the number of events at event time ti, while ni is the number at risk immediately before that time.

Log-rank test

The log-rank test compares time-to-event experience between groups across the observed follow-up. It is particularly useful when the question concerns whether the survival distributions differ rather than whether a single fixed-time proportion differs.

In PROCLAIM, the registry reports the log-rank method for overall survival and progression-free survival, and also reports it for objective response rate. The page preserves that registry-reported methodology rather than replacing it with an inferred method.

Hazard ratio

A hazard ratio summarizes the relative instantaneous event rate between two groups under a time-to-event model. An HR below 1 indicates a lower estimated hazard for the numerator group, while an HR above 1 indicates a higher estimated hazard.

Interpretation of the hazard ratio
HR < 1  →  lower estimated instantaneous event rate in the first group

A hazard ratio is not an absolute risk difference, not a probability of survival, and not necessarily a constant relative difference in cumulative event probability at every time point.

Fisher exact test

Fisher's exact test evaluates a two-group comparison for a categorical outcome using the exact distribution of the observed table under the null hypothesis. It can be especially useful when expected cell counts are small, although the ClinicalTrials.gov record does not report the cell counts for the PROCLAIM relapse or swallowing analyses.

Confidence intervals

A 95% confidence interval describes the uncertainty associated with an estimated parameter under the statistical model and sampling framework. For the PROCLAIM overall-survival HR, the interval is 0.79–1.20. For PFS, it is 0.71–1.04.

Because both intervals include 1.00, neither interval excludes the null value for a hazard ratio. This is consistent with the corresponding P-values of 0.831 for OS and 0.130 for PFS.

11. Statistical Methods Explained

Why was a log-rank test used for overall survival?

Overall survival records both whether an event occurred and when it occurred. A simple comparison of the proportion dead at a single time point would discard much of that information and would not naturally handle participants whose follow-up ends before death. The log-rank test uses the ordering of event times across the groups and accommodates right-censored observations.

What does an overall-survival HR of 0.98 mean?

Under the hazard-ratio interpretation, 0.98 means that the estimated instantaneous rate of death for Arm A relative to Arm B was 0.98. Equivalently, 1 − 0.98 = 0.02, so the point estimate corresponds to a 2% lower estimated hazard in Arm A. That arithmetic describes the point estimate only; it does not establish a 2% reduction in absolute mortality.

Why does the 95% CI matter more than the point estimate alone?

A point estimate is only one summary of the observed comparison. The 95% CI of 0.79–1.20 shows that considerable uncertainty surrounds the OS estimate. It includes values below 1 and above 1, so the data are not precise enough to isolate a narrow range of relative hazard differences.

Why doesn't P = 0.831 mean there is an 83.1% probability that the treatments are equivalent?

A P-value is calculated under a null hypothesis and describes the extremeness of the observed data, or data more extreme, under that hypothesis. It is not the posterior probability that the null hypothesis is true and it does not establish equivalence. A separate equivalence or non-inferiority design would require prespecified margins and a corresponding hypothesis-testing framework.

Why use Fisher exact testing for relapse categories?

The relapse outcomes are categorical: a participant with objective progression can be classified according to a specified first site of disease failure. Fisher's exact test provides an exact two-group test for a categorical contingency table. The ClinicalTrials.gov record does not provide the underlying counts, so the test results cannot be translated into unreported effect sizes.

Why is randomization important to the statistical analysis?

Randomization creates the treatment groups through a prespecified allocation mechanism rather than allowing participants or investigators to choose treatment. This supports the causal interpretation of between-group efficacy comparisons, subject to the trial's conduct, follow-up, endpoint definitions, and analysis assumptions.

12. Censoring and Analysis Populations

The overall-survival analysis used all randomized participants. Arm A had 124 participants censored and Arm B had 117 participants censored. The PFS analysis also used all randomized participants, with 99 censored in Arm A and 87 in Arm B.

EndpointAnalysis populationArm A censoredArm B censored
Overall SurvivalAll randomized participants124117
Progression-free SurvivalAll randomized participants9987
Objective Response RateAll randomized participantsNot reportedNot reported
First Site of Disease FailureAll randomized participants with objective PDNot reportedNot reported
Swallowing DiaryAll randomized participants with at least one post-baseline swallowing diary scoreNot reportedNot reported

Censoring is not equivalent to an unsuccessful outcome. For example, a participant who remains alive at the data cut-off contributes survival information up to the last known time they were observed alive. The validity of standard survival analysis depends on assumptions about the relationship between censoring and the event process; the ClinicalTrials.gov record does not report a detailed missing-data or censoring sensitivity analysis.

13. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized arm. The safety measure is presented as the number of affected participants divided by the number at risk.

Safety measureArm AArm B
Serious adverse events134 / 283145 / 272
Arm definitionPemetrexed + Cisplatin and TRTEtoposide + Cisplatin and TRT
Serious adverse events: affected / at risk
Arm A
134 / 283
Arm B
145 / 272

The displayed percentages in the graphic are simple visual representations of the registry-reported affected/at-risk counts and are not additional reported trial estimates. The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or P-value for serious adverse events.

Safety interpretation: efficacy and safety answer different statistical questions. The randomized survival analysis evaluates time-to-event efficacy outcomes, whereas serious adverse events are participant-level categorical safety outcomes. They should not be collapsed into a single numerical benefit-risk measure.

14. What the Primary Hazard Ratio Does — and Does Not — Mean

Point estimate

The OS hazard ratio of 0.98 means that the estimated instantaneous rate of death in Arm A was 0.98 times that in Arm B under the reported analysis. Expressed as a simple relative-hazard calculation, this is a 2% lower estimated hazard for Arm A at the point estimate.

It does not mean that 2% fewer participants died, that survival probability increased by 2 percentage points, or that every participant experienced the same relative difference.

Confidence interval

The 95% CI of 0.79–1.20 communicates uncertainty around the estimated hazard ratio. The interval includes the null value of 1.00, as well as values representing lower and higher hazards for Arm A relative to Arm B.

P-value

The P-value of 0.831 is evidence about the statistical compatibility of the observed result with the null hypothesis used by the log-rank test. It is not a measure of effect size and should not be interpreted as a probability that one treatment is better, worse, or equivalent to the other.

Absolute outcomes

Hazard ratios are relative measures. A complete clinical interpretation normally benefits from absolute survival estimates at meaningful time points and median survival when available. Those quantities are not included in the registry-reported PROCLAIM statistical-analysis data and therefore are not reported here.

15. Interpreting the Secondary PFS Result

The PFS estimate of 0.86 is numerically farther below 1 than the OS estimate of 0.98. A simple interpretation of the point estimate is that the estimated instantaneous rate of progression or death was 14% lower in Arm A. However, the 95% CI of 0.71–1.04 includes 1.00 and the P-value is 0.130.

This is a useful example of why effect size, precision, and hypothesis testing should be read together. The point estimate alone suggests a relative difference; the confidence interval shows uncertainty around that estimate; and the P-value provides the result of the specified hypothesis test. None of these quantities should be interpreted in isolation.

16. Endpoint-Specific Statistical Map

EndpointTypeMethodEffect measureP-value
Overall SurvivalTime-to-eventLog-rankHR 0.98 (95% CI 0.79–1.20)0.831
Progression-free SurvivalTime-to-eventLog-rankHR 0.86 (95% CI 0.71–1.04)0.130
Objective Response RateTime-to-event in registry-reported registry classificationLog-rankNot reported0.458
Relapse: radiation treatment fieldBinaryFisher exactNot reported0.132
Relapse: inside thorax, outside radiation fieldBinaryFisher exactNot reported0.337
Relapse: distant diseaseBinaryFisher exactNot reported0.457
Swallowing diary score ≥4BinaryFisher exactNot reported0.150

The statistical map illustrates an important feature of clinical-trial reporting: not every endpoint is summarized with the same effect measure. Time-to-event outcomes naturally support hazard ratios and survival estimates, while categorical outcomes can be compared using exact tests. A P-value alone does not replace an effect measure.

17. Multiplicity and the Superiority Framework

The ClinicalTrials.gov record identifies Overall Survival as the single registered primary endpoint and identify the hypothesis type as superiority. Seven statistical analyses are posted in total: one primary analysis and six secondary analyses in the ClinicalTrials.gov record.

FeatureRegistry information
Registered primary endpoints1
Primary endpointOverall Survival
Primary hypothesisSuperiority
Statistical analyses posted7
Primary analyses with estimate + CI1
Non-inferiority marginNot reported in the ClinicalTrials.gov record
Interim-analysis methodNot reported in the ClinicalTrials.gov record
Bayesian methodNot reported in the ClinicalTrials.gov record

Because the primary endpoint is superiority, the interpretation of the OS result is based on the reported superiority analysis rather than a non-inferiority framework. A non-inferiority analysis would require a prespecified margin defining the largest clinically acceptable loss of efficacy; no such margin is included in the ClinicalTrials.gov record.

Multiplicity caution: seven statistical analyses are posted, but the ClinicalTrials.gov record does not specify a multiplicity-adjustment strategy. The secondary P-values should therefore be interpreted as endpoint-specific reported test results rather than as evidence of a particular familywise error-control procedure that has not been documented in the ClinicalTrials.gov record.

18. Missing Data and Imputation

The ClinicalTrials.gov record does not report an explicit missing-data or imputation method. This matters particularly for the swallowing diary endpoint, whose analysis population requires at least one post-baseline swallowing diary score.

For time-to-event endpoints, censoring is part of the endpoint analysis rather than a conventional imputation of a missing event time. The registry explicitly describes censoring for overall survival and provides censored-participant counts for OS and PFS. The ClinicalTrials.gov record does not describe sensitivity analyses under alternative missing-data assumptions.

OS

Participants not known to have died by the data cut-off were censored at the last contact date known to be alive.

PFS

The analysis includes all randomized participants and reports censored counts, but the ClinicalTrials.gov record does not provide a fuller missing-data strategy.

Swallowing diary

The analysis requires at least one post-baseline swallowing diary score.

Other endpoints

No additional imputation procedures are specified in the ClinicalTrials.gov record.

19. Randomization and Causal Interpretation

Randomization is the foundation of the main comparative inference in PROCLAIM. With two randomized parallel arms, the treatment groups are intended to differ systematically in treatment assignment while balancing prognostic factors in expectation.

That design does not make every observed difference automatically causal. Causal interpretation still depends on adherence to the randomized design, follow-up, outcome ascertainment, censoring, and prespecified analysis. For the primary OS endpoint, the use of all randomized participants maintains the treatment assignment framework specified in the ClinicalTrials.gov record.

Randomized comparison
Treatment assignment → observed time-to-event outcomes → between-group statistical comparison

The strength of the randomized comparison comes from assigning treatment before the outcome is observed, rather than from the P-value alone.

20. Why Time-to-Event Analysis Is Central Here

Both the primary OS endpoint and the secondary PFS endpoint are time-to-event outcomes. This changes the statistical problem compared with a simple binary endpoint. The analysis needs to preserve information about when an event occurs and how long each participant remains under observation.

For OS, the event is death from any cause. For PFS, the event is measured progressive disease or death from any cause. A participant without an event at the relevant data cut-off contributes follow-up information until censoring.

Event timing

The analysis uses the timing of events rather than only whether an event eventually occurred.

Censoring

Participants can contribute partial follow-up without being classified as having experienced the event.

Kaplan-Meier

Provides an estimate of the event-free survival function over time.

Log-rank

Provides a formal comparison of the time-to-event experience between groups.

21. Important Limitations and Interpretation Issues

22. Why This Trial Matters Statistically

PROCLAIM is a useful teaching case because its registry record connects randomized treatment assignment with time-to-event analysis, categorical exact testing, censoring, confidence intervals, and multiple secondary endpoints.

ConceptHow it appears in PROCLAIM
RandomizationParticipants were randomized to two parallel treatment arms.
Superiority testingThe primary hypothesis type is superiority.
Kaplan-Meier estimationOverall survival was summarized using Kaplan-Meier estimates.
Hazard ratioOS HR 0.98; PFS HR 0.86.
Confidence interval95% two-sided CIs are reported for the OS and PFS hazard ratios.
Log-rank testingUsed for the primary OS analysis and reported secondary analyses.
Fisher exact testingUsed for specified relapse locations and the swallowing diary endpoint.
Censoring124 vs 117 OS participants and 99 vs 87 PFS participants were censored by arm.
Analysis populationsSome secondary analyses use restricted populations rather than all randomized participants.
MultiplicitySeven statistical analyses are posted, while the ClinicalTrials.gov record does not specify a multiplicity procedure.

23. Statistical Methods Explained: A Deeper View

Why isn't a hazard ratio the same as a risk ratio?

A risk ratio compares probabilities over a specified period. A hazard ratio compares instantaneous event rates under a time-to-event model. The two measures answer different questions and need not have the same numerical value.

Why can a confidence interval include 1 even when the point estimate is below 1?

The point estimate is based on the observed data, while the confidence interval incorporates sampling uncertainty. For PFS, the point estimate is 0.86, but the 95% CI extends from 0.71 to 1.04. The observed estimate is below 1, while plausible parameter values under the stated confidence framework include values above 1.

Why are censored patients still useful?

A censored participant provides information about remaining event-free up to the censoring time. Treating that participant as if an event occurred at censoring would incorrectly introduce an event that was not observed.

Why are OS and PFS not interchangeable?

OS uses death from any cause as the event. PFS uses measured progressive disease or death from any cause. A participant can therefore experience progression before death, meaning the two endpoints measure different points in the disease course.

Why can the analysis populations differ?

An endpoint can require information that is not needed for another endpoint. The swallowing analysis requires at least one post-baseline diary score, while relapse analyses require objective PD. These eligibility conditions define different analysis populations and should be stated explicitly.

Why should secondary P-values not be treated as rankings?

Each P-value corresponds to a particular statistical question. A smaller P-value does not automatically indicate a larger or more important treatment effect, particularly when the underlying effect estimates are not reported. Multiple endpoints also create a broader multiplicity question that cannot be resolved from individual P-values alone.

24. Reported Results Summary

EndpointArm A vs Arm B95% CIP-valueMethod
Overall SurvivalHR 0.980.79–1.200.831Log-rank
Progression-free SurvivalHR 0.860.71–1.040.130Log-rank
Objective Response RateEffect estimate not reportedTwo-sided0.458Log-rank
Relapse: radiation treatment fieldEffect estimate not reportedTwo-sided0.132Fisher exact
Relapse: inside thorax, outside radiation fieldEffect estimate not reportedTwo-sided0.337Fisher exact
Relapse: distant diseaseEffect estimate not reportedTwo-sided0.457Fisher exact
Swallowing diary score ≥4Effect estimate not reportedTwo-sided0.150Fisher exact

The two endpoints with reported effect estimates are both time-to-event outcomes. Their point estimates are below 1, but their two-sided 95% confidence intervals include 1.00. The other five posted analyses supply P-values but not effect estimates in the ClinicalTrials.gov record, so their magnitude cannot be reconstructed without introducing information outside the permitted trial record.

25. Clinical Interpretation vs Statistical Interpretation

Statistical interpretation

The primary OS analysis reported an HR of 0.98 with a 95% CI of 0.79–1.20 and P = 0.831. The secondary PFS analysis reported an HR of 0.86 with a 95% CI of 0.71–1.04 and P = 0.130.

What remains unreported here

The ClinicalTrials.gov record does not provide median OS, median PFS, time-specific survival percentages, detailed response percentages, baseline characteristics, or formal multiplicity and interim-analysis specifications.

This distinction is important. A statistical analysis should describe what the reported estimates establish, what they leave uncertain, and which quantities are simply not available in the ClinicalTrials.gov record. Filling those gaps from memory or from an outside publication would change the evidentiary basis of the page.

26. Related Tutorials

Learn more about the methods used in this trial:

27. Related Calculators

28. Sources

Continue with the statistical methods behind PROCLAIM

Explore the survival-analysis, hypothesis-testing, confidence-interval, and categorical-data methods that appear in the trial's registered statistical analyses.

29. Record Summary

PROCLAIM provides a compact example of how a randomized phase 3 oncology trial can combine several statistical frameworks. Its registered primary endpoint was overall survival, analyzed with Kaplan-Meier estimation and a log-rank comparison summarized by a hazard ratio. The reported OS estimate was 0.98 with a 95% CI of 0.79–1.20 and P = 0.831. PFS was analyzed similarly, with an HR of 0.86, 95% CI 0.71–1.04, and P = 0.130. Secondary analyses used both log-rank and Fisher exact methods.

The most important statistical lesson is that these numbers should be interpreted together rather than independently. The hazard ratio communicates relative event rates, the confidence interval communicates precision, and the P-value addresses the specified hypothesis test. For categorical endpoints where only P-values are reported, the magnitude of the treatment difference cannot be inferred. The analysis populations and censoring rules also matter: the primary OS analysis used all randomized participants, while some secondary analyses required objective progression or post-baseline diary data.

Clinical Biostats methodology: This page separates the numerical results reported by the registry from statistical explanation. Where the registry-reported PROCLAIM record does not provide an estimate, confidence interval, baseline table, median survival, multiplicity procedure, or interim-analysis specification, those quantities are not reconstructed from outside information.