← Clinical Trials
Renal Cell Carcinoma Phase 3 Time-to-Event Analysis NCT01865747

METEOR: Complete Statistical Analysis of Cabozantinib in Metastatic Renal Cell Carcinoma

An independent statistical analysis of the randomized phase 3 METEOR trial comparing cabozantinib (XL184) with everolimus (Afinitor) in subjects with metastatic renal cell carcinoma, focusing on progression-free survival, overall survival, objective response rate, and the statistical methods used for their comparison.

Trial status: COMPLETED  ·  Enrollment: 658  ·  Primary completion: 2015-05-22
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical results on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

METEOR was a randomized, parallel-group, open-label phase 3 trial comparing cabozantinib tablets with everolimus tablets in subjects with metastatic renal cell carcinoma. The registry reports 658 enrolled subjects, two treatment arms, one registered primary time-to-event endpoint, and formal statistical analyses for progression-free survival, overall survival, and objective response rate.

658
Enrolled
Randomized trial
2
Treatment arms
Parallel design
0.58
PFS HR
95% CI 0.45–0.74
0.66
OS HR
95% CI 0.53–0.83
FeatureMETEOR
Trial nameMETEOR
Brief titleA Study of Cabozantinib (XL184) vs Everolimus in Subjects With Metastatic Renal Cell Carcinoma
PhasePhase 3
ConditionRenal Cell Carcinoma
AllocationRandomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment658
InterventionsCabozantinib tablets (drug); Everolimus (Afinitor) tablets (drug)
Lead sponsorExelixis
Sponsor typeIndustry
StatusCompleted
Start2013-06
Primary completion2015-05-22
ClinicalTrials.govNCT01865747

2. Clinical Question

The central statistical question was whether randomized assignment to cabozantinib, compared with everolimus, was associated with a difference in progression-free survival in subjects with metastatic renal cell carcinoma. The registry identifies progression-free survival as the single registered primary endpoint and classifies the formal hypothesis as one of superiority.

Population

Subjects with metastatic renal cell carcinoma enrolled in the randomized phase 3 METEOR trial.

Intervention

Cabozantinib (XL184) tablets.

Comparator

Everolimus (Afinitor) tablets.

Primary question

Does cabozantinib produce a superior progression-free survival outcome compared with everolimus?

3. Trial Design

01
Randomize658 subjects
02
Two armsCabozantinib vs everolimus
03
Follow-upTime-to-event outcomes
04
Primary analysisPFS
05
Secondary analysesOS and ORR
ARM A

Cabozantinib

  • Cabozantinib (XL184) tablets
  • Randomized treatment assignment
  • Compared with everolimus for PFS, OS, and ORR analyses
ARM B

Everolimus

  • Everolimus (Afinitor) tablets
  • Randomized treatment assignment
  • Comparator arm for PFS, OS, and ORR analyses
Allocation
Randomized allocation was used to establish the treatment comparison.
Design model
Parallel-group design with two treatment arms.
Masking
None. The registry classifies the study as unmasked.
Primary purpose
Treatment.

Because this was an open-label randomized comparison, the statistical interpretation depends particularly strongly on how objective time-to-event endpoints were defined and analyzed. Randomization establishes the treatment comparison, while the survival-analysis methods translate follow-up information into an estimate of the difference between treatment groups.

4. Endpoints

EndpointRegistry definition / time frameTypeAnalysis
Progression-free Survival (PFS) PFS is measured from the date of randomization until the date of first documented disease progression or date of death from any cause as determined by the Independent Radiology Committee (IRC) per RECIST 1.1, assessed for up to 17 months.. The primary analysis of PFS is the time from randomization to date of first documented tumor progression as determined by investigator (per RECIST 1.1 criteria) or death due to any cause, whichever occurred first. Time-to-event Stratified log-rank test; hazard ratio
Overall Survival (OS) OS was measured from the time of randomization until 320 deaths, approximately 28 months. Time-to-event Stratified log-rank test; hazard ratio
Objective Response Rate (ORR) ORR was assessed at 8 weeks post-randomization, every 8 weeks for 12 months, and every 12 weeks until date of disease progression or death, up to May 2015 (approximately 21 months). Binary Cochran-Mantel-Haenszel test

The registry identifies one primary endpoint: progression-free survival. Overall survival and objective response rate are reported as secondary endpoints. The distinction matters because the inferential role of a result depends not only on its numerical value but also on where the endpoint sits in the prespecified trial structure.

Endpoint definition matters. For PFS, an event could occur through either investigator-determined tumor progression under RECIST 1.1 or death from any cause, whichever occurred first. This makes PFS a composite time-to-event endpoint rather than a measure of tumor progression alone.

5. Statistical Methodology

Stratified log-rank testing

The registry reports a log-rank test for the primary PFS analysis and for the secondary OS analysis. For PFS, the analysis text states that the log-rank test was stratified by the Memorial Sloan-Kettering Cancer Center (MSKCC) group and the number of prior VEGFR TKIs. For OS, the same two stratification concepts are reported: MSKCC risk group and number of prior VEGFR TKIs.

A log-rank test compares the observed pattern of event occurrence between randomized groups across follow-up. In a time-to-event trial, it uses the ordering of event times while accounting for the changing number of participants still at risk. Stratification allows the comparison to be performed within prespecified strata rather than treating all subjects as though the stratification factors were irrelevant.

Conceptual survival comparison
Observed events  ↔  Expected events under the null hypothesis

The log-rank framework asks whether the event experience over follow-up differs systematically between randomized treatment groups, conditional on the analysis structure.

Hazard ratio as the effect measure

The registry reports the hazard ratio (HR) as the effect measure for both PFS and OS. A hazard ratio compares the estimated instantaneous event rates between treatment groups over the analyzed follow-up.

Interpretation
HR = hazard in cabozantinib group ÷ hazard in everolimus group

An HR below 1 indicates a lower estimated instantaneous event rate in the cabozantinib group relative to the everolimus group under the analysis model.

Cochran-Mantel-Haenszel analysis

The secondary ORR analysis used the Cochran-Mantel-Haenszel test. This is a categorical-data method that can compare treatment groups while accounting for stratification. The registry identifies the analysis as a superiority test in the ITT population.

Unlike PFS and OS, ORR is binary: a participant either meets the prespecified response definition or does not. That difference in endpoint structure explains why a categorical-data method was used for ORR while a survival-analysis method was used for PFS and OS.

Intention-to-treat analysis

The OS analysis explicitly used the Intent to Treat (ITT) population, consisting of all 658 randomized subjects: 330 assigned to cabozantinib and 328 assigned to everolimus. The ORR analysis also used the ITT population, with the same randomized group sizes. ITT analysis preserves the randomized treatment comparison rather than redefining groups according to treatment received after randomization.

Kaplan-Meier estimation

The registered PFS definition states that a Kaplan-Meier analysis was performed to estimate the median duration. Kaplan-Meier estimation is appropriate for time-to-event data because it can incorporate participants whose event has not occurred by the end of available follow-up.

Kaplan-Meier survival estimate
S(t) = ∏ti ≤ t (1 − di/ni)

Here, di represents events at an event time and ni represents the number at risk immediately before that time. The ClinicalTrials.gov record does not report the resulting median PFS, so this page does not introduce a median value.

6. Analysis Populations and Stratification

EndpointAnalysis populationRandomized group sizesImportant analysis features
Primary PFS First 375 randomized subjects 187 cabozantinib; 188 everolimus Stratified log-rank test; stratified by MSKCC group and number of prior VEGFR TKIs
OS ITT population 330 cabozantinib; 328 everolimus Second interim analysis; cutoff 31 December 2015; stratified log-rank test
ORR ITT population 330 cabozantinib; 328 everolimus IRC response assessment per RECIST 1.1; Cochran-Mantel-Haenszel test

The difference between the PFS and OS analysis populations is an important statistical feature. The prespecified primary PFS analysis was based on the first 375 randomized subjects, whereas the reported OS analysis used all 658 randomized subjects in the ITT population. These are therefore not interchangeable analyses, even though both compare cabozantinib with everolimus.

Do not silently pool analysis populations. A treatment effect estimated in 375 participants cannot be described as though it were estimated in all 658 randomized subjects. The analysis population is part of the definition of the result.

7. Results: Progression-Free Survival

PFS was the single registered primary endpoint. The primary analysis was based on the first 375 randomized subjects: 187 assigned to cabozantinib and 188 assigned to everolimus. The reported method was a log-rank test, stratified by MSKCC group and number of prior VEGFR TKIs.

Primary PFS treatment effect

HR 0.58

95% CI: 0.45–0.74   ·   P < 0.0001

Cabozantinib vs everolimus  ·  Superiority hypothesis

Primary PFS elementReported value
Analysis populationFirst 375 randomized subjects
Cabozantinib187 subjects
Everolimus188 subjects
MethodLog-rank test
StratificationMSKCC group and number of prior VEGFR TKIs
Effect measureHazard ratio
Hazard ratio0.58
95% confidence interval0.45–0.74
P-value<0.0001
HypothesisSuperiority
Clinical Biostats interpretation

An HR of 0.58 means that, under the time-to-event analysis, the estimated instantaneous rate of progression or death in the cabozantinib group was approximately 58% of that in the everolimus group. Equivalently, 0.58 corresponds to an estimated 42% lower instantaneous hazard relative to everolimus.

The HR does not mean that 42% of patients avoided progression, that every patient experienced exactly a 42% reduction in risk, or that the probability of progression was reduced by 42% at every individual time point. A hazard ratio is a relative time-to-event measure, not an absolute risk difference.

The 95% CI of 0.45–0.74 describes statistical uncertainty around the estimated hazard ratio under the analysis framework. It does not describe the range of effects that individual patients experienced.

The P-value of <0.0001 addresses the strength of evidence against the null hypothesis within the specified statistical test. It is not a measure of effect size, clinical importance, or the probability that the treatment is effective.

Because the analysis is based on time-to-event data, interpretation also depends on censoring and on the assumptions underlying the hazard-ratio representation. The ClinicalTrials.gov record identifies the stratified log-rank analysis and HR but do not provide additional model-diagnostic information that would permit a separate assessment of the proportional-hazards assumption.

Why the PFS analysis is statistically distinctive

The primary PFS analysis combines several important design features: randomization, a prespecified subset of the randomized population, a time-to-event endpoint, stratification, a log-rank comparison, and a hazard-ratio effect measure. Each component answers a different statistical question.

Randomization

Creates the treatment groups whose outcomes are compared, reducing the role of measured and unmeasured baseline differences in the treatment comparison.

Time-to-event endpoint

Uses not only whether progression or death occurred, but also when the event occurred and how long participants remained event-free.

Stratification

Accounts for the prespecified MSKCC group and number of prior VEGFR TKIs in the log-rank comparison.

Hazard ratio

Provides a relative measure of the event rate between the two randomized treatment groups.

8. Results: Overall Survival

Overall survival was a secondary endpoint. The registry states that OS was measured from randomization until 320 deaths, approximately 28 months. The second interim analysis used the full ITT population of 658 randomized subjects, with a cutoff date of 31 December 2015.

Secondary OS treatment effect

HR 0.66

95% CI: 0.53–0.83   ·   P = 0.0003

Cabozantinib vs everolimus  ·  Superiority hypothesis

OS analysis elementReported value
Analysis populationITT population
Cabozantinib330 subjects
Everolimus328 subjects
Interim analysisSecond interim analysis
Cutoff date31 December 2015
Event target / time frame320 deaths, approximately 28 months
MethodLog-rank test
StratificationMSKCC risk group and number of prior VEGFR TKIs
Effect measureHazard ratio
Hazard ratio0.66
95% confidence interval0.53–0.83
P-value0.0003
Clinical Biostats interpretation

An HR of 0.66 means that the estimated instantaneous rate of death in the cabozantinib group was approximately 66% of that in the everolimus group under the reported analysis. Equivalently, 0.66 corresponds to an estimated 34% lower instantaneous hazard relative to everolimus.

This is not the same as saying that 34% fewer patients died, that survival probability increased by 34%, or that each participant had exactly a 34% reduction in the probability of death. The hazard ratio summarizes a relative time-to-event comparison.

The 95% CI of 0.53–0.83 provides a measure of precision around the estimated HR. It indicates that the reported estimate is not a single exact population quantity known without uncertainty; the interval reflects statistical uncertainty under the analysis framework.

The P-value of 0.0003 quantifies evidence against the null hypothesis for the reported test. It does not quantify the magnitude of the treatment effect, the clinical value of the treatment, or the probability that the null hypothesis is true.

The OS analysis was explicitly an interim analysis. Interim analyses require careful interpretation because the timing of an analysis relative to accumulating events can affect statistical inference. The ClinicalTrials.gov record identifies the analysis as a second interim analysis but do not provide an alpha-spending boundary or numerical information-allocation schedule, so none is inferred here.

Why OS and PFS should not be treated as interchangeable

PFS and OS are both time-to-event outcomes, but they define different events. PFS ends at the first documented tumor progression or death, whichever occurs first. OS ends at death from any cause. A participant can therefore experience a PFS event without having experienced an OS event.

This distinction is important when interpreting the two HRs. The PFS HR of 0.58 and OS HR of 0.66 describe different event processes. One should not be interpreted as a direct surrogate for the other solely because both are reported as hazard ratios.

9. Results: Objective Response Rate

ORR was a secondary binary endpoint. The registry states that ORR was assessed at 8 weeks post-randomization, every 8 weeks for 12 months, and every 12 weeks thereafter until the date of disease progression as recorded in the registry wording.

The ORR analysis used the ITT population: all randomized participants, consisting of 330 cabozantinib and 328 everolimus. Response was determined by an Independent Radiology Committee (IRC) according to RECIST 1.1.

Secondary ORR comparison

P < 0.0001

Cochran-Mantel-Haenszel test  ·  ITT population

Cabozantinib vs everolimus  ·  Superiority hypothesis

ORR analysis elementReported value
Analysis populationITT population
Cabozantinib330 randomized subjects
Everolimus328 randomized subjects
Response assessmentIndependent Radiology Committee per RECIST 1.1
Endpoint typeBinary
MethodCochran-Mantel-Haenszel test
P-value<0.0001
HypothesisSuperiority
Clinical Biostats interpretation

The reported P < 0.0001 indicates strong evidence against the null hypothesis under the reported Cochran-Mantel-Haenszel comparison of ORR. It does not, by itself, tell us the magnitude of the difference in response rates.

That distinction is especially important here because the ClinicalTrials.gov record provides the formal ORR test and its P-value but do not provide the arm-specific response percentages or a confidence interval for the response-rate difference. Therefore, this page does not invent an ORR effect estimate that is absent from the ClinicalTrials.gov record.

ORR also answers a different question from PFS. A response endpoint classifies tumor response, whereas PFS incorporates the timing of progression or death. Consequently, the ORR P-value should not be interpreted as evidence for a particular magnitude of PFS benefit.

10. Safety Results

The ClinicalTrials.gov record reports serious adverse events by randomized treatment arm using affected participants over participants at risk. These data should be read separately from the efficacy analyses because safety and efficacy answer different questions and may use different analysis populations or definitions.

Safety measureCabozantinib (XL184)Everolimus (Afinitor)
Serious adverse events131 / 331 affected / at risk139 / 322 affected / at risk

The registry-derived ClinicalTrials.gov record do not provide a formal statistical comparison, confidence interval, or P-value for serious adverse events. Accordingly, the figures are presented descriptively rather than converted into an inferred hypothesis test.

Why denominators matter: 131/331 and 139/322 are affected-participant counts divided by the registry-reported at-risk counts. They should not be silently replaced with the randomized sample sizes used in the ITT efficacy analyses.

11. Interim Analysis and Statistical Inference

The OS analysis was explicitly identified as the second interim analysis, with a cutoff date of 31 December 2015 and an endpoint time frame defined around 320 deaths, approximately 28 months.

Interim analysis is a design feature rather than merely a date label. When investigators examine accumulating outcome data before the trial is complete, the statistical design must account for the fact that multiple opportunities to evaluate the evidence can alter the probability of observing a nominally significant result by chance.

What the registry establishes

The ClinicalTrials.gov record establishes that the reported OS result came from a second interim analysis and that the analysis included the full ITT population.

What the registry data do not establish

The ClinicalTrials.gov record does not provide an alpha-spending function, interim boundary, or numerical allocation of type I error. These details are therefore not reconstructed here.

This distinction is important because a nominal P-value such as 0.0003 should be interpreted in the context of the prespecified interim-analysis framework. The existence of an interim analysis does not by itself invalidate the result; rather, it means that the timing and rules governing interim inference are part of the statistical design.

12. Stratified Analysis

Both time-to-event analyses used stratification. The PFS log-rank test was stratified by MSKCC group and number of prior VEGFR TKIs. The OS log-rank test was stratified by MSKCC risk group and number of prior VEGFR TKIs.

Why stratify?
Treatment comparison within prespecified strata → combined treatment comparison

Stratification allows the survival comparison to account for specified baseline or disease-history categories rather than ignoring them completely in the primary test.

Stratification does not mean that the trial produced a separate primary conclusion for every MSKCC category or every level of prior VEGFR TKI exposure. It means that these factors were incorporated into the reported comparison. The ClinicalTrials.gov record does not provide subgroup-specific hazard ratios or interaction tests, so no subgroup treatment-effect conclusions are added.

13. Statistical Methods Explained

Why was a log-rank test used for PFS?

PFS is a time-to-event endpoint. Some participants experience progression or death during follow-up, while others may not experience the event during the period in which they are observed. The log-rank test is designed to compare the event-time distributions of two groups while incorporating the timing of observed events and censoring.

For METEOR, the reported PFS comparison was additionally stratified by MSKCC group and number of prior VEGFR TKIs. The resulting P-value therefore belongs to the specified stratified comparison, not to an unadjusted comparison invented after the fact.

What does an HR of 0.58 mean?

An HR of 0.58 means that the estimated instantaneous rate of progression or death was 0.58 times that of the comparator under the reported time-to-event analysis. Expressed as a relative difference, 1 − 0.58 = 0.42, so the estimate corresponds to a 42% lower instantaneous hazard relative to everolimus.

It does not mean a 42% reduction in every patient's probability of progression, and it does not provide an absolute percentage of patients who benefited.

Why does the confidence interval matter?

The PFS HR was estimated as 0.58 with a 95% CI of 0.45–0.74. The interval conveys statistical uncertainty around the point estimate. A narrower interval generally represents greater precision than a wider interval, although precision also depends on the underlying design and amount of information.

The confidence interval should not be interpreted as the range of individual patient responses. It concerns uncertainty in the estimated treatment effect at the population-analysis level.

Why doesn't the P-value measure effect size?

A P-value measures how compatible the observed data are with a specified null hypothesis under the statistical test. It depends on both the magnitude of the observed difference and the amount of statistical information available. Consequently, a small P-value does not by itself establish that an effect is large or clinically important.

For METEOR, the PFS result provides both pieces of information separately: the HR 0.58 describes the estimated relative treatment effect, while P < 0.0001 describes the strength of evidence against the tested null hypothesis.

Why use the ITT population for OS and ORR?

The ITT principle retains participants in the groups to which they were randomized. For METEOR, the OS and ORR analyses used all 658 randomized subjects, with 330 assigned to cabozantinib and 328 assigned to everolimus.

This preserves the treatment comparison generated by randomization. It also means that the analysis is not restricted only to participants who completed treatment or followed an ideal treatment course.

Why was a Cochran-Mantel-Haenszel test used for ORR?

ORR is a binary endpoint rather than a time-to-event endpoint. The Cochran-Mantel-Haenszel test is a categorical-data method that can compare groups while accounting for stratification. In METEOR, the registry reports this method for the ITT ORR analysis and reports a superiority hypothesis.

Why should PFS and OS results not be combined into one number?

PFS and OS measure different events. PFS counts the first documented progression or death, whereas OS counts death from any cause. Their hazard ratios therefore summarize different underlying event processes. Reporting both provides complementary information rather than two measurements of exactly the same outcome.

14. Reading the Hazard Ratios Together

EndpointHR95% CIP-valueStatistical role
PFS0.580.45–0.74<0.0001Primary endpoint
OS0.660.53–0.830.0003Secondary endpoint; second interim analysis

The two HRs are both below 1, but they should be interpreted independently because their event definitions and analysis roles differ. The PFS estimate is based on the first 375 randomized subjects, while the OS estimate is based on all 658 randomized subjects. The OS analysis also has an explicitly identified second-interim-analysis context.

A useful statistical reading

The most defensible summary is not simply that "both hazard ratios were below 1." A fuller reading identifies what event was analyzed, which population was analyzed, how the comparison was performed, what the HR estimates, how precise the estimates are, and what inferential framework generated the P-values.

For PFS, the HR of 0.58 corresponds to a 42% lower estimated instantaneous hazard relative to everolimus. For OS, the HR of 0.66 corresponds to a 34% lower estimated instantaneous hazard. These are relative hazard statements, not absolute survival differences.

15. Censoring and Time-to-Event Interpretation

Time-to-event analysis differs from a simple comparison of proportions because not every participant necessarily has the event observed during the relevant follow-up. Kaplan-Meier methods allow information from participants who remain event-free at the end of their observable follow-up to contribute until their censoring time.

This is one reason a hazard ratio cannot be translated directly into an absolute percentage reduction in events. The HR summarizes the relative event rate over time within a survival-analysis framework, whereas an absolute event probability would require a specified time point and corresponding survival estimates.

Censoring

A participant without an observed event can still contribute follow-up information until the censoring point.

Event timing

An event occurring earlier and an event occurring later are not treated as equivalent observations in a time-to-event analysis.

Hazard

The hazard describes an instantaneous event rate conditional on remaining event-free up to a particular time.

Survival probability

A survival probability describes the probability of remaining event-free through a specified time point and is conceptually different from the HR.

16. Limitations and Interpretation Issues

17. Why This Trial Matters Statistically

METEOR is a useful statistical teaching case because its registry record brings together several foundational clinical-trial concepts without requiring them to be treated as interchangeable. It has randomized treatment allocation, a time-to-event primary endpoint, stratified survival analysis, a hazard-ratio effect measure, an ITT population for secondary analyses, a binary response endpoint, a Cochran-Mantel-Haenszel comparison, and an interim OS analysis.

ConceptHow it appears in METEOR
RandomizationRandomized allocation to cabozantinib or everolimus.
Parallel-group designTwo treatment arms were evaluated in parallel.
Time-to-event endpointPFS was the single registered primary endpoint.
Kaplan-Meier estimationThe registered PFS definition states that Kaplan-Meier analysis was used to estimate median duration.
Hazard ratioReported for both PFS and OS.
Log-rank testUsed for the primary PFS and secondary OS comparisons.
Stratified analysisPFS and OS log-rank tests were stratified by MSKCC group/risk group and number of prior VEGFR TKIs.
Intention-to-treat analysisUsed for OS and ORR in all 658 randomized subjects.
Binary endpoint analysisORR was analyzed with a Cochran-Mantel-Haenszel test.
Interim analysisThe reported OS analysis was the second interim analysis with a 31 December 2015 cutoff.
Confidence intervals95% two-sided confidence intervals were reported for the PFS and OS HRs.
Superiority testingThe reported PFS, OS, and ORR analyses used superiority hypotheses.

The most important lesson is methodological: a clinical-trial result is not just a number. The meaning of a number depends on the endpoint definition, analysis population, statistical method, stratification, timing of the analysis, effect measure, and uncertainty interval surrounding the estimate.

18. Related Tutorials

Learn more about the methods used in this trial:

19. Related Calculators

20. Sources

Continue through the Clinical Biostats statistical learning pathway

Explore the underlying survival-analysis, categorical-data, confidence-interval, and clinical-trial methods used to interpret randomized evidence.

21. Record Summary

METEOR provides a compact example of how several clinical-trial statistical methods fit together. The trial randomized 658 subjects in a two-arm, parallel, unmasked phase 3 design comparing cabozantinib with everolimus. Its registered primary endpoint was PFS, analyzed in the first 375 randomized subjects using a stratified log-rank test, producing an HR of 0.58 with a 95% CI of 0.45–0.74 and P < 0.0001. The secondary OS analysis used the full ITT population at the second interim analysis and produced an HR of 0.66 with a 95% CI of 0.53–0.83 and P = 0.0003. ORR was analyzed as a binary endpoint in the ITT population using a Cochran-Mantel-Haenszel test, with P < 0.0001.

The statistical interpretation should remain tied to the specific analysis that generated each estimate. PFS and OS measure different events; the primary PFS and secondary OS analyses used different populations; the OS result came from an interim analysis; and the ORR record supplies a P-value without an arm-specific response estimate in the ClinicalTrials.gov record. Keeping those distinctions visible is essential for a precise reading of the evidence.

Clinical Biostats methodology: A trial-results page should separate reported evidence from statistical interpretation. Effect estimates, confidence intervals, P-values, endpoint definitions, analysis populations, and design features should be interpreted together rather than reduced to a single numerical conclusion.