Introduction
Clinical trials often measure more than one outcome. One endpoint may occur earlier and be easier to observe, while another may be more directly meaningful to patients.
In oncology, for example, investigators may measure progression-free survival (PFS) before overall survival (OS).
PFS can be observed sooner because progression may occur before death. OS, however, is generally a more direct measure of whether patients remain alive.
This creates an important statistical question: Can PFS reliably serve as a surrogate endpoint for OS?
What Is a Surrogate Endpoint?
Let:
- \(S\) denote a surrogate endpoint.
- \(T\) denote the clinically meaningful true endpoint.
- \(Z\) denote treatment assignment.
A surrogate endpoint is an intermediate outcome intended to substitute for, or provide reliable evidence about, the effect of treatment on the true clinical endpoint.
For survival studies, the surrogate might be:
- Progression-free survival
- Disease-free survival
- Event-free survival
- Time to progression
- Biochemical progression
- A molecular or imaging-based time-to-event endpoint
The true endpoint might be:
- Overall survival
- Time to death
- A patient-centered clinical outcome
Why Use a Surrogate Endpoint?
The primary motivation is often speed.
Suppose overall survival requires several years of follow-up. If an intermediate endpoint occurs much earlier, using it may allow a trial to reach an informative analysis sooner.
Other potential advantages include:
- Shorter follow-up
- Smaller or faster trials
- Earlier treatment-development decisions
- Earlier evidence about biological activity
- Reduced duration of clinical development
However, these advantages come with a major statistical requirement: the surrogate must reliably predict the treatment effect on the true clinical outcome.
Surrogate Endpoint vs. Prognostic Endpoint
One of the most important distinctions is between a prognostic variable and a surrogate endpoint.
A prognostic variable predicts patient outcome.
A surrogate endpoint must do more: it should capture information about how treatment changes the clinically meaningful outcome.
| Concept | Question |
|---|---|
| Prognostic factor | Does the variable predict patient outcome? |
| Predictive factor | Does the variable identify patients who benefit differently from treatment? |
| Surrogate endpoint | Does treatment's effect on the surrogate reliably represent treatment's effect on the true endpoint? |
The Basic Survival-Endpoint Framework
Suppose a randomized clinical trial compares treatment \(Z=1\) with control \(Z=0\).
Let:
For an oncology trial, one possible example is:
The treatment may reduce the hazard of progression and death.
The central surrogate question is whether the treatment effect on \(S\) adequately predicts the treatment effect on \(T\).
Hazard Ratios for Surrogate and True Endpoints
Suppose the treatment effect on PFS is summarized by:
and the treatment effect on OS is:
If treatment substantially improves PFS but has little or no effect on OS, then PFS may be a poor surrogate despite being strongly associated with survival at the individual patient level.
Individual-Level Surrogacy
Individual-level surrogacy asks whether patients who have better outcomes on the surrogate also tend to have better outcomes on the true endpoint.
Conceptually:
For example, patients who remain progression-free longer may tend to survive longer.
This is useful evidence, but it is not sufficient by itself.
Trial-Level Surrogacy
Trial-level surrogacy asks a different question: Do treatments that produce larger effects on the surrogate also produce larger effects on the true endpoint?
Suppose several randomized trials compare treatment and control. For trial \(j\), let:
represent the treatment effect on the surrogate and:
represent the treatment effect on the true endpoint.
Trial-level surrogacy examines the relationship:
If the treatment effects on PFS and OS move together across trials, that is evidence supporting trial-level surrogacy.
Why Trial-Level Surrogacy Is So Important
Imagine that PFS and OS are strongly correlated across individual patients.
That does not guarantee that an intervention improving PFS will improve OS.
A treatment could delay radiographic progression while simultaneously causing toxicity that offsets or reverses the survival benefit.
Alternatively, subsequent therapies after progression could substantially alter the relationship between PFS and OS.
The Causal Perspective
A useful conceptual framework separates three quantities:
- Treatment assignment
- Effect on the surrogate
- Effect on the true endpoint
The causal pathway might be represented as:
A valid surrogate should capture a sufficiently large portion of the causal effect of treatment on the true endpoint.
The Prentice Framework
One influential framework for surrogate endpoint evaluation was proposed by Prentice.
The basic idea is that, conditional on the surrogate, treatment should provide little additional information about the true endpoint.
Conceptually, the surrogate would satisfy a relationship of the form:
where \(T\) is the true endpoint, \(Z\) is treatment assignment, and \(S\) is the surrogate.
In words: once the surrogate is known, treatment assignment should no longer provide important additional information about the true endpoint.
Four Classical Prentice Conditions
The traditional framework is often described through four requirements.
| Condition | Interpretation |
|---|---|
| Treatment affects the true endpoint | The intervention must have an effect on the clinically meaningful outcome. |
| Treatment affects the surrogate | The intervention must change the intermediate endpoint. |
| Surrogate predicts the true endpoint | The surrogate must be associated with the true clinical outcome. |
| True endpoint is independent of treatment given surrogate | Once the surrogate is accounted for, treatment should add little information about the true endpoint. |
These criteria provide a useful conceptual starting point, but modern validation usually goes further.
Why Correlation Alone Is Not Enough
Suppose a dataset shows a strong correlation between PFS and OS.
For example:
This may appear impressive.
But it does not establish that PFS is a valid surrogate.
The reason is that correlation measures association among observed outcomes, whereas surrogacy is fundamentally about whether treatment effects on one endpoint predict treatment effects on another.
| Analysis | What It Tells You |
|---|---|
| Patient-level correlation | Whether patients with favorable surrogate outcomes tend to have favorable true outcomes |
| Trial-level correlation | Whether treatment effects on the surrogate predict treatment effects on the true endpoint |
| Causal validation | Whether the surrogate captures the treatment effect on the true endpoint |
A Simple Multi-Trial Example
Suppose five randomized oncology trials evaluate similar therapies.
| Trial | PFS HR | OS HR |
|---|---|---|
| 1 | 0.70 | 0.78 |
| 2 | 0.65 | 0.72 |
| 3 | 0.82 | 0.90 |
| 4 | 0.95 | 0.98 |
| 5 | 0.55 | 0.68 |
The pattern suggests that stronger treatment effects on PFS tend to accompany stronger treatment effects on OS.
A formal analysis would typically work on the logarithmic hazard-ratio scale:
and then model:
The strength of this relationship provides evidence about trial-level surrogacy.
Why the Log Hazard-Ratio Scale Is Natural
Hazard ratios are multiplicative. A hazard ratio of 0.70 represents a 30% relative reduction in hazard, whereas 1.30 represents a 30% increase.
Taking logarithms makes the treatment effect additive:
For example:
and:
This scale is convenient for meta-analysis and regression-based surrogate validation.
Surrogacy Measures
Several statistical measures have been proposed to quantify surrogacy.
Common approaches include:
- Correlation coefficients
- Coefficient of determination \(R^2\)
- Trial-level \(R^2\)
- Individual-level \(R^2\)
- Surrogate threshold effect
- Meta-analytic validation models
- Information-theoretic and causal measures
No single number should automatically determine whether an endpoint is valid.
Trial-Level \(R^2\)
Suppose a regression of treatment effects on OS against treatment effects on PFS produces:
This means that approximately 80% of the variation in observed treatment effects on OS is explained by the treatment effects on PFS in the fitted trial-level model.
A high \(R^2\) supports surrogacy.
However, it does not mean that PFS perfectly predicts OS in every future trial.
Surrogate Threshold Effect
A particularly useful concept is the surrogate threshold effect (STE).
The STE is the minimum treatment effect on the surrogate required before the analysis predicts a statistically significant beneficial effect on the true endpoint.
For example, suppose a meta-analytic validation model indicates that the treatment must produce a PFS hazard ratio below:
before a benefit in OS can be predicted with the desired level of confidence.
The STE therefore provides a clinically interpretable way of translating surrogate performance into a treatment-effect threshold.
A Surrogate Threshold Is Not a Magic Cutoff
The STE depends on:
- The collection of validation trials
- The statistical model
- The uncertainty of treatment-effect estimates
- The target confidence level
- The heterogeneity among trials
- The relationship between surrogate and true endpoint
Therefore, an STE should be interpreted as a model-based prediction threshold, not as a universal biological law.
Meta-Analysis for Surrogate Validation
Surrogate validation is often strongest when multiple randomized trials are available.
Each trial contributes two treatment effects:
and:
The collection of paired effects can then be analyzed using a meta-analytic framework.
This is fundamentally different from pooling only individual patient-level observations.
Why Multiple Trials Matter
A single randomized trial can show that treatment improves both PFS and OS.
But this does not establish that PFS will predict OS in future trials.
Validation requires evidence that the relationship is reproducible across different treatment effects, populations, studies, and clinical settings.
Hierarchical Modeling
Treatment-effect estimates have different precision across trials. A large trial may estimate its hazard ratio much more precisely than a small trial.
A hierarchical model can account for this structure.
One conceptual model is:
where:
- \(\alpha\) is the intercept.
- \(\beta\) describes the average relationship between surrogate and true treatment effects.
- \(u_j\) represents between-trial variation.
- \(\varepsilon_j\) represents residual uncertainty.
More sophisticated models can incorporate the estimated standard errors of the trial-specific treatment effects and measurement error.
Measurement Error Matters
The estimated hazard ratio in a trial is not the true treatment effect. It is an estimate with uncertainty.
For trial \(j\):
and similarly:
Ignoring this estimation error can distort the estimated relationship between surrogate and true treatment effects.
Restricted Mean Survival Time and Surrogates
Surrogate validation does not have to rely exclusively on hazard ratios.
An alternative treatment-effect measure is the restricted mean survival time (RMST).
For a survival function \(S(t)\), the RMST through time \(\tau\) is:
A treatment effect can be expressed as:
One could therefore investigate whether treatment effects on a surrogate RMST measure predict treatment effects on the true endpoint's RMST.
Example: PFS as a Candidate Surrogate for OS
Consider a randomized oncology trial evaluating a new therapy.
The primary endpoint is PFS, while OS is a key secondary endpoint. Suppose the results are:
| Endpoint | Hazard Ratio | Interpretation |
|---|---|---|
| PFS | 0.62 | Strong reduction in progression/death hazard |
| OS | 0.91 | Little evidence of an overall survival benefit |
Can we conclude that PFS is not a valid surrogate?
Not from this trial alone.
Several explanations are possible:
- Post-progression treatments may dilute the OS effect.
- The trial may have insufficient power for OS.
- The treatment may delay progression without extending life.
- Subsequent therapies may improve survival in the control group.
- The relationship between PFS and OS may vary across treatment classes.
This illustrates why surrogate validation should be based on evidence across multiple trials whenever possible.
When PFS Can Fail as a Surrogate
There are several mechanisms through which PFS can improve without producing a corresponding OS benefit.
Effective Subsequent Therapy
If patients in the control group receive highly effective rescue therapy after progression, an experimental treatment may substantially improve PFS while having little effect on OS.
Treatment-Related Toxicity
A treatment could delay progression but cause serious toxicity that offsets some or all of the survival benefit.
Non-Proportional Hazards
The hazard ratio may not adequately summarize treatment effects when survival curves cross or hazards vary substantially over time.
Endpoint Measurement Differences
PFS can depend on imaging schedules, assessment rules, censoring conventions, and blinded independent review.
Differences in measurement can weaken the relationship between PFS and OS.
Progression-Free Survival Is a Composite Endpoint
PFS commonly measures time from randomization to either progression or death.
Conceptually:
This creates an important distinction from OS, where the event is death.
A treatment can affect progression and death differently.
Informative Censoring and Surrogate Analysis
Survival endpoints rely on censoring assumptions.
Standard Kaplan-Meier and Cox analyses generally require assumptions concerning the relationship between censoring and the event process.
If censoring differs systematically across treatment groups or depends on unobserved future outcomes, treatment-effect estimates can be biased.
Surrogate validation therefore inherits many of the challenges of ordinary survival analysis.
Competing Risks and Surrogate Endpoints
In some settings, the surrogate and true endpoint can be affected by competing events.
For example, non-cancer death may compete with cancer progression.
This can complicate interpretation of PFS and disease-specific endpoints.
The appropriate analysis depends on the estimand and the clinical question.
Surrogacy and the Estimand Framework
A surrogate analysis should begin with a clearly defined treatment effect.
For example:
- Effect of treatment on PFS
- Effect of treatment on OS
- Effect under treatment policy
- Effect in a specified population
- Effect through a specified follow-up time
If the estimand for the surrogate differs fundamentally from the estimand for the true endpoint, their relationship can become difficult to interpret.
Surrogate Validation Is Context Specific
A surrogate is rarely valid in an abstract universal sense.
Instead, validity may depend on:
- Disease
- Stage
- Line of therapy
- Treatment mechanism
- Patient population
- Background therapy
- Subsequent therapy
- Endpoint definition
- Follow-up duration
A surrogate validated in one disease setting should not automatically be assumed valid in another.
Individual-Level vs. Trial-Level Surrogacy
| Feature | Individual Level | Trial Level |
|---|---|---|
| Unit of analysis | Patient | Clinical trial |
| Main question | Do patients with favorable surrogate outcomes have favorable true outcomes? | Do treatment effects on the surrogate predict treatment effects on the true endpoint? |
| Typical measure | Correlation or association | Regression/correlation of treatment effects |
| Primary concern | Prognostic association | Predictive treatment-effect relationship |
| Evidence for substitution | Helpful | Usually more directly relevant |
A Common Ecological Fallacy
The difference between patient-level and trial-level evidence creates an important statistical warning.
A strong patient-level relationship does not imply a strong trial-level relationship.
Conversely, treatment effects can show a strong trial-level relationship even when individual-level associations are more complicated.
The two levels answer different questions and should not be substituted for one another.
Surrogate Endpoint Validation Using R
Suppose a dataset contains one row per randomized trial with:
triallog_hr_pfsse_pfslog_hr_osse_os
A basic exploratory plot can examine the relationship between treatment effects.
library(ggplot2)
ggplot(
trials,
aes(
x = log_hr_pfs,
y = log_hr_os
)
) +
geom_point(size = 3) +
geom_smooth(
method = "lm",
se = TRUE
) +
labs(
x = "Log hazard ratio for PFS",
y = "Log hazard ratio for OS",
title = "Trial-Level Surrogate Relationship"
)
This is an exploratory analysis rather than a complete surrogate-validation model.
Calculate the Trial-Level Regression
fit <- lm( log_hr_os ~ log_hr_pfs, data = trials ) summary(fit)
The coefficient of determination can be extracted with:
summary(fit)$r.squared
A high value supports a strong trial-level relationship, but uncertainty and prediction performance must also be evaluated.
Weighted Regression
Trials differ in the precision of their OS treatment-effect estimates. A simple weighted analysis can use inverse variance weighting.
fit_w <- lm( log_hr_os ~ log_hr_pfs, data = trials, weights = 1 / se_os^2 ) summary(fit_w)
This is useful as an exploratory model, but formal surrogate validation may require a hierarchical model that accounts for uncertainty in both treatment effects.
Predicting the OS Effect
Suppose the fitted model is:
A new trial with estimated PFS effect \(\theta_S\) would have predicted OS effect:
The important point is that the prediction should include an uncertainty interval.
Prediction Intervals Matter
Suppose the predicted OS hazard ratio is:
That alone does not establish reliable prediction.
If the prediction interval is very wide, the model may be unable to exclude an OS hazard ratio close to 1.0 or even above 1.0.
Validation vs. Qualification
The terms validation and qualification can have different meanings depending on the regulatory and methodological context.
In general, surrogate evaluation should establish a body of evidence showing that the endpoint reliably predicts the clinically meaningful outcome for the intended use.
A surrogate that is useful for decision-making in one development context may not automatically be suitable as the primary basis for another regulatory claim.
Surrogate Endpoints in Randomized Trials
Randomization is particularly important because surrogate validation is concerned with treatment effects.
Without randomization, an observed association between the surrogate and true endpoint can be confounded.
For example, patients with better baseline health may both:
- Remain progression-free longer
- Survive longer
This association does not establish that changing PFS through treatment will necessarily change OS.
What Happens When the Surrogate Is Imperfect?
An imperfect surrogate can still be useful.
The relevant question is often not whether the surrogate is mathematically perfect, but whether it provides sufficiently reliable information for the intended decision.
For example, a surrogate might:
- Predict the direction of the true treatment effect well.
- Predict the magnitude moderately well.
- Provide useful information with shorter follow-up.
The acceptable degree of uncertainty depends on the clinical and regulatory context.
Surrogacy and Treatment Mechanism
A treatment's mechanism can affect whether a surrogate is credible.
If an intervention primarily prevents tumor progression without affecting mechanisms associated with mortality, PFS may be less informative about OS.
Conversely, if progression is tightly linked to subsequent mortality and the treatment acts through that pathway, PFS may have stronger surrogate properties.
Mechanistic plausibility therefore complements statistical validation.
Why Surrogacy Can Change Over Time
The relationship between PFS and OS can change as treatment landscapes evolve.
For example, improvements in:
- Salvage therapy
- Immunotherapy
- Targeted therapy
- Supportive care
- Sequencing strategies
can alter the relationship between progression and death.
Consequently, surrogate validation should be periodically reassessed when the clinical environment changes substantially.
Common Mistake: "PFS Was Significant, So OS Must Be Better"
A statistically significant PFS result does not logically imply a statistically significant OS result.
For example:
can coexist with:
The two endpoints represent different event processes.
Common Mistake: "PFS and OS Are Correlated"
Even if longer PFS is strongly associated with longer OS, this is not sufficient to establish surrogacy.
The critical question concerns the relationship between randomized treatment effects.
Common Mistake: Ignoring Subsequent Therapy
Subsequent therapy can substantially alter OS.
If control patients receive effective post-progression therapy, the OS difference between randomized groups may become much smaller than the PFS difference.
This does not necessarily mean the treatment failed biologically. It means that PFS and OS may be measuring different aspects of the treatment strategy.
Common Mistake: Using Too Few Trials
A surrogate validation analysis based on only a handful of trials may produce an apparently strong correlation that is unstable.
With few trials:
- Regression coefficients can be unstable.
- \(R^2\) can be highly variable.
- Prediction intervals can be wide.
- One influential trial can dominate the relationship.
External validation and sensitivity analyses are therefore important.
Common Mistake: Mixing Different Treatment Classes
Suppose a validation dataset combines:
- Cytotoxic chemotherapy
- Targeted therapy
- Immunotherapy
- Cell therapy
The PFS-to-OS relationship may differ substantially among these mechanisms.
Pooling them indiscriminately can obscure important heterogeneity.
Heterogeneity Across Trials
Trial-level surrogate validation should consider heterogeneity in:
- Patient population
- Control treatment
- Experimental mechanism
- Line of therapy
- Follow-up duration
- Subsequent therapy
- Endpoint definitions
A high average correlation may conceal clinically important differences among subgroups of trials.
Sensitivity Analyses
A robust validation program should consider sensitivity analyses such as:
- Removing influential trials
- Restricting to randomized Phase III trials
- Restricting to a specific disease setting
- Restricting to a treatment class
- Using alternative treatment-effect measures
- Using alternative censoring assumptions
- Comparing fixed-effect and random-effects approaches
The purpose is to determine whether the surrogate relationship is robust rather than driven by a particular modeling decision.
Surrogate Endpoints and Multiplicity
Trials often evaluate PFS and OS as multiple endpoints.
The statistical interpretation depends on the prespecified testing strategy.
For example, a protocol may specify a hierarchical sequence in which PFS must demonstrate benefit before formal testing of OS.
Alternatively, the endpoints may be analyzed using a multiplicity adjustment.
Surrogate validation does not eliminate the need to address multiplicity.
Surrogate Endpoints and Missing Data
Missing data can affect surrogate evaluation at both the patient and trial levels.
Examples include:
- Missing tumor assessments
- Incomplete progression assessments
- Loss to follow-up
- Incomplete survival follow-up
The statistical analysis should follow the prespecified estimand and missing data strategy.
A Practical Validation Workflow
A Compact Mathematical Summary
Let \(S\) be the surrogate and \(T\) the true endpoint. At the individual level, investigate:
At the trial level, investigate:
where:
A strong surrogate should demonstrate both meaningful patient-level association and, especially for treatment substitution, a strong and reproducible trial-level relationship.
Example Interpretation
Suppose a validation program finds:
| Measure | Result |
|---|---|
| Individual-level association | Strong |
| Trial-level \(R^2\) | 0.82 |
| Prediction interval | Moderately narrow |
| Surrogate threshold effect | Clinically attainable |
| Sensitivity analyses | Relationship remains consistent |
This would provide substantially stronger evidence for surrogacy than a single trial showing a significant association between PFS and OS.
But What If the Trial-Level \(R^2\) Is Low?
Suppose instead:
This means that only a relatively small fraction of variation in trial-level OS treatment effects is explained by PFS treatment effects in the fitted model.
That would provide weak evidence for using PFS as a substitute for OS.
The appropriate conclusion would not necessarily be that PFS is useless. Instead, it may still be valuable as:
- An early efficacy endpoint
- A secondary endpoint
- A measure of disease control
- A component of the overall evidence package
But it would be weaker evidence for replacing OS as the principal measure of clinical benefit.
Surrogate Endpoint vs. Co-Primary Endpoint
A surrogate can also be used alongside the true endpoint rather than replacing it.
For example, a study might use:
- PFS as an earlier endpoint
- OS as a key confirmatory endpoint
This can preserve the clinical importance of OS while allowing earlier assessment of treatment activity.
Surrogate Endpoint vs. Intermediate Endpoint
The terms "intermediate endpoint" and "surrogate endpoint" are sometimes used interchangeably, but they are not conceptually identical.
An intermediate endpoint is simply an outcome occurring between treatment and the ultimate clinical outcome.
A surrogate endpoint carries an additional claim: the intermediate endpoint is sufficiently validated to stand in for the clinical endpoint for a specified purpose.
Why Surrogate Validation Is Hard
The difficulty is fundamentally causal.
Treatment can influence several pathways simultaneously:
If the surrogate captures only one of these pathways, it may fail to represent the total treatment effect on survival.
The Most Important Clinical Example
The classic oncology example is:
PFS is attractive because it can often be observed substantially earlier than OS.
But the validity of PFS as an OS surrogate depends on the disease and treatment context.
A PFS improvement should therefore not automatically be translated into an equivalent OS benefit.
Practical Checklist
Before treating a survival endpoint as a surrogate, ask:
- Is there a clear biological rationale?
- Does treatment consistently affect the surrogate?
- Does the surrogate predict the true endpoint at the patient level?
- Do treatment effects on the surrogate predict treatment effects on the true endpoint?
- How many randomized trials support the relationship?
- How large is the trial-level \(R^2\)?
- How wide are prediction intervals?
- Is the relationship consistent across treatment classes?
- Is the relationship robust to influential-trial analyses?
- Could subsequent therapy alter the surrogate-to-true-endpoint relationship?
- Could treatment toxicity break the relationship?
- Is the surrogate being used in the same clinical setting in which it was validated?
- Does the surrogate support the specific regulatory or clinical decision being made?
Common Statistical Pitfalls
- Confusing association with surrogacy. A surrogate can be strongly prognostic without being a valid surrogate.
- Using only patient-level correlation. Treatment-effect relationships across randomized trials are critical.
- Ignoring uncertainty in treatment effects. Trial-specific hazard ratios are estimates, not known constants.
- Using too few validation trials. Small numbers of trials can produce unstable relationships.
- Ignoring heterogeneity. Surrogacy can vary by disease, treatment class, and clinical setting.
- Ignoring subsequent therapy. Post-progression treatment can materially alter OS.
- Assuming a significant PFS result guarantees OS benefit. The endpoints measure different event processes.
- Using a high \(R^2\) without examining prediction error. A strong fitted relationship does not guarantee accurate prediction in a new trial.
- Assuming a surrogate is universally valid. Validation is context dependent.
What Should Be Reported in a Surrogate Validation Analysis?
A transparent analysis should report:
- Definition of the surrogate endpoint
- Definition of the true endpoint
- Clinical setting
- Treatment classes represented
- Number of randomized trials
- Trial eligibility criteria
- Patient-level association measures
- Trial-level treatment effects
- Statistical model
- Estimated trial-level association
- \(R^2\) or related surrogacy measures
- Prediction intervals
- Surrogate threshold effect when appropriate
- Between-trial heterogeneity
- Sensitivity analyses
- Limitations of extrapolation
What a Strong Conclusion Looks Like
A strong surrogate analysis should avoid statements such as:
"PFS is highly correlated with OS, therefore PFS is a validated surrogate."
A more defensible conclusion describes the evidence and its scope.
For example:
The Key Difference Between Prediction and Substitution
A surrogate can be useful for predicting the eventual true endpoint without being sufficiently reliable to replace it.
These are different claims.
| Claim | Strength |
|---|---|
| The surrogate is prognostic | Weakest |
| The surrogate predicts the true endpoint | Stronger |
| Treatment effects on surrogate predict treatment effects on true endpoint | Much stronger |
| Surrogate can replace the true endpoint for a defined decision | Strongest |
Bottom Line
Key Takeaways
- A surrogate endpoint is an intermediate outcome intended to represent treatment effects on a clinically meaningful endpoint.
- PFS is a common candidate surrogate for OS in oncology.
- Prognostic association is not the same as surrogate validity.
- Individual-level and trial-level surrogacy answer different questions.
- Trial-level surrogacy examines whether treatment effects on the surrogate predict treatment effects on the true endpoint.
- Meta-analytic validation across randomized trials is particularly important.
- \(R^2\) measures can summarize the strength of a surrogate relationship but should not be interpreted without prediction uncertainty.
- The surrogate threshold effect can provide a clinically interpretable treatment-effect threshold.
- Subsequent therapy, treatment toxicity, competing risks, and changing standards of care can weaken surrogate relationships.
- Surrogate validity is generally context specific.
References
Prentice, R.L. (1989). Surrogate endpoints in clinical trials: definition and operational
criteria. Statistics in Medicine, 8, 431–440.
Buyse, M., Molenberghs, G., Burzykowski, T., Renard, D. & Geys, H. (2000). The validation of surrogate endpoints in meta-analyses of randomized
experiments. Biostatistics, 1(1), 49–67.
Burzykowski, T., Molenberghs, G. & Buyse, M. (eds.) (2005). The Evaluation of Surrogate Endpoints.
Springer.
Daniels, M.J. & Hughes, M.D. (1997). Meta-analysis for the evaluation of potential surrogate markers. Statistics in Medicine, 16, 1965–1982.
Buyse, M., Burzykowski, T., Carroll, K., Geys, H., Michiels, S., Sargent, D.J.,
Miller, L.L., El-Helw, L., Saad, E.D., Sweeney, C., Burghardt, M.C.,
Quinaux, E., Bogaerts, J., Thirion, P. & Piedbois, P. (2007). Progression-free survival is a surrogate for survival in advanced
colorectal cancer. Journal of Clinical Oncology, 25(33), 5218–5224.
Fleming, T.R. & DeMets, D.L. (1996). Surrogate end points in clinical trials: are we being misled? Annals of Internal Medicine, 125, 605–613.
Ciani, O., Buyse, M., Garside, R., Peters, J., Taylor, R.S., Sargent, D.J. &
Tappenden, P. (2017). Comparison of treatment effect sizes associated with surrogate and
final patient relevant outcomes in randomised controlled trials. BMJ, 356, j374.