Tutorials › Pharmacometrics › External Validation of Population PK Models
Pharmacokinetics · Population PK

External Validation of Population PK Models

Learn how to evaluate whether a population pharmacokinetic model predicts drug concentrations adequately in an independent dataset—and how prediction-based diagnostics can reveal bias, imprecision, and limitations that may not be apparent during model development.

Intermediate Population PK Model Evaluation Model Validation
01 · The big picture

1. What Is External Validation?

External validation evaluates a population PK model using data that were not used to develop the model. The purpose is to determine whether the model can adequately predict observations in a new population, study, dosing regimen, clinical setting, or other dataset that is meaningfully independent of the development data.

This distinction is important because a model can describe its development dataset well without necessarily generalizing to new observations. External validation therefore focuses on predictive performance beyond the data used for model construction.

Development population PK data PK model fixed effects variability covariates Validation independent data Observed concentrations are compared with model-based predictions

External validation asks whether a previously developed population PK model predicts an independent dataset adequately.

Core idea: external validation is not simply another goodness-of-fit exercise. The key question is whether a model developed elsewhere can make useful predictions for new observations without being refitted to those observations.
02 · Why validate?

2. Why Is External Validation Important?

Population PK models are often intended for use beyond the specific subjects and study conditions from which they were developed. External validation provides evidence about whether the model's predictive behavior transfers to another dataset.

QuestionWhat external validation can assess
Does the model predict concentrations without systematic bias?Agreement between observed concentrations and model-based predictions
Is predictive variability reasonable?Whether the observed distribution is consistent with model expectations
Does the model generalize?Performance in a new population or study setting
Are important covariate effects transferable?Whether predictions remain appropriate when patient characteristics differ
Can the model support prospective use?Whether predictions are sufficiently reliable for the intended application

External validation is especially informative when the validation population differs from the development population in characteristics such as age, body size, renal function, disease state, concomitant medications, sampling design, or dosing regimen.

03 · Independence

3. What Makes a Validation Dataset Independent?

The strongest external validation uses observations that were not involved in model development. The exact meaning of independence depends on how the model was developed and how the validation data were selected.

Independent subjects

A common approach is to evaluate the model in a completely separate group of individuals. This avoids evaluating the model on subjects whose data contributed to parameter estimation.

Independent study

A stronger test of generalizability may involve data from another clinical study. Differences in study design, sampling schedule, dosing, and patient characteristics can reveal limitations that are difficult to detect internally.

New clinical setting

External validation can also evaluate whether the model transfers to a different clinical setting, such as a new treatment population or a therapeutic drug-monitoring environment.

Important distinction: a random subset of the development dataset is useful for some forms of internal validation, but it is not equivalent to demonstrating performance in a genuinely external dataset.
04 · Prediction types

4. What Predictions Are Used in External Validation?

Population PK models can generate different types of predictions. The choice of prediction should match the question being asked and the information available at the time of prediction.

PredictionDefinitionTypical use
Population prediction, \(PRED\)Prediction based on typical population parameters and available covariates, without using an individual's observed concentrations to estimate their random effectsAssess population-level predictive performance
Individual prediction, \(IPRED\)Prediction incorporating individual-specific information, typically through empirical Bayes estimatesAssess how well the model describes observations after individual information is incorporated
Simulation-based predictionConcentrations simulated repeatedly from the complete model, including appropriate variabilityAssess whether the observed data resemble the distribution predicted by the model

For a genuinely prospective validation question, population predictions are particularly important because they assess what the model predicts before the individual's validation observations are used to update their random effects.

05 · Observed versus predicted

5. Observed Versus Predicted Concentrations

A basic external-validation diagnostic compares observed concentrations with model predictions.

\[ \text{Observed concentration} \quad \text{vs.} \quad \text{Predicted concentration} \]

If the model is well calibrated, points should generally lie around the identity line:

\[ y=x \]

Systematic departures from this relationship can indicate prediction bias. For example, the model might consistently underpredict concentrations at high values or overpredict concentrations at low values.

Population prediction Observed concentration Identity line

Observed-versus-predicted plots can reveal systematic bias, changes in variability, and regions where model predictions depart from observations.

However, visual agreement alone is not sufficient. A model may appear acceptable on a scatterplot while still showing clinically meaningful bias in a particular concentration range or patient subgroup.

06 · Prediction error

6. Quantifying Prediction Error

Prediction errors can be calculated by comparing an observed concentration \(C_i\) with its corresponding prediction \(\widehat{C}_i\).

\[ PE_i=C_i-\widehat{C}_i \]

A relative form expresses the discrepancy relative to the prediction:

\[ RPE_i=\frac{C_i-\widehat{C}_i}{\widehat{C}_i}\times100\% \]

Summary measures can then describe the magnitude and direction of prediction error.

MeasureInterpretation
Mean prediction errorIndicates the average direction of prediction bias
Mean absolute prediction errorDescribes average absolute prediction discrepancy
Median prediction errorProvides a robust summary when prediction errors are skewed
Median absolute prediction errorSummarizes typical absolute error while reducing sensitivity to extreme observations

No single summary statistic completely characterizes predictive performance. Error summaries should therefore be interpreted alongside graphical diagnostics and the scientific purpose of the model.

07 · Bias

7. Detecting Systematic Bias

One of the most important goals of external validation is identifying systematic prediction bias.

Suppose a model systematically predicts concentrations below the observations. The resulting prediction errors will tend to be positive:

\[ C_i-\widehat{C}_i>0 \]

Conversely, systematic overprediction produces predominantly negative errors.

Bias can also depend on the magnitude of the prediction. For example, a model might perform well at low concentrations but increasingly underpredict high concentrations. This type of pattern can indicate an inadequately specified structural model, residual error model, covariate relationship, or another aspect of the model.

Look for patterns, not just averages: an average prediction error close to zero does not guarantee good predictions. Positive and negative errors can cancel each other while substantial systematic errors remain in particular regions of the data.
08 · Time-dependent diagnostics

8. Prediction Errors Versus Time

Population PK predictions should also be examined as a function of time after dose.

A model may reproduce overall concentration levels while missing a specific part of the concentration-time profile. For example, systematic deviations might occur during:

  • the absorption phase;
  • the distribution phase;
  • the terminal elimination phase; or
  • a period surrounding a dosing event.

Plotting prediction errors against time can therefore provide information that is not visible in an overall observed-versus-predicted plot.

Practical question: does the model miss the same part of the concentration-time profile repeatedly? If so, the issue may concern the structural model rather than random prediction noise.
09 · Simulation-based validation

9. Visual Predictive Checks in External Validation

A visual predictive check (VPC) compares observed concentrations with the distribution of concentrations simulated from the model.

The model is used to generate many simulated datasets under the validation study design. Prediction intervals or percentile bands are then constructed from the simulated concentrations and compared with the observed data.

Time after dose Concentration Simulation interval Observed median

Conceptual VPC: observed concentrations are compared with prediction intervals and summary percentiles generated from model simulations.

A VPC can reveal whether the model reproduces the central tendency and variability of the validation data over time. If observed concentrations systematically fall outside the expected simulation envelope, the model may not adequately describe the validation population or study conditions.

10 · Normalized prediction

10. Normalized Prediction Distribution Errors

Simulation-based diagnostics can also be transformed into normalized prediction distribution errors (NPDE). The underlying idea is to determine where each observed concentration falls within the distribution of concentrations predicted by the model simulations.

If the model is adequate and the simulation procedure is correctly specified, the resulting NPDE distribution should be broadly consistent with the theoretical distribution expected under the model.

NPDE can therefore provide a complementary assessment to conventional VPCs and prediction-error plots, particularly when the goal is to evaluate the full predictive distribution rather than only its median or selected percentiles.

Interpretation: NPDE is a model-based diagnostic. It should not be treated as a standalone pass/fail test. Its meaning depends on the simulation design, model assumptions, validation data, and the scientific question.
11 · Generalizability

11. Evaluating Covariate Generalizability

A population PK model may perform differently in a validation population because the distribution of important covariates differs from that in the development population.

CovariatePotential validation question
Body weightDoes the model appropriately predict concentrations across the validation weight range?
Renal functionDoes the clearance relationship remain adequate in subjects with different renal function?
AgeDoes the model extrapolate appropriately to older or younger subjects?
Disease statusDoes the model transfer to a population with different disease characteristics?
Concomitant medicationDoes the model account adequately for relevant drug interactions?

A model should not automatically be considered invalid simply because the validation population differs from the development population. In fact, meaningful differences can make external validation particularly informative. The important question is whether the model's predictions remain adequate across those differences.

12 · Individual predictions

12. A Caution About Individual Predictions and Shrinkage

Individual predictions often use empirical Bayes estimates of individual random effects. These estimates depend on the observed concentrations for the individual being evaluated.

Consequently, \(IPRED\) can provide useful information about how the model describes observations after individual information has been incorporated, but it answers a different question from a prospective population prediction.

This distinction is especially important when evaluating external predictive performance:

DiagnosticQuestion
PREDHow well does the model predict the validation data from population parameters and covariates?
IPREDHow well does the model describe observations after individual-specific information has been incorporated?

When individual random effects show substantial shrinkage, the apparent agreement between IPRED and observations may be partly driven by the model pulling individual estimates toward the population mean. Therefore, excellent IPRED diagnostics do not by themselves establish strong prospective predictive performance.

13 · Validation workflow

13. A Practical External Validation Workflow

  1. Freeze the model. Clearly define the model structure, parameter estimates, covariate relationships, variability models, and implementation to be evaluated.
  2. Define the validation dataset. Confirm that the data were not used in model development and characterize how the validation population differs from the development population.
  3. Reproduce the model implementation. Verify units, dosing records, infusion information, covariates, event records, and parameter transformations.
  4. Generate predictions. Produce population predictions using the model without refitting the model to the validation data.
  5. Compare observations and predictions. Examine observed-versus-predicted concentrations and prediction errors.
  6. Evaluate time-dependent behavior. Examine prediction errors and observed-versus-predicted relationships across time and relevant concentration ranges.
  7. Perform simulation-based diagnostics. Where appropriate, simulate the validation study design and construct VPCs or related diagnostics.
  8. Assess important subgroups. Examine performance across clinically relevant covariate ranges and patient subgroups.
  9. Evaluate clinical relevance. Determine whether any observed prediction errors matter for the intended application of the model.
  10. Document limitations. Describe differences between development and validation populations, sampling limitations, extrapolation, and uncertainty in the validation assessment.
14 · Worked example

14. Worked Example: External Validation of a Population PK Model

Suppose a population PK model was developed using 120 subjects with an oral drug. The final model describes clearance using body weight and renal function.

A separate study containing 40 new subjects is available for external validation. The validation study includes concentration samples collected after the same dose, but the subjects are somewhat heavier and have a wider range of renal function than the development population.

Step 1: Apply the frozen model

The previously estimated population parameters and covariate relationships are applied to the validation subjects. No model parameters are re-estimated using the validation concentrations.

Step 2: Generate population predictions

For each observed concentration \(C_i\), obtain the corresponding population prediction \(\widehat{C}_i\).

Suppose one validation observation is:

\[ C_i=8.0\text{ mg/L} \]

and the model predicts:

\[ \widehat{C}_i=7.2\text{ mg/L} \]

Step 3: Calculate prediction error

\[ PE_i=8.0-7.2=0.8\text{ mg/L} \]

Step 4: Calculate relative prediction error

\[ RPE_i=\frac{8.0-7.2}{7.2}\times100\%\approx11.1\% \]

This individual observation is therefore approximately 11% higher than its population prediction.

Step 5: Examine the complete validation dataset

Suppose the observed-versus-predicted plot shows reasonable agreement overall, but concentrations above 10 mg/L are consistently underpredicted. A VPC also shows that the upper portion of the observed concentration distribution lies above the model's simulated prediction interval.

The important conclusion is not simply that the model has a particular average prediction error. The pattern suggests that predictive performance may deteriorate at higher concentrations and should be investigated in relation to model structure, covariates, variability, dose, and the intended use of the model.

Worked-example lesson: external validation is strongest when multiple complementary diagnostics tell a consistent story. A single numerical summary should not replace inspection of prediction bias, variability, time dependence, and clinically relevant subgroups.
15 · Interpretation

15. How Should External Validation Results Be Interpreted?

External validation does not generally produce a universal numerical threshold separating a "valid" model from an "invalid" model. The interpretation depends on the purpose of the model and the consequences of prediction error.

For example, a model intended to describe population-level exposure in a simulation exercise may tolerate different levels of prediction error than a model intended to support individualized dose selection.

FindingPossible interpretation
Predictions show little systematic biasSupports adequate calibration in the validation setting
Errors increase with concentrationMay indicate model misspecification or inadequate handling of concentration-dependent behavior
Errors vary systematically over timeMay indicate structural or absorption/disposition misspecification
One subgroup performs poorlyMay indicate inadequate covariate relationships or limited extrapolation
Observed variability exceeds simulated variabilityMay indicate underestimation of between-subject or residual variability
Validation population lies outside development rangePredictions may involve extrapolation and require additional caution

The appropriate interpretation should connect the statistical diagnostics to the scientific purpose of the model rather than relying on a single pass/fail criterion.

16 · Common mistakes

16. Common External Validation Mistakes

  • Refitting the model before evaluating predictions. This removes much of the value of external validation because the validation data are being used to adapt the model.
  • Calling an internal split an external validation. A randomly held-out subset can be useful for internal validation, but it does not necessarily test transfer to an independent study or population.
  • Using IPRED as the only evidence of predictive performance. Individual predictions incorporate information from the observed concentrations themselves.
  • Looking only at goodness-of-fit plots. A visually acceptable scatterplot does not establish that the model reproduces the full predictive distribution.
  • Ignoring the validation population. Differences in covariate distributions, dosing, disease, or sampling can explain differences in predictive performance.
  • Using one summary statistic as a pass/fail criterion. Prediction error, VPCs, time trends, subgroup behavior, and clinical relevance provide complementary information.
  • Ignoring implementation errors. Differences in units, dosing times, infusion duration, covariate coding, or model code can create apparent validation failures.
  • Confusing lack of evidence with evidence of adequate performance. A small validation dataset may simply have limited ability to reveal important discrepancies.
Validation principle: before concluding that a model does not generalize, verify that the validation dataset, dosing records, covariates, units, and model implementation are all consistent with the original model specification.
17 · Internal versus external

17. External Validation Versus Internal Validation

External validation is one part of a broader model evaluation strategy. Internal validation methods can identify overfitting and assess model stability, whereas external validation evaluates performance in independent data.

ApproachPrimary question
Goodness-of-fit diagnosticsDoes the model describe the development observations reasonably?
Bootstrap validationHow stable are model estimates and performance under resampling of the development data?
Cross-validationHow does predictive performance behave when observations are repeatedly held out from model development?
External validationDoes the model predict an independent dataset adequately?
Simulation-based diagnosticsCan the model reproduce the observed distribution and variability under the validation design?

These approaches answer related but distinct questions. Strong model evaluation typically combines several forms of evidence rather than treating any one diagnostic as definitive.

18 · Prospective use

18. External Validation and Prospective Prediction

The strongest practical motivation for external validation is often prospective model use. A population PK model may ultimately be used to predict exposure for patients who were not part of the original model development dataset.

In that setting, the validation exercise should resemble the intended application as closely as practical.

For example, if the model will be used to predict concentrations using dose and patient covariates before therapeutic drug monitoring results are available, the validation should emphasize predictions generated without using those future observations.

If the intended application involves Bayesian forecasting after one or more concentrations are collected, the validation design should instead reflect that workflow and evaluate the model's performance under the corresponding information structure.

Key distinction: the appropriate validation design depends on the prediction task. A model can have different predictive performance before and after individual concentration data are incorporated.
19 · Clinical relevance

19. Statistical Validation Is Not the Same as Clinical Utility

A model can show statistically reasonable predictive performance while still requiring additional evaluation before it is used for clinical decisions.

Clinical utility depends on how prediction errors affect the decision being made. For example, an error that is negligible for estimating population exposure may be important if the same prediction is used to select an individualized dose near a narrow therapeutic window.

Therefore, external validation should ultimately be connected to the model's intended purpose:

  • population exposure characterization;
  • dose optimization;
  • therapeutic drug monitoring;
  • clinical trial simulation;
  • label or regulatory support;
  • precision dosing; or
  • other pharmacometric applications.

The more consequential the intended decision, the more important it is to evaluate prediction performance in the context of that decision rather than relying solely on generic statistical diagnostics.

20 · Practical checklist

20. External Validation Checklist

StepQuestion
1. Model definitionIs the model fully specified and frozen before validation?
2. Data independenceWere the validation observations excluded from model development?
3. Population comparisonHow does the validation population differ from the development population?
4. ImplementationAre doses, times, units, covariates, and model code implemented correctly?
5. Population predictionsAre predictions generated without using validation concentrations to estimate individual effects?
6. Prediction errorsIs systematic bias apparent in the validation observations?
7. Time trendsDoes prediction performance change across the concentration-time profile?
8. Simulation diagnosticsDoes the model reproduce the observed distribution and variability?
9. SubgroupsDoes predictive performance remain reasonable across clinically important covariates?
10. Intended useAre observed discrepancies important for the model's actual application?

21. Key Takeaways

  • External validation evaluates a population PK model using data that were not used to develop the model.
  • The central question is whether the model generalizes to new observations, subjects, studies, or clinical settings.
  • Population predictions are particularly important when evaluating prospective predictive performance.
  • Observed-versus-predicted plots can reveal systematic bias but should not be used as the only diagnostic.
  • Prediction-error summaries quantify the direction and magnitude of discrepancies between observations and predictions.
  • Prediction errors should also be examined against time, concentration, and clinically important covariates.
  • VPCs and related simulation-based diagnostics assess whether the validation observations are consistent with the distribution predicted by the model.
  • IPRED and PRED answer different questions because individual predictions incorporate information from observed concentrations.
  • External validation should account for differences between development and validation populations, including covariate distributions and study design.
  • A validation dataset should generally be evaluated using a frozen model rather than refitting the model to the validation observations.
  • There is no universal numerical threshold that determines whether every population PK model has been externally validated successfully.
  • The appropriate standard depends on the scientific question, prediction task, and intended clinical or pharmacometric use of the model.
Next step

Where to Go Next

A natural progression is to study bootstrap validation of population PK models, followed by visual predictive checks, prediction-corrected visual predictive checks, normalized prediction distribution errors, and the interpretation of eta and epsilon shrinkage.

These methods complement external validation by providing different perspectives on model stability, predictive distribution, residual behavior, and the ability of a population PK model to reproduce observed data.

← Back to Pharmacokinetics Tutorials