1. What Are Goodness-of-Fit Diagnostics?
Goodness-of-fit (GOF) diagnostics are graphical and numerical tools used to evaluate how well a population pharmacokinetic model describes the observed concentration data.
In population PK, the objective is not simply to obtain parameter estimates. A model should provide an adequate representation of the observed data while appropriately describing between-subject variability and residual unexplained variability.
GOF diagnostics help identify situations in which the model systematically misses observations. They can reveal problems with the structural model, covariate relationships, variability models, residual error model, or individual subject predictions.
2. Observed Versus Predicted Concentrations
The most familiar GOF diagnostic compares observed concentrations with model-predicted concentrations. Let \(y_{ij}\) denote the observed concentration for subject \(i\) at time \(j\), and let \(\hat{y}_{ij}\) denote the corresponding prediction.
If the model adequately describes the data, observations should generally be distributed around the identity line:
Observed-versus-predicted plots ask whether the model predictions track the observed concentrations without obvious systematic departures.
Important patterns include curvature, systematic deviations at low or high concentrations, increasing spread, and clusters of observations that deviate from the overall relationship.
3. Population Predictions Versus Individual Predictions
Population PK models commonly distinguish between a population prediction and an individual prediction.
The population prediction, often denoted PRED, represents the expected concentration based on typical population parameters and covariates. The individual prediction, often denoted IPRED, incorporates the estimated individual random effects.
| Prediction | Common notation | What it represents |
|---|---|---|
| Population prediction | PRED | Prediction based on typical population parameters and observed covariates |
| Individual prediction | IPRED | Prediction incorporating the subject-specific estimated random effects |
| Observation | DV | Measured dependent variable, such as drug concentration |
Comparing DV with PRED primarily evaluates the model's ability to describe the population-level concentration data. Comparing DV with IPRED asks a somewhat different question: after accounting for estimated individual variability, does the model reproduce each subject's observations reasonably well?
4. What Are Residuals?
A residual measures the discrepancy between an observed concentration and a model prediction. A simple residual can be written as:
In population PK, several residual definitions are commonly examined because the scale and variance of concentration measurements may change across the prediction range.
4.1 Raw residuals
Raw residuals are simply observed minus predicted values:
They retain the original concentration units and can be useful for identifying large discrepancies, but their variance may depend strongly on concentration.
4.2 Weighted residuals
Weighted residuals attempt to account for the expected observation variance. A generic representation is:
The precise implementation depends on the residual error model and estimation framework.
For a reasonable model, residuals should generally fluctuate around zero without obvious systematic trends.
5. Residuals Versus Predictions
Plotting residuals against predictions is one of the most useful ways to investigate the residual error model.
A desirable residual plot has residuals scattered around zero without strong curvature, trends, or systematic changes in spread.
Potential warning patterns include:
- Curvature: may indicate structural model misspecification.
- Funnel-shaped spread: may indicate an inappropriate residual error model.
- Systematic positive or negative regions: may indicate bias in prediction.
- Clusters: may reflect subjects, occasions, dose groups, or other unmodeled structure.
6. Residuals Versus Time
Residuals can also be plotted against time after dose. This is particularly important in PK because different parts of the concentration-time profile contain information about different processes.
For example, systematic residual patterns during the early post-dose period may suggest problems with absorption or distribution, whereas patterns during the terminal phase may indicate an issue with elimination or the structural model.
| Pattern | Possible interpretation |
|---|---|
| Early positive residuals followed by negative residuals | Potential mismatch in the predicted absorption or distribution profile |
| Persistent terminal-phase residual trend | Potential misspecification of elimination or disposition |
| Increasing residual variance with time | May indicate an inappropriate error structure or changing measurement variability |
| Occasion-specific pattern | May suggest interoccasion variability or an unmodeled occasion effect |
These interpretations are hypotheses rather than automatic diagnoses. A residual pattern should be evaluated alongside the structural model, sampling design, parameter estimates, and other diagnostics.
7. Why Should Residuals Be Centered Around Zero?
If predictions are systematically too low, residuals will tend to be positive. If predictions are systematically too high, residuals will tend to be negative.
The important concept is not that every residual should be close to zero. Random variability naturally produces positive and negative residuals. The diagnostic question is whether there is a systematic pattern in their location or spread.
8. Goodness of Fit and the Residual Error Model
The residual error model describes the difference between the model's expected concentration and the observed concentration after accounting for the modeled PK processes and variability.
Common residual error structures include additive, proportional, and combined models.
| Error model | Conceptual form | Typical implication |
|---|---|---|
| Additive | \(DV=PRED+\epsilon\) | Constant absolute error variance |
| Proportional | \(DV=PRED(1+\epsilon)\) | Error scales with prediction magnitude |
| Combined | \(DV=PRED(1+\epsilon_1)+\epsilon_2\) | Allows both proportional and additive components |
If residual spread changes substantially with the magnitude of the prediction, the residual diagnostics may suggest that the selected error model does not adequately represent the observation process.
9. Goodness of Fit and Between-Subject Variability
Population PK models commonly represent individual parameters as deviations from typical population values. For example, clearance may be modeled as:
where \(CL_i\) is the individual's clearance, \(CL_{pop}\) is the typical population clearance, and \(\eta_{CL,i}\) represents the individual's random effect.
Diagnostics involving estimated random effects can help identify whether important relationships remain unexplained.
For example, plots of ETA versus covariates can be informative during covariate model development. A strong systematic relationship between an ETA and body weight, renal function, age, or another clinically meaningful covariate may suggest that the covariate relationship has not been adequately represented.
10. Why Shrinkage Matters for Diagnostic Interpretation
Individual random effects are estimated from the available data for each subject. When an individual's data contain limited information, the estimated random effect may be pulled toward the population mean of zero. This behavior is commonly referred to as shrinkage.
Shrinkage is important because some diagnostics involving individual ETAs or individual predictions can become difficult to interpret when shrinkage is substantial.
For example, an ETA-versus-covariate plot may appear less informative when individual ETAs have substantial shrinkage. Similarly, strong agreement between DV and IPRED should not automatically be interpreted as evidence of a highly informative model if individual predictions are strongly influenced by the observations being evaluated.
Consequently, GOF diagnostics should be interpreted together with shrinkage, sampling density, and the amount of information available for individual parameter estimation.
11. A Practical Set of GOF Plots
A common population PK diagnostic workflow examines several plots rather than relying on one diagnostic.
| Diagnostic | Main question |
|---|---|
| DV versus PRED | Does the model reproduce the population-level observations? |
| DV versus IPRED | How well do individual predictions reproduce observations? |
| RES versus PRED | Are residuals unbiased across the prediction range? |
| WRES versus PRED | Are standardized residuals appropriately distributed across predictions? |
| RES or WRES versus TIME | Are there systematic temporal patterns? |
| ETA versus covariates | Could important systematic relationships remain unexplained? |
| Distribution of residuals | Are residuals broadly consistent with the assumed observation model? |
The precise diagnostic set should depend on the model, study design, sampling schedule, endpoint, and scientific objective.
12. What Patterns Are We Looking For?
There is no universal plot that proves a population PK model is correct. Instead, investigators generally look for the absence of important systematic patterns.
| Diagnostic feature | Generally desirable pattern | Potential concern |
|---|---|---|
| DV vs PRED | Observations distributed around the identity line | Curvature or systematic departures |
| DV vs IPRED | Close correspondence without obvious systematic bias | Persistent subject-level discrepancies |
| WRES vs PRED | Approximately centered around zero with no major trend | Trend, curvature, or changing spread |
| WRES vs TIME | No substantial temporal structure | Time-dependent patterns |
| ETA vs covariate | No important unexplained systematic relationship | Strong covariate-related structure |
13. How Should Outlying Observations Be Interpreted?
GOF plots often identify observations that are substantially different from the model prediction. These observations deserve investigation, but an outlier should not automatically be deleted.
Potential explanations include:
- Data-entry or transcription errors.
- Incorrect dose or dosing-time information.
- Incorrect sampling-time information.
- Assay problems.
- Unexpected clinical or physiological circumstances.
- True biological variability not represented by the model.
- Model misspecification.
The appropriate response depends on the evidence. Removing an observation solely because it produces a poor fit can distort parameter estimation and hide genuine model deficiencies.
14. GOF Diagnostics for Structural Model Development
Goodness-of-fit diagnostics are especially useful when comparing candidate structural models.
Suppose a dataset is initially described with a one-compartment model. If residuals show systematic deviations during distribution and terminal phases, a two-compartment model may provide a more appropriate representation of the observed concentration-time behavior.
Similarly, systematic deviations following extravascular dosing may suggest that the absorption model is inadequate.
| Observed pattern | Possible model-development question |
|---|---|
| Early-time curvature | Is the absorption model sufficiently flexible? |
| Distribution-phase bias | Is a multi-compartment structural model needed? |
| Terminal-phase bias | Is elimination adequately represented? |
| Concentration-dependent residual spread | Is the residual error model appropriate? |
| Covariate-related ETA pattern | Is an important covariate relationship missing? |
These diagnostics generate hypotheses for model refinement. They do not determine the correct model in isolation.
15. Worked Example: Interpreting a GOF Diagnostic Set
Consider a hypothetical population PK model with 200 subjects and 2,400 observed concentrations. The investigator evaluates DV versus PRED, DV versus IPRED, and WRES versus PRED.
Step 1: Examine DV versus PRED
Suppose most observations are distributed around the identity line, but the highest concentrations tend to fall below the line.
This suggests that the model may systematically overpredict concentrations in the upper prediction range.
Step 2: Examine DV versus IPRED
Suppose DV versus IPRED looks substantially tighter, with most observations near the identity line.
This indicates that subject-specific random effects allow the model to account for substantial individual differences. It does not by itself establish that the population model is adequate.
Step 3: Examine WRES versus PRED
Suppose WRES is approximately centered around zero at low and moderate predictions but becomes increasingly negative at high predictions.
The pattern is consistent with systematic overprediction at high concentrations and suggests that the issue should be investigated further.
Step 4: Investigate possible explanations
- Check whether the structural model adequately describes the peak.
- Examine whether the residual error model is appropriate.
- Review the dose and sampling information.
- Evaluate whether an important covariate affects exposure.
- Compare the pattern with alternative structural models.
Step 5: Reassess after model refinement
After modifying the model, the same diagnostics should be reconsidered. Improvement should be evaluated across the diagnostic set rather than by selecting whichever individual plot looks most favorable.
16. What GOF Diagnostics Cannot Prove
Goodness-of-fit plots are powerful, but they have important limitations.
- A good fit does not prove the model is biologically true.
- A good fit does not prove the parameters are uniquely identified.
- A poor fit does not identify the cause automatically.
- Individual predictions can be influenced by the same observations used to estimate them.
- Sparse sampling can hide structural model deficiencies.
- Visual agreement does not replace evaluation of parameter uncertainty and plausibility.
- GOF diagnostics alone do not establish predictive performance in new data.
For these reasons, GOF diagnostics are one component of a broader model evaluation strategy.
17. GOF Versus Predictive Diagnostics
A model can describe the data used for estimation reasonably well while still having limitations when applied to new subjects or new observations.
This distinction motivates complementary diagnostics such as:
- Visual predictive checks (VPCs)
- Prediction-corrected VPCs
- Bootstrap validation
- External validation
- Normalized prediction distribution errors (NPDE)
GOF diagnostics primarily ask whether the model adequately describes the observations used in model development. Predictive diagnostics ask additional questions about how well the model predicts data under specified conditions.
18. A Practical GOF Diagnostic Workflow
- Start with the observed data. Understand dose history, sampling times, concentration ranges, BLQ observations, and study design.
- Review the structural model. Confirm that the proposed absorption and disposition model is scientifically plausible.
- Review DV versus PRED. Look for population-level bias, curvature, and concentration-dependent discrepancies.
- Review DV versus IPRED. Determine whether individual random effects substantially improve the correspondence between predictions and observations.
- Review residual diagnostics. Examine RES, WRES, or other appropriate residuals against predictions and time.
- Examine random effects. Review ETA distributions and ETA-versus- covariate relationships while considering shrinkage.
- Investigate unusual observations. Check data quality and scientific explanations before considering exclusions.
- Compare plausible model alternatives. Use diagnostics together with parameter estimates, uncertainty, objective-function-based criteria when appropriate, and scientific plausibility.
- Perform predictive diagnostics. Use VPC, pcVPC, bootstrap, external validation, NPDE, or other appropriate methods when needed.
- Document the reasoning. Record why the final model was selected and how diagnostic findings influenced model development.
19. How to Report GOF Diagnostics
A population PK analysis should make it possible for readers to understand how model adequacy was assessed.
A useful diagnostic summary can describe:
- The observed-versus-predicted relationships.
- The residual diagnostics examined.
- Whether systematic trends were identified.
- How individual predictions behaved relative to observations.
- Whether important ETA-versus-covariate relationships remained.
- Any influential or unusual observations and how they were investigated.
- How GOF findings contributed to structural, variability, or covariate model decisions.
- Which complementary predictive diagnostics were performed.
The objective is not merely to display plots. The diagnostic assessment should explain what the plots showed and how those findings affected model development.
20. Key Takeaways
- Goodness-of-fit diagnostics evaluate how well a population PK model describes observed concentration data.
- DV versus PRED primarily evaluates population-level agreement between observations and predictions.
- DV versus IPRED evaluates the correspondence after incorporating individual random effects.
- Residual-versus-prediction plots help identify bias, curvature, and inappropriate variance patterns.
- Residual-versus-time plots can reveal systematic discrepancies across different parts of the PK profile.
- ETA diagnostics can reveal potentially unexplained relationships with covariates, but they should be interpreted with shrinkage and sampling information in mind.
- Outlying observations should be investigated scientifically rather than removed simply because they fit poorly.
- GOF diagnostics can suggest structural-model, covariate-model, or residual-error-model problems, but they do not identify the correct solution automatically.
- A visually good fit does not prove biological truth, parameter identifiability, or predictive performance.
- GOF diagnostics should be combined with parameter assessment and predictive diagnostics such as VPCs, pcVPCs, bootstrap validation, external validation, and NPDE when appropriate.
Where to Go Next
A natural next step is to study Visual Predictive Checks for Population PK, which extend model evaluation from pointwise goodness of fit to comparison of the observed data with distributions of simulated data generated from the population model.
Related topics include prediction-corrected VPCs, bootstrap validation, external validation, normalized prediction distribution errors, ETA shrinkage, epsilon shrinkage, and covariate diagnostics.