Introduction
A companion diagnostic, commonly abbreviated CDx, is an in vitro diagnostic test whose information is essential for the safe and effective use of a corresponding therapeutic product.
The statistical validation of such a test is therefore different from simply demonstrating that a laboratory assay produces repeatable numbers.
The central question is whether the diagnostic can reliably identify the biomarker-defined population for which a particular therapeutic decision is intended.
That requires evidence across multiple dimensions.
- Analytical performance
- Precision and reproducibility
- Accuracy or agreement
- Analytical sensitivity
- Analytical specificity
- Reportable range
- Interference and cross-reactivity
- Specimen stability
- Robustness
- Clinical concordance
- Invalid and indeterminate results
- Predefined statistical acceptance criteria
What Makes a Companion Diagnostic Different?
A conventional laboratory assay may be developed primarily to measure a biological characteristic.
A companion diagnostic has a more consequential role.
Its result may determine whether a patient:
- Receives a treatment
- Does not receive a treatment
- Receives a particular dose or regimen
- Enters a biomarker-defined clinical trial
- Continues treatment based on a molecular characteristic
Consequently, an analytical error can translate into a clinically meaningful classification error.
For example, a false-negative CDx result could classify a biomarker-positive patient as negative and potentially exclude that patient from a treatment intended for the biomarker-defined population.
A false-positive result can create a different problem by identifying a patient as eligible when the therapeutic evidence does not support that classification.
The Validation Framework
A useful way to organize CDx validation is into two major evidence streams:
Analytical Validation vs. Clinical Validation
| Dimension | Analytical Validation | Clinical Validation |
|---|---|---|
| Primary question | Does the assay measure reliably? | Does the result support the intended clinical use? |
| Typical data | Replicates, controls, specimens, concentrations | Clinical specimens and clinical outcome or comparator data |
| Examples | Precision, LoD, linearity, interference | Clinical concordance, PPA, NPA |
| Primary unit | Assay result | Patient classification |
| Statistical emphasis | Variance, bias, agreement, detection probability | Classification agreement and clinical performance |
The First Statistical Principle: Define the Intended Use
Statistical validation should begin with the intended-use statement.
Consider three hypothetical assays:
| Assay | Intended Measurement | Potential Validation Focus |
|---|---|---|
| DNA mutation assay | Presence or absence of a specific variant | Accuracy, LoD, precision, variant-level agreement |
| RNA expression assay | Expression above a predefined threshold | Precision, cutoff performance, reproducibility |
| Protein IHC assay | Biomarker expression category | Reader agreement, categorical agreement, reproducibility |
The statistical methods cannot be selected independently of the assay's intended use.
Analytical Accuracy
Analytical accuracy concerns how closely the assay result agrees with an appropriate reference or comparator, depending on the measurement and validation design.
For a continuous measurement, one might evaluate:
- Bias
- Mean difference
- Regression
- Limits of agreement
- Correlation as a supplemental descriptive measure
For a qualitative assay, the analysis often focuses on classification agreement.
Suppose a CDx classifies specimens as positive or negative.
The results can be summarized in a two-by-two table.
Counts are simulated for teaching purposes. In a real validation report, the reference method and acceptance criteria should be prespecified.
The 2 × 2 Table
For a binary comparator, the fundamental table is:
| Comparator Positive | Comparator Negative | |
|---|---|---|
| CDx Positive | True positive / positive agreement cell | False positive / negative-discordance cell |
| CDx Negative | False negative / negative-discordance cell | True negative / negative agreement cell |
Let:
represent the four cells.
Then classical sensitivity and specificity are:
However, when neither method can legitimately be treated as a perfect reference standard, the preferred terminology may be positive percent agreement and negative percent agreement rather than sensitivity and specificity.
Positive Percent Agreement
Positive percent agreement, or PPA, is commonly calculated as:
This asks:
PPA is therefore numerically identical to sensitivity in the two-by-two framework when the comparator-positive category is used as the denominator.
The distinction is interpretive: the comparator is not necessarily assumed to represent absolute biological truth.
Negative Percent Agreement
Negative percent agreement, or NPA, is:
It asks:
Overall Percent Agreement
Overall percent agreement can be calculated as:
OPA is useful as a descriptive summary, but it can hide asymmetric performance.
For example, an assay could have 98% OPA while having materially different positive and negative agreement.
Confidence Intervals Matter
A point estimate such as:
does not tell the complete statistical story.
Suppose the estimate is based on only 20 comparator-positive specimens.
The uncertainty can be substantial.
A much larger validation dataset may produce the same point estimate with a much narrower confidence interval.
Therefore, CDx validation reports should generally present the point estimate together with an appropriate confidence interval.
Which Confidence Interval?
For binomial proportions, commonly used approaches include:
- Exact binomial confidence intervals
- Wilson confidence intervals
- Other prespecified interval methods appropriate to the design
The ordinary Wald interval is often undesirable when sample sizes are small or proportions are near 0% or 100%.
For a proportion:
the interval method should be selected prospectively and consistently.
Illustrative PPA Calculation
Suppose a validation study includes:
Then:
The important reporting statement is not simply "PPA was 92%."
A complete result should include the numerator, denominator, percentage, and confidence interval.
Sample Size for Agreement Studies
Sample-size planning should reflect the performance claim being evaluated.
For example, suppose the goal is to demonstrate that PPA exceeds a minimum acceptable value:
A study may be designed around a one-sided hypothesis such as:
versus:
The required number of positive specimens depends on:
- The minimum acceptable performance
- The anticipated true performance
- Type I error
- Desired power
- The confidence-bound strategy
- Expected evaluable specimen availability
Analytical Precision
Precision describes the closeness of repeated measurements under specified conditions.
Precision can be divided into components such as:
- Repeatability
- Within-run precision
- Between-run precision
- Between-day precision
- Between-operator precision
- Between-instrument precision
- Between-site reproducibility
The appropriate components depend on the technology and intended use.
Continuous Measurement Precision
For a continuous assay result \(Y\), variance can be decomposed into multiple sources.
More complex studies may include additional variance components for operator, instrument, day, lot, site, or other factors.
A variance-components model can therefore be useful when a CDx is intended to operate across multiple laboratories or platforms.
Coefficient of Variation
For positive continuous measurements, the coefficient of variation is often useful:
For example, if repeated measurements have:
then:
CV is especially useful when variability is naturally proportional to the magnitude of the measurement.
It is not universally appropriate for every assay, particularly where values can be near zero or where the measurement scale does not support a meaningful ratio interpretation.
Precision Study Design
| Factor | Illustrative Levels | Purpose |
|---|---|---|
| Specimen | Negative, low positive, moderate positive, high positive | Evaluate performance across clinically relevant concentrations |
| Operator | Multiple operators | Assess operator-related variability |
| Day | Multiple nonconsecutive days | Assess temporal variability |
| Run | Multiple runs | Assess run-to-run variation |
| Lot | Multiple reagent lots | Assess lot-to-lot variability |
| Instrument | Multiple instruments where applicable | Assess platform variability |
Reproducibility
Reproducibility asks whether the assay produces consistent results when conditions vary in ways representative of intended use.
For a multicenter CDx, this can be especially important.
A study might deliberately include:
- Multiple laboratories
- Multiple operators
- Multiple reagent lots
- Multiple instruments
- Multiple testing days
The statistical analysis should distinguish the experimental design from a simple collection of repeated observations.
Limit of Detection
The limit of detection, or LoD, is the lowest amount or concentration of analyte that can be reliably detected under the defined conditions of the assay.
LoD is especially important for low-frequency molecular biomarkers.
Consider a mutation with variant allele frequency:
The assay may detect it in some replicates but not others.
Therefore, LoD is fundamentally a probability-of-detection problem rather than simply the concentration at which the mean signal becomes nonzero.
Probability of Detection
Suppose an assay is tested at several concentrations.
| Concentration | Detected / Tested | Detection Rate |
|---|---|---|
| 0.05% | 4 / 20 | 20% |
| 0.10% | 12 / 20 | 60% |
| 0.15% | 18 / 20 | 90% |
| 0.20% | 20 / 20 | 100% |
These data illustrate why the LoD should be evaluated across concentrations rather than inferred from one successful measurement.
Illustrative data only. A real LoD study should use a protocol-defined estimation method and acceptance criterion appropriate to the assay.
LoD and Hit Rates
For qualitative assays, one useful representation is:
The detection probability generally increases as analyte concentration increases.
The validation team may define an LoD at a specified detection probability, provided that the definition is appropriate to the assay and prespecified.
LoQ Is Different From LoD
The limit of quantitation, or LoQ, concerns the lowest concentration at which the analyte can be quantified with predefined performance characteristics.
LoD and LoQ should not be treated as interchangeable.
| Concept | Question |
|---|---|
| LoD | Can the analyte be reliably detected? |
| LoQ | Can the analyte be quantified with acceptable performance? |
Analytical Specificity
Analytical specificity concerns whether the assay detects the intended target without inappropriate interference from other substances or biological features.
Depending on the technology, studies may evaluate:
- Potential interferents
- Cross-reactive targets
- Closely related sequences
- Endogenous substances
- Exogenous substances
- Common medications
- Matrix effects
Interference Studies
Suppose a potentially interfering substance is present at a clinically relevant concentration.
A useful comparison is:
For a qualitative assay, the focus may instead be on whether the classification changes.
For example:
| Specimen | Without Interferent | With Interferent | Interpretation |
|---|---|---|---|
| Low positive | Positive | Positive | No observed classification effect |
| Near cutoff | Positive | Negative | Potential interference requiring investigation |
| Negative | Negative | Negative | No observed false-positive effect |
Reportable Range
The reportable range describes the range over which the assay can provide results that meet the defined performance requirements.
For quantitative assays, this may involve:
- Linearity
- Accuracy
- Precision
- Recovery
- Bias
For categorical molecular assays, the relevant concept may instead involve the range of analyte concentrations or variant frequencies over which the assay maintains acceptable classification performance.
Linearity
For quantitative measurements, linearity asks whether the observed result changes appropriately with the true or expected concentration.
A simple model is:
where:
- \(X_i\) = expected or reference concentration
- \(Y_i\) = observed assay result
- \(\beta_0\) = intercept
- \(\beta_1\) = slope
A slope near 1 and intercept near 0 can support agreement, but regression coefficients should not be interpreted in isolation.
Bias
Bias represents systematic deviation from a reference or expected value.
For paired continuous results:
The mean difference is:
A mean difference close to zero suggests little average bias, although the distribution of individual differences must also be considered.
Agreement for Continuous Results
When two methods measure the same continuous quantity, Bland-Altman-style analysis can be informative.
The difference is plotted against the average:
Approximate limits of agreement can then be expressed as:
The actual analysis should account for the study design and assumptions rather than mechanically applying the formula.
Agreement plots should be interpreted relative to clinically acceptable differences, not merely according to whether the confidence limits look narrow.
Cutoff Determination
Many companion diagnostics convert a continuous or semi-quantitative assay result into a binary classification.
For example:
where \(c\) is the classification cutoff.
Cutoff selection is therefore a major statistical issue.
Cutoff Validation Principles
A cutoff should not simply be selected because it produces the most favorable validation results in a convenient dataset.
Important considerations include:
- Clinical rationale
- Analytical characteristics
- Biological plausibility
- Prespecified development strategy
- Independent or appropriately controlled evaluation
- Expected clinical consequences of misclassification
ROC Curves
When a continuous biomarker is used to classify patients, receiver operating characteristic analysis may be useful during development.
The ROC curve displays:
across possible thresholds.
The area under the curve, or AUC, summarizes discrimination over the range of thresholds.
However, the AUC does not automatically establish that a selected cutoff is clinically appropriate.
Clinical Concordance
Clinical validation asks a different question from analytical validation.
The central issue is whether the diagnostic identifies the population in which the therapeutic benefit demonstrated in clinical development applies.
For example, suppose a pivotal clinical trial enrolled patients based on biomarker status determined using an investigational assay.
A commercial CDx may need to demonstrate appropriate concordance with that clinical-trial classification.
Bridging a Clinical-Trial Assay to a CDx
A common development problem is that the assay used during clinical development differs from the final commercial diagnostic.
The statistical bridge may involve comparing:
- Clinical-trial assay results
- Candidate CDx results
- Reference or orthogonal assay results
- Archived clinical specimens
The key question is whether patients classified as biomarker-positive or negative by the final assay correspond appropriately to the population supporting the therapeutic indication.
Positive and Negative Percent Agreement in Bridging
Suppose the clinical-trial assay identifies 150 positive specimens and the candidate CDx identifies 143 of those as positive.
If there are 7 discordant positive specimens:
The analysis should then investigate the discordant specimens rather than treating the aggregate percentage as the end of the analysis.
Discordant Analysis
Discordant specimens can provide important information.
For every clinically important discordance, consider whether the difference could arise from:
- Specimen quality
- Low analyte concentration
- Heterogeneity
- Pre-analytical handling
- Assay failure
- Cutoff differences
- Biological differences
- True disagreement between measurement technologies
A good validation report should not simply list discordances without investigation.
Borderline Specimens
Specimens near the assay cutoff are particularly informative.
Suppose a continuous assay has a positive cutoff of:
A specimen with:
may be classified positive while another measurement of the same specimen could be negative.
This makes near-cutoff precision and reproducibility especially important.
The figure is illustrative. Actual cutoff evaluation should follow the assay's intended-use and validation strategy.
Clinical Consequences of Misclassification
A useful validation framework distinguishes statistical performance from clinical consequences.
| Error | Diagnostic Result | Potential Clinical Consequence |
|---|---|---|
| False positive | CDx positive when target is absent | Potential treatment of a patient outside the evidence-supported population |
| False negative | CDx negative when target is present | Potential exclusion from an indicated or investigational therapy |
| Invalid result | No interpretable result | Potential delay, repeat testing, or unresolved treatment decision |
The relative importance of these errors depends on the clinical context.
Predictive Values
Positive predictive value is:
Negative predictive value is:
Unlike PPA and NPA, predictive values depend strongly on the prevalence of the biomarker in the tested population.
Prevalence Effects
Consider an assay with:
If biomarker prevalence is low, the number of false positives can be large relative to true positives.
Therefore, the same assay can have substantially different PPV and NPV in different populations.
Invalid Rates
A CDx validation should explicitly evaluate invalid, indeterminate, or uninterpretable results when such outcomes are possible.
For example:
An assay with excellent PPA among valid results may still have an important clinical limitation if a substantial fraction of specimens produce no valid result.
First-Pass Success Rate
Another useful metric can be the percentage of specimens producing an interpretable result on the initial testing attempt.
This can be particularly relevant for assays involving complex specimen processing.
Missing Data and Unevaluable Specimens
Diagnostic validation studies can lose specimens because of:
- Insufficient specimen quantity
- Degraded material
- Failed extraction
- Instrument failure
- Invalid control results
- Missing comparator results
- Inadequate specimen quality
The analysis should distinguish:
- Not tested
- Test attempted but invalid
- Valid negative
- Valid positive
- Unevaluable for a particular endpoint
These categories should not be silently collapsed into a single "missing" category.
Repeat Testing
Repeated testing after an initial invalid result introduces another important statistical issue.
Suppose:
- 100 specimens are initially tested.
- 10 produce invalid results.
- 8 of the 10 become valid after repeat testing.
Reporting only the final valid results could make the assay appear to have 100% valid performance among the retained specimens.
The initial invalid rate remains clinically relevant.
Robustness
Robustness examines whether modest, realistic variations in operating conditions materially affect assay performance.
Examples include:
- Incubation time
- Temperature
- Input amount
- Extraction conditions
- Storage duration
- Reagent handling
- Instrument settings
The purpose is not to intentionally break the assay.
The purpose is to determine whether reasonable deviations within the intended operating environment cause unacceptable performance changes.
Specimen Stability
Specimen stability can be evaluated at clinically relevant storage conditions.
For example:
| Condition | Time | Primary Question |
|---|---|---|
| Room temperature | Defined interval | Does expected handling affect classification? |
| Refrigerated | Defined interval | Is refrigerated storage acceptable? |
| Frozen | Defined interval | Does long-term storage preserve performance? |
| Freeze-thaw | Multiple cycles | Does repeated handling materially affect results? |
Lot-to-Lot Consistency
If reagent lots can change over time, the validation strategy should evaluate whether lot-to-lot variation remains within predefined requirements.
For a quantitative assay, this may involve estimating:
For a qualitative assay, classification agreement may be more appropriate.
Reader Agreement
Some CDx technologies involve subjective interpretation.
Immunohistochemistry is a common example where trained readers may assign categorical scores.
Possible analyses include:
- Raw agreement
- Positive agreement
- Negative agreement
- Inter-reader agreement
- Intra-reader agreement
- Kappa statistics where appropriate
Kappa Statistics
For categorical assessments, Cohen's kappa is:
where:
- \(P_o\) = observed agreement
- \(P_e\) = agreement expected by chance under the specified model
Kappa can be useful, but it should not replace clinically meaningful agreement measures.
In situations with highly imbalanced prevalence, kappa can behave in ways that make interpretation difficult.
Multi-Category CDx Results
Not all CDx classifications are binary.
An assay might classify specimens as:
- Negative
- Low positive
- Intermediate
- High positive
For ordered categories, weighted agreement methods may be appropriate.
The analysis should preserve the ordinal structure rather than automatically reducing everything to positive versus negative.
NGS-Based Companion Diagnostics
Next-generation sequencing introduces additional statistical considerations.
Potential validation dimensions include:
- Variant detection
- Variant type
- Allele fraction
- Coverage
- Read quality
- Sample quality
- False-positive variant calls
- False-negative variant calls
- Pipeline reproducibility
- Reference genome and bioinformatics configuration
The statistical unit may differ depending on whether the endpoint is:
- Specimen-level classification
- Variant-level detection
- Patient-level biomarker status
Variant-Level vs. Patient-Level Agreement
| Unit | Example Question |
|---|---|
| Patient | Is this patient's biomarker status correctly classified? |
| Specimen | Does the assay correctly classify this tested specimen? |
| Variant | Does the assay detect this specific molecular alteration? |
| Call | Does the pipeline generate the correct individual result? |
Nested Data and Correlation
When multiple measurements come from the same specimen or patient, the data are clustered.
For example:
Observations within the same hierarchy are not necessarily independent.
Possible analytical approaches include:
- Mixed-effects models
- Variance-component models
- Generalized estimating equations
- Cluster-aware confidence intervals
- Specimen-level summaries
The method should follow the estimand and experimental design.
Analytical Sensitivity vs. Clinical Sensitivity
These terms are often confused.
| Term | Meaning |
|---|---|
| Analytical sensitivity | Ability to detect the analyte at low concentrations or low abundance |
| Clinical sensitivity | Ability to identify patients in the clinically relevant positive population |
A test can have excellent analytical sensitivity while still having imperfect clinical sensitivity because clinical specimens introduce biological heterogeneity and pre-analytical variation.
Pre-Analytical Variables
The assay does not operate in isolation.
The full diagnostic workflow can be represented as:
Validation should consider the portions of this pathway included in the intended use.
Sample Quality
For tissue-based diagnostics, sample quality can materially influence performance.
Potential factors include:
- Tumor content
- Necrosis
- DNA or RNA integrity
- Fixation quality
- Specimen age
- Section thickness
- Extraction yield
These variables may need to be incorporated into analytical validation or eligibility criteria.
Statistical Acceptance Criteria
A validation study should not be designed around vague goals such as "demonstrate good performance."
Instead, each critical performance characteristic should have a predefined acceptance criterion.
| Performance Characteristic | Example Acceptance Framework |
|---|---|
| PPA | Lower confidence bound exceeds a prespecified minimum |
| NPA | Lower confidence bound exceeds a prespecified minimum |
| Precision | Variance or agreement remains within predefined limits |
| LoD | Detection probability meets the predefined criterion |
| Interference | No clinically meaningful classification or measurement effect |
| Stability | Performance remains within predefined limits over intended conditions |
Point Estimate vs. Confidence-Bound Criterion
Suppose the acceptance criterion is:
A study might observe:
But if the lower confidence bound is below 90%, the prespecified success rule may not be met.
Therefore:
The inferential rule must be defined in advance.
Multiple Performance Characteristics
A CDx validation program can contain many endpoints.
For example:
- PPA
- NPA
- Precision
- LoD
- Linearity
- Interference
- Stability
- Reproducibility
Not every performance characteristic should automatically be treated as a formal hypothesis test requiring multiplicity adjustment.
The statistical framework should distinguish:
- Descriptive performance characterization
- Formal acceptance criteria
- Primary validation claims
- Secondary supporting analyses
Study Design Matrix
A useful planning tool is a validation matrix.
| Question | Study Type | Typical Data | Statistical Output |
|---|---|---|---|
| Can the assay detect low-level target? | LoD | Replicate low-concentration specimens | Detection probability |
| Is repeated testing consistent? | Precision | Replicate results | SD, CV, agreement |
| Does the assay agree with comparator? | Agreement | Paired specimens | PPA, NPA, OPA, CI |
| Does performance vary by operator/site? | Reproducibility | Multifactorial replicates | Variance components/agreement |
| Does storage affect results? | Stability | Paired stored and reference specimens | Difference/agreement |
| Do substances interfere? | Interference | Spiked specimens | Classification or quantitative effect |
Validation Hierarchy
A useful conceptual hierarchy is:
Clinical Cutoff vs. Analytical Cutoff
A subtle but important distinction is the difference between an analytical detection threshold and a clinical classification threshold.
For example, an assay might reliably detect a molecular signal at:
while the clinically validated positive classification cutoff may be:
These are not contradictory.
The assay can detect lower levels than are required for the clinical classification.
Precision Near the Clinical Cutoff
Near-cutoff precision is often more consequential than precision far from the cutoff.
Suppose:
and repeated measurements are:
The assay may have excellent precision in absolute terms but still produce classification switching around the decision boundary.
Therefore, validation should include specimens representative of the decision region.
Classification Concordance
For a binary CDx, classification concordance can be expressed as:
But concordance should be accompanied by the directional components.
| Measure | Why It Matters |
|---|---|
| PPA | Captures agreement in the positive population |
| NPA | Captures agreement in the negative population |
| OPA | Provides overall descriptive agreement |
| Discordant analysis | Identifies the causes and clinical relevance of disagreement |
Reference Method Limitations
A comparator method is not automatically a perfect gold standard.
This is particularly important for biomarkers where biological truth cannot be directly observed.
If two assays use different technologies, a disagreement may reflect:
- Technical error by one assay
- Technical error by the other assay
- Differences in analytical sensitivity
- Different target definitions
- Biological heterogeneity
Therefore, an orthogonal method or adjudication strategy may sometimes be useful for discordant cases.
Orthogonal Testing
An orthogonal assay uses a different analytical principle to investigate a result.
For example, an NGS finding could potentially be investigated using an independent molecular method where appropriate.
The purpose is not necessarily to declare one assay correct automatically.
Instead, orthogonal evidence can help determine whether a discordance reflects a genuine biological signal or an analytical limitation.
Clinical Specimen Diversity
A strong validation program should reflect the intended clinical population and specimen environment.
Depending on the indication, this may include variation in:
- Disease stage
- Specimen type
- Tumor content
- Biomarker abundance
- Demographic and clinical characteristics
- Specimen age
- Collection site
The precise sampling strategy depends on the intended-use population and development program.
Rare Biomarkers
Rare biomarker populations create a special statistical challenge.
Suppose the biomarker prevalence is only 5%.
A general population sample of 1,000 patients would contain approximately:
positive patients on average.
If the precision of PPA is important, the study may therefore require intentional enrichment or another appropriate specimen-acquisition strategy, provided that the resulting analysis remains valid for the intended claim.
Enrichment Does Not Change the Mathematics of PPA
If the study intentionally enriches positive specimens, PPA and NPA remain defined using their respective denominators.
However, predictive values estimated from the enriched sample would generally not represent the predictive values expected in a routine population unless appropriately adjusted.
Statistical Analysis Plan
The statistical analysis plan for a CDx validation study should define at least:
- Analysis populations
- Primary endpoints
- Secondary endpoints
- Specimen eligibility
- Handling of invalid results
- Handling of repeat testing
- Comparator definition
- Cutoff definition
- Confidence interval methodology
- Missing-data handling
- Outlier handling
- Acceptance criteria
- Subgroup analyses
- Discordant-result investigations
Analysis Populations
A useful framework may distinguish:
| Population | Description |
|---|---|
| All specimens | All specimens entering the validation workflow |
| Tested specimens | Specimens for which testing was attempted |
| Valid specimens | Specimens producing interpretable results |
| Paired-evaluable specimens | Specimens with interpretable results from both candidate and comparator methods |
Different endpoints may use different denominators.
That is acceptable as long as the definitions are explicit.
Subgroup Analysis
CDx performance can vary across clinically relevant specimen characteristics.
Potential subgroups may include:
- Specimen type
- Tumor content
- Biomarker abundance
- Site
- Operator
- Instrument
- Reagent lot
Subgroup analyses should generally be prespecified where they are intended to support a validation claim.
Small Subgroups
A subgroup with only a few positive specimens can produce an apparently extreme PPA.
For example:
does not mean the subgroup's true performance is known with certainty.
The confidence interval is essential.
Bootstrapping and Resampling
For some complex validation statistics, resampling methods may provide useful uncertainty estimates.
However, resampling should respect the data hierarchy.
If five replicates come from one specimen, randomly resampling individual replicates as though they were independent patient observations can be misleading.
A cluster-aware bootstrap may instead resample at the specimen level.
Power and Precision
Validation studies are often framed around acceptance criteria rather than traditional superiority hypotheses.
Consequently, sample-size planning may focus on the width of a confidence interval.
For example, the study team may require:
where \(W\) is the maximum acceptable uncertainty.
This is particularly intuitive for agreement estimates.
Example of Precision-Based Planning
Suppose investigators expect:
and want a reasonably narrow confidence interval around the estimate.
The number of positive specimens required will generally be driven much more by the desired precision than by the total number of specimens alone.
This illustrates a useful planning principle:
Validation of a Quantitative Biomarker
Suppose the CDx reports a quantitative biomarker value.
A comprehensive analytical validation might examine:
- Accuracy
- Precision
- Linearity
- Reportable range
- Analytical sensitivity
- Analytical specificity
- Interference
- Stability
- Carryover
- Robustness
The final clinical classification may then be based on a prespecified cutoff.
Validation of a Qualitative Biomarker
For a qualitative CDx, the core package may emphasize:
- Positive agreement
- Negative agreement
- Overall agreement
- Precision
- Reproducibility
- LoD
- Interference
- Cross-reactivity
- Invalid rate
- Clinical concordance
The statistical methods should match the categorical nature of the endpoint.
Example CDx Validation Dataset
SPECIMEN CDX COMP RESULT 001 + + Concordant Positive 002 + + Concordant Positive 003 + - Discordant 004 - - Concordant Negative 005 - + Discordant 006 - - Concordant Negative 007 + + Concordant Positive 008 - - Concordant Negative
From the table, the analysis can derive the four classification cells.
In real validation programming, the data structure would generally contain additional variables for specimen source, assay run, lot, operator, site, instrument, biomarker level, validity status, and other protocol-defined attributes.
R Implementation: PPA and NPA
tab <- table( CDx = data$CDX, Comparator = data$COMP ) TP <- tab["Positive", "Positive"] FP <- tab["Positive", "Negative"] FN <- tab["Negative", "Positive"] TN <- tab["Negative", "Negative"] PPA <- TP / (TP + FN) NPA <- TN / (TN + FP) OPA <- (TP + TN) / sum(tab) results <- data.frame( PPA = PPA, NPA = NPA, OPA = OPA )
In production programming, explicit handling of missing cells, zero denominators, factor levels, and invalid specimens should be added.
R Implementation: Wilson Confidence Interval
wilson_ci <- function(x, n, z = 1.96) {
p <- x / n
denominator <- 1 + z^2 / n
center <- (
p + z^2 / (2 * n)
) / denominator
half_width <- (
z * sqrt(
p * (1 - p) / n +
z^2 / (4 * n^2)
)
) / denominator
c(
lower = center - half_width,
upper = center + half_width
)
}
wilson_ci(TP, TP + FN)
For a regulated validation package, the chosen confidence-interval method should be specified in the statistical analysis plan and implemented consistently.
R Implementation: LoD Detection Probability
If replicate results are coded as detected versus not detected, a logistic model can be useful for exploratory estimation of the detection curve.
lod_model <- glm( detected ~ log10(concentration), data = lod_data, family = binomial() ) summary(lod_model) pred <- predict( lod_model, type = "response" )
The exact LoD estimation approach should be selected according to the assay, study design, protocol, and intended claim.
R Implementation: Bland-Altman Data
data$mean_method <- ( data$reference + data$candidate ) / 2 data$difference <- ( data$candidate - data$reference ) mean_difference <- mean( data$difference, na.rm = TRUE ) sd_difference <- sd( data$difference, na.rm = TRUE ) upper_loa <- mean_difference + 1.96 * sd_difference lower_loa <- mean_difference - 1.96 * sd_difference
The resulting values can be plotted against the paired means.
Validation Programming QC
Statistical validation requires more than checking whether the final percentages look plausible.
Independent QC should verify:
- Specimen inclusion
- Comparator classification
- Candidate classification
- Cutoff derivation
- Invalid handling
- Repeat-testing logic
- Numerators
- Denominators
- Confidence intervals
- Precision estimates
- LoD derivations
- Subgroup definitions
- Figure values
Patient-Level Traceability
Every summary statistic should be traceable back to specimen-level records.
For example:
Common Statistical Mistakes
- Calling a comparator the gold standard without justification. Agreement with a comparator is not necessarily equivalent to truth.
- Reporting only overall agreement. PPA and NPA can reveal asymmetric performance hidden by OPA.
- Ignoring confidence intervals. A point estimate does not quantify uncertainty.
- Treating replicates as independent patients. Repeated measurements within a specimen are correlated.
- Ignoring invalid results. Invalid rates can have direct clinical relevance.
- Using post hoc cutoff optimization. This can create optimistic estimates of performance.
- Confusing LoD with clinical sensitivity. Analytical detection capability is not the same as patient-level diagnostic performance.
- Using correlation as a substitute for agreement. High correlation can coexist with systematic bias.
- Ignoring rare positive populations. A large total sample does not guarantee precise PPA.
- Ignoring discordant specimens. Discordances can reveal important assay limitations.
- Pooling heterogeneous specimen types without justification. Different matrices can behave differently.
- Using inappropriate denominators. The denominator must correspond to the defined estimand.
Regulatory Perspective
Companion diagnostic development is closely linked to the therapeutic development program.
The statistical validation strategy should therefore be consistent with:
- The intended use of the diagnostic
- The biomarker definition
- The therapeutic indication
- The clinical-trial enrollment strategy
- The clinical evidence supporting the treatment
- The analytical characteristics of the final assay
In the United States, FDA guidance concerning in vitro companion diagnostic devices emphasizes the close relationship between the diagnostic and the corresponding therapeutic product.
The FDA has also issued guidance addressing analytical validation considerations for certain next-generation sequencing-based in vitro diagnostics.
CDx and the Drug Development Program
The diagnostic and therapeutic programs should generally be considered as an integrated development strategy.
The relationship can be represented as:
Bridging Studies
A bridging study may be necessary when the assay used to select patients in clinical development differs from the final commercial CDx.
The statistical question is whether the final assay reproduces the clinically meaningful classification sufficiently well.
Potential endpoints include:
- PPA
- NPA
- OPA
- Classification consistency
- Discordance rates
- Clinical outcome concordance where applicable
Clinical Outcome Concordance
In some programs, the strongest question is not merely:
but rather:
This is a fundamentally different level of evidence.
A diagnostic can agree extremely well with another assay while neither assay adequately identifies the population that benefits from treatment.
Diagnostic Performance vs. Clinical Utility
| Question | Evidence Type |
|---|---|
| Does the assay detect the biomarker? | Analytical validation |
| Does it reproduce the result? | Precision / reproducibility |
| Does it agree with another method? | Method comparison |
| Does it classify the clinical population consistently? | Clinical concordance / bridging |
| Does biomarker-positive treatment produce benefit? | Clinical trial evidence |
Subgroup Performance
A CDx can have excellent overall performance while exhibiting weaker performance in a clinically important subgroup.
For example, performance could differ by:
- Low versus high biomarker abundance
- Specimen type
- Fresh versus archived tissue
- Primary versus metastatic tissue
- Different collection sites
Subgroup analyses should therefore be planned around plausible sources of heterogeneity.
Interaction Between Biomarker Level and Error
Diagnostic error is often not constant across the measurement range.
For example:
The probability of correct classification may be highest for very strong positives and strong negatives, while specimens near the cutoff may have greater uncertainty.
Visualizing Performance Across the Measurement Range
Illustrative only. The actual relationship between biomarker level and classification performance depends on assay technology and biological context.
Clinical Specimen Testing Strategy
A practical validation program should include specimens spanning the relevant range of biomarker status.
For a binary biomarker, this may include:
- Strong positives
- Moderate positives
- Near-cutoff positives
- Near-cutoff negatives
- Strong negatives
This design helps distinguish broad assay performance from decision-boundary performance.
Precision and Reproducibility Near LoD
Low-level specimens can simultaneously challenge:
- Analytical sensitivity
- Precision
- Classification stability
Therefore, LoD studies and precision studies may provide complementary evidence.
Carryover
For some assays, carryover from a high-positive specimen into a subsequent negative specimen can create false-positive results.
A typical design may alternate high-positive and negative specimens.
The key question is whether the negative specimens remain negative under the specified conditions.
Cross-Reactivity
Cross-reactivity evaluates whether related analytes or sequences produce inappropriate positive signals.
This can be particularly important for molecular assays targeting closely related variants or organisms.
Analytical Validation Summary Matrix
| Characteristic | Typical Statistical Question |
|---|---|
| Accuracy | How close is the result to an appropriate reference or comparator? |
| Precision | How variable are repeated results? |
| Reproducibility | Does performance remain consistent across intended sources of variation? |
| LoD | At what low concentration is detection sufficiently reliable? |
| Linearity | Does response track concentration appropriately? |
| Specificity | Does the assay avoid inappropriate detection? |
| Interference | Do relevant substances alter results materially? |
| Stability | Does specimen handling preserve performance? |
| Robustness | Does modest operational variation affect performance? |
Clinical Validation Summary Matrix
| Characteristic | Typical Statistical Question |
|---|---|
| PPA | How often does the CDx classify comparator-positive specimens as positive? |
| NPA | How often does the CDx classify comparator-negative specimens as negative? |
| OPA | What proportion of paired results agree overall? |
| Discordance | Where and why do methods disagree? |
| Clinical concordance | Does the CDx reproduce clinically meaningful biomarker classification? |
| Clinical utility | Does the diagnostic support the treatment decision for which it is intended? |
Validation Reporting
A final validation report should allow an independent reviewer to understand:
- What was tested
- Why it was tested
- How specimens were selected
- How results were classified
- Which specimens were excluded
- How invalids were handled
- How repeat testing was handled
- Which statistical methods were used
- Which acceptance criteria were prespecified
- Whether the criteria were met
- What limitations remain
Traceability From Requirement to Result
A strong validation package can be organized as:
Example Validation Table
| Endpoint | Result | 95% CI | Acceptance Criterion | Conclusion |
|---|---|---|---|---|
| PPA | 95.2% | 91.8–97.4% | Lower bound ≥90% | Met |
| NPA | 97.1% | 94.6–98.5% | Lower bound ≥95% | Met |
| Invalid rate | 1.4% | — | ≤3% | Met |
| Repeatability CV | 4.8% | — | ≤7% | Met |
| LoD | 0.12% VAF | — | ≤0.20% VAF | Met |
The numerical values above are illustrative rather than regulatory requirements.
How to Interpret a Validation Failure
Not every failure means that the entire assay is unusable.
A failed acceptance criterion should trigger investigation.
Potential explanations include:
- Unexpected specimen variability
- Insufficient sample size
- Incorrect design assumptions
- Assay instability
- Unexpected interference
- Lot effects
- Operator effects
- Inappropriate cutoff
- Data-processing error
The appropriate response depends on the specific failure and its clinical importance.
Statistical Validation Is Not the Same as Statistical Significance
A common misconception is that validation means obtaining statistically significant p-values.
That is often not the central objective.
The real question may be whether performance is sufficiently precise and reliable to support a predefined claim.
For example:
can be more informative than:
for an assay-performance claim.
What Makes a Strong CDx Statistical Package?
A strong package has five characteristics.
1. Prespecified
The endpoints, estimands, acceptance criteria, and analysis methods are defined before examining the final validation results.
2. Representative
The specimens and operating conditions represent the intended use.
3. Quantitative
Performance is expressed with appropriate estimates and uncertainty.
4. Traceable
Every summary result can be traced to the underlying specimen-level data.
5. Clinically connected
Analytical performance is ultimately connected to the clinical decision the diagnostic is intended to support.
Practical CDx Validation Checklist
Common Reporting Language
A concise statistical conclusion might state:
What the Final Statistical Report Should Contain
| Section | Content |
|---|---|
| Objective | Intended use and validation claims |
| Study design | Specimens, operators, sites, lots, instruments, conditions |
| Analysis populations | Definitions of evaluable, valid, invalid, and paired specimens |
| Methods | Statistical procedures and confidence intervals |
| Results | Performance estimates and uncertainty |
| Discordance | Detailed investigation of discrepant specimens |
| Subgroups | Clinically relevant performance stratification |
| Limitations | Known limitations and residual uncertainty |
| Conclusion | Whether predefined validation criteria were satisfied |
Final Perspective
Companion diagnostic statistical validation sits at the intersection of laboratory science, biostatistics, clinical development, and regulatory science.
The central statistical challenge is not simply calculating a sensitivity, specificity, or coefficient of variation.
The deeper question is whether the evidence establishes that the diagnostic can make the intended clinical classification reliably enough for its purpose.
That requires a chain of evidence.
Each link answers a different question.
An assay may be analytically precise but clinically irrelevant.
An assay may agree well with a comparator but perform poorly near the therapeutic decision threshold.
An assay may have excellent PPA but an unacceptable invalid rate.
A large study may produce a precise overall estimate while providing too few positive specimens to establish a reliable PPA.
These distinctions are why companion diagnostic validation requires careful statistical planning before the laboratory study begins.
References
U.S. Food and Drug Administration. In Vitro Companion Diagnostic
Devices: Guidance for Industry and Food and Drug Administration Staff.
2014.
U.S. Food and Drug Administration. Considerations for Design,
Development, and Analytical Validation of Next Generation Sequencing
(NGS)-Based In Vitro Diagnostics (IVDs) Intended to Aid in the Diagnosis of
Cancer.
U.S. Food and Drug Administration. Statistical Guidance on Reporting
Results from Studies Evaluating Diagnostic Tests.
Clinical and Laboratory Standards Institute. Evaluation of Precision
of Quantitative Measurement Procedures. CLSI EP05.
Clinical and Laboratory Standards Institute. Evaluation of Detection
Capability for Clinical Laboratory Measurement Procedures. CLSI
EP17.
Clinical and Laboratory Standards Institute. Measurement Procedure
Comparison and Bias Estimation Using Patient Samples. CLSI EP09.
Clinical and Laboratory Standards Institute. Evaluation of
Qualitative, Binary Output Examination Performance. CLSI EP12.
Bland, J.M. and Altman, D.G. Statistical methods for assessing
agreement between two methods of clinical measurement. The Lancet, 1986.
McHugh, M.L. Interrater reliability: the kappa statistic. Biochemia Medica, 2012.