Tutorials › Biostatistics › Companion Diagnostic Statistical Validation

Diagnostic Statistical Validation

Companion Diagnostic Statistical Validation

A practical guide to statistically validating companion diagnostics, including analytical accuracy, precision, reproducibility, sensitivity, specificity, limit of detection, reportable range, interference, robustness, clinical concordance, positive and negative percent agreement, confidence intervals, sample-size planning, discordant results, and regulatory reporting.

Advanced 24 min read

What You'll Learn

  • How companion diagnostic validation differs from ordinary assay verification
  • How to design analytical accuracy, precision, and reproducibility studies
  • How to estimate sensitivity, specificity, LoD, and reportable range
  • How to evaluate agreement with a comparator or reference method
  • How clinical concordance differs from analytical accuracy
  • How to build a statistically defensible CDx validation package

Introduction

A companion diagnostic, commonly abbreviated CDx, is an in vitro diagnostic test whose information is essential for the safe and effective use of a corresponding therapeutic product.

The statistical validation of such a test is therefore different from simply demonstrating that a laboratory assay produces repeatable numbers.

The central question is whether the diagnostic can reliably identify the biomarker-defined population for which a particular therapeutic decision is intended.

That requires evidence across multiple dimensions.

  • Analytical performance
  • Precision and reproducibility
  • Accuracy or agreement
  • Analytical sensitivity
  • Analytical specificity
  • Reportable range
  • Interference and cross-reactivity
  • Specimen stability
  • Robustness
  • Clinical concordance
  • Invalid and indeterminate results
  • Predefined statistical acceptance criteria
Key idea: CDx validation is not one statistical test. It is a structured body of evidence demonstrating that the assay performs adequately for its intended clinical use.

What Makes a Companion Diagnostic Different?

A conventional laboratory assay may be developed primarily to measure a biological characteristic.

A companion diagnostic has a more consequential role.

Its result may determine whether a patient:

  • Receives a treatment
  • Does not receive a treatment
  • Receives a particular dose or regimen
  • Enters a biomarker-defined clinical trial
  • Continues treatment based on a molecular characteristic

Consequently, an analytical error can translate into a clinically meaningful classification error.

For example, a false-negative CDx result could classify a biomarker-positive patient as negative and potentially exclude that patient from a treatment intended for the biomarker-defined population.

A false-positive result can create a different problem by identifying a patient as eligible when the therapeutic evidence does not support that classification.

Validation must follow intended use. The appropriate statistical design depends on what the assay is intended to measure, the specimen type, the technology, the biomarker, the clinical decision, the comparator or reference method, and the role of the assay in the therapeutic development program.

The Validation Framework

A useful way to organize CDx validation is into two major evidence streams:

1
Analytical validation Demonstrates that the assay measures the intended analyte or biomarker reliably under defined laboratory conditions.
2
Clinical validation or clinical concordance Demonstrates that the assay's result is appropriately associated with the clinical or therapeutic context for which the diagnostic is intended.
3
Integrated validation Connects analytical performance to the clinical decision supported by the diagnostic.

Analytical Validation vs. Clinical Validation

Dimension Analytical Validation Clinical Validation
Primary question Does the assay measure reliably? Does the result support the intended clinical use?
Typical data Replicates, controls, specimens, concentrations Clinical specimens and clinical outcome or comparator data
Examples Precision, LoD, linearity, interference Clinical concordance, PPA, NPA
Primary unit Assay result Patient classification
Statistical emphasis Variance, bias, agreement, detection probability Classification agreement and clinical performance

The First Statistical Principle: Define the Intended Use

Statistical validation should begin with the intended-use statement.

Consider three hypothetical assays:

Assay Intended Measurement Potential Validation Focus
DNA mutation assay Presence or absence of a specific variant Accuracy, LoD, precision, variant-level agreement
RNA expression assay Expression above a predefined threshold Precision, cutoff performance, reproducibility
Protein IHC assay Biomarker expression category Reader agreement, categorical agreement, reproducibility

The statistical methods cannot be selected independently of the assay's intended use.

Analytical Accuracy

Analytical accuracy concerns how closely the assay result agrees with an appropriate reference or comparator, depending on the measurement and validation design.

For a continuous measurement, one might evaluate:

  • Bias
  • Mean difference
  • Regression
  • Limits of agreement
  • Correlation as a supplemental descriptive measure

For a qualitative assay, the analysis often focuses on classification agreement.

Suppose a CDx classifies specimens as positive or negative.

The results can be summarized in a two-by-two table.

Figure 1. Analytical Classification Agreement
Illustrative 2 × 2 comparison between the candidate companion diagnostic and a comparator classification.
Concordant classification
Discordant classification

Counts are simulated for teaching purposes. In a real validation report, the reference method and acceptance criteria should be prespecified.

The 2 × 2 Table

For a binary comparator, the fundamental table is:

Comparator Positive Comparator Negative
CDx Positive True positive / positive agreement cell False positive / negative-discordance cell
CDx Negative False negative / negative-discordance cell True negative / negative agreement cell

Let:

$$ TP,\;FP,\;FN,\;TN $$

represent the four cells.

Then classical sensitivity and specificity are:

$$ Sensitivity= \frac{TP}{TP+FN} $$
$$ Specificity= \frac{TN}{TN+FP} $$

However, when neither method can legitimately be treated as a perfect reference standard, the preferred terminology may be positive percent agreement and negative percent agreement rather than sensitivity and specificity.

Positive Percent Agreement

Positive percent agreement, or PPA, is commonly calculated as:

$$ PPA= \frac{TP}{TP+FN}\times100\% $$

This asks:

$$ \text{Among comparator-positive specimens, how often is the CDx positive?} $$

PPA is therefore numerically identical to sensitivity in the two-by-two framework when the comparator-positive category is used as the denominator.

The distinction is interpretive: the comparator is not necessarily assumed to represent absolute biological truth.

Negative Percent Agreement

Negative percent agreement, or NPA, is:

$$ NPA= \frac{TN}{TN+FP}\times100\% $$

It asks:

$$ \text{Among comparator-negative specimens, how often is the CDx negative?} $$

Overall Percent Agreement

Overall percent agreement can be calculated as:

$$ OPA= \frac{TP+TN}{TP+FP+FN+TN}\times100\% $$

OPA is useful as a descriptive summary, but it can hide asymmetric performance.

For example, an assay could have 98% OPA while having materially different positive and negative agreement.

Best practice: Do not rely on OPA alone for a clinically consequential binary CDx. Report the relevant directional agreement measures and their confidence intervals.

Confidence Intervals Matter

A point estimate such as:

$$ PPA=95\% $$

does not tell the complete statistical story.

Suppose the estimate is based on only 20 comparator-positive specimens.

The uncertainty can be substantial.

A much larger validation dataset may produce the same point estimate with a much narrower confidence interval.

Therefore, CDx validation reports should generally present the point estimate together with an appropriate confidence interval.

Which Confidence Interval?

For binomial proportions, commonly used approaches include:

  • Exact binomial confidence intervals
  • Wilson confidence intervals
  • Other prespecified interval methods appropriate to the design

The ordinary Wald interval is often undesirable when sample sizes are small or proportions are near 0% or 100%.

For a proportion:

$$ \hat p=\frac{x}{n} $$

the interval method should be selected prospectively and consistently.

Illustrative PPA Calculation

Suppose a validation study includes:

$$ TP=92 \qquad FN=8 $$

Then:

$$ PPA= \frac{92}{92+8} = 0.92 = 92\% $$

The important reporting statement is not simply "PPA was 92%."

A complete result should include the numerator, denominator, percentage, and confidence interval.

Sample Size for Agreement Studies

Sample-size planning should reflect the performance claim being evaluated.

For example, suppose the goal is to demonstrate that PPA exceeds a minimum acceptable value:

$$ PPA\ge p_0 $$

A study may be designed around a one-sided hypothesis such as:

$$ H_0:p\le p_0 $$

versus:

$$ H_A:p>p_0 $$

The required number of positive specimens depends on:

  • The minimum acceptable performance
  • The anticipated true performance
  • Type I error
  • Desired power
  • The confidence-bound strategy
  • Expected evaluable specimen availability
Do not calculate sample size only from the total number of patients. For a binary CDx, the number of evaluable positive specimens may determine precision for PPA, while the number of negative specimens determines precision for NPA.

Analytical Precision

Precision describes the closeness of repeated measurements under specified conditions.

Precision can be divided into components such as:

  • Repeatability
  • Within-run precision
  • Between-run precision
  • Between-day precision
  • Between-operator precision
  • Between-instrument precision
  • Between-site reproducibility

The appropriate components depend on the technology and intended use.

Continuous Measurement Precision

For a continuous assay result \(Y\), variance can be decomposed into multiple sources.

$$ \sigma^2_{total} = \sigma^2_{within} + \sigma^2_{between} $$

More complex studies may include additional variance components for operator, instrument, day, lot, site, or other factors.

A variance-components model can therefore be useful when a CDx is intended to operate across multiple laboratories or platforms.

Coefficient of Variation

For positive continuous measurements, the coefficient of variation is often useful:

$$ CV(\%)= \frac{SD}{Mean}\times100 $$

For example, if repeated measurements have:

$$ Mean=100 \qquad SD=5 $$

then:

$$ CV=5\% $$

CV is especially useful when variability is naturally proportional to the magnitude of the measurement.

It is not universally appropriate for every assay, particularly where values can be near zero or where the measurement scale does not support a meaningful ratio interpretation.

Precision Study Design

Factor Illustrative Levels Purpose
Specimen Negative, low positive, moderate positive, high positive Evaluate performance across clinically relevant concentrations
Operator Multiple operators Assess operator-related variability
Day Multiple nonconsecutive days Assess temporal variability
Run Multiple runs Assess run-to-run variation
Lot Multiple reagent lots Assess lot-to-lot variability
Instrument Multiple instruments where applicable Assess platform variability

Reproducibility

Reproducibility asks whether the assay produces consistent results when conditions vary in ways representative of intended use.

For a multicenter CDx, this can be especially important.

A study might deliberately include:

  • Multiple laboratories
  • Multiple operators
  • Multiple reagent lots
  • Multiple instruments
  • Multiple testing days

The statistical analysis should distinguish the experimental design from a simple collection of repeated observations.

Repeated measurements are not automatically independent. If the same specimen is tested repeatedly, those observations share biological and specimen-level characteristics. Treating every replicate as an independent patient-level observation can produce artificially narrow confidence intervals.

Limit of Detection

The limit of detection, or LoD, is the lowest amount or concentration of analyte that can be reliably detected under the defined conditions of the assay.

LoD is especially important for low-frequency molecular biomarkers.

Consider a mutation with variant allele frequency:

$$ VAF=0.5\% $$

The assay may detect it in some replicates but not others.

Therefore, LoD is fundamentally a probability-of-detection problem rather than simply the concentration at which the mean signal becomes nonzero.

Probability of Detection

Suppose an assay is tested at several concentrations.

Concentration Detected / Tested Detection Rate
0.05% 4 / 20 20%
0.10% 12 / 20 60%
0.15% 18 / 20 90%
0.20% 20 / 20 100%

These data illustrate why the LoD should be evaluated across concentrations rather than inferred from one successful measurement.

Figure 2. Illustrative Probability of Detection Near the LoD
Simulated detection probabilities across increasing biomarker concentrations. The curve illustrates the transition from unreliable to highly reliable detection.

Illustrative data only. A real LoD study should use a protocol-defined estimation method and acceptance criterion appropriate to the assay.

LoD and Hit Rates

For qualitative assays, one useful representation is:

$$ P(\text{detected}\mid concentration) $$

The detection probability generally increases as analyte concentration increases.

The validation team may define an LoD at a specified detection probability, provided that the definition is appropriate to the assay and prespecified.

LoQ Is Different From LoD

The limit of quantitation, or LoQ, concerns the lowest concentration at which the analyte can be quantified with predefined performance characteristics.

LoD and LoQ should not be treated as interchangeable.

Concept Question
LoD Can the analyte be reliably detected?
LoQ Can the analyte be quantified with acceptable performance?

Analytical Specificity

Analytical specificity concerns whether the assay detects the intended target without inappropriate interference from other substances or biological features.

Depending on the technology, studies may evaluate:

  • Potential interferents
  • Cross-reactive targets
  • Closely related sequences
  • Endogenous substances
  • Exogenous substances
  • Common medications
  • Matrix effects

Interference Studies

Suppose a potentially interfering substance is present at a clinically relevant concentration.

A useful comparison is:

$$ Difference= Result_{interferent}-Result_{control} $$

For a qualitative assay, the focus may instead be on whether the classification changes.

For example:

Specimen Without Interferent With Interferent Interpretation
Low positive Positive Positive No observed classification effect
Near cutoff Positive Negative Potential interference requiring investigation
Negative Negative Negative No observed false-positive effect

Reportable Range

The reportable range describes the range over which the assay can provide results that meet the defined performance requirements.

For quantitative assays, this may involve:

  • Linearity
  • Accuracy
  • Precision
  • Recovery
  • Bias

For categorical molecular assays, the relevant concept may instead involve the range of analyte concentrations or variant frequencies over which the assay maintains acceptable classification performance.

Linearity

For quantitative measurements, linearity asks whether the observed result changes appropriately with the true or expected concentration.

A simple model is:

$$ Y_i=\beta_0+\beta_1X_i+\epsilon_i $$

where:

  • \(X_i\) = expected or reference concentration
  • \(Y_i\) = observed assay result
  • \(\beta_0\) = intercept
  • \(\beta_1\) = slope

A slope near 1 and intercept near 0 can support agreement, but regression coefficients should not be interpreted in isolation.

Correlation is not agreement. A correlation coefficient can be close to 1 even when the candidate assay has substantial systematic bias. Agreement analyses should therefore be appropriate to the measurement scale and intended use.

Bias

Bias represents systematic deviation from a reference or expected value.

For paired continuous results:

$$ d_i=Y_i-X_i $$

The mean difference is:

$$ \bar d= \frac{1}{n} \sum_{i=1}^{n}d_i $$

A mean difference close to zero suggests little average bias, although the distribution of individual differences must also be considered.

Agreement for Continuous Results

When two methods measure the same continuous quantity, Bland-Altman-style analysis can be informative.

The difference is plotted against the average:

$$ M_i=\frac{X_i+Y_i}{2} $$
$$ D_i=Y_i-X_i $$

Approximate limits of agreement can then be expressed as:

$$ \bar D\pm1.96SD_D $$

The actual analysis should account for the study design and assumptions rather than mechanically applying the formula.

Figure 3. Illustrative Method-Agreement Plot
Simulated paired measurements comparing a candidate assay with a comparator. The horizontal line represents mean difference and the dashed lines represent illustrative limits of agreement.

Agreement plots should be interpreted relative to clinically acceptable differences, not merely according to whether the confidence limits look narrow.

Cutoff Determination

Many companion diagnostics convert a continuous or semi-quantitative assay result into a binary classification.

For example:

$$ CDx= \begin{cases} Positive,&X\ge c\\ Negative,&X

where \(c\) is the classification cutoff.

Cutoff selection is therefore a major statistical issue.

Cutoff Validation Principles

A cutoff should not simply be selected because it produces the most favorable validation results in a convenient dataset.

Important considerations include:

  • Clinical rationale
  • Analytical characteristics
  • Biological plausibility
  • Prespecified development strategy
  • Independent or appropriately controlled evaluation
  • Expected clinical consequences of misclassification
Cutoff overfitting is a major risk. If the cutoff is repeatedly optimized against the same validation specimens used to claim performance, the apparent diagnostic performance may be optimistically biased.

ROC Curves

When a continuous biomarker is used to classify patients, receiver operating characteristic analysis may be useful during development.

The ROC curve displays:

$$ Sensitivity \quad\text{vs.}\quad 1-Specificity $$

across possible thresholds.

The area under the curve, or AUC, summarizes discrimination over the range of thresholds.

However, the AUC does not automatically establish that a selected cutoff is clinically appropriate.

Clinical Concordance

Clinical validation asks a different question from analytical validation.

The central issue is whether the diagnostic identifies the population in which the therapeutic benefit demonstrated in clinical development applies.

For example, suppose a pivotal clinical trial enrolled patients based on biomarker status determined using an investigational assay.

A commercial CDx may need to demonstrate appropriate concordance with that clinical-trial classification.

Bridging a Clinical-Trial Assay to a CDx

A common development problem is that the assay used during clinical development differs from the final commercial diagnostic.

The statistical bridge may involve comparing:

  • Clinical-trial assay results
  • Candidate CDx results
  • Reference or orthogonal assay results
  • Archived clinical specimens

The key question is whether patients classified as biomarker-positive or negative by the final assay correspond appropriately to the population supporting the therapeutic indication.

Positive and Negative Percent Agreement in Bridging

Suppose the clinical-trial assay identifies 150 positive specimens and the candidate CDx identifies 143 of those as positive.

If there are 7 discordant positive specimens:

$$ PPA= \frac{143}{143+7} = 95.3\% $$

The analysis should then investigate the discordant specimens rather than treating the aggregate percentage as the end of the analysis.

Discordant Analysis

Discordant specimens can provide important information.

For every clinically important discordance, consider whether the difference could arise from:

  • Specimen quality
  • Low analyte concentration
  • Heterogeneity
  • Pre-analytical handling
  • Assay failure
  • Cutoff differences
  • Biological differences
  • True disagreement between measurement technologies

A good validation report should not simply list discordances without investigation.

Borderline Specimens

Specimens near the assay cutoff are particularly informative.

Suppose a continuous assay has a positive cutoff of:

$$ X=10 $$

A specimen with:

$$ X=10.01 $$

may be classified positive while another measurement of the same specimen could be negative.

This makes near-cutoff precision and reproducibility especially important.

Figure 4. Illustrative Classification Stability Around a Cutoff
Simulated assay measurements near a predefined cutoff. Borderline specimens can be disproportionately important when evaluating classification robustness.

The figure is illustrative. Actual cutoff evaluation should follow the assay's intended-use and validation strategy.

Clinical Consequences of Misclassification

A useful validation framework distinguishes statistical performance from clinical consequences.

Error Diagnostic Result Potential Clinical Consequence
False positive CDx positive when target is absent Potential treatment of a patient outside the evidence-supported population
False negative CDx negative when target is present Potential exclusion from an indicated or investigational therapy
Invalid result No interpretable result Potential delay, repeat testing, or unresolved treatment decision

The relative importance of these errors depends on the clinical context.

Predictive Values

Positive predictive value is:

$$ PPV= \frac{TP}{TP+FP} $$

Negative predictive value is:

$$ NPV= \frac{TN}{TN+FN} $$

Unlike PPA and NPA, predictive values depend strongly on the prevalence of the biomarker in the tested population.

Do not confuse agreement with predictive value. PPA and NPA describe agreement with a comparator-positive or comparator-negative group. PPV and NPV describe the probability of the reference classification conditional on the CDx result and therefore depend on the underlying population composition.

Prevalence Effects

Consider an assay with:

$$ Sensitivity=95\% \qquad Specificity=95\% $$

If biomarker prevalence is low, the number of false positives can be large relative to true positives.

Therefore, the same assay can have substantially different PPV and NPV in different populations.

Invalid Rates

A CDx validation should explicitly evaluate invalid, indeterminate, or uninterpretable results when such outcomes are possible.

For example:

$$ Invalid\ Rate= \frac{Number\ of\ invalid\ results} {Number\ of\ specimens\ tested} \times100\% $$

An assay with excellent PPA among valid results may still have an important clinical limitation if a substantial fraction of specimens produce no valid result.

First-Pass Success Rate

Another useful metric can be the percentage of specimens producing an interpretable result on the initial testing attempt.

$$ First\ Pass\ Success= \frac{Valid\ first\ attempts} {All\ first\ attempts} \times100\% $$

This can be particularly relevant for assays involving complex specimen processing.

Missing Data and Unevaluable Specimens

Diagnostic validation studies can lose specimens because of:

  • Insufficient specimen quantity
  • Degraded material
  • Failed extraction
  • Instrument failure
  • Invalid control results
  • Missing comparator results
  • Inadequate specimen quality

The analysis should distinguish:

  • Not tested
  • Test attempted but invalid
  • Valid negative
  • Valid positive
  • Unevaluable for a particular endpoint

These categories should not be silently collapsed into a single "missing" category.

Repeat Testing

Repeated testing after an initial invalid result introduces another important statistical issue.

Suppose:

  • 100 specimens are initially tested.
  • 10 produce invalid results.
  • 8 of the 10 become valid after repeat testing.

Reporting only the final valid results could make the assay appear to have 100% valid performance among the retained specimens.

The initial invalid rate remains clinically relevant.

Always preserve the testing pathway. The denominator used for analytical performance should be explicitly defined. Repeated specimens should not be allowed to disappear from the analysis merely because a later attempt generated a valid result.

Robustness

Robustness examines whether modest, realistic variations in operating conditions materially affect assay performance.

Examples include:

  • Incubation time
  • Temperature
  • Input amount
  • Extraction conditions
  • Storage duration
  • Reagent handling
  • Instrument settings

The purpose is not to intentionally break the assay.

The purpose is to determine whether reasonable deviations within the intended operating environment cause unacceptable performance changes.

Specimen Stability

Specimen stability can be evaluated at clinically relevant storage conditions.

For example:

Condition Time Primary Question
Room temperature Defined interval Does expected handling affect classification?
Refrigerated Defined interval Is refrigerated storage acceptable?
Frozen Defined interval Does long-term storage preserve performance?
Freeze-thaw Multiple cycles Does repeated handling materially affect results?

Lot-to-Lot Consistency

If reagent lots can change over time, the validation strategy should evaluate whether lot-to-lot variation remains within predefined requirements.

For a quantitative assay, this may involve estimating:

$$ Difference_{lot} = Result_{Lot\,B}-Result_{Lot\,A} $$

For a qualitative assay, classification agreement may be more appropriate.

Reader Agreement

Some CDx technologies involve subjective interpretation.

Immunohistochemistry is a common example where trained readers may assign categorical scores.

Possible analyses include:

  • Raw agreement
  • Positive agreement
  • Negative agreement
  • Inter-reader agreement
  • Intra-reader agreement
  • Kappa statistics where appropriate

Kappa Statistics

For categorical assessments, Cohen's kappa is:

$$ \kappa= \frac{P_o-P_e}{1-P_e} $$

where:

  • \(P_o\) = observed agreement
  • \(P_e\) = agreement expected by chance under the specified model

Kappa can be useful, but it should not replace clinically meaningful agreement measures.

In situations with highly imbalanced prevalence, kappa can behave in ways that make interpretation difficult.

Multi-Category CDx Results

Not all CDx classifications are binary.

An assay might classify specimens as:

  • Negative
  • Low positive
  • Intermediate
  • High positive

For ordered categories, weighted agreement methods may be appropriate.

The analysis should preserve the ordinal structure rather than automatically reducing everything to positive versus negative.

NGS-Based Companion Diagnostics

Next-generation sequencing introduces additional statistical considerations.

Potential validation dimensions include:

  • Variant detection
  • Variant type
  • Allele fraction
  • Coverage
  • Read quality
  • Sample quality
  • False-positive variant calls
  • False-negative variant calls
  • Pipeline reproducibility
  • Reference genome and bioinformatics configuration

The statistical unit may differ depending on whether the endpoint is:

  • Specimen-level classification
  • Variant-level detection
  • Patient-level biomarker status
Define the unit of analysis before calculating agreement. Counting thousands of variants from a small number of patients does not create thousands of independent clinical specimens. Variant-level and specimen-level analyses answer different questions.

Variant-Level vs. Patient-Level Agreement

Unit Example Question
Patient Is this patient's biomarker status correctly classified?
Specimen Does the assay correctly classify this tested specimen?
Variant Does the assay detect this specific molecular alteration?
Call Does the pipeline generate the correct individual result?

Nested Data and Correlation

When multiple measurements come from the same specimen or patient, the data are clustered.

For example:

$$ Patient \rightarrow Specimen \rightarrow Replicate $$

Observations within the same hierarchy are not necessarily independent.

Possible analytical approaches include:

  • Mixed-effects models
  • Variance-component models
  • Generalized estimating equations
  • Cluster-aware confidence intervals
  • Specimen-level summaries

The method should follow the estimand and experimental design.

Analytical Sensitivity vs. Clinical Sensitivity

These terms are often confused.

Term Meaning
Analytical sensitivity Ability to detect the analyte at low concentrations or low abundance
Clinical sensitivity Ability to identify patients in the clinically relevant positive population

A test can have excellent analytical sensitivity while still having imperfect clinical sensitivity because clinical specimens introduce biological heterogeneity and pre-analytical variation.

Pre-Analytical Variables

The assay does not operate in isolation.

The full diagnostic workflow can be represented as:

1
Patient selection
2
Specimen collection
3
Transport and storage
4
Sample preparation
5
Analytical measurement
6
Result interpretation
7
Clinical classification

Validation should consider the portions of this pathway included in the intended use.

Sample Quality

For tissue-based diagnostics, sample quality can materially influence performance.

Potential factors include:

  • Tumor content
  • Necrosis
  • DNA or RNA integrity
  • Fixation quality
  • Specimen age
  • Section thickness
  • Extraction yield

These variables may need to be incorporated into analytical validation or eligibility criteria.

Statistical Acceptance Criteria

A validation study should not be designed around vague goals such as "demonstrate good performance."

Instead, each critical performance characteristic should have a predefined acceptance criterion.

Performance Characteristic Example Acceptance Framework
PPA Lower confidence bound exceeds a prespecified minimum
NPA Lower confidence bound exceeds a prespecified minimum
Precision Variance or agreement remains within predefined limits
LoD Detection probability meets the predefined criterion
Interference No clinically meaningful classification or measurement effect
Stability Performance remains within predefined limits over intended conditions

Point Estimate vs. Confidence-Bound Criterion

Suppose the acceptance criterion is:

$$ PPA\ge90\% $$

A study might observe:

$$ PPA=94\% $$

But if the lower confidence bound is below 90%, the prespecified success rule may not be met.

Therefore:

$$ \text{Observed estimate} \neq \text{Evidence of meeting the acceptance criterion} $$

The inferential rule must be defined in advance.

Multiple Performance Characteristics

A CDx validation program can contain many endpoints.

For example:

  • PPA
  • NPA
  • Precision
  • LoD
  • Linearity
  • Interference
  • Stability
  • Reproducibility

Not every performance characteristic should automatically be treated as a formal hypothesis test requiring multiplicity adjustment.

The statistical framework should distinguish:

  • Descriptive performance characterization
  • Formal acceptance criteria
  • Primary validation claims
  • Secondary supporting analyses
Do not apply multiplicity mechanically. The appropriate treatment of multiple endpoints depends on the validation objective, regulatory strategy, and prespecified decision framework.

Study Design Matrix

A useful planning tool is a validation matrix.

Question Study Type Typical Data Statistical Output
Can the assay detect low-level target? LoD Replicate low-concentration specimens Detection probability
Is repeated testing consistent? Precision Replicate results SD, CV, agreement
Does the assay agree with comparator? Agreement Paired specimens PPA, NPA, OPA, CI
Does performance vary by operator/site? Reproducibility Multifactorial replicates Variance components/agreement
Does storage affect results? Stability Paired stored and reference specimens Difference/agreement
Do substances interfere? Interference Spiked specimens Classification or quantitative effect

Validation Hierarchy

A useful conceptual hierarchy is:

A
Can the assay detect the analyte?
B
Can it measure or classify the analyte reproducibly?
C
Does it agree with an appropriate comparator?
D
Does it reproduce across intended operating conditions?
E
Does its classification support the intended clinical use?

Clinical Cutoff vs. Analytical Cutoff

A subtle but important distinction is the difference between an analytical detection threshold and a clinical classification threshold.

For example, an assay might reliably detect a molecular signal at:

$$ VAF=0.1\% $$

while the clinically validated positive classification cutoff may be:

$$ VAF=1.0\% $$

These are not contradictory.

The assay can detect lower levels than are required for the clinical classification.

Precision Near the Clinical Cutoff

Near-cutoff precision is often more consequential than precision far from the cutoff.

Suppose:

$$ Cutoff=10 $$

and repeated measurements are:

$$ 9.7,\;10.2,\;9.8,\;10.1,\;9.9 $$

The assay may have excellent precision in absolute terms but still produce classification switching around the decision boundary.

Therefore, validation should include specimens representative of the decision region.

Classification Concordance

For a binary CDx, classification concordance can be expressed as:

$$ Concordance= \frac{\text{Number of concordant classifications}} {\text{Total evaluable classifications}} $$

But concordance should be accompanied by the directional components.

Measure Why It Matters
PPA Captures agreement in the positive population
NPA Captures agreement in the negative population
OPA Provides overall descriptive agreement
Discordant analysis Identifies the causes and clinical relevance of disagreement

Reference Method Limitations

A comparator method is not automatically a perfect gold standard.

This is particularly important for biomarkers where biological truth cannot be directly observed.

If two assays use different technologies, a disagreement may reflect:

  • Technical error by one assay
  • Technical error by the other assay
  • Differences in analytical sensitivity
  • Different target definitions
  • Biological heterogeneity

Therefore, an orthogonal method or adjudication strategy may sometimes be useful for discordant cases.

Orthogonal Testing

An orthogonal assay uses a different analytical principle to investigate a result.

For example, an NGS finding could potentially be investigated using an independent molecular method where appropriate.

The purpose is not necessarily to declare one assay correct automatically.

Instead, orthogonal evidence can help determine whether a discordance reflects a genuine biological signal or an analytical limitation.

Clinical Specimen Diversity

A strong validation program should reflect the intended clinical population and specimen environment.

Depending on the indication, this may include variation in:

  • Disease stage
  • Specimen type
  • Tumor content
  • Biomarker abundance
  • Demographic and clinical characteristics
  • Specimen age
  • Collection site

The precise sampling strategy depends on the intended-use population and development program.

Rare Biomarkers

Rare biomarker populations create a special statistical challenge.

Suppose the biomarker prevalence is only 5%.

A general population sample of 1,000 patients would contain approximately:

$$ 1000\times0.05=50 $$

positive patients on average.

If the precision of PPA is important, the study may therefore require intentional enrichment or another appropriate specimen-acquisition strategy, provided that the resulting analysis remains valid for the intended claim.

Enrichment Does Not Change the Mathematics of PPA

If the study intentionally enriches positive specimens, PPA and NPA remain defined using their respective denominators.

However, predictive values estimated from the enriched sample would generally not represent the predictive values expected in a routine population unless appropriately adjusted.

Important distinction: Enrichment can be highly useful for estimating agreement efficiently, but prevalence-dependent measures require careful interpretation when the study population does not reflect the intended-use population.

Statistical Analysis Plan

The statistical analysis plan for a CDx validation study should define at least:

  • Analysis populations
  • Primary endpoints
  • Secondary endpoints
  • Specimen eligibility
  • Handling of invalid results
  • Handling of repeat testing
  • Comparator definition
  • Cutoff definition
  • Confidence interval methodology
  • Missing-data handling
  • Outlier handling
  • Acceptance criteria
  • Subgroup analyses
  • Discordant-result investigations

Analysis Populations

A useful framework may distinguish:

Population Description
All specimens All specimens entering the validation workflow
Tested specimens Specimens for which testing was attempted
Valid specimens Specimens producing interpretable results
Paired-evaluable specimens Specimens with interpretable results from both candidate and comparator methods

Different endpoints may use different denominators.

That is acceptable as long as the definitions are explicit.

Subgroup Analysis

CDx performance can vary across clinically relevant specimen characteristics.

Potential subgroups may include:

  • Specimen type
  • Tumor content
  • Biomarker abundance
  • Site
  • Operator
  • Instrument
  • Reagent lot

Subgroup analyses should generally be prespecified where they are intended to support a validation claim.

Small Subgroups

A subgroup with only a few positive specimens can produce an apparently extreme PPA.

For example:

$$ 5/5=100\% $$

does not mean the subgroup's true performance is known with certainty.

The confidence interval is essential.

Bootstrapping and Resampling

For some complex validation statistics, resampling methods may provide useful uncertainty estimates.

However, resampling should respect the data hierarchy.

If five replicates come from one specimen, randomly resampling individual replicates as though they were independent patient observations can be misleading.

A cluster-aware bootstrap may instead resample at the specimen level.

Power and Precision

Validation studies are often framed around acceptance criteria rather than traditional superiority hypotheses.

Consequently, sample-size planning may focus on the width of a confidence interval.

For example, the study team may require:

$$ Width(CI)\le W $$

where \(W\) is the maximum acceptable uncertainty.

This is particularly intuitive for agreement estimates.

Example of Precision-Based Planning

Suppose investigators expect:

$$ PPA\approx95\% $$

and want a reasonably narrow confidence interval around the estimate.

The number of positive specimens required will generally be driven much more by the desired precision than by the total number of specimens alone.

This illustrates a useful planning principle:

Plan the number of informative specimens, not just the total number of specimens. For PPA, positive specimens are informative about positive agreement. For NPA, negative specimens are informative about negative agreement.

Validation of a Quantitative Biomarker

Suppose the CDx reports a quantitative biomarker value.

A comprehensive analytical validation might examine:

  1. Accuracy
  2. Precision
  3. Linearity
  4. Reportable range
  5. Analytical sensitivity
  6. Analytical specificity
  7. Interference
  8. Stability
  9. Carryover
  10. Robustness

The final clinical classification may then be based on a prespecified cutoff.

Validation of a Qualitative Biomarker

For a qualitative CDx, the core package may emphasize:

  • Positive agreement
  • Negative agreement
  • Overall agreement
  • Precision
  • Reproducibility
  • LoD
  • Interference
  • Cross-reactivity
  • Invalid rate
  • Clinical concordance

The statistical methods should match the categorical nature of the endpoint.

Example CDx Validation Dataset

SPECIMEN   CDX   COMP   RESULT
001        +     +      Concordant Positive
002        +     +      Concordant Positive
003        +     -      Discordant
004        -     -      Concordant Negative
005        -     +      Discordant
006        -     -      Concordant Negative
007        +     +      Concordant Positive
008        -     -      Concordant Negative

From the table, the analysis can derive the four classification cells.

In real validation programming, the data structure would generally contain additional variables for specimen source, assay run, lot, operator, site, instrument, biomarker level, validity status, and other protocol-defined attributes.

R Implementation: PPA and NPA

tab <- table(
  CDx = data$CDX,
  Comparator = data$COMP
)

TP <- tab["Positive", "Positive"]
FP <- tab["Positive", "Negative"]
FN <- tab["Negative", "Positive"]
TN <- tab["Negative", "Negative"]

PPA <- TP / (TP + FN)
NPA <- TN / (TN + FP)
OPA <- (TP + TN) / sum(tab)

results <- data.frame(
  PPA = PPA,
  NPA = NPA,
  OPA = OPA
)

In production programming, explicit handling of missing cells, zero denominators, factor levels, and invalid specimens should be added.

R Implementation: Wilson Confidence Interval

wilson_ci <- function(x, n, z = 1.96) {

  p <- x / n

  denominator <- 1 + z^2 / n

  center <- (
    p + z^2 / (2 * n)
  ) / denominator

  half_width <- (
    z * sqrt(
      p * (1 - p) / n +
      z^2 / (4 * n^2)
    )
  ) / denominator

  c(
    lower = center - half_width,
    upper = center + half_width
  )
}

wilson_ci(TP, TP + FN)

For a regulated validation package, the chosen confidence-interval method should be specified in the statistical analysis plan and implemented consistently.

R Implementation: LoD Detection Probability

If replicate results are coded as detected versus not detected, a logistic model can be useful for exploratory estimation of the detection curve.

lod_model <- glm(
  detected ~ log10(concentration),
  data = lod_data,
  family = binomial()
)

summary(lod_model)

pred <- predict(
  lod_model,
  type = "response"
)

The exact LoD estimation approach should be selected according to the assay, study design, protocol, and intended claim.

R Implementation: Bland-Altman Data

data$mean_method <- (
  data$reference +
  data$candidate
) / 2

data$difference <- (
  data$candidate -
  data$reference
)

mean_difference <- mean(
  data$difference,
  na.rm = TRUE
)

sd_difference <- sd(
  data$difference,
  na.rm = TRUE
)

upper_loa <-
  mean_difference +
  1.96 * sd_difference

lower_loa <-
  mean_difference -
  1.96 * sd_difference

The resulting values can be plotted against the paired means.

Validation Programming QC

Statistical validation requires more than checking whether the final percentages look plausible.

Independent QC should verify:

  • Specimen inclusion
  • Comparator classification
  • Candidate classification
  • Cutoff derivation
  • Invalid handling
  • Repeat-testing logic
  • Numerators
  • Denominators
  • Confidence intervals
  • Precision estimates
  • LoD derivations
  • Subgroup definitions
  • Figure values

Patient-Level Traceability

Every summary statistic should be traceable back to specimen-level records.

For example:

1
Start with the reported PPA.
2
Trace the numerator to positive concordant specimens.
3
Trace the denominator to all eligible comparator-positive specimens.
4
Verify inclusion and exclusion decisions.
5
Independently reproduce the percentage and confidence interval.

Common Statistical Mistakes

  1. Calling a comparator the gold standard without justification. Agreement with a comparator is not necessarily equivalent to truth.
  2. Reporting only overall agreement. PPA and NPA can reveal asymmetric performance hidden by OPA.
  3. Ignoring confidence intervals. A point estimate does not quantify uncertainty.
  4. Treating replicates as independent patients. Repeated measurements within a specimen are correlated.
  5. Ignoring invalid results. Invalid rates can have direct clinical relevance.
  6. Using post hoc cutoff optimization. This can create optimistic estimates of performance.
  7. Confusing LoD with clinical sensitivity. Analytical detection capability is not the same as patient-level diagnostic performance.
  8. Using correlation as a substitute for agreement. High correlation can coexist with systematic bias.
  9. Ignoring rare positive populations. A large total sample does not guarantee precise PPA.
  10. Ignoring discordant specimens. Discordances can reveal important assay limitations.
  11. Pooling heterogeneous specimen types without justification. Different matrices can behave differently.
  12. Using inappropriate denominators. The denominator must correspond to the defined estimand.

Regulatory Perspective

Companion diagnostic development is closely linked to the therapeutic development program.

The statistical validation strategy should therefore be consistent with:

  • The intended use of the diagnostic
  • The biomarker definition
  • The therapeutic indication
  • The clinical-trial enrollment strategy
  • The clinical evidence supporting the treatment
  • The analytical characteristics of the final assay

In the United States, FDA guidance concerning in vitro companion diagnostic devices emphasizes the close relationship between the diagnostic and the corresponding therapeutic product.

The FDA has also issued guidance addressing analytical validation considerations for certain next-generation sequencing-based in vitro diagnostics.

Regulatory principle: The statistical analysis should support a clearly defined regulatory claim. A technically impressive collection of experiments is not necessarily a complete validation package if the evidence does not establish performance for the intended use.

CDx and the Drug Development Program

The diagnostic and therapeutic programs should generally be considered as an integrated development strategy.

The relationship can be represented as:

1
Biomarker hypothesis
2
Clinical assay development
3
Biomarker-defined clinical enrollment
4
Therapeutic efficacy evidence
5
Final CDx analytical validation
6
Clinical bridging/concordance
7
Integrated diagnostic and therapeutic labeling strategy

Bridging Studies

A bridging study may be necessary when the assay used to select patients in clinical development differs from the final commercial CDx.

The statistical question is whether the final assay reproduces the clinically meaningful classification sufficiently well.

Potential endpoints include:

  • PPA
  • NPA
  • OPA
  • Classification consistency
  • Discordance rates
  • Clinical outcome concordance where applicable

Clinical Outcome Concordance

In some programs, the strongest question is not merely:

$$ CDx\;result \leftrightarrow Comparator\;result $$

but rather:

$$ CDx\;result \leftrightarrow Treatment\;benefit $$

This is a fundamentally different level of evidence.

A diagnostic can agree extremely well with another assay while neither assay adequately identifies the population that benefits from treatment.

Analytical agreement does not prove clinical utility. Clinical utility requires evidence connecting the diagnostic classification to the clinical decision or treatment effect for which the diagnostic is intended.

Diagnostic Performance vs. Clinical Utility

Question Evidence Type
Does the assay detect the biomarker? Analytical validation
Does it reproduce the result? Precision / reproducibility
Does it agree with another method? Method comparison
Does it classify the clinical population consistently? Clinical concordance / bridging
Does biomarker-positive treatment produce benefit? Clinical trial evidence

Subgroup Performance

A CDx can have excellent overall performance while exhibiting weaker performance in a clinically important subgroup.

For example, performance could differ by:

  • Low versus high biomarker abundance
  • Specimen type
  • Fresh versus archived tissue
  • Primary versus metastatic tissue
  • Different collection sites

Subgroup analyses should therefore be planned around plausible sources of heterogeneity.

Interaction Between Biomarker Level and Error

Diagnostic error is often not constant across the measurement range.

For example:

$$ P(\text{correct classification}) = f(\text{biomarker abundance}) $$

The probability of correct classification may be highest for very strong positives and strong negatives, while specimens near the cutoff may have greater uncertainty.

Visualizing Performance Across the Measurement Range

Figure 5. Illustrative Performance Across Biomarker Abundance
Simulated classification performance across increasing biomarker levels. Performance is intentionally shown as strongest away from the decision boundary and less certain near the cutoff.

Illustrative only. The actual relationship between biomarker level and classification performance depends on assay technology and biological context.

Clinical Specimen Testing Strategy

A practical validation program should include specimens spanning the relevant range of biomarker status.

For a binary biomarker, this may include:

  • Strong positives
  • Moderate positives
  • Near-cutoff positives
  • Near-cutoff negatives
  • Strong negatives

This design helps distinguish broad assay performance from decision-boundary performance.

Precision and Reproducibility Near LoD

Low-level specimens can simultaneously challenge:

  • Analytical sensitivity
  • Precision
  • Classification stability

Therefore, LoD studies and precision studies may provide complementary evidence.

Carryover

For some assays, carryover from a high-positive specimen into a subsequent negative specimen can create false-positive results.

A typical design may alternate high-positive and negative specimens.

The key question is whether the negative specimens remain negative under the specified conditions.

Cross-Reactivity

Cross-reactivity evaluates whether related analytes or sequences produce inappropriate positive signals.

This can be particularly important for molecular assays targeting closely related variants or organisms.

Analytical Validation Summary Matrix

Characteristic Typical Statistical Question
Accuracy How close is the result to an appropriate reference or comparator?
Precision How variable are repeated results?
Reproducibility Does performance remain consistent across intended sources of variation?
LoD At what low concentration is detection sufficiently reliable?
Linearity Does response track concentration appropriately?
Specificity Does the assay avoid inappropriate detection?
Interference Do relevant substances alter results materially?
Stability Does specimen handling preserve performance?
Robustness Does modest operational variation affect performance?

Clinical Validation Summary Matrix

Characteristic Typical Statistical Question
PPA How often does the CDx classify comparator-positive specimens as positive?
NPA How often does the CDx classify comparator-negative specimens as negative?
OPA What proportion of paired results agree overall?
Discordance Where and why do methods disagree?
Clinical concordance Does the CDx reproduce clinically meaningful biomarker classification?
Clinical utility Does the diagnostic support the treatment decision for which it is intended?

Validation Reporting

A final validation report should allow an independent reviewer to understand:

  • What was tested
  • Why it was tested
  • How specimens were selected
  • How results were classified
  • Which specimens were excluded
  • How invalids were handled
  • How repeat testing was handled
  • Which statistical methods were used
  • Which acceptance criteria were prespecified
  • Whether the criteria were met
  • What limitations remain

Traceability From Requirement to Result

A strong validation package can be organized as:

1
Intended-use requirement
2
Performance characteristic
3
Study design
4
Statistical endpoint
5
Acceptance criterion
6
Observed result
7
Confidence interval / uncertainty
8
Conclusion and clinical interpretation

Example Validation Table

Endpoint Result 95% CI Acceptance Criterion Conclusion
PPA 95.2% 91.8–97.4% Lower bound ≥90% Met
NPA 97.1% 94.6–98.5% Lower bound ≥95% Met
Invalid rate 1.4% — ≤3% Met
Repeatability CV 4.8% — ≤7% Met
LoD 0.12% VAF — ≤0.20% VAF Met

The numerical values above are illustrative rather than regulatory requirements.

How to Interpret a Validation Failure

Not every failure means that the entire assay is unusable.

A failed acceptance criterion should trigger investigation.

Potential explanations include:

  • Unexpected specimen variability
  • Insufficient sample size
  • Incorrect design assumptions
  • Assay instability
  • Unexpected interference
  • Lot effects
  • Operator effects
  • Inappropriate cutoff
  • Data-processing error

The appropriate response depends on the specific failure and its clinical importance.

Statistical Validation Is Not the Same as Statistical Significance

A common misconception is that validation means obtaining statistically significant p-values.

That is often not the central objective.

The real question may be whether performance is sufficiently precise and reliable to support a predefined claim.

For example:

$$ Lower\ Confidence\ Bound \ge Minimum\ Acceptable\ Performance $$

can be more informative than:

$$ p<0.05 $$

for an assay-performance claim.

What Makes a Strong CDx Statistical Package?

A strong package has five characteristics.

1. Prespecified

The endpoints, estimands, acceptance criteria, and analysis methods are defined before examining the final validation results.

2. Representative

The specimens and operating conditions represent the intended use.

3. Quantitative

Performance is expressed with appropriate estimates and uncertainty.

4. Traceable

Every summary result can be traced to the underlying specimen-level data.

5. Clinically connected

Analytical performance is ultimately connected to the clinical decision the diagnostic is intended to support.

Practical CDx Validation Checklist

1
Define the intended use.
2
Define the biomarker and clinical classification.
3
Define the specimen type and intended workflow.
4
Define the analytical performance characteristics.
5
Define comparator/reference methods where applicable.
6
Define the clinical cutoff.
7
Prespecify statistical acceptance criteria.
8
Calculate sample size based on the informative endpoint.
9
Evaluate positive and negative agreement.
10
Evaluate precision and reproducibility.
11
Evaluate LoD and relevant analytical sensitivity.
12
Evaluate interference, specificity, and robustness.
13
Evaluate invalid and repeat-testing rates.
14
Investigate discordant and borderline specimens.
15
Connect analytical performance to clinical concordance.
16
Perform independent statistical programming QC.

Common Reporting Language

A concise statistical conclusion might state:

Example: The companion diagnostic demonstrated high agreement with the prespecified comparator method. Positive percent agreement and negative percent agreement were estimated with two-sided 95% confidence intervals using the prespecified binomial interval method. Analytical precision was evaluated across relevant specimen levels and sources of variation. Additional studies evaluated analytical sensitivity, specificity, interference, specimen stability, and reproducibility. Discordant and invalid results were investigated according to the predefined analysis plan. All numerical values in this example are illustrative.

What the Final Statistical Report Should Contain

Section Content
Objective Intended use and validation claims
Study design Specimens, operators, sites, lots, instruments, conditions
Analysis populations Definitions of evaluable, valid, invalid, and paired specimens
Methods Statistical procedures and confidence intervals
Results Performance estimates and uncertainty
Discordance Detailed investigation of discrepant specimens
Subgroups Clinically relevant performance stratification
Limitations Known limitations and residual uncertainty
Conclusion Whether predefined validation criteria were satisfied

Final Perspective

Companion diagnostic statistical validation sits at the intersection of laboratory science, biostatistics, clinical development, and regulatory science.

The central statistical challenge is not simply calculating a sensitivity, specificity, or coefficient of variation.

The deeper question is whether the evidence establishes that the diagnostic can make the intended clinical classification reliably enough for its purpose.

That requires a chain of evidence.

$$ \text{Analytical Capability} \rightarrow \text{Reproducibility} \rightarrow \text{Agreement} \rightarrow \text{Clinical Concordance} \rightarrow \text{Clinical Decision} $$

Each link answers a different question.

An assay may be analytically precise but clinically irrelevant.

An assay may agree well with a comparator but perform poorly near the therapeutic decision threshold.

An assay may have excellent PPA but an unacceptable invalid rate.

A large study may produce a precise overall estimate while providing too few positive specimens to establish a reliable PPA.

These distinctions are why companion diagnostic validation requires careful statistical planning before the laboratory study begins.

Bottom line: Companion diagnostic statistical validation is a structured evidence program, not a single statistical analysis. Analytical validation establishes that the assay can measure or classify the biomarker reliably. Precision and reproducibility establish consistency across repeated and varied conditions. LoD, specificity, interference, stability, and robustness define important operating characteristics. Agreement and clinical concordance establish how the final diagnostic relates to comparator classifications and the clinical development population. Throughout the program, confidence intervals, appropriate denominators, prespecified acceptance criteria, specimen-level traceability, and careful handling of invalid and discordant results are essential. The ultimate objective is to demonstrate that the companion diagnostic is fit for its intended clinical role.

References

U.S. Food and Drug Administration. In Vitro Companion Diagnostic Devices: Guidance for Industry and Food and Drug Administration Staff. 2014.

U.S. Food and Drug Administration. Considerations for Design, Development, and Analytical Validation of Next Generation Sequencing (NGS)-Based In Vitro Diagnostics (IVDs) Intended to Aid in the Diagnosis of Cancer.

U.S. Food and Drug Administration. Statistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests.

Clinical and Laboratory Standards Institute. Evaluation of Precision of Quantitative Measurement Procedures. CLSI EP05.

Clinical and Laboratory Standards Institute. Evaluation of Detection Capability for Clinical Laboratory Measurement Procedures. CLSI EP17.

Clinical and Laboratory Standards Institute. Measurement Procedure Comparison and Bias Estimation Using Patient Samples. CLSI EP09.

Clinical and Laboratory Standards Institute. Evaluation of Qualitative, Binary Output Examination Performance. CLSI EP12.

Bland, J.M. and Altman, D.G. Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet, 1986.

McHugh, M.L. Interrater reliability: the kappa statistic. Biochemia Medica, 2012.