Tutorials › Biostatistics › Maximum Tolerated Dose (MTD) Determination

Phase I & Dose-Escalation Design

Maximum Tolerated Dose (MTD) Determination: A Worked Example

A practical guide to determining the maximum tolerated dose in Phase I clinical trials, including dose-limiting toxicities, the traditional 3+3 design, dose-escalation rules, MTD selection, statistical interpretation, and modern model-based alternatives.

Advanced 18 min read

What You'll Learn

  • What the MTD represents in a Phase I dose-escalation study
  • How dose-limiting toxicities (DLTs) define the escalation problem
  • How the traditional 3+3 design makes dose-escalation decisions
  • How to work through a complete dose-escalation example
  • Why the observed MTD is not the same thing as the true biological toxicity threshold
  • How CRM, BOIN, mTPI, and other model-based designs differ from 3+3

Introduction

Early-phase clinical trials must answer a fundamental question before a new treatment can be studied at larger doses: how much drug can be administered with an acceptable level of toxicity?

In a conventional Phase I oncology trial, patients are assigned to progressively higher dose levels while investigators monitor for dose-limiting toxicities (DLTs). The goal is not usually to estimate efficacy. Instead, the study is primarily concerned with characterizing safety and selecting a dose or dose range suitable for subsequent development.

The term maximum tolerated dose (MTD) generally refers to the highest dose associated with an acceptable probability of severe or clinically important toxicity, according to the rules specified in the protocol.

Key idea: MTD determination is a dose-escalation problem, not simply a search for the highest dose at which no adverse events occur. The design must balance information about toxicity against the ethical requirement to avoid exposing patients to doses that are excessively toxic.

What Is a Dose-Limiting Toxicity?

A dose-limiting toxicity, or DLT, is a prespecified toxicity that is considered sufficiently serious to prevent further escalation or otherwise influence the dose-escalation decision.

The protocol must define the DLT criteria before the study begins. Depending on the disease, drug, treatment schedule, and therapeutic context, DLT criteria may involve hematologic toxicity, nonhematologic toxicity, organ toxicity, prolonged treatment interruption, hospitalization, or other clinically important events.

For statistical purposes, the essential feature is that each evaluable patient contributes a binary outcome during the specified DLT observation period:

$$ Y_i = \begin{cases} 1, & \text{DLT observed}\\ 0, & \text{DLT not observed} \end{cases} $$

If the true probability of a DLT at dose \(d\) is \(p(d)\), then for \(n\) patients treated at that dose:

$$ X_d\sim\operatorname{Binomial}(n,p(d)) $$

where \(X_d\) is the number of patients experiencing a DLT.

The Target Toxicity Rate

Modern statistical dose-escalation designs are often formulated around a target DLT probability. Let

$$ \phi = \text{target probability of a DLT}. $$

For example, a protocol might define a target toxicity probability of 25%:

$$ \phi=0.25. $$

The conceptual objective is then to identify the dose whose toxicity probability is closest to the target:

$$ d^*=\arg\min_d |p(d)-\phi|. $$

This formulation is particularly natural for model-based designs such as the continual reassessment method (CRM) and interval-based designs such as BOIN. The traditional 3+3 design is different: it does not explicitly estimate \(p(d)\) or select the dose closest to a numerical target.

Important distinction: The phrase “MTD” can refer to different operational definitions. In a 3+3 study it is usually a dose selected by the protocol's escalation rules. In a model-based design, it is commonly defined as the dose whose estimated DLT probability is closest to a prespecified target.

Why Dose Escalation Is Necessary

Before the first patient is treated, the dose-toxicity relationship is usually uncertain. Preclinical studies, pharmacology, toxicology, and prior clinical experience may provide useful information, but the probability of severe toxicity in humans is not known with certainty.

Consequently, a Phase I trial typically evaluates a sequence of dose levels:

1
Start at an appropriate initial dose: the starting dose is selected using the protocol's preclinical and clinical safety rationale.
2
Observe toxicity: patients are monitored during the prespecified DLT evaluation period.
3
Make a dose decision: the observed DLT experience determines whether to escalate, remain at the current dose, de-escalate, or stop.
4
Repeat: the process continues until the protocol's stopping and dose-selection criteria are satisfied.

The Traditional 3+3 Design

The 3+3 design is one of the most widely recognized traditional dose-escalation designs in oncology. It enrolls patients in cohorts of three and uses the number of DLTs observed within a cohort to determine the next action.

A basic version of the design can be summarized as follows.

DLTs in first 3 patientsDecision
0 DLTsEscalate to the next dose
1 DLTExpand the cohort at the same dose to 6 patients
≥2 DLTsStop escalation; the current dose is too toxic under the design

When a cohort is expanded to six patients, the conventional rules are:

Total DLTs among 6 patientsDecision
≤1 DLTEscalate to the next dose
≥2 DLTsStop escalation; select the previous dose as the MTD, subject to the protocol's final criteria
Why “3+3”? The name comes directly from the basic cohort structure: three patients are treated initially, and a cohort is expanded to six when exactly one DLT is observed.

Formalizing the 3+3 Decision Rules

Let \(d_j\) denote dose level \(j\), and let \(X_j\) be the number of DLTs observed among the initial three patients at that dose.

The initial decision rule can be written:

$$ X_j=0 \Rightarrow d_{j+1}\text{ is evaluated} $$
$$ X_j=1 \Rightarrow \text{treat 3 additional patients at }d_j $$
$$ X_j\ge2 \Rightarrow \text{stop escalation}. $$

After six patients have been evaluated following a one-DLT initial cohort:

$$ X_j\le1 \Rightarrow \text{escalate} $$
$$ X_j\ge2 \Rightarrow \text{stop escalation}. $$

These rules are intentionally simple. No likelihood model is fit, and no formal posterior distribution is required to determine the next dose.

A Complete Worked Example

Consider a hypothetical Phase I oncology study evaluating an investigational agent at six planned dose levels. The protocol uses a conventional 3+3 design.

Dose LevelDoseInitial CohortObserved DLTsAction
110 mg30Escalate
220 mg30Escalate
340 mg31Expand to 6
340 mg6 total1 totalEscalate
460 mg31Expand to 6
460 mg6 total2 totalStop escalation

The key question is now: which dose is selected as the MTD?

Because escalation was stopped at 60 mg after two DLTs among six patients, the conventional 3+3 decision is to select the previous dose, 40 mg, as the MTD, provided the protocol's final dose-selection and safety criteria are satisfied.

$$ \boxed{\text{Selected MTD}=40\text{ mg}} $$

Step 1: Evaluate 10 mg

Three patients receive 10 mg. Suppose none experiences a DLT:

$$ X_1=0. $$

The 3+3 rule therefore permits escalation:

$$ 0\text{ DLTs among 3}\Rightarrow\text{escalate to 20 mg}. $$

Step 2: Evaluate 20 mg

Three patients receive 20 mg, again with no DLTs:

$$ X_2=0. $$

The trial escalates to 40 mg.

Step 3: Evaluate 40 mg

Three patients receive 40 mg. One patient experiences a protocol-defined DLT:

$$ X_3=1. $$

The trial does not immediately escalate. Instead, the cohort is expanded from three to six patients at 40 mg.

$$ n_3=6,\qquad X_3=1. $$

Because only one DLT has been observed among six patients, escalation is permitted under the 3+3 rules.

Interpretation: One DLT among six patients corresponds to an observed DLT proportion of \(1/6\approx16.7\%\). However, the 3+3 decision is not based on whether 16.7% is statistically “below” a target toxicity rate. It follows the prespecified cohort rule.

Step 4: Evaluate 60 mg

Three patients receive 60 mg. One patient experiences a DLT:

$$ X_4=1. $$

The cohort must therefore be expanded to six patients.

Suppose one additional DLT occurs among the next three patients. The total is then:

$$ X_4=2,\qquad n_4=6. $$

The 3+3 rule now requires escalation to stop:

$$ 2\text{ DLTs among 6}\Rightarrow\text{stop escalation}. $$

The previous dose, 40 mg, is therefore selected as the MTD under this simplified conventional 3+3 scenario.

The Complete Decision Algorithm

Current DoseDLTsPatientsDecision
Any dose0/33Escalate
Any dose1/33Expand to 6
Any dose≥2/33Stop escalation
Expanded dose≤1/66Escalate
Expanded dose≥2/66Stop escalation; select previous dose

Observed DLT Rate Versus True DLT Probability

A common misconception is that an observed DLT proportion such as \(1/6\) is the true toxicity probability. It is not.

If six patients are treated at a dose and one experiences a DLT, the observed rate is:

$$ \hat p=\frac{1}{6}=0.1667. $$

But the underlying probability \(p\) is unknown. The six-patient sample provides only an estimate of it.

For example, if:

$$ X\sim\operatorname{Binomial}(6,p), $$

then many values of \(p\) can produce one observed DLT with non-negligible probability. This is one reason why the 3+3 design can provide relatively imprecise estimates of the dose-toxicity curve.

Statistical point: The MTD is a decision under uncertainty. A small Phase I cohort cannot precisely estimate the toxicity probability at every dose level.

What Does “Maximum Tolerated” Actually Mean?

The word maximum can be misleading. MTD does not necessarily mean the highest dose that can physically be given to humans, nor does it mean the highest dose at which no serious adverse events occur.

Instead, the MTD is defined relative to a protocol-specific notion of acceptable toxicity. The definition depends on:

  • the DLT definition;
  • the DLT assessment window;
  • the dose levels investigated;
  • the dose-escalation algorithm;
  • the target or acceptable toxicity level, when applicable; and
  • the final dose-selection rules.

MTD Is Not Necessarily the Recommended Phase II Dose

Another important distinction is between the MTD and the recommended Phase II dose (RP2D).

The dose ultimately selected for later development may incorporate substantially more information than DLT counts alone. Pharmacokinetics, pharmacodynamics, cumulative toxicity, dose intensity, exposure-response relationships, efficacy signals, tolerability, and dose-schedule considerations may all influence dose selection.

ConceptPrimary Question
MTDWhat is the highest dose meeting the protocol's toxicity criterion?
Target toxicity doseWhich dose has a DLT probability closest to the prespecified target?
RP2DWhich dose or regimen is appropriate for subsequent clinical development?
Important: An MTD-based dose selection strategy should not automatically be interpreted as proof that the selected dose is the biologically optimal or clinically optimal dose.

Why the 3+3 Design Is Statistically Limited

The 3+3 design is simple and operationally familiar, but it has several important statistical limitations.

  • It does not explicitly model the dose-toxicity relationship.
  • It does not formally target a prespecified DLT probability.
  • Its decisions depend heavily on small cohorts.
  • It may treat patients at doses that a model-based design would consider excessively toxic.
  • It can produce relatively imprecise estimates of toxicity probabilities.
  • The final selected dose may have a true DLT probability substantially different from the intended clinical target.

These limitations motivated a large family of model-based and model-assisted dose-escalation designs.

Model-Based MTD Determination

A model-based design treats the observed DLT data as information about the underlying dose-toxicity relationship.

Suppose the study has doses \(d_1,\ldots,d_K\), with true DLT probabilities \(p_1,\ldots,p_K\). A model-based approach attempts to learn about these probabilities as patients accumulate.

The target may be expressed as:

$$ \phi=\text{target DLT probability}. $$

The dose selected as the MTD is then often the dose whose estimated toxicity probability is closest to \(\phi\):

$$ \hat d_{\text{MTD}} = \arg\min_{d_j} \left|\hat p_j-\phi\right|. $$

This is fundamentally different from the 3+3 approach.

Continual Reassessment Method (CRM)

The continual reassessment method (CRM) is a Bayesian model-based dose-escalation design. Rather than using only fixed cohort rules, CRM continuously updates a model for the dose-toxicity relationship as DLT information becomes available.

A simple one-parameter formulation may specify:

$$ p(d_j)=F(d_j,\theta), $$

where \(F\) is a prespecified monotonic dose-toxicity model and \(\theta\) is an unknown model parameter.

After each cohort, the accumulated DLT data update the distribution of \(\theta\), and the posterior toxicity probabilities are used to determine the next dose.

Conceptual advantage: CRM uses information across dose levels. Evidence collected at one dose can influence the estimated toxicity probabilities at other doses through the dose-toxicity model.

Bayesian Optimal Interval (BOIN)

The Bayesian optimal interval (BOIN) design is a model-assisted approach that uses prespecified toxicity probability intervals to determine whether to escalate, remain at the current dose, or de-escalate.

Let \(\phi\) denote the target toxicity probability. The design defines an interval around \(\phi\):

$$ \lambda_e < \phi < \lambda_d, $$

where the exact boundaries are determined during design calibration.

Conceptually:

Observed evidenceTypical BOIN action
Evidence that toxicity is below the targetEscalate
Evidence consistent with the target intervalStay at the current dose
Evidence that toxicity is above the targetDe-escalate

Unlike the fixed 3+3 rule, BOIN explicitly incorporates the target toxicity probability into the dose-escalation decision.

3+3 Versus CRM and BOIN

Feature3+3CRMBOIN
Model-basedNoYesModel-assisted
Explicit target DLT rateNoYesYes
Uses information across dosesLimitedYesYes
Can de-escalateUnder protocol-specific rulesYesYes
Statistical complexityLowHigherModerate
Common cohort structure3 or 6FlexibleFlexible
Primary dose-selection conceptProtocol ruleEstimated target doseEstimated target interval

Safety Boundaries and Dose Skipping

Modern dose-escalation designs generally incorporate safety constraints rather than allowing the statistical model to dictate unrestricted escalation.

For example, a protocol may prohibit escalation beyond a dose level that has demonstrated excessive toxicity, require a minimum number of patients at a dose before escalation, or restrict escalation to adjacent dose levels.

A
Escalation boundary: prevents progression when the observed toxicity is too high.
B
De-escalation boundary: moves treatment toward a safer dose when evidence indicates excessive toxicity.
C
Stopping boundary: terminates escalation when the study has sufficient evidence that additional doses should not be evaluated.

How the Dose-Toxicity Curve Changes the Interpretation

The fundamental statistical assumption in most dose-escalation methods is that toxicity does not decrease as dose increases. In other words, the dose-toxicity relationship is generally assumed to be monotonic:

$$ d_i

The curve may be shallow at low doses and steep at higher doses. It may also be uncertain, particularly when few patients have been treated at each dose.

This means that observing zero DLTs at a low dose does not imply that the next dose is safe with certainty. It merely provides evidence consistent with a relatively low toxicity probability.

A Numerical Illustration of Uncertainty

Suppose six patients receive a dose and zero DLTs are observed:

$$ X=0,\qquad n=6. $$

The observed DLT rate is:

$$ \hat p=0. $$

But the true DLT probability is not necessarily zero. Even when no DLTs are observed, the underlying probability remains uncertain.

This is an important practical lesson in Phase I trials: absence of observed toxicity is not evidence of absence of toxicity.

MTD Determination as a Sequential Decision Problem

Dose escalation is sequential because each new cohort depends on information generated by previous cohorts.

Let \(D_t\) represent the dose assigned to cohort \(t\), and let \(Y_t\) represent the observed DLT outcomes. Then the next dose can be viewed abstractly as:

$$ D_{t+1}=g(D_1,Y_1,\ldots,D_t,Y_t), $$

where \(g(\cdot)\) is the dose-escalation rule specified by the design.

For 3+3, \(g(\cdot)\) is a discrete rule based on DLT counts. For CRM, it depends on the updated dose-toxicity model. For BOIN, it depends on the observed toxicity evidence relative to calibrated decision boundaries.

Stopping the Dose-Escalation Trial

A Phase I study does not necessarily continue until a DLT is observed at every dose. A trial may stop because:

  • a prespecified dose has been identified as sufficiently characterized;
  • the highest planned dose has been reached;
  • toxicity becomes excessive;
  • a protocol-defined stopping rule is triggered;
  • enrollment or feasibility considerations require termination; or
  • additional information indicates that further escalation is not appropriate.
Do not equate “highest tested dose” with MTD. If no dose reaches the protocol's toxicity criterion, the study may simply have not identified an MTD within the investigated dose range.

What Happens If No DLT Is Observed?

Suppose a study evaluates several dose levels and observes no DLTs at any dose.

It would be incorrect to conclude:

$$ p(d)=0\quad\text{for all doses}. $$

The correct conclusion is that no DLT was observed in the evaluated patients during the prespecified assessment period. The available data may still be compatible with a nonzero toxicity probability.

The protocol may therefore permit continued escalation, evaluation of additional doses, or selection of the highest tested dose depending on its prespecified rules.

What Happens If Toxicity Appears Earlier Than Expected?

If multiple DLTs occur at a low dose, the design should protect subsequent patients by preventing automatic escalation.

For example, in a simple 3+3 implementation:

$$ X_j\ge2\text{ among the first 3 patients} \Rightarrow\text{stop escalation}. $$

The study may then select a lower dose, expand an earlier dose, or terminate, depending on the specific protocol.

Common Mistakes in MTD Analysis

1. Treating the observed DLT rate as the true DLT probability

An observed proportion such as \(2/6=33.3\%\) is an estimate, not a known population probability.

2. Assuming the MTD is always the highest tested dose

The highest tested dose may be below, near, or above the clinically desired toxicity level. Selection depends on the prespecified design rules.

3. Assuming no DLT means zero risk

A small cohort can observe zero DLTs even when the underlying toxicity probability is meaningful.

4. Confusing MTD with RP2D

The recommended dose for later development can incorporate PK, PD, efficacy, cumulative toxicity, and regimen-specific considerations beyond the MTD decision.

5. Ignoring the DLT evaluation window

A DLT observed outside the predefined evaluation period may be handled differently depending on the protocol. Delayed or cumulative toxicity can therefore be important when interpreting MTD results.

Practical MTD Workflow

A statistical programming or clinical-trial team can conceptualize the workflow as:

1
Define the DLT: specify exactly which toxicities count and during what assessment window.
2
Define dose levels: establish the planned dose sequence and escalation restrictions.
3
Select the design: 3+3, CRM, BOIN, mTPI, or another prespecified dose-escalation method.
4
Collect DLT outcomes: evaluate each cohort according to the protocol.
5
Apply the decision rule: escalate, remain, de-escalate, or stop.
6
Select the dose: apply the design's final MTD or target-dose rule.
7
Integrate additional evidence: consider whether the selected dose is appropriate for subsequent development.

Summary of the Worked Example

DosePatientsDLTsObserved Rate3+3 Decision
10 mg300%Escalate
20 mg300%Escalate
40 mg6116.7%Escalate
60 mg6233.3%Stop escalation

Under the simplified conventional 3+3 rules used in this example, the previous dose is selected:

$$ \boxed{\text{MTD}=40\text{ mg}}. $$

Key Takeaways

  • MTD determination is a sequential dose-escalation problem.
  • DLTs are protocol-defined toxicities that drive dose-escalation decisions.
  • The traditional 3+3 design uses small cohorts and simple DLT-count rules.
  • With 0 DLTs among 3 patients, the conventional design escalates.
  • With 1 DLT among 3 patients, the cohort is expanded to 6.
  • With ≥2 DLTs among 3, or ≥2 among 6 after expansion, escalation stops under the conventional rule.
  • The observed DLT rate is an estimate and should not be confused with the true toxicity probability.
  • CRM explicitly models dose-toxicity relationships and targets a prespecified toxicity probability.
  • BOIN uses calibrated toxicity intervals to make escalation and de-escalation decisions.
  • The MTD and RP2D are not necessarily identical.

Final Perspective

The statistical problem underlying MTD determination is deceptively simple: identify a dose that provides an acceptable balance between increasing exposure and increasing toxicity. However, the information available in a Phase I trial is sparse, sequential, and uncertain.

The traditional 3+3 design addresses the problem with an intentionally simple decision algorithm. Modern approaches such as CRM and BOIN instead use the accumulating toxicity information more efficiently and explicitly define the target toxicity probability.

For contemporary Phase I development, the choice of dose-escalation design should therefore be made prospectively and should reflect the drug, disease, treatment schedule, anticipated toxicity profile, patient population, and objectives of the study.

Bottom line: The MTD is not simply “the highest dose without toxicity.” It is the dose selected by a prespecified statistical and clinical decision framework using the observed dose-toxicity relationship and an explicit definition of acceptable toxicity.

Related Concepts

MTD determination connects naturally to several other Phase I and clinical pharmacology topics:

  • Dose-escalation and dose-finding designs
  • Continual Reassessment Method (CRM)
  • Bayesian Optimal Interval (BOIN)
  • Modified Toxicity Probability Interval (mTPI)
  • Time-to-Event CRM (TITE-CRM)
  • Pharmacokinetic dose proportionality
  • Exposure-response analysis
  • PK/PD-guided dose selection
  • Recommended Phase II dose (RP2D)
  • First-in-human dose selection