Introduction
Early-phase clinical trials must answer a fundamental question before a new treatment can be studied at larger doses: how much drug can be administered with an acceptable level of toxicity?
In a conventional Phase I oncology trial, patients are assigned to progressively higher dose levels while investigators monitor for dose-limiting toxicities (DLTs). The goal is not usually to estimate efficacy. Instead, the study is primarily concerned with characterizing safety and selecting a dose or dose range suitable for subsequent development.
The term maximum tolerated dose (MTD) generally refers to the highest dose associated with an acceptable probability of severe or clinically important toxicity, according to the rules specified in the protocol.
What Is a Dose-Limiting Toxicity?
A dose-limiting toxicity, or DLT, is a prespecified toxicity that is considered sufficiently serious to prevent further escalation or otherwise influence the dose-escalation decision.
The protocol must define the DLT criteria before the study begins. Depending on the disease, drug, treatment schedule, and therapeutic context, DLT criteria may involve hematologic toxicity, nonhematologic toxicity, organ toxicity, prolonged treatment interruption, hospitalization, or other clinically important events.
For statistical purposes, the essential feature is that each evaluable patient contributes a binary outcome during the specified DLT observation period:
If the true probability of a DLT at dose \(d\) is \(p(d)\), then for \(n\) patients treated at that dose:
where \(X_d\) is the number of patients experiencing a DLT.
The Target Toxicity Rate
Modern statistical dose-escalation designs are often formulated around a target DLT probability. Let
For example, a protocol might define a target toxicity probability of 25%:
The conceptual objective is then to identify the dose whose toxicity probability is closest to the target:
This formulation is particularly natural for model-based designs such as the continual reassessment method (CRM) and interval-based designs such as BOIN. The traditional 3+3 design is different: it does not explicitly estimate \(p(d)\) or select the dose closest to a numerical target.
Why Dose Escalation Is Necessary
Before the first patient is treated, the dose-toxicity relationship is usually uncertain. Preclinical studies, pharmacology, toxicology, and prior clinical experience may provide useful information, but the probability of severe toxicity in humans is not known with certainty.
Consequently, a Phase I trial typically evaluates a sequence of dose levels:
The Traditional 3+3 Design
The 3+3 design is one of the most widely recognized traditional dose-escalation designs in oncology. It enrolls patients in cohorts of three and uses the number of DLTs observed within a cohort to determine the next action.
A basic version of the design can be summarized as follows.
| DLTs in first 3 patients | Decision |
|---|---|
| 0 DLTs | Escalate to the next dose |
| 1 DLT | Expand the cohort at the same dose to 6 patients |
| ≥2 DLTs | Stop escalation; the current dose is too toxic under the design |
When a cohort is expanded to six patients, the conventional rules are:
| Total DLTs among 6 patients | Decision |
|---|---|
| ≤1 DLT | Escalate to the next dose |
| ≥2 DLTs | Stop escalation; select the previous dose as the MTD, subject to the protocol's final criteria |
Formalizing the 3+3 Decision Rules
Let \(d_j\) denote dose level \(j\), and let \(X_j\) be the number of DLTs observed among the initial three patients at that dose.
The initial decision rule can be written:
After six patients have been evaluated following a one-DLT initial cohort:
These rules are intentionally simple. No likelihood model is fit, and no formal posterior distribution is required to determine the next dose.
A Complete Worked Example
Consider a hypothetical Phase I oncology study evaluating an investigational agent at six planned dose levels. The protocol uses a conventional 3+3 design.
| Dose Level | Dose | Initial Cohort | Observed DLTs | Action |
|---|---|---|---|---|
| 1 | 10 mg | 3 | 0 | Escalate |
| 2 | 20 mg | 3 | 0 | Escalate |
| 3 | 40 mg | 3 | 1 | Expand to 6 |
| 3 | 40 mg | 6 total | 1 total | Escalate |
| 4 | 60 mg | 3 | 1 | Expand to 6 |
| 4 | 60 mg | 6 total | 2 total | Stop escalation |
The key question is now: which dose is selected as the MTD?
Because escalation was stopped at 60 mg after two DLTs among six patients, the conventional 3+3 decision is to select the previous dose, 40 mg, as the MTD, provided the protocol's final dose-selection and safety criteria are satisfied.
Step 1: Evaluate 10 mg
Three patients receive 10 mg. Suppose none experiences a DLT:
The 3+3 rule therefore permits escalation:
Step 2: Evaluate 20 mg
Three patients receive 20 mg, again with no DLTs:
The trial escalates to 40 mg.
Step 3: Evaluate 40 mg
Three patients receive 40 mg. One patient experiences a protocol-defined DLT:
The trial does not immediately escalate. Instead, the cohort is expanded from three to six patients at 40 mg.
Because only one DLT has been observed among six patients, escalation is permitted under the 3+3 rules.
Step 4: Evaluate 60 mg
Three patients receive 60 mg. One patient experiences a DLT:
The cohort must therefore be expanded to six patients.
Suppose one additional DLT occurs among the next three patients. The total is then:
The 3+3 rule now requires escalation to stop:
The previous dose, 40 mg, is therefore selected as the MTD under this simplified conventional 3+3 scenario.
The Complete Decision Algorithm
| Current Dose | DLTs | Patients | Decision |
|---|---|---|---|
| Any dose | 0/3 | 3 | Escalate |
| Any dose | 1/3 | 3 | Expand to 6 |
| Any dose | ≥2/3 | 3 | Stop escalation |
| Expanded dose | ≤1/6 | 6 | Escalate |
| Expanded dose | ≥2/6 | 6 | Stop escalation; select previous dose |
Observed DLT Rate Versus True DLT Probability
A common misconception is that an observed DLT proportion such as \(1/6\) is the true toxicity probability. It is not.
If six patients are treated at a dose and one experiences a DLT, the observed rate is:
But the underlying probability \(p\) is unknown. The six-patient sample provides only an estimate of it.
For example, if:
then many values of \(p\) can produce one observed DLT with non-negligible probability. This is one reason why the 3+3 design can provide relatively imprecise estimates of the dose-toxicity curve.
What Does “Maximum Tolerated” Actually Mean?
The word maximum can be misleading. MTD does not necessarily mean the highest dose that can physically be given to humans, nor does it mean the highest dose at which no serious adverse events occur.
Instead, the MTD is defined relative to a protocol-specific notion of acceptable toxicity. The definition depends on:
- the DLT definition;
- the DLT assessment window;
- the dose levels investigated;
- the dose-escalation algorithm;
- the target or acceptable toxicity level, when applicable; and
- the final dose-selection rules.
MTD Is Not Necessarily the Recommended Phase II Dose
Another important distinction is between the MTD and the recommended Phase II dose (RP2D).
The dose ultimately selected for later development may incorporate substantially more information than DLT counts alone. Pharmacokinetics, pharmacodynamics, cumulative toxicity, dose intensity, exposure-response relationships, efficacy signals, tolerability, and dose-schedule considerations may all influence dose selection.
| Concept | Primary Question |
|---|---|
| MTD | What is the highest dose meeting the protocol's toxicity criterion? |
| Target toxicity dose | Which dose has a DLT probability closest to the prespecified target? |
| RP2D | Which dose or regimen is appropriate for subsequent clinical development? |
Why the 3+3 Design Is Statistically Limited
The 3+3 design is simple and operationally familiar, but it has several important statistical limitations.
- It does not explicitly model the dose-toxicity relationship.
- It does not formally target a prespecified DLT probability.
- Its decisions depend heavily on small cohorts.
- It may treat patients at doses that a model-based design would consider excessively toxic.
- It can produce relatively imprecise estimates of toxicity probabilities.
- The final selected dose may have a true DLT probability substantially different from the intended clinical target.
These limitations motivated a large family of model-based and model-assisted dose-escalation designs.
Model-Based MTD Determination
A model-based design treats the observed DLT data as information about the underlying dose-toxicity relationship.
Suppose the study has doses \(d_1,\ldots,d_K\), with true DLT probabilities \(p_1,\ldots,p_K\). A model-based approach attempts to learn about these probabilities as patients accumulate.
The target may be expressed as:
The dose selected as the MTD is then often the dose whose estimated toxicity probability is closest to \(\phi\):
This is fundamentally different from the 3+3 approach.
Continual Reassessment Method (CRM)
The continual reassessment method (CRM) is a Bayesian model-based dose-escalation design. Rather than using only fixed cohort rules, CRM continuously updates a model for the dose-toxicity relationship as DLT information becomes available.
A simple one-parameter formulation may specify:
where \(F\) is a prespecified monotonic dose-toxicity model and \(\theta\) is an unknown model parameter.
After each cohort, the accumulated DLT data update the distribution of \(\theta\), and the posterior toxicity probabilities are used to determine the next dose.
Bayesian Optimal Interval (BOIN)
The Bayesian optimal interval (BOIN) design is a model-assisted approach that uses prespecified toxicity probability intervals to determine whether to escalate, remain at the current dose, or de-escalate.
Let \(\phi\) denote the target toxicity probability. The design defines an interval around \(\phi\):
where the exact boundaries are determined during design calibration.
Conceptually:
| Observed evidence | Typical BOIN action |
|---|---|
| Evidence that toxicity is below the target | Escalate |
| Evidence consistent with the target interval | Stay at the current dose |
| Evidence that toxicity is above the target | De-escalate |
Unlike the fixed 3+3 rule, BOIN explicitly incorporates the target toxicity probability into the dose-escalation decision.
3+3 Versus CRM and BOIN
| Feature | 3+3 | CRM | BOIN |
|---|---|---|---|
| Model-based | No | Yes | Model-assisted |
| Explicit target DLT rate | No | Yes | Yes |
| Uses information across doses | Limited | Yes | Yes |
| Can de-escalate | Under protocol-specific rules | Yes | Yes |
| Statistical complexity | Low | Higher | Moderate |
| Common cohort structure | 3 or 6 | Flexible | Flexible |
| Primary dose-selection concept | Protocol rule | Estimated target dose | Estimated target interval |
Safety Boundaries and Dose Skipping
Modern dose-escalation designs generally incorporate safety constraints rather than allowing the statistical model to dictate unrestricted escalation.
For example, a protocol may prohibit escalation beyond a dose level that has demonstrated excessive toxicity, require a minimum number of patients at a dose before escalation, or restrict escalation to adjacent dose levels.
How the Dose-Toxicity Curve Changes the Interpretation
The fundamental statistical assumption in most dose-escalation methods is that toxicity does not decrease as dose increases. In other words, the dose-toxicity relationship is generally assumed to be monotonic:
The curve may be shallow at low doses and steep at higher doses. It may also be uncertain, particularly when few patients have been treated at each dose.
This means that observing zero DLTs at a low dose does not imply that the next dose is safe with certainty. It merely provides evidence consistent with a relatively low toxicity probability.
A Numerical Illustration of Uncertainty
Suppose six patients receive a dose and zero DLTs are observed:
The observed DLT rate is:
But the true DLT probability is not necessarily zero. Even when no DLTs are observed, the underlying probability remains uncertain.
This is an important practical lesson in Phase I trials: absence of observed toxicity is not evidence of absence of toxicity.
MTD Determination as a Sequential Decision Problem
Dose escalation is sequential because each new cohort depends on information generated by previous cohorts.
Let \(D_t\) represent the dose assigned to cohort \(t\), and let \(Y_t\) represent the observed DLT outcomes. Then the next dose can be viewed abstractly as:
where \(g(\cdot)\) is the dose-escalation rule specified by the design.
For 3+3, \(g(\cdot)\) is a discrete rule based on DLT counts. For CRM, it depends on the updated dose-toxicity model. For BOIN, it depends on the observed toxicity evidence relative to calibrated decision boundaries.
Stopping the Dose-Escalation Trial
A Phase I study does not necessarily continue until a DLT is observed at every dose. A trial may stop because:
- a prespecified dose has been identified as sufficiently characterized;
- the highest planned dose has been reached;
- toxicity becomes excessive;
- a protocol-defined stopping rule is triggered;
- enrollment or feasibility considerations require termination; or
- additional information indicates that further escalation is not appropriate.
What Happens If No DLT Is Observed?
Suppose a study evaluates several dose levels and observes no DLTs at any dose.
It would be incorrect to conclude:
The correct conclusion is that no DLT was observed in the evaluated patients during the prespecified assessment period. The available data may still be compatible with a nonzero toxicity probability.
The protocol may therefore permit continued escalation, evaluation of additional doses, or selection of the highest tested dose depending on its prespecified rules.
What Happens If Toxicity Appears Earlier Than Expected?
If multiple DLTs occur at a low dose, the design should protect subsequent patients by preventing automatic escalation.
For example, in a simple 3+3 implementation:
The study may then select a lower dose, expand an earlier dose, or terminate, depending on the specific protocol.
Common Mistakes in MTD Analysis
1. Treating the observed DLT rate as the true DLT probability
An observed proportion such as \(2/6=33.3\%\) is an estimate, not a known population probability.
2. Assuming the MTD is always the highest tested dose
The highest tested dose may be below, near, or above the clinically desired toxicity level. Selection depends on the prespecified design rules.
3. Assuming no DLT means zero risk
A small cohort can observe zero DLTs even when the underlying toxicity probability is meaningful.
4. Confusing MTD with RP2D
The recommended dose for later development can incorporate PK, PD, efficacy, cumulative toxicity, and regimen-specific considerations beyond the MTD decision.
5. Ignoring the DLT evaluation window
A DLT observed outside the predefined evaluation period may be handled differently depending on the protocol. Delayed or cumulative toxicity can therefore be important when interpreting MTD results.
Practical MTD Workflow
A statistical programming or clinical-trial team can conceptualize the workflow as:
Summary of the Worked Example
| Dose | Patients | DLTs | Observed Rate | 3+3 Decision |
|---|---|---|---|---|
| 10 mg | 3 | 0 | 0% | Escalate |
| 20 mg | 3 | 0 | 0% | Escalate |
| 40 mg | 6 | 1 | 16.7% | Escalate |
| 60 mg | 6 | 2 | 33.3% | Stop escalation |
Under the simplified conventional 3+3 rules used in this example, the previous dose is selected:
Key Takeaways
- MTD determination is a sequential dose-escalation problem.
- DLTs are protocol-defined toxicities that drive dose-escalation decisions.
- The traditional 3+3 design uses small cohorts and simple DLT-count rules.
- With 0 DLTs among 3 patients, the conventional design escalates.
- With 1 DLT among 3 patients, the cohort is expanded to 6.
- With ≥2 DLTs among 3, or ≥2 among 6 after expansion, escalation stops under the conventional rule.
- The observed DLT rate is an estimate and should not be confused with the true toxicity probability.
- CRM explicitly models dose-toxicity relationships and targets a prespecified toxicity probability.
- BOIN uses calibrated toxicity intervals to make escalation and de-escalation decisions.
- The MTD and RP2D are not necessarily identical.
Final Perspective
The statistical problem underlying MTD determination is deceptively simple: identify a dose that provides an acceptable balance between increasing exposure and increasing toxicity. However, the information available in a Phase I trial is sparse, sequential, and uncertain.
The traditional 3+3 design addresses the problem with an intentionally simple decision algorithm. Modern approaches such as CRM and BOIN instead use the accumulating toxicity information more efficiently and explicitly define the target toxicity probability.
For contemporary Phase I development, the choice of dose-escalation design should therefore be made prospectively and should reflect the drug, disease, treatment schedule, anticipated toxicity profile, patient population, and objectives of the study.
Related Concepts
MTD determination connects naturally to several other Phase I and clinical pharmacology topics:
- Dose-escalation and dose-finding designs
- Continual Reassessment Method (CRM)
- Bayesian Optimal Interval (BOIN)
- Modified Toxicity Probability Interval (mTPI)
- Time-to-Event CRM (TITE-CRM)
- Pharmacokinetic dose proportionality
- Exposure-response analysis
- PK/PD-guided dose selection
- Recommended Phase II dose (RP2D)
- First-in-human dose selection