This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.
1. Trial at a Glance
SCOUT is a phase 1/2, non-randomized, open-label treatment study of larotrectinib in children with solid tumors harboring NTRK fusion. The study uses a parallel design with 6 arms and an enrollment of 154 participants.
| Feature | SCOUT |
|---|---|
| Study acronym | SCOUT |
| NCT ID | NCT02637687 |
| Phase | Phase 1/2 |
| Condition | Solid Tumors Harboring NTRK Fusion |
| Intervention | Larotrectinib (Vitrakvi, BAY2757556) |
| Allocation | Non-randomized |
| Design model | Parallel |
| Masking | None |
| Primary purpose | Treatment |
| Enrollment | 154 |
| Number of arms | 6 |
| Study status | Active, not recruiting |
| Lead sponsor | Bayer |
| Sponsor type | Industry |
2. Clinical Question
The SCOUT study addresses two related statistical questions. In phase 1, the central question is whether dose cohorts of larotrectinib can be evaluated for treatment-emergent adverse events and dose-limiting toxicity during the first 28-day cycle. In phase 2, the principal efficacy question is whether children with NTRK gene fusions experience an objective tumor response as assessed by an independent radiology review committee.
Population
Children with solid tumors harboring NTRK fusion.
Intervention
Larotrectinib, identified in the registry as Vitrakvi and BAY2757556.
Comparator
There is no randomized comparator. The allocation is explicitly non-randomized.
Primary questions
How frequently do dose-limiting toxicities and treatment-emergent adverse events occur, and what is the overall response rate in the phase 2 population?
The absence of randomization is central to interpreting the study statistically. The design is capable of describing safety and response under treatment, but it does not create the same randomized counterfactual comparison that would arise from assigning otherwise comparable participants to larotrectinib and a control treatment.
3. Trial Design
A useful distinction is between phase 1 safety evaluation and phase 2 efficacy evaluation. They use different primary endpoints because the statistical question changes. Early dose evaluation focuses on toxicity, particularly dose-limiting toxicity during Cycle 1. The phase 2 endpoint instead asks whether tumors meet predefined criteria for complete or partial response.
4. Endpoints
| Endpoint | Time frame | Statistical role |
|---|---|---|
| Phase 1: Number of participants in an assigned dose cohort with treatment emergent adverse events (TEAEs) by grade assessed by NCI-CTCAE v 4.03 who experience a DLT | From Day 1 to Day 28 of Cycle 1 (1 Cycle=28 days) | Primary safety / dose-limiting toxicity endpoint |
| Phase 1: Number of participants with TEAEs | From first dose of larotrectinib up to 93 months | Primary safety endpoint |
| Phase 1: Severity of TEAEs | From first dose of larotrectinib up to 93 months | Primary safety endpoint |
| Phase 2: Overall response rate (ORR) by IRRC | From first dose of Larotrectinib to disease progression or subsequent therapy or surgical intervention or death, up to 76 months | Primary efficacy endpoint |
The registry defines dose-limiting toxicity as DLT and identifies NCI-CTCAE v 4.03 as the grading framework for treatment-emergent adverse events. For the phase 2 efficacy endpoint, the registry defines ORR as the proportion of participants with a best overall response of complete response or partial response.
The phase 2 response assessment is particularly important because it is not simply a measurement of any tumor shrinkage. The endpoint requires a best overall response of CR or PR, and the response determination is assigned by an independent radiology review committee using RECIST 1.1, RANO, or INRC as appropriate to tumor type. The population is further defined by the presence of NTRK gene fusions.
5. Primary Endpoint Structure
| Study phase | Endpoint family | What is being measured |
|---|---|---|
| Phase 1 | Dose-limiting toxicity | Participants in an assigned dose cohort with TEAEs by grade who experience a DLT during Day 1 through Day 28 of Cycle 1. |
| Phase 1 | Adverse events | Number of participants with TEAEs from first dose through up to 93 months. |
| Phase 1 | Adverse-event severity | Severity of TEAEs from first dose through up to 93 months. |
| Phase 2 | Overall response rate | Proportion with CR or PR by independent radiology review, assessed from first dose through disease progression, subsequent therapy, surgical intervention, or death, up to 76 months. |
This structure illustrates why a single statistical summary would be inadequate for SCOUT. DLT is a short-window safety endpoint; TEAEs are longitudinal safety outcomes; severity is an ordinal or categorical characterization of adverse events; and ORR is a binary response endpoint defined from tumor-assessment criteria. Each requires a different statistical description.
6. Planned Analysis
No formal statistical analyses are posted to ClinicalTrials.gov for the registered primary endpoints. The registry therefore provides the endpoint definitions and assessment windows, but not numerical estimates, confidence intervals, or p-values for those endpoints.
Phase 1: Dose-limiting toxicity
For the phase 1 DLT endpoint, the natural primary summary is the number and proportion of participants within each assigned dose cohort who experience a DLT during Day 1 through Day 28 of Cycle 1. Because the endpoint is explicitly defined within an assigned dose cohort and a fixed first-cycle window, cohort-level safety summaries are more directly aligned with the question than a time-to-event model.
The statistical interpretation depends on the cohort definition and evaluability rules. The registry does not provide a posted statistical analysis specifying an alternative estimator or inferential procedure for this endpoint.
Phase 1: Treatment-emergent adverse events
For TEAEs, a conventional analysis would summarize the number and proportion of participants experiencing at least one event, together with event severity and clinically relevant event categories. Because the registry specifies follow-up from first dose through up to 93 months, the duration of observation is an important part of interpretation: participants may have different lengths of follow-up, and simple event proportions do not necessarily represent incidence rates over equal exposure time.
Phase 1: Severity of TEAEs
Severity is distinct from whether an adverse event occurred. A useful statistical presentation would preserve the ordered grading information rather than collapsing all adverse events into a single yes/no outcome. The registry identifies NCI-CTCAE v 4.03 for grading the DLT-related endpoint and separately specifies severity of TEAEs as a primary endpoint.
Phase 2: Overall response rate
ORR is naturally analyzed as a binomial proportion: the numerator is the number of participants whose best overall response is CR or PR, and the denominator is the applicable phase 2 analysis population. A confidence interval for the proportion would quantify uncertainty around the observed response rate.
The registry's response definition is based on independent radiology review and tumor-type-appropriate criteria: RECIST 1.1, RANO, or INRC.
7. Statistical Methodology
Descriptive analysis of safety
Safety endpoints in an early-phase treatment study are commonly summarized descriptively. For a binary event such as whether a participant experienced a TEAE, the basic quantities are counts and proportions. Severity can then be summarized across the applicable grading categories.
For DLT, the short assessment window is particularly meaningful. The registry defines the endpoint from Day 1 through Day 28 of Cycle 1, with one cycle defined as 28 days. That makes the endpoint a fixed-window safety measure rather than a general statement about all toxicity observed over the entire treatment period.
Binomial proportions
ORR is fundamentally a binomial endpoint because each participant can be classified according to whether their best overall response meets the CR-or-PR definition. If \(x\) participants respond among \(n\) participants, the observed response proportion is \(x/n\).
Here, \(p\) represents the underlying probability that a participant in the defined analysis population has a best overall response of CR or PR under the endpoint definition.
For a small or moderate response sample, exact or appropriately constructed binomial confidence intervals can be preferable to relying automatically on a normal approximation. The registry does not state which confidence-interval method would be used.
Independent radiology review
The phase 2 ORR endpoint is determined by an independent radiology review committee. Statistically, this matters because response classification is not simply the investigator's subjective assessment. The independent review creates a specified assessment process for determining whether the participant's best overall response satisfies the CR or PR definition.
The criteria also vary according to tumor type. RECIST 1.1, RANO, or INRC may be used as appropriate. That means the response classification framework is linked to the characteristics of the tumor being assessed rather than assuming that one radiologic criterion applies identically to every tumor type.
Time-to-event considerations
The registry's phase 2 endpoint includes an assessment period extending from first dose until disease progression, subsequent therapy, surgical intervention, or death, up to 76 months. Those events define the endpoint's assessment context, but the primary endpoint itself is ORR rather than a reported time-to-event measure.
This distinction is important. A participant may have a response and later experience progression, while another participant may never meet the CR-or-PR definition. ORR reduces the tumor-assessment history to a binary response classification and therefore cannot by itself describe the duration or timing of the response.
8. Statistical Methods Explained
Why is DLT evaluated during the first 28 days?
The registry defines the DLT endpoint specifically from Day 1 to Day 28 of Cycle 1, with one cycle equal to 28 days. This creates a standardized early safety window. From a statistical perspective, a fixed observation window makes participants more directly comparable for the particular DLT question than an endpoint based on an unrestricted amount of follow-up.
Why summarize DLT by dose cohort?
The endpoint is explicitly defined as the number of participants in an assigned dose cohort who experience a DLT. A cohort-level analysis therefore retains information about the relationship between the assigned dose cohort and early toxicity. Combining all cohorts into a single percentage could obscure differences between dose levels.
Why is ORR analyzed as a proportion?
ORR asks whether a participant's best overall response is CR or PR. Each participant therefore contributes a response classification to the numerator or denominator. The resulting proportion is easy to interpret, while a confidence interval expresses uncertainty around the observed proportion.
Why does independent radiology review matter statistically?
Response is determined by an independent radiology review committee rather than relying solely on the treatment team. This creates a defined assessment process for the primary efficacy endpoint and helps separate response classification from the clinical team's knowledge of treatment exposure.
Why is a non-randomized design important when interpreting ORR?
Because allocation is non-randomized, an observed response proportion describes outcomes among participants treated in the study. It does not, by itself, estimate the causal difference between larotrectinib and a concurrently randomized control treatment. Without randomization, differences in patient characteristics, disease characteristics, or other factors can affect comparisons with external populations.
Why are TEAEs different from DLTs?
TEAE is a broad treatment-emergent safety concept, whereas DLT is a more specific endpoint focused on toxicity that meets the study's dose-limiting definition during the first 28-day cycle. A participant can therefore contribute to the broader TEAE assessment without necessarily having a DLT.
Why does the endpoint's stopping rule matter for ORR?
The registry defines the ORR assessment period as extending from first dose to disease progression or subsequent therapy or surgical intervention or death, up to 76 months. These events can affect whether a later tumor assessment contributes to the participant's best overall response classification. The timing and rules for assessment therefore matter to the denominator and to the resulting response estimate.
9. Interpreting an ORR Without a Comparator
A response rate is often intuitively attractive because it can be expressed as a simple percentage. Statistically, however, its meaning depends on the population and endpoint definition. In SCOUT, the relevant population is participants with tumors harboring NTRK gene fusions, and response is determined by independent radiology review using the appropriate tumor-specific response criteria.
An ORR would represent the proportion of participants in the defined analysis population whose best overall response was complete response or partial response. It is a population-level summary of tumor response under the specified assessment framework.
ORR would not mean that the same proportion of participants were cured, remained progression-free for a particular length of time, or survived for a particular duration. It also would not establish a causal treatment effect against a control because the study allocation is non-randomized.
A confidence interval around ORR would describe statistical uncertainty around the estimated response proportion. A wider interval would indicate less precision, while a narrower interval would indicate greater precision under the selected statistical framework.
If a formal hypothesis test were used, its p-value would quantify evidence against a specified null hypothesis under the statistical model. It would not measure the magnitude of tumor response. The response proportion itself is the effect estimate; the confidence interval describes its precision.
10. Interpreting the Safety Endpoints
Safety requires more than one number because the SCOUT registry separates the occurrence and severity of TEAEs from the phase 1 DLT endpoint. The statistical question is therefore multidimensional: whether an event occurred, how severe it was, and whether it met the criteria for a dose-limiting toxicity during the defined first-cycle window.
Occurrence
The number of participants with TEAEs describes how frequently treatment-emergent adverse events occurred in the study population.
Severity
Severity adds information about the clinical grade or seriousness of the observed adverse events rather than simply counting any event.
DLT
DLT focuses specifically on qualifying toxicities within Day 1 through Day 28 of Cycle 1.
Follow-up
The TEAE endpoint extends from first dose up to 93 months, which is a substantially broader observation period than the DLT window.
These different windows should not be conflated. A 28-day DLT assessment answers an early dose-limiting safety question, whereas the TEAE endpoint can capture events over a much longer period. The two endpoints therefore provide complementary rather than interchangeable summaries.
11. Analysis Population and Follow-Up
The registry reports an overall enrollment of 154 participants. It does not provide, in the ClinicalTrials.gov record, numerical analysis populations for each individual primary endpoint or a posted statistical analysis specifying the exact evaluability rules for every endpoint.
| Feature | Registry information | Statistical implication |
|---|---|---|
| Overall enrollment | 154 | The enrolled population provides the overall study scale, but the applicable denominator for a specific endpoint may differ depending on endpoint-specific eligibility and assessment. |
| Phase 1 DLT window | Day 1 to Day 28 of Cycle 1 | Early toxicity is assessed within a defined fixed window. |
| Phase 1 TEAE window | First dose up to 93 months | Longer observation introduces varying exposure and follow-up considerations. |
| Phase 2 ORR window | First dose to disease progression or subsequent therapy or surgical intervention or death, up to 76 months | Response classification depends on the prespecified assessment context and stopping events. |
For an endpoint such as ORR, the denominator is especially important. The meaningful question is not simply how many responses occurred, but how many participants were eligible for the response analysis under the study's endpoint rules. The registry does not post the endpoint-specific numerical denominator or a formal statistical analysis.
12. Non-Randomized Design and Causal Interpretation
Randomization is a mechanism for balancing both measured and unmeasured prognostic factors between treatment groups in expectation. SCOUT is explicitly non-randomized, so that mechanism is not present.
A response estimate describes the treated study population. A causal treatment effect requires a valid comparison with a counterfactual outcome under another treatment or condition.
This does not make the response endpoint statistically uninformative. It changes the question it can answer. A non-randomized phase 1/2 study can characterize safety, describe response, evaluate the behavior of a therapy in a defined molecular population, and provide evidence for subsequent development. It cannot by itself provide the same randomized estimate of comparative efficacy that a parallel randomized controlled trial would provide.
The distinction also affects interpretation of external comparisons. If an observed ORR were compared with a historical response rate, differences in eligibility criteria, tumor types, prior treatment, assessment practices, follow-up, and patient characteristics could contribute to the observed difference. Such a comparison requires assumptions beyond the response proportion itself.
13. Six-Arm Parallel Structure
The registry lists 6 arms within a parallel design, while the intervention is larotrectinib. Because allocation is non-randomized and the ClinicalTrials.gov record does not provide separate numerical enrollment or outcome results for each arm, the overall enrollment should not be interpreted as six equally sized groups.
For statistical reporting, arm-specific results are most useful when the denominators are explicit. A single study-wide response proportion can hide important differences in the composition of the participants contributing to that estimate. Conversely, comparing small arm-specific proportions without accounting for their precision can make random variation appear more meaningful than it is.
14. Timing of the Study
Study start
The registry lists December 16, 2015 as the start date for SCOUT.
Primary completion
The registry lists July 20, 2024 as the primary completion date.
Active, not recruiting
The ClinicalTrials.gov record currently identifies the study status as active, not recruiting.
The timing of follow-up is relevant to the statistical interpretation of the endpoints. DLT has a fixed 28-day Cycle 1 window, whereas TEAEs can be assessed through up to 93 months and the phase 2 ORR assessment extends through a defined event-based period up to 76 months. These are fundamentally different observation structures.
15. Missing Data and Assessment Issues
The registry specifies several events that can terminate the phase 2 ORR assessment period: disease progression, subsequent therapy, surgical intervention, or death. These events are statistically relevant because response assessment is not conducted in an unlimited observational window.
For a response endpoint, missing tumor assessments can be consequential. If participants are not assessed after treatment begins, the statistical analysis must follow prespecified rules for determining whether they are evaluable and how their lack of an assessable response is handled. The ClinicalTrials.gov record does not post those detailed imputation or missing-data rules.
The same principle applies to safety. Participants may have different durations of treatment and follow-up, particularly when the TEAE window extends to 93 months. A simple proportion of participants with an event is therefore not automatically equivalent to an incidence rate or a time-to-first-event estimate.
16. What This Design Can and Cannot Establish
Safety characterization
The phase 1 endpoints can describe DLTs, TEAEs, and their severity within the defined study population and observation periods.
Response characterization
The phase 2 endpoint can describe the proportion of participants achieving CR or PR under the independent radiology review framework.
Molecularly defined population
The ORR definition specifically concerns participants who express NTRK gene fusions, linking the efficacy question to the molecular characteristic in the registry endpoint.
Comparative efficacy
Because allocation is non-randomized and no comparator is listed, the study does not provide a randomized treatment-versus-control efficacy contrast.
This distinction is one of the most important statistical lessons from SCOUT. An efficacy endpoint can be precisely defined without being a randomized comparative endpoint. Statistical precision and causal identification are different properties: a very precisely estimated response proportion can still be a descriptive treatment outcome rather than a randomized estimate of treatment effect.
17. Important Limitations and Interpretation Issues
- Non-randomized allocation: participants were not randomized to treatment groups, so the study does not have the causal protection provided by randomization.
- No randomized comparator: the registry identifies larotrectinib as the intervention and does not identify a comparator treatment.
- Arm-specific information: the registry lists 6 arms but does not provide separate numerical enrollment or primary-endpoint results for each arm in the posted statistical analysis information.
- Endpoint-specific denominators: the overall enrollment of 154 should not automatically be assumed to be the denominator for every primary endpoint.
- Different follow-up windows: DLT is evaluated during Day 1 through Day 28 of Cycle 1, while TEAEs can extend to 93 months and ORR assessment can extend up to 76 months.
- Response assessment: ORR is dependent on tumor-response criteria and independent radiology review, with RECIST 1.1, RANO, or INRC used as appropriate to tumor type.
- Missing assessment information: the registry does not post detailed imputation rules or a complete statistical analysis plan for the primary endpoints.
- Descriptive versus causal interpretation: an observed response proportion should not be interpreted as a randomized treatment effect.
18. Why This Trial Matters Statistically
SCOUT is a useful statistical teaching case because it combines early-phase safety methodology with a molecularly defined efficacy endpoint. The study also illustrates why the design of a clinical trial determines the interpretation of its numerical results.
| Concept | How it appears in SCOUT |
|---|---|
| Phase 1 safety analysis | DLT, TEAE occurrence, and TEAE severity are registered primary endpoints. |
| Fixed-window safety assessment | DLT is assessed from Day 1 through Day 28 of Cycle 1. |
| Longitudinal safety | TEAEs are assessed from first dose up to 93 months. |
| Binary efficacy endpoint | ORR classifies participants according to whether their best overall response is CR or PR. |
| Independent assessment | Phase 2 ORR is determined by an independent radiology review committee. |
| Multiple response criteria | RECIST 1.1, RANO, or INRC is used as appropriate to tumor type. |
| Non-randomized design | Allocation is explicitly non-randomized, so response estimates are not randomized treatment effects. |
| Parallel design | The registry lists 6 arms in a parallel study model. |
| Endpoint-specific follow-up | Different primary endpoints use different observation windows and stopping events. |
| Confidence intervals | A binomial confidence interval would quantify uncertainty around an ORR estimate if reported. |
19. Statistical Interpretation vs Clinical Interpretation
Statistical interpretation
The study is structured to describe early dose-limiting toxicity, treatment-emergent adverse events, adverse-event severity, and overall response. The appropriate statistical summaries differ because these endpoints represent different types of data.
Clinical interpretation
The clinical meaning of an observed response or adverse-event rate depends on the characteristics of the pediatric population, the NTRK-fusion tumor context, the treatment exposure, the assessment criteria, and the duration of follow-up.
Keeping these two interpretations separate is especially important in a non-randomized study. Statistical analysis can describe the observed pattern with considerable precision, but precision does not substitute for a randomized comparator when the question is comparative causal efficacy.
20. A Framework for Reading Future SCOUT Results
If formal numerical results are posted for the primary endpoints, the most informative statistical reading would begin with the denominator and endpoint definition before examining any p-value.
| Question | Why it matters |
|---|---|
| How many participants were included in the endpoint analysis? | Defines the denominator and determines how much information contributes to the estimate. |
| How many participants experienced the event? | Provides the numerator for a binary safety or response proportion. |
| What was the confidence interval? | Shows the statistical precision of the estimated proportion. |
| Which response criteria were used? | Determines how CR and PR were classified. |
| Was the analysis independently reviewed? | For ORR, the registry specifies independent radiology review. |
| How long were participants followed? | Safety and response observation windows differ substantially in SCOUT. |
| Was there a comparator? | The registered design is non-randomized and does not identify a comparator treatment. |
This sequence prevents a common statistical error: interpreting a numerical estimate before establishing exactly what population, endpoint definition, assessment window, and denominator produced it.
21. Sources
- ClinicalTrials.gov: NCT02637687 — SCOUT.
Continue through the Clinical Biostats knowledge graph
Explore broader statistical concepts and clinical-trial methodology through the Clinical Biostats educational library.
22. Record Summary
SCOUT is a phase 1/2, non-randomized, unmasked, parallel treatment study of larotrectinib in children with solid tumors harboring NTRK fusion. The registry reports an enrollment of 154 participants and 6 arms. Its primary endpoints span two statistical domains: early and longitudinal safety in phase 1, and independently reviewed overall response rate in phase 2.
The phase 1 DLT endpoint is confined to Day 1 through Day 28 of Cycle 1, making it a fixed-window dose-safety measure. The broader TEAE endpoints extend from first dose up to 93 months. The phase 2 ORR endpoint extends from first dose through disease progression or subsequent therapy or surgical intervention or death, up to 76 months, and defines response as CR or PR according to RECIST 1.1, RANO, or INRC as appropriate to tumor type.
From a statistical perspective, the most important feature is the non-randomized design. ORR can be summarized as a binomial response proportion and its uncertainty can be quantified with a confidence interval, but that estimate should remain a description of treatment outcomes in the study population rather than being interpreted as a randomized comparative treatment effect.