← Clinical Trials
NTRK-Fusion Solid Tumors Phase 1/2 Active, Not Recruiting NCT02637687

SCOUT: Complete Statistical Analysis of Larotrectinib in NTRK-Fusion Tumors

A detailed statistical review of the phase 1/2 SCOUT study evaluating the safety and efficacy of larotrectinib in children with tumors harboring NTRK fusion.

Start date: December 16, 2015  ·  Primary completion: July 20, 2024  ·  Sponsor: Bayer
Official registry

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

SCOUT is a phase 1/2, non-randomized, open-label treatment study of larotrectinib in children with solid tumors harboring NTRK fusion. The study uses a parallel design with 6 arms and an enrollment of 154 participants.

154
Enrollment
Clinical study population
6
Arms
Parallel design
1/2
Phase
Phase 1 and Phase 2
2015
Study start
December 16, 2015
FeatureSCOUT
Study acronymSCOUT
NCT IDNCT02637687
PhasePhase 1/2
ConditionSolid Tumors Harboring NTRK Fusion
InterventionLarotrectinib (Vitrakvi, BAY2757556)
AllocationNon-randomized
Design modelParallel
MaskingNone
Primary purposeTreatment
Enrollment154
Number of arms6
Study statusActive, not recruiting
Lead sponsorBayer
Sponsor typeIndustry

2. Clinical Question

The SCOUT study addresses two related statistical questions. In phase 1, the central question is whether dose cohorts of larotrectinib can be evaluated for treatment-emergent adverse events and dose-limiting toxicity during the first 28-day cycle. In phase 2, the principal efficacy question is whether children with NTRK gene fusions experience an objective tumor response as assessed by an independent radiology review committee.

Population

Children with solid tumors harboring NTRK fusion.

Intervention

Larotrectinib, identified in the registry as Vitrakvi and BAY2757556.

Comparator

There is no randomized comparator. The allocation is explicitly non-randomized.

Primary questions

How frequently do dose-limiting toxicities and treatment-emergent adverse events occur, and what is the overall response rate in the phase 2 population?

The absence of randomization is central to interpreting the study statistically. The design is capable of describing safety and response under treatment, but it does not create the same randomized counterfactual comparison that would arise from assigning otherwise comparable participants to larotrectinib and a control treatment.

3. Trial Design

01
Enroll 154 participants
02
Assign Non-randomized allocation
03
Treat Larotrectinib
04
Assess Safety and tumor response
05
Review Independent radiology review for ORR
Allocation
Non-randomized. Treatment assignment was not based on random allocation.
Masking
None. The registry describes the study as unmasked.
Design model
Parallel. The registry lists 6 study arms.
Purpose
Treatment. The study evaluates larotrectinib for tumors harboring NTRK fusion.

A useful distinction is between phase 1 safety evaluation and phase 2 efficacy evaluation. They use different primary endpoints because the statistical question changes. Early dose evaluation focuses on toxicity, particularly dose-limiting toxicity during Cycle 1. The phase 2 endpoint instead asks whether tumors meet predefined criteria for complete or partial response.

4. Endpoints

EndpointTime frameStatistical role
Phase 1: Number of participants in an assigned dose cohort with treatment emergent adverse events (TEAEs) by grade assessed by NCI-CTCAE v 4.03 who experience a DLT From Day 1 to Day 28 of Cycle 1 (1 Cycle=28 days) Primary safety / dose-limiting toxicity endpoint
Phase 1: Number of participants with TEAEs From first dose of larotrectinib up to 93 months Primary safety endpoint
Phase 1: Severity of TEAEs From first dose of larotrectinib up to 93 months Primary safety endpoint
Phase 2: Overall response rate (ORR) by IRRC From first dose of Larotrectinib to disease progression or subsequent therapy or surgical intervention or death, up to 76 months Primary efficacy endpoint

The registry defines dose-limiting toxicity as DLT and identifies NCI-CTCAE v 4.03 as the grading framework for treatment-emergent adverse events. For the phase 2 efficacy endpoint, the registry defines ORR as the proportion of participants with a best overall response of complete response or partial response.

The phase 2 response assessment is particularly important because it is not simply a measurement of any tumor shrinkage. The endpoint requires a best overall response of CR or PR, and the response determination is assigned by an independent radiology review committee using RECIST 1.1, RANO, or INRC as appropriate to tumor type. The population is further defined by the presence of NTRK gene fusions.

5. Primary Endpoint Structure

Study phaseEndpoint familyWhat is being measured
Phase 1 Dose-limiting toxicity Participants in an assigned dose cohort with TEAEs by grade who experience a DLT during Day 1 through Day 28 of Cycle 1.
Phase 1 Adverse events Number of participants with TEAEs from first dose through up to 93 months.
Phase 1 Adverse-event severity Severity of TEAEs from first dose through up to 93 months.
Phase 2 Overall response rate Proportion with CR or PR by independent radiology review, assessed from first dose through disease progression, subsequent therapy, surgical intervention, or death, up to 76 months.

This structure illustrates why a single statistical summary would be inadequate for SCOUT. DLT is a short-window safety endpoint; TEAEs are longitudinal safety outcomes; severity is an ordinal or categorical characterization of adverse events; and ORR is a binary response endpoint defined from tumor-assessment criteria. Each requires a different statistical description.

6. Planned Analysis

No formal statistical analyses are posted to ClinicalTrials.gov for the registered primary endpoints. The registry therefore provides the endpoint definitions and assessment windows, but not numerical estimates, confidence intervals, or p-values for those endpoints.

Phase 1: Dose-limiting toxicity

For the phase 1 DLT endpoint, the natural primary summary is the number and proportion of participants within each assigned dose cohort who experience a DLT during Day 1 through Day 28 of Cycle 1. Because the endpoint is explicitly defined within an assigned dose cohort and a fixed first-cycle window, cohort-level safety summaries are more directly aligned with the question than a time-to-event model.

Typical cohort-level summary
DLT proportion = participants with DLT / participants evaluable for the cohort endpoint

The statistical interpretation depends on the cohort definition and evaluability rules. The registry does not provide a posted statistical analysis specifying an alternative estimator or inferential procedure for this endpoint.

Phase 1: Treatment-emergent adverse events

For TEAEs, a conventional analysis would summarize the number and proportion of participants experiencing at least one event, together with event severity and clinically relevant event categories. Because the registry specifies follow-up from first dose through up to 93 months, the duration of observation is an important part of interpretation: participants may have different lengths of follow-up, and simple event proportions do not necessarily represent incidence rates over equal exposure time.

Phase 1: Severity of TEAEs

Severity is distinct from whether an adverse event occurred. A useful statistical presentation would preserve the ordered grading information rather than collapsing all adverse events into a single yes/no outcome. The registry identifies NCI-CTCAE v 4.03 for grading the DLT-related endpoint and separately specifies severity of TEAEs as a primary endpoint.

Phase 2: Overall response rate

ORR is naturally analyzed as a binomial proportion: the numerator is the number of participants whose best overall response is CR or PR, and the denominator is the applicable phase 2 analysis population. A confidence interval for the proportion would quantify uncertainty around the observed response rate.

Conceptual ORR calculation
ORR = (number with CR or PR) / (number in the defined response analysis population)

The registry's response definition is based on independent radiology review and tumor-type-appropriate criteria: RECIST 1.1, RANO, or INRC.

Important distinction: ORR is a response proportion, not a time-to-event endpoint. It does not by itself measure how long a response lasts, how long participants remain progression-free, or how long participants survive. Those questions require different endpoints and corresponding analyses.

7. Statistical Methodology

Descriptive analysis of safety

Safety endpoints in an early-phase treatment study are commonly summarized descriptively. For a binary event such as whether a participant experienced a TEAE, the basic quantities are counts and proportions. Severity can then be summarized across the applicable grading categories.

For DLT, the short assessment window is particularly meaningful. The registry defines the endpoint from Day 1 through Day 28 of Cycle 1, with one cycle defined as 28 days. That makes the endpoint a fixed-window safety measure rather than a general statement about all toxicity observed over the entire treatment period.

Binomial proportions

ORR is fundamentally a binomial endpoint because each participant can be classified according to whether their best overall response meets the CR-or-PR definition. If \(x\) participants respond among \(n\) participants, the observed response proportion is \(x/n\).

Binomial model
X ∼ Binomial(n, p)

Here, \(p\) represents the underlying probability that a participant in the defined analysis population has a best overall response of CR or PR under the endpoint definition.

For a small or moderate response sample, exact or appropriately constructed binomial confidence intervals can be preferable to relying automatically on a normal approximation. The registry does not state which confidence-interval method would be used.

Independent radiology review

The phase 2 ORR endpoint is determined by an independent radiology review committee. Statistically, this matters because response classification is not simply the investigator's subjective assessment. The independent review creates a specified assessment process for determining whether the participant's best overall response satisfies the CR or PR definition.

The criteria also vary according to tumor type. RECIST 1.1, RANO, or INRC may be used as appropriate. That means the response classification framework is linked to the characteristics of the tumor being assessed rather than assuming that one radiologic criterion applies identically to every tumor type.

Time-to-event considerations

The registry's phase 2 endpoint includes an assessment period extending from first dose until disease progression, subsequent therapy, surgical intervention, or death, up to 76 months. Those events define the endpoint's assessment context, but the primary endpoint itself is ORR rather than a reported time-to-event measure.

This distinction is important. A participant may have a response and later experience progression, while another participant may never meet the CR-or-PR definition. ORR reduces the tumor-assessment history to a binary response classification and therefore cannot by itself describe the duration or timing of the response.

8. Statistical Methods Explained

Why is DLT evaluated during the first 28 days?

The registry defines the DLT endpoint specifically from Day 1 to Day 28 of Cycle 1, with one cycle equal to 28 days. This creates a standardized early safety window. From a statistical perspective, a fixed observation window makes participants more directly comparable for the particular DLT question than an endpoint based on an unrestricted amount of follow-up.

Why summarize DLT by dose cohort?

The endpoint is explicitly defined as the number of participants in an assigned dose cohort who experience a DLT. A cohort-level analysis therefore retains information about the relationship between the assigned dose cohort and early toxicity. Combining all cohorts into a single percentage could obscure differences between dose levels.

Why is ORR analyzed as a proportion?

ORR asks whether a participant's best overall response is CR or PR. Each participant therefore contributes a response classification to the numerator or denominator. The resulting proportion is easy to interpret, while a confidence interval expresses uncertainty around the observed proportion.

Why does independent radiology review matter statistically?

Response is determined by an independent radiology review committee rather than relying solely on the treatment team. This creates a defined assessment process for the primary efficacy endpoint and helps separate response classification from the clinical team's knowledge of treatment exposure.

Why is a non-randomized design important when interpreting ORR?

Because allocation is non-randomized, an observed response proportion describes outcomes among participants treated in the study. It does not, by itself, estimate the causal difference between larotrectinib and a concurrently randomized control treatment. Without randomization, differences in patient characteristics, disease characteristics, or other factors can affect comparisons with external populations.

Why are TEAEs different from DLTs?

TEAE is a broad treatment-emergent safety concept, whereas DLT is a more specific endpoint focused on toxicity that meets the study's dose-limiting definition during the first 28-day cycle. A participant can therefore contribute to the broader TEAE assessment without necessarily having a DLT.

Why does the endpoint's stopping rule matter for ORR?

The registry defines the ORR assessment period as extending from first dose to disease progression or subsequent therapy or surgical intervention or death, up to 76 months. These events can affect whether a later tumor assessment contributes to the participant's best overall response classification. The timing and rules for assessment therefore matter to the denominator and to the resulting response estimate.

9. Interpreting an ORR Without a Comparator

A response rate is often intuitively attractive because it can be expressed as a simple percentage. Statistically, however, its meaning depends on the population and endpoint definition. In SCOUT, the relevant population is participants with tumors harboring NTRK gene fusions, and response is determined by independent radiology review using the appropriate tumor-specific response criteria.

What an ORR would mean

An ORR would represent the proportion of participants in the defined analysis population whose best overall response was complete response or partial response. It is a population-level summary of tumor response under the specified assessment framework.

What an ORR would not mean

ORR would not mean that the same proportion of participants were cured, remained progression-free for a particular length of time, or survived for a particular duration. It also would not establish a causal treatment effect against a control because the study allocation is non-randomized.

Why the confidence interval matters

A confidence interval around ORR would describe statistical uncertainty around the estimated response proportion. A wider interval would indicate less precision, while a narrower interval would indicate greater precision under the selected statistical framework.

Why a p-value is not an effect size

If a formal hypothesis test were used, its p-value would quantify evidence against a specified null hypothesis under the statistical model. It would not measure the magnitude of tumor response. The response proportion itself is the effect estimate; the confidence interval describes its precision.

10. Interpreting the Safety Endpoints

Safety requires more than one number because the SCOUT registry separates the occurrence and severity of TEAEs from the phase 1 DLT endpoint. The statistical question is therefore multidimensional: whether an event occurred, how severe it was, and whether it met the criteria for a dose-limiting toxicity during the defined first-cycle window.

Occurrence

The number of participants with TEAEs describes how frequently treatment-emergent adverse events occurred in the study population.

Severity

Severity adds information about the clinical grade or seriousness of the observed adverse events rather than simply counting any event.

DLT

DLT focuses specifically on qualifying toxicities within Day 1 through Day 28 of Cycle 1.

Follow-up

The TEAE endpoint extends from first dose up to 93 months, which is a substantially broader observation period than the DLT window.

These different windows should not be conflated. A 28-day DLT assessment answers an early dose-limiting safety question, whereas the TEAE endpoint can capture events over a much longer period. The two endpoints therefore provide complementary rather than interchangeable summaries.

11. Analysis Population and Follow-Up

The registry reports an overall enrollment of 154 participants. It does not provide, in the ClinicalTrials.gov record, numerical analysis populations for each individual primary endpoint or a posted statistical analysis specifying the exact evaluability rules for every endpoint.

FeatureRegistry informationStatistical implication
Overall enrollment 154 The enrolled population provides the overall study scale, but the applicable denominator for a specific endpoint may differ depending on endpoint-specific eligibility and assessment.
Phase 1 DLT window Day 1 to Day 28 of Cycle 1 Early toxicity is assessed within a defined fixed window.
Phase 1 TEAE window First dose up to 93 months Longer observation introduces varying exposure and follow-up considerations.
Phase 2 ORR window First dose to disease progression or subsequent therapy or surgical intervention or death, up to 76 months Response classification depends on the prespecified assessment context and stopping events.

For an endpoint such as ORR, the denominator is especially important. The meaningful question is not simply how many responses occurred, but how many participants were eligible for the response analysis under the study's endpoint rules. The registry does not post the endpoint-specific numerical denominator or a formal statistical analysis.

12. Non-Randomized Design and Causal Interpretation

Randomization is a mechanism for balancing both measured and unmeasured prognostic factors between treatment groups in expectation. SCOUT is explicitly non-randomized, so that mechanism is not present.

Core causal distinction
Observed response proportion ≠ randomized treatment effect

A response estimate describes the treated study population. A causal treatment effect requires a valid comparison with a counterfactual outcome under another treatment or condition.

This does not make the response endpoint statistically uninformative. It changes the question it can answer. A non-randomized phase 1/2 study can characterize safety, describe response, evaluate the behavior of a therapy in a defined molecular population, and provide evidence for subsequent development. It cannot by itself provide the same randomized estimate of comparative efficacy that a parallel randomized controlled trial would provide.

The distinction also affects interpretation of external comparisons. If an observed ORR were compared with a historical response rate, differences in eligibility criteria, tumor types, prior treatment, assessment practices, follow-up, and patient characteristics could contribute to the observed difference. Such a comparison requires assumptions beyond the response proportion itself.

13. Six-Arm Parallel Structure

The registry lists 6 arms within a parallel design, while the intervention is larotrectinib. Because allocation is non-randomized and the ClinicalTrials.gov record does not provide separate numerical enrollment or outcome results for each arm, the overall enrollment should not be interpreted as six equally sized groups.

Number of arms
6 arms are listed in the registry.
Allocation
Non-randomized allocation means arm assignment does not arise from randomization.
Design model
Parallel design indicates distinct treatment-arm pathways rather than a crossover design.
Intervention
Larotrectinib is the registered drug intervention.

For statistical reporting, arm-specific results are most useful when the denominators are explicit. A single study-wide response proportion can hide important differences in the composition of the participants contributing to that estimate. Conversely, comparing small arm-specific proportions without accounting for their precision can make random variation appear more meaningful than it is.

14. Timing of the Study

December 16, 2015

Study start

The registry lists December 16, 2015 as the start date for SCOUT.

July 20, 2024

Primary completion

The registry lists July 20, 2024 as the primary completion date.

Current registry status

Active, not recruiting

The ClinicalTrials.gov record currently identifies the study status as active, not recruiting.

The timing of follow-up is relevant to the statistical interpretation of the endpoints. DLT has a fixed 28-day Cycle 1 window, whereas TEAEs can be assessed through up to 93 months and the phase 2 ORR assessment extends through a defined event-based period up to 76 months. These are fundamentally different observation structures.

15. Missing Data and Assessment Issues

The registry specifies several events that can terminate the phase 2 ORR assessment period: disease progression, subsequent therapy, surgical intervention, or death. These events are statistically relevant because response assessment is not conducted in an unlimited observational window.

For a response endpoint, missing tumor assessments can be consequential. If participants are not assessed after treatment begins, the statistical analysis must follow prespecified rules for determining whether they are evaluable and how their lack of an assessable response is handled. The ClinicalTrials.gov record does not post those detailed imputation or missing-data rules.

Interpretation caution: an ORR estimate should always be read together with its analysis population and assessment rules. The same observed number of responses can produce different response proportions if the denominator or evaluability definition differs.

The same principle applies to safety. Participants may have different durations of treatment and follow-up, particularly when the TEAE window extends to 93 months. A simple proportion of participants with an event is therefore not automatically equivalent to an incidence rate or a time-to-first-event estimate.

16. What This Design Can and Cannot Establish

Safety characterization

The phase 1 endpoints can describe DLTs, TEAEs, and their severity within the defined study population and observation periods.

Response characterization

The phase 2 endpoint can describe the proportion of participants achieving CR or PR under the independent radiology review framework.

Molecularly defined population

The ORR definition specifically concerns participants who express NTRK gene fusions, linking the efficacy question to the molecular characteristic in the registry endpoint.

Comparative efficacy

Because allocation is non-randomized and no comparator is listed, the study does not provide a randomized treatment-versus-control efficacy contrast.

This distinction is one of the most important statistical lessons from SCOUT. An efficacy endpoint can be precisely defined without being a randomized comparative endpoint. Statistical precision and causal identification are different properties: a very precisely estimated response proportion can still be a descriptive treatment outcome rather than a randomized estimate of treatment effect.

17. Important Limitations and Interpretation Issues

18. Why This Trial Matters Statistically

SCOUT is a useful statistical teaching case because it combines early-phase safety methodology with a molecularly defined efficacy endpoint. The study also illustrates why the design of a clinical trial determines the interpretation of its numerical results.

ConceptHow it appears in SCOUT
Phase 1 safety analysis DLT, TEAE occurrence, and TEAE severity are registered primary endpoints.
Fixed-window safety assessment DLT is assessed from Day 1 through Day 28 of Cycle 1.
Longitudinal safety TEAEs are assessed from first dose up to 93 months.
Binary efficacy endpoint ORR classifies participants according to whether their best overall response is CR or PR.
Independent assessment Phase 2 ORR is determined by an independent radiology review committee.
Multiple response criteria RECIST 1.1, RANO, or INRC is used as appropriate to tumor type.
Non-randomized design Allocation is explicitly non-randomized, so response estimates are not randomized treatment effects.
Parallel design The registry lists 6 arms in a parallel study model.
Endpoint-specific follow-up Different primary endpoints use different observation windows and stopping events.
Confidence intervals A binomial confidence interval would quantify uncertainty around an ORR estimate if reported.

19. Statistical Interpretation vs Clinical Interpretation

Statistical interpretation

The study is structured to describe early dose-limiting toxicity, treatment-emergent adverse events, adverse-event severity, and overall response. The appropriate statistical summaries differ because these endpoints represent different types of data.

Clinical interpretation

The clinical meaning of an observed response or adverse-event rate depends on the characteristics of the pediatric population, the NTRK-fusion tumor context, the treatment exposure, the assessment criteria, and the duration of follow-up.

Keeping these two interpretations separate is especially important in a non-randomized study. Statistical analysis can describe the observed pattern with considerable precision, but precision does not substitute for a randomized comparator when the question is comparative causal efficacy.

20. A Framework for Reading Future SCOUT Results

If formal numerical results are posted for the primary endpoints, the most informative statistical reading would begin with the denominator and endpoint definition before examining any p-value.

QuestionWhy it matters
How many participants were included in the endpoint analysis? Defines the denominator and determines how much information contributes to the estimate.
How many participants experienced the event? Provides the numerator for a binary safety or response proportion.
What was the confidence interval? Shows the statistical precision of the estimated proportion.
Which response criteria were used? Determines how CR and PR were classified.
Was the analysis independently reviewed? For ORR, the registry specifies independent radiology review.
How long were participants followed? Safety and response observation windows differ substantially in SCOUT.
Was there a comparator? The registered design is non-randomized and does not identify a comparator treatment.

This sequence prevents a common statistical error: interpreting a numerical estimate before establishing exactly what population, endpoint definition, assessment window, and denominator produced it.

21. Sources

Continue through the Clinical Biostats knowledge graph

Explore broader statistical concepts and clinical-trial methodology through the Clinical Biostats educational library.

22. Record Summary

SCOUT is a phase 1/2, non-randomized, unmasked, parallel treatment study of larotrectinib in children with solid tumors harboring NTRK fusion. The registry reports an enrollment of 154 participants and 6 arms. Its primary endpoints span two statistical domains: early and longitudinal safety in phase 1, and independently reviewed overall response rate in phase 2.

The phase 1 DLT endpoint is confined to Day 1 through Day 28 of Cycle 1, making it a fixed-window dose-safety measure. The broader TEAE endpoints extend from first dose up to 93 months. The phase 2 ORR endpoint extends from first dose through disease progression or subsequent therapy or surgical intervention or death, up to 76 months, and defines response as CR or PR according to RECIST 1.1, RANO, or INRC as appropriate to tumor type.

From a statistical perspective, the most important feature is the non-randomized design. ORR can be summarized as a binomial response proportion and its uncertainty can be quantified with a confidence interval, but that estimate should remain a description of treatment outcomes in the study population rather than being interpreted as a randomized comparative treatment effect.

Clinical Biostats methodology: The statistical story of a trial begins with its design and endpoint definitions. For SCOUT, understanding the distinction between fixed-window DLT, long-term TEAE assessment, independently reviewed ORR, and non-randomized allocation is essential before interpreting any future numerical results.