← Clinical Trials
RET-Altered Cancers Phase 1/2 Completed NCT03037385

ARROW: Complete Statistical Analysis of Pralsetinib in RET-Altered Cancers

An independent statistical review of the phase 1/2 ARROW trial evaluating the highly selective RET inhibitor pralsetinib in participants with thyroid cancer, non-small cell lung cancer, and other advanced solid tumors.

Trial period: 2017-03-17 to 2024-03-21  ·  Industry-sponsored  ·  590 enrolled
Scope of this record

This page separates registry-reported trial information from statistical interpretation. ClinicalTrials.gov provides the official trial registry record. The registry reports results for the registered primary endpoints, but the ClinicalTrials.gov record contains no formal statistical-analysis records with treatment-effect estimates, confidence intervals, or p-values.

Registry record: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

ARROW was a completed, non-randomized, unmasked, parallel phase 1/2 study of pralsetinib in participants with thyroid cancer, non-small cell lung cancer, and other advanced solid tumors, including RET-altered cancers. The registry reports 590 participants enrolled across 2 arms.

590
Enrolled
Total participants
2
Arms
Parallel design
1/2
Phase
Phase 1/2
3
Primary endpoints
Registry-defined
FeatureARROW
Trial nameARROW
NCT IDNCT03037385
PhasePhase 1/2
StatusCOMPLETED
AllocationNON_RANDOMIZED
Design modelPARALLEL
MaskingNONE
Primary purposeTREATMENT
Enrollment590.0
Arms2
InterventionPralsetinib (BLU-667)
Lead sponsorHoffmann-La Roche
Sponsor typeINDUSTRY

2. Clinical Question

The ARROW program addressed the development of pralsetinib in participants with thyroid cancer, non-small cell lung cancer, and other advanced solid tumors, with the ClinicalTrials.gov record specifically identifying several RET-altered disease populations.

Population

Participants with thyroid cancer, non-small cell lung cancer, and other advanced solid tumors, including RET-altered non-small cell lung cancer, medullary thyroid cancer, RET-altered papillary thyroid cancer, RET-altered colon cancer, and RET-altered solid tumors.

Intervention

Pralsetinib (BLU-667), the study drug evaluated in the ARROW program.

Comparator

The ClinicalTrials.gov record does not identify a randomized comparator treatment. The study allocation was non-randomized.

Primary questions

What dose has acceptable toxicity for further development, what adverse-event and serious-adverse-event burden is observed, and what overall response rate is observed during phase 2 evaluation?

The non-randomized design is fundamental to interpreting the study. Unlike a randomized comparative trial, ARROW does not use random assignment to create two exchangeable treatment groups for a causal treatment-versus-control comparison. Its primary statistical questions therefore center on dose selection, observed safety, and response activity rather than a randomized estimate of comparative efficacy.

3. Trial Design

01
Enroll590 participants
02
Phase 1Dose escalation
03
Dose selectionMTD / RP2D
04
Phase 2Response evaluation
05
SafetyAE / SAE assessment
Allocation
NON_RANDOMIZED. Participants were not assigned through a randomized treatment-allocation mechanism.
Design model
PARALLEL. The registry identifies 2 arms within a parallel study design.
Masking
NONE. The registry identifies the study as unmasked.
Primary purpose
TREATMENT. The study evaluates pralsetinib as an active treatment intervention.
DOSE GROUP · SAFETY DATA

Pralsetinib ≤ 300 mg QD

  • Serious adverse events: 26/37 affected/at risk
  • Registry safety data identify this as one of the reported pralsetinib dose categories.
DOSE GROUP · SAFETY DATA

Pralsetinib 400 mg QD

  • Serious adverse events: 381/540 affected/at risk
  • Registry safety data identify this as one of the reported pralsetinib dose categories.

BID dosing schedule

The registry safety data report 7/9 participants affected by serious adverse events in the pralsetinib BID dosing schedule category.

All doses

Across all pralsetinib doses, serious adverse events affected 416/590 participants according to the ClinicalTrials.gov record.

Important design distinction: the dose-specific safety categories should not be interpreted as randomized treatment arms. The ClinicalTrials.gov record identifies the overall study as non-randomized and report serious adverse events by dose category, but they do not provide a randomized comparison between those dose groups.

4. Trial Timeline and Registry Status

2017-03-17

Study start

The ARROW trial began on 2017-03-17.

2024-03-21

Primary completion

The registry reports primary completion on 2024-03-21.

Completed

Current trial status in the ClinicalTrials.gov record

the ClinicalTrials.gov record identifies ARROW as COMPLETED, with 590 participants enrolled.

the ClinicalTrials.gov record does not provide a separate enrollment-period endpoint, a data-cutoff date for each efficacy analysis, or a publication-specific analysis population. Those quantities are therefore not introduced here.

5. Primary Endpoints

EndpointTime frameRegistry definitionStatistical type
Phase 1: Maximum Tolerated Dose (MTD) and Recommended Phase 2 Dose (RP2D) of Pralsetinib Up to approximately 30.8 months MTD was defined as the highest tolerated dose of pralsetinib without causing dose limiting toxicities (DLTs). DLT was defined as any Grade ≥3 adverse event occurring during Cycle 1 during Phase 1 (dose escalation) that is not clearly caused by something other than pralsetinib. RP2D was defined as the highest dose with acceptable toxicity as determined from dose-escalation phase. Other / unclear
Phase 1 and Phase 2: Number of Participants With AEs and Serious AEs (SAEs) From Cycle 1 Day 1 up to 30 days after the final dose of study drug (up to approximately 6.7 years) An AE was any untoward medical occurrence associated with the use of a drug in humans, whether or not considered drug related. A SAE is a significant safety event as defined by the registry. Binary
Phase 2: Overall Response Rate (ORR) Up to approximately 79.8 months ORR was defined as the percentage of participants with a confirmed complete response (CR) or partial response (PR) for at least two assessments with at least 28 days apart and no disease progression (PD) in between, using RECIST v1.1. Binary

These endpoints represent three different statistical questions. The MTD/RP2D endpoint is a dose-selection problem. The AE/SAE endpoint is an event-count problem. ORR is a binary tumor-response endpoint. Treating all three as if they were conventional randomized time-to-event efficacy endpoints would misrepresent the design and the registry definitions.

6. Statistical Methodology

Dose escalation and dose selection

The phase 1 MTD/RP2D endpoint is fundamentally a dose-escalation and toxicity-assessment problem. The registry defines MTD through the occurrence of dose-limiting toxicities during Cycle 1 and defines RP2D as the highest dose with acceptable toxicity determined from the dose-escalation phase.

Dose-selection concept
MTD = highest tolerated dose without protocol-defined DLTs

The statistical objective is not to estimate a treatment-versus-control effect. Instead, dose levels are evaluated in relation to observed toxicity, with the selected dose carried forward as the recommended phase 2 dose.

This distinction matters because an MTD is a decision threshold rather than a conventional continuous or binary treatment-effect estimate. The observed dose-toxicity pattern, the number of participants evaluated at each dose, and the rules governing escalation or de-escalation determine how much evidence supports a selected dose.

Binary safety endpoints

The registered AE/SAE endpoint is a binary participant-level outcome: whether a participant experienced an adverse event or serious adverse event during the specified observation period. A natural descriptive analysis is therefore the number and percentage of affected participants, together with the denominator at risk.

Binary-event framework
Observed proportion = number of participants with event / number of participants at risk

For example, the ClinicalTrials.gov record reports serious adverse events as 416/590 across all pralsetinib doses. This is a descriptive safety proportion, not a randomized estimate of relative treatment effect.

Overall response rate

ORR is defined in the registry as the percentage of participants achieving a confirmed complete response or partial response for at least two assessments at least 28 days apart, without disease progression between those assessments. The primary statistical summary for such an endpoint is ordinarily the observed response proportion accompanied by an appropriate confidence interval.

Because ARROW is non-randomized, an ORR does not by itself quantify the causal effect of pralsetinib relative to a concurrent randomized control group. It describes the response experience within the evaluated study population under the registry's response definition.

Confidence intervals for binary endpoints

For a binary endpoint such as ORR or the proportion of participants with an SAE, a confidence interval quantifies statistical uncertainty around the observed proportion. The interval is particularly important when sample sizes are small because the observed percentage can be unstable.

Clinical Biostats interpretation

The registry-reported ARROW data provide counts for serious adverse events but do not provide confidence intervals or formal statistical-analysis records. The appropriate interpretation is therefore descriptive. A reported proportion should not be treated as a precise estimate without considering the denominator and uncertainty around that estimate.

Time-frame-specific safety assessment

The AE/SAE endpoint has a long observation window: from Cycle 1 Day 1 through 30 days after the final dose, up to approximately 6.7 years. This means the endpoint is not simply a snapshot of early toxicity. It can accumulate events over a participant's study exposure and follow-up period.

When interpreting such an endpoint, the observation period is essential. Two studies can report the same percentage of participants with an adverse event while having different follow-up durations, making the percentages difficult to compare directly without understanding the underlying exposure time and event definitions.

7. Results and Registry Reporting

Important reporting limitation: the registry-reported ARROW trial data indicate that results are posted for all 3 registered primary endpoints, but no formal statistical analyses were posted. The results below are limited to the numerical registry information actually reported.

Primary Endpoint 1: MTD and RP2D

Registry result status

Posted

Results are reported for the phase 1 MTD/RP2D endpoint.

The ClinicalTrials.gov record does not provide a numerical MTD result, a numerical RP2D result, a DLT count by dose, or a formal statistical-analysis record.

The registry definition makes the endpoint a dose-selection decision based on tolerability. The key statistical information that would ordinarily support interpretation includes the number of participants evaluated at each dose, the number and type of DLTs during Cycle 1, the rules for determining dose escalation, and the dose ultimately selected as the RP2D. Those numerical details are not present in the ClinicalTrials.gov record and are not reported here.

Clinical Biostats interpretation

An MTD or RP2D is not analogous to a hazard ratio or an odds ratio. It represents a dose-selection conclusion based on the observed toxicity pattern and the prespecified escalation framework. Without the dose-level DLT results and the formal dose-selection analysis, the ClinicalTrials.gov record supports identification of the endpoint and its statistical purpose, but not an independent numerical estimate of the selected dose.

The absence of a reported confidence interval is not evidence that the dose-selection process was imprecise; it means the ClinicalTrials.gov record does not contain such an interval.

Primary Endpoint 2: Number of Participants With AEs and SAEs

Serious adverse events across all pralsetinib doses

416 / 590

Participants affected / participants at risk

The ClinicalTrials.gov record identifies this as the all-dose safety result.

Pralsetinib dose categorySAE affectedAt risk
Pralsetinib ≤ 300 mg QD2637
Pralsetinib 400 mg QD381540
Pralsetinib BID Dosing Schedule79
Pralsetinib All Doses416590

The denominators show why dose-specific safety counts require care. The categories represent different exposure groups within a non-randomized development program rather than randomized comparator arms. Differences in the number of participants and exposure patterns can therefore affect the raw event proportions.

Clinical Biostats interpretation

The 416/590 figure is a descriptive count of participants affected by serious adverse events across all pralsetinib doses in the ClinicalTrials.gov record. It does not estimate the excess SAE risk attributable to pralsetinib relative to an untreated or randomized control population.

A proportion also does not communicate the timing, duration, severity distribution, recurrence, or exposure-adjusted incidence of individual events. The ClinicalTrials.gov record does not provide those additional analyses.

No p-value or confidence interval is reported in the registry-reported statistical-analysis field, so no formal hypothesis test or precision claim should be attached to the count.

Primary Endpoint 3: Overall Response Rate

Registry result status

Posted

Results are reported for the phase 2 ORR endpoint.

The ClinicalTrials.gov record does not provide the numerical ORR, response count, confidence interval, or p-value.

The registry definition is specific: a participant must have a confirmed complete response or partial response for at least two assessments separated by at least 28 days, with no disease progression between those assessments. This confirmation rule reduces the chance that a transient or uncertain radiologic change is counted as a confirmed response.

Clinical Biostats interpretation

For ORR, the natural statistical result would be an observed proportion of participants meeting the registry response definition, accompanied by a confidence interval. In a non-randomized study, that estimate describes response activity in the study population; it does not establish a randomized causal treatment effect.

A p-value, if one were reported, would address a prespecified hypothesis about the response proportion or a comparison—not the magnitude of the response itself. The registry-reported ARROW data contain no formal p-value or confidence interval for ORR.

8. Safety Results

The ClinicalTrials.gov record provides serious adverse event counts by pralsetinib dose category. They do not provide a complete adverse-event table by event term, grade, seriousness category, relatedness, or timing. The safety analysis below therefore remains deliberately limited to the reported serious-adverse-event counts.

Safety categoryAffectedAt risk
Pralsetinib ≤ 300 mg QD2637
Pralsetinib 400 mg QD381540
Pralsetinib BID Dosing Schedule79
Pralsetinib All Doses416590

The all-dose denominator of 590 matches the total enrollment reported in the ClinicalTrials.gov record. That correspondence provides a useful internal consistency check on the ClinicalTrials.gov record, but it does not establish that every participant had identical exposure duration or identical opportunity for an SAE.

Do not compare the dose-group percentages as if they were randomized treatment effects. The study is explicitly non-randomized, and the ClinicalTrials.gov record does not provide a randomized comparator. A higher or lower observed SAE proportion in one dose category could reflect dose, exposure duration, participant characteristics, or other differences in the populations evaluated at those doses.

9. What the Registry Endpoints Mean Statistically

MTD

A dose-selection endpoint based on whether dose-limiting toxicities occur under the protocol-defined evaluation window.

RP2D

The highest dose judged to have acceptable toxicity during the dose-escalation phase.

AE / SAE

A participant-level safety outcome observed over the specified treatment and follow-up period.

ORR

The percentage of participants meeting the registry's confirmed CR-or-PR response definition.

The endpoints are therefore complementary rather than interchangeable. MTD/RP2D addresses which dose should proceed. Safety endpoints address what adverse events occurred. ORR addresses how many participants achieved a confirmed tumor response. None of these, by itself, is a randomized estimate of comparative survival benefit.

10. Statistical Methods Explained

Why is MTD different from an ordinary efficacy endpoint?

MTD is a decision endpoint rather than a conventional comparative effect measure. The registry defines it as the highest tolerated dose without protocol-defined DLTs. The statistical evidence comes from the observed dose-toxicity experience and the prespecified escalation rules, not from comparing two randomized treatment groups.

Why does the definition of DLT specify Cycle 1?

The registry defines DLT in the context of Cycle 1 during phase 1 dose escalation. Restricting the formal DLT window creates a standardized period for making the dose-escalation decision. It does not mean that adverse events after Cycle 1 are unimportant; those events remain relevant to the broader safety endpoint.

Why is ORR treated as a binary endpoint?

Each participant ultimately either meets the registry definition of a confirmed CR or PR or does not. That makes ORR naturally expressible as a proportion. The statistical analysis would ordinarily summarize the number responding, divide by the relevant analysis denominator, and quantify uncertainty with a confidence interval.

Why does confirmation require two assessments at least 28 days apart?

The confirmation requirement reduces the likelihood that a single assessment is counted as a response when it is not sustained or when progression occurs before confirmation. It therefore makes the endpoint more specific to confirmed tumor response under the registered RECIST v1.1 framework.

Why does non-randomization change the interpretation?

Randomization is what ordinarily balances measured and unmeasured prognostic factors probabilistically between treatment groups. ARROW is explicitly non-randomized. Consequently, an observed response rate or safety proportion describes the study population but cannot, without a suitable comparator and additional assumptions, be interpreted as a causal difference attributable to pralsetinib.

Why should dose-specific SAE percentages not be treated as dose-response evidence?

The ClinicalTrials.gov record identifies different dose categories, but they do not establish randomized assignment, equal follow-up, or comparable participant characteristics across those categories. A difference in observed SAE proportions can therefore have several explanations. A formal dose-toxicity analysis would need the underlying participant-level or sufficiently detailed aggregate data and the prespecified dose-escalation framework.

What would a confidence interval add to the ORR result?

A confidence interval would show the statistical uncertainty around the observed response proportion. A high response percentage with a wide interval can provide less precise information than a somewhat lower percentage estimated with much greater precision. The registry-reported ARROW data do not provide the ORR estimate or its confidence interval, so neither can be reconstructed without additional data.

11. Planned Analysis Framework for the Posted Results

Because the ClinicalTrials.gov record contains posted outcomes but no formal statistical analyses, the most defensible approach is to distinguish what the registry reports from the analysis that would normally accompany each endpoint.

EndpointWhat is reported in the ClinicalTrials.gov recordTypical statistical treatment
MTD / RP2D Endpoint is posted; definition is reported. Summarize DLT occurrence by dose during the defined escalation window and apply the protocol's dose-selection rules to identify the highest tolerated dose and recommended phase 2 dose.
AEs / SAEs Serious-adverse-event counts are reported by dose category and across all doses. Summarize participant counts and percentages, with event definitions and observation windows clearly specified. Confidence intervals may be used for proportions when appropriate.
ORR Endpoint is posted; registry definition is reported, but the numerical ORR is not included in the ClinicalTrials.gov record. Calculate the proportion of participants with confirmed CR or PR under RECIST v1.1 and report an appropriate confidence interval. Comparative hypothesis testing would require a defined comparator or prespecified reference value.

This distinction prevents a common error in registry-based trial summaries: treating the existence of a posted outcome measure as evidence that a formal comparative statistical analysis was performed or that its numerical result is available in the ClinicalTrials.gov record.

12. Confidence Intervals, P-Values, and Effect Size

No confidence intervals or p-values are included in the posted results. That has important implications for interpretation.

Three different statistical questions
Estimate → How large is the observed effect or proportion?
Confidence interval → How precisely is it estimated?
P-value → How compatible are the data with a specified null hypothesis?

These quantities are related but answer different questions. A p-value is not a measure of effect size, and a confidence interval is not a prediction interval for individual participants.

For ORR, the effect size would ordinarily be the response proportion. For safety, the descriptive effect size could be the proportion of participants experiencing an AE or SAE. For dose selection, the key result is the selected dose under the protocol-defined toxicity rules.

Because no formal analysis estimates are reported, assigning a confidence interval or p-value would require assumptions about the analysis population, missing responses, denominator, statistical method, and null hypothesis. Those assumptions would go beyond the ClinicalTrials.gov record.

13. Analysis Population and Denominator Issues

The denominator is particularly important for ARROW because the study is phase 1/2 and non-randomized. The overall enrollment is 590.0, while the registry-reported SAE categories include 37 participants in the ≤300 mg QD category, 540 in the 400 mg QD category, and 9 in the BID dosing category.

QuantitySupplied valueInterpretive role
Total enrollment590.0Overall study enrollment in the ClinicalTrials.gov record
All-dose SAE denominator590Denominator reported for the all-dose SAE result
≤300 mg QD denominator37Dose-category safety denominator
400 mg QD denominator540Dose-category safety denominator
BID dosing denominator9Dose-category safety denominator

The ClinicalTrials.gov record does not explain the exact relationship among dose categories beyond the reported counts, nor do they provide a participant-level accounting of exposure duration. Therefore, these categories should be presented as registry-reported safety groupings rather than reconstructed as mutually exclusive randomized cohorts.

14. Statistical Interpretation of the Non-Randomized Design

Causal inference

In a randomized trial, the treatment assignment mechanism provides the foundation for interpreting differences between groups as causal, subject to the usual assumptions. ARROW does not have that feature: its allocation is explicitly non-randomized.

What the study can describe

The study can characterize dose selection, observed adverse events, serious adverse events, and tumor-response activity in the participants evaluated under its protocol. These are clinically informative descriptive quantities.

What the ClinicalTrials.gov record cannot establish

The ClinicalTrials.gov record does not support a randomized estimate of how much better or worse pralsetinib performs than a concurrent control treatment. Such a conclusion would require a suitable comparator and an analysis designed to address the resulting comparative question.

15. Limitations

16. Design Topics Not Established by the Supplied Data

Several statistical topics are often important in clinical-trial analysis, but they should not be attributed to ARROW without evidence in the ClinicalTrials.gov record.

TopicWhat can be concluded from the ClinicalTrials.gov record
Non-inferiority marginNot reported; the study is non-randomized and the ClinicalTrials.gov record does not identify a non-inferiority design.
CrossoverNot reported in the ClinicalTrials.gov record.
Factorial designNot reported. The design model is identified as parallel.
Multiplicity adjustmentNot reported in the ClinicalTrials.gov record.
Interim efficacy analysisNot reported in the ClinicalTrials.gov record.
Missing-data or imputation strategyNot reported in the ClinicalTrials.gov record.
Stratification factorsNot reported in the ClinicalTrials.gov record.
Bayesian methodsNot reported in the ClinicalTrials.gov record.

This is an important methodological discipline. The absence of a listed method in the ClinicalTrials.gov recordset is not evidence that the investigators did or did not use a particular technique elsewhere in the protocol or statistical analysis plan. It simply means that the method is not established by the information used for this page.

17. Why This Trial Matters Statistically

ARROW is a useful teaching case because it illustrates a different statistical structure from the randomized phase 3 trials that dominate clinical-trial evidence summaries. Its central questions involve dose escalation, toxicity-based dose selection, binary safety outcomes, and confirmed tumor response within a non-randomized development program.

ConceptHow it appears in ARROW
Dose escalationPhase 1 evaluation of pralsetinib with MTD and RP2D as a registered primary endpoint.
DLTProtocol-defined Grade ≥3 adverse event during Cycle 1 in phase 1 that is not clearly caused by something other than pralsetinib.
Binary endpointsAE/SAE and ORR are registered as binary-type outcomes.
Response confirmationORR requires confirmed CR or PR over at least two assessments at least 28 days apart with no progression between assessments.
Non-randomized designObserved outcomes are not based on randomized treatment allocation.
Denominator interpretationSafety results are reported with different denominators by dose category.
Long follow-upThe AE/SAE endpoint extends up to approximately 6.7 years.
Endpoint heterogeneityDose selection, safety, and tumor response answer different statistical questions.
Statistical reporting discipline

The trial also illustrates why statistical interpretation should begin with the design. A response percentage can look straightforward, but its meaning depends on whether participants were randomized, what denominator was used, how response was confirmed, and what follow-up was available. Likewise, a safety percentage is inseparable from the exposure and observation period used to define it.

18. A Practical Reading Framework for ARROW

First: identify the design

ARROW is phase 1/2, non-randomized, parallel, and unmasked. That immediately limits the type of comparative causal inference that can be made.

Second: identify the endpoint

MTD/RP2D, AE/SAE, and ORR are different endpoint classes and require different statistical summaries.

Third: inspect the denominator

The overall safety denominator is 590, while the registry-reported dose-specific safety denominators are 37, 540, and 9.

Fourth: separate evidence from inference

The registry states that results are posted, but the registry-reported statistical-analysis field contains no formal analyses.

This framework is broadly applicable to early-phase oncology trials. The statistical sophistication of a study is not determined solely by whether it reports a p-value. In dose-escalation studies, the most consequential statistical decision may instead be the rule that translates observed toxicity into a recommended dose.

19. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

20. Related Statistical Calculators

21. Sources

Continue exploring clinical-trial statistics

Connect the trial's dose-finding, binary-endpoint, safety, and response concepts to deeper statistical tutorials and analysis tools.

22. Record Summary

ARROW is a completed phase 1/2, non-randomized, unmasked, parallel study of pralsetinib in participants with thyroid cancer, non-small cell lung cancer, and other advanced solid tumors, including multiple RET-altered disease populations. Its three registered primary endpoints address dose selection through MTD/RP2D, participant-level adverse events and serious adverse events, and confirmed overall response rate.

The ClinicalTrials.gov record reports 590 participants enrolled and provide serious-adverse-event counts of 26/37 for pralsetinib ≤ 300 mg QD, 381/540 for pralsetinib 400 mg QD, 7/9 for the BID dosing schedule, and 416/590 across all pralsetinib doses. Results are posted for all three primary endpoints, but the registry-reported statistical-analysis field contains no formal analyses. Accordingly, this page does not infer missing ORR estimates, MTD or RP2D values, confidence intervals, p-values, comparative treatment effects, or other unreported statistical quantities.

Clinical Biostats methodology: A rigorous trial-results page should distinguish registry facts from statistical interpretation. For ARROW, the most important statistical lesson is that endpoint class and study design determine what can legitimately be inferred from the reported numbers. A non-randomized dose-development study should not be analyzed as though it were a randomized comparative efficacy trial.