This page separates registry-reported trial information from statistical interpretation. ClinicalTrials.gov provides the official trial registry record. The registry reports results for the registered primary endpoints, but the ClinicalTrials.gov record contains no formal statistical-analysis records with treatment-effect estimates, confidence intervals, or p-values.
1. Trial at a Glance
ARROW was a completed, non-randomized, unmasked, parallel phase 1/2 study of pralsetinib in participants with thyroid cancer, non-small cell lung cancer, and other advanced solid tumors, including RET-altered cancers. The registry reports 590 participants enrolled across 2 arms.
| Feature | ARROW |
|---|---|
| Trial name | ARROW |
| NCT ID | NCT03037385 |
| Phase | Phase 1/2 |
| Status | COMPLETED |
| Allocation | NON_RANDOMIZED |
| Design model | PARALLEL |
| Masking | NONE |
| Primary purpose | TREATMENT |
| Enrollment | 590.0 |
| Arms | 2 |
| Intervention | Pralsetinib (BLU-667) |
| Lead sponsor | Hoffmann-La Roche |
| Sponsor type | INDUSTRY |
2. Clinical Question
The ARROW program addressed the development of pralsetinib in participants with thyroid cancer, non-small cell lung cancer, and other advanced solid tumors, with the ClinicalTrials.gov record specifically identifying several RET-altered disease populations.
Population
Participants with thyroid cancer, non-small cell lung cancer, and other advanced solid tumors, including RET-altered non-small cell lung cancer, medullary thyroid cancer, RET-altered papillary thyroid cancer, RET-altered colon cancer, and RET-altered solid tumors.
Intervention
Pralsetinib (BLU-667), the study drug evaluated in the ARROW program.
Comparator
The ClinicalTrials.gov record does not identify a randomized comparator treatment. The study allocation was non-randomized.
Primary questions
What dose has acceptable toxicity for further development, what adverse-event and serious-adverse-event burden is observed, and what overall response rate is observed during phase 2 evaluation?
The non-randomized design is fundamental to interpreting the study. Unlike a randomized comparative trial, ARROW does not use random assignment to create two exchangeable treatment groups for a causal treatment-versus-control comparison. Its primary statistical questions therefore center on dose selection, observed safety, and response activity rather than a randomized estimate of comparative efficacy.
3. Trial Design
Pralsetinib ≤ 300 mg QD
- Serious adverse events: 26/37 affected/at risk
- Registry safety data identify this as one of the reported pralsetinib dose categories.
Pralsetinib 400 mg QD
- Serious adverse events: 381/540 affected/at risk
- Registry safety data identify this as one of the reported pralsetinib dose categories.
BID dosing schedule
The registry safety data report 7/9 participants affected by serious adverse events in the pralsetinib BID dosing schedule category.
All doses
Across all pralsetinib doses, serious adverse events affected 416/590 participants according to the ClinicalTrials.gov record.
4. Trial Timeline and Registry Status
Study start
The ARROW trial began on 2017-03-17.
Primary completion
The registry reports primary completion on 2024-03-21.
Current trial status in the ClinicalTrials.gov record
the ClinicalTrials.gov record identifies ARROW as COMPLETED, with 590 participants enrolled.
the ClinicalTrials.gov record does not provide a separate enrollment-period endpoint, a data-cutoff date for each efficacy analysis, or a publication-specific analysis population. Those quantities are therefore not introduced here.
5. Primary Endpoints
| Endpoint | Time frame | Registry definition | Statistical type |
|---|---|---|---|
| Phase 1: Maximum Tolerated Dose (MTD) and Recommended Phase 2 Dose (RP2D) of Pralsetinib | Up to approximately 30.8 months | MTD was defined as the highest tolerated dose of pralsetinib without causing dose limiting toxicities (DLTs). DLT was defined as any Grade ≥3 adverse event occurring during Cycle 1 during Phase 1 (dose escalation) that is not clearly caused by something other than pralsetinib. RP2D was defined as the highest dose with acceptable toxicity as determined from dose-escalation phase. | Other / unclear |
| Phase 1 and Phase 2: Number of Participants With AEs and Serious AEs (SAEs) | From Cycle 1 Day 1 up to 30 days after the final dose of study drug (up to approximately 6.7 years) | An AE was any untoward medical occurrence associated with the use of a drug in humans, whether or not considered drug related. A SAE is a significant safety event as defined by the registry. | Binary |
| Phase 2: Overall Response Rate (ORR) | Up to approximately 79.8 months | ORR was defined as the percentage of participants with a confirmed complete response (CR) or partial response (PR) for at least two assessments with at least 28 days apart and no disease progression (PD) in between, using RECIST v1.1. | Binary |
These endpoints represent three different statistical questions. The MTD/RP2D endpoint is a dose-selection problem. The AE/SAE endpoint is an event-count problem. ORR is a binary tumor-response endpoint. Treating all three as if they were conventional randomized time-to-event efficacy endpoints would misrepresent the design and the registry definitions.
6. Statistical Methodology
Dose escalation and dose selection
The phase 1 MTD/RP2D endpoint is fundamentally a dose-escalation and toxicity-assessment problem. The registry defines MTD through the occurrence of dose-limiting toxicities during Cycle 1 and defines RP2D as the highest dose with acceptable toxicity determined from the dose-escalation phase.
The statistical objective is not to estimate a treatment-versus-control effect. Instead, dose levels are evaluated in relation to observed toxicity, with the selected dose carried forward as the recommended phase 2 dose.
This distinction matters because an MTD is a decision threshold rather than a conventional continuous or binary treatment-effect estimate. The observed dose-toxicity pattern, the number of participants evaluated at each dose, and the rules governing escalation or de-escalation determine how much evidence supports a selected dose.
Binary safety endpoints
The registered AE/SAE endpoint is a binary participant-level outcome: whether a participant experienced an adverse event or serious adverse event during the specified observation period. A natural descriptive analysis is therefore the number and percentage of affected participants, together with the denominator at risk.
For example, the ClinicalTrials.gov record reports serious adverse events as 416/590 across all pralsetinib doses. This is a descriptive safety proportion, not a randomized estimate of relative treatment effect.
Overall response rate
ORR is defined in the registry as the percentage of participants achieving a confirmed complete response or partial response for at least two assessments at least 28 days apart, without disease progression between those assessments. The primary statistical summary for such an endpoint is ordinarily the observed response proportion accompanied by an appropriate confidence interval.
Because ARROW is non-randomized, an ORR does not by itself quantify the causal effect of pralsetinib relative to a concurrent randomized control group. It describes the response experience within the evaluated study population under the registry's response definition.
Confidence intervals for binary endpoints
For a binary endpoint such as ORR or the proportion of participants with an SAE, a confidence interval quantifies statistical uncertainty around the observed proportion. The interval is particularly important when sample sizes are small because the observed percentage can be unstable.
The registry-reported ARROW data provide counts for serious adverse events but do not provide confidence intervals or formal statistical-analysis records. The appropriate interpretation is therefore descriptive. A reported proportion should not be treated as a precise estimate without considering the denominator and uncertainty around that estimate.
Time-frame-specific safety assessment
The AE/SAE endpoint has a long observation window: from Cycle 1 Day 1 through 30 days after the final dose, up to approximately 6.7 years. This means the endpoint is not simply a snapshot of early toxicity. It can accumulate events over a participant's study exposure and follow-up period.
When interpreting such an endpoint, the observation period is essential. Two studies can report the same percentage of participants with an adverse event while having different follow-up durations, making the percentages difficult to compare directly without understanding the underlying exposure time and event definitions.
7. Results and Registry Reporting
Primary Endpoint 1: MTD and RP2D
Registry result status
Results are reported for the phase 1 MTD/RP2D endpoint.
The ClinicalTrials.gov record does not provide a numerical MTD result, a numerical RP2D result, a DLT count by dose, or a formal statistical-analysis record.
The registry definition makes the endpoint a dose-selection decision based on tolerability. The key statistical information that would ordinarily support interpretation includes the number of participants evaluated at each dose, the number and type of DLTs during Cycle 1, the rules for determining dose escalation, and the dose ultimately selected as the RP2D. Those numerical details are not present in the ClinicalTrials.gov record and are not reported here.
An MTD or RP2D is not analogous to a hazard ratio or an odds ratio. It represents a dose-selection conclusion based on the observed toxicity pattern and the prespecified escalation framework. Without the dose-level DLT results and the formal dose-selection analysis, the ClinicalTrials.gov record supports identification of the endpoint and its statistical purpose, but not an independent numerical estimate of the selected dose.
The absence of a reported confidence interval is not evidence that the dose-selection process was imprecise; it means the ClinicalTrials.gov record does not contain such an interval.
Primary Endpoint 2: Number of Participants With AEs and SAEs
Serious adverse events across all pralsetinib doses
Participants affected / participants at risk
The ClinicalTrials.gov record identifies this as the all-dose safety result.
| Pralsetinib dose category | SAE affected | At risk |
|---|---|---|
| Pralsetinib ≤ 300 mg QD | 26 | 37 |
| Pralsetinib 400 mg QD | 381 | 540 |
| Pralsetinib BID Dosing Schedule | 7 | 9 |
| Pralsetinib All Doses | 416 | 590 |
The denominators show why dose-specific safety counts require care. The categories represent different exposure groups within a non-randomized development program rather than randomized comparator arms. Differences in the number of participants and exposure patterns can therefore affect the raw event proportions.
The 416/590 figure is a descriptive count of participants affected by serious adverse events across all pralsetinib doses in the ClinicalTrials.gov record. It does not estimate the excess SAE risk attributable to pralsetinib relative to an untreated or randomized control population.
A proportion also does not communicate the timing, duration, severity distribution, recurrence, or exposure-adjusted incidence of individual events. The ClinicalTrials.gov record does not provide those additional analyses.
No p-value or confidence interval is reported in the registry-reported statistical-analysis field, so no formal hypothesis test or precision claim should be attached to the count.
Primary Endpoint 3: Overall Response Rate
Registry result status
Results are reported for the phase 2 ORR endpoint.
The ClinicalTrials.gov record does not provide the numerical ORR, response count, confidence interval, or p-value.
The registry definition is specific: a participant must have a confirmed complete response or partial response for at least two assessments separated by at least 28 days, with no disease progression between those assessments. This confirmation rule reduces the chance that a transient or uncertain radiologic change is counted as a confirmed response.
For ORR, the natural statistical result would be an observed proportion of participants meeting the registry response definition, accompanied by a confidence interval. In a non-randomized study, that estimate describes response activity in the study population; it does not establish a randomized causal treatment effect.
A p-value, if one were reported, would address a prespecified hypothesis about the response proportion or a comparison—not the magnitude of the response itself. The registry-reported ARROW data contain no formal p-value or confidence interval for ORR.
8. Safety Results
The ClinicalTrials.gov record provides serious adverse event counts by pralsetinib dose category. They do not provide a complete adverse-event table by event term, grade, seriousness category, relatedness, or timing. The safety analysis below therefore remains deliberately limited to the reported serious-adverse-event counts.
| Safety category | Affected | At risk |
|---|---|---|
| Pralsetinib ≤ 300 mg QD | 26 | 37 |
| Pralsetinib 400 mg QD | 381 | 540 |
| Pralsetinib BID Dosing Schedule | 7 | 9 |
| Pralsetinib All Doses | 416 | 590 |
The all-dose denominator of 590 matches the total enrollment reported in the ClinicalTrials.gov record. That correspondence provides a useful internal consistency check on the ClinicalTrials.gov record, but it does not establish that every participant had identical exposure duration or identical opportunity for an SAE.
9. What the Registry Endpoints Mean Statistically
MTD
A dose-selection endpoint based on whether dose-limiting toxicities occur under the protocol-defined evaluation window.
RP2D
The highest dose judged to have acceptable toxicity during the dose-escalation phase.
AE / SAE
A participant-level safety outcome observed over the specified treatment and follow-up period.
ORR
The percentage of participants meeting the registry's confirmed CR-or-PR response definition.
The endpoints are therefore complementary rather than interchangeable. MTD/RP2D addresses which dose should proceed. Safety endpoints address what adverse events occurred. ORR addresses how many participants achieved a confirmed tumor response. None of these, by itself, is a randomized estimate of comparative survival benefit.
10. Statistical Methods Explained
Why is MTD different from an ordinary efficacy endpoint?
MTD is a decision endpoint rather than a conventional comparative effect measure. The registry defines it as the highest tolerated dose without protocol-defined DLTs. The statistical evidence comes from the observed dose-toxicity experience and the prespecified escalation rules, not from comparing two randomized treatment groups.
Why does the definition of DLT specify Cycle 1?
The registry defines DLT in the context of Cycle 1 during phase 1 dose escalation. Restricting the formal DLT window creates a standardized period for making the dose-escalation decision. It does not mean that adverse events after Cycle 1 are unimportant; those events remain relevant to the broader safety endpoint.
Why is ORR treated as a binary endpoint?
Each participant ultimately either meets the registry definition of a confirmed CR or PR or does not. That makes ORR naturally expressible as a proportion. The statistical analysis would ordinarily summarize the number responding, divide by the relevant analysis denominator, and quantify uncertainty with a confidence interval.
Why does confirmation require two assessments at least 28 days apart?
The confirmation requirement reduces the likelihood that a single assessment is counted as a response when it is not sustained or when progression occurs before confirmation. It therefore makes the endpoint more specific to confirmed tumor response under the registered RECIST v1.1 framework.
Why does non-randomization change the interpretation?
Randomization is what ordinarily balances measured and unmeasured prognostic factors probabilistically between treatment groups. ARROW is explicitly non-randomized. Consequently, an observed response rate or safety proportion describes the study population but cannot, without a suitable comparator and additional assumptions, be interpreted as a causal difference attributable to pralsetinib.
Why should dose-specific SAE percentages not be treated as dose-response evidence?
The ClinicalTrials.gov record identifies different dose categories, but they do not establish randomized assignment, equal follow-up, or comparable participant characteristics across those categories. A difference in observed SAE proportions can therefore have several explanations. A formal dose-toxicity analysis would need the underlying participant-level or sufficiently detailed aggregate data and the prespecified dose-escalation framework.
What would a confidence interval add to the ORR result?
A confidence interval would show the statistical uncertainty around the observed response proportion. A high response percentage with a wide interval can provide less precise information than a somewhat lower percentage estimated with much greater precision. The registry-reported ARROW data do not provide the ORR estimate or its confidence interval, so neither can be reconstructed without additional data.
11. Planned Analysis Framework for the Posted Results
Because the ClinicalTrials.gov record contains posted outcomes but no formal statistical analyses, the most defensible approach is to distinguish what the registry reports from the analysis that would normally accompany each endpoint.
| Endpoint | What is reported in the ClinicalTrials.gov record | Typical statistical treatment |
|---|---|---|
| MTD / RP2D | Endpoint is posted; definition is reported. | Summarize DLT occurrence by dose during the defined escalation window and apply the protocol's dose-selection rules to identify the highest tolerated dose and recommended phase 2 dose. |
| AEs / SAEs | Serious-adverse-event counts are reported by dose category and across all doses. | Summarize participant counts and percentages, with event definitions and observation windows clearly specified. Confidence intervals may be used for proportions when appropriate. |
| ORR | Endpoint is posted; registry definition is reported, but the numerical ORR is not included in the ClinicalTrials.gov record. | Calculate the proportion of participants with confirmed CR or PR under RECIST v1.1 and report an appropriate confidence interval. Comparative hypothesis testing would require a defined comparator or prespecified reference value. |
This distinction prevents a common error in registry-based trial summaries: treating the existence of a posted outcome measure as evidence that a formal comparative statistical analysis was performed or that its numerical result is available in the ClinicalTrials.gov record.
12. Confidence Intervals, P-Values, and Effect Size
No confidence intervals or p-values are included in the posted results. That has important implications for interpretation.
Confidence interval → How precisely is it estimated?
P-value → How compatible are the data with a specified null hypothesis?
These quantities are related but answer different questions. A p-value is not a measure of effect size, and a confidence interval is not a prediction interval for individual participants.
For ORR, the effect size would ordinarily be the response proportion. For safety, the descriptive effect size could be the proportion of participants experiencing an AE or SAE. For dose selection, the key result is the selected dose under the protocol-defined toxicity rules.
Because no formal analysis estimates are reported, assigning a confidence interval or p-value would require assumptions about the analysis population, missing responses, denominator, statistical method, and null hypothesis. Those assumptions would go beyond the ClinicalTrials.gov record.
13. Analysis Population and Denominator Issues
The denominator is particularly important for ARROW because the study is phase 1/2 and non-randomized. The overall enrollment is 590.0, while the registry-reported SAE categories include 37 participants in the ≤300 mg QD category, 540 in the 400 mg QD category, and 9 in the BID dosing category.
| Quantity | Supplied value | Interpretive role |
|---|---|---|
| Total enrollment | 590.0 | Overall study enrollment in the ClinicalTrials.gov record |
| All-dose SAE denominator | 590 | Denominator reported for the all-dose SAE result |
| ≤300 mg QD denominator | 37 | Dose-category safety denominator |
| 400 mg QD denominator | 540 | Dose-category safety denominator |
| BID dosing denominator | 9 | Dose-category safety denominator |
The ClinicalTrials.gov record does not explain the exact relationship among dose categories beyond the reported counts, nor do they provide a participant-level accounting of exposure duration. Therefore, these categories should be presented as registry-reported safety groupings rather than reconstructed as mutually exclusive randomized cohorts.
14. Statistical Interpretation of the Non-Randomized Design
In a randomized trial, the treatment assignment mechanism provides the foundation for interpreting differences between groups as causal, subject to the usual assumptions. ARROW does not have that feature: its allocation is explicitly non-randomized.
The study can characterize dose selection, observed adverse events, serious adverse events, and tumor-response activity in the participants evaluated under its protocol. These are clinically informative descriptive quantities.
The ClinicalTrials.gov record does not support a randomized estimate of how much better or worse pralsetinib performs than a concurrent control treatment. Such a conclusion would require a suitable comparator and an analysis designed to address the resulting comparative question.
15. Limitations
- Non-randomized allocation: the trial does not provide the treatment-group balance created by randomization, limiting causal interpretation of observed response and safety outcomes.
- No formal statistical-analysis records reported: results are posted, but no formal statistical analyses were posted. Numerical estimates, confidence intervals, and p-values are therefore not available for reconstruction.
- Incomplete endpoint result detail: the ClinicalTrials.gov record identifies that MTD/RP2D and ORR results are posted, but do not provide their numerical results.
- Safety denominator differences: the dose-category safety denominators differ substantially, and the ClinicalTrials.gov record does not provide enough detail to determine exposure-adjusted event rates.
- Long observation window: the AE/SAE endpoint extends from Cycle 1 Day 1 through 30 days after the final dose, up to approximately 6.7 years. Simple participant proportions do not capture differences in exposure time.
- ORR is not survival: response rate measures confirmed tumor response under the registry definition and does not directly measure duration of survival.
- No randomized comparator: differences among dose categories cannot be interpreted as randomized dose effects.
- Unreported statistical details: the ClinicalTrials.gov record does not identify the confidence-interval method, missing-data strategy, formal hypothesis-testing framework, multiplicity procedure, or interim-analysis method for the posted results.
- Population heterogeneity: the ClinicalTrials.gov record includes multiple disease settings, including RET-altered non-small cell lung cancer, medullary thyroid cancer, RET-altered papillary thyroid cancer, RET-altered colon cancer, and other RET-altered solid tumors. An overall study summary may therefore combine clinically distinct populations.
16. Design Topics Not Established by the Supplied Data
Several statistical topics are often important in clinical-trial analysis, but they should not be attributed to ARROW without evidence in the ClinicalTrials.gov record.
| Topic | What can be concluded from the ClinicalTrials.gov record |
|---|---|
| Non-inferiority margin | Not reported; the study is non-randomized and the ClinicalTrials.gov record does not identify a non-inferiority design. |
| Crossover | Not reported in the ClinicalTrials.gov record. |
| Factorial design | Not reported. The design model is identified as parallel. |
| Multiplicity adjustment | Not reported in the ClinicalTrials.gov record. |
| Interim efficacy analysis | Not reported in the ClinicalTrials.gov record. |
| Missing-data or imputation strategy | Not reported in the ClinicalTrials.gov record. |
| Stratification factors | Not reported in the ClinicalTrials.gov record. |
| Bayesian methods | Not reported in the ClinicalTrials.gov record. |
This is an important methodological discipline. The absence of a listed method in the ClinicalTrials.gov recordset is not evidence that the investigators did or did not use a particular technique elsewhere in the protocol or statistical analysis plan. It simply means that the method is not established by the information used for this page.
17. Why This Trial Matters Statistically
ARROW is a useful teaching case because it illustrates a different statistical structure from the randomized phase 3 trials that dominate clinical-trial evidence summaries. Its central questions involve dose escalation, toxicity-based dose selection, binary safety outcomes, and confirmed tumor response within a non-randomized development program.
| Concept | How it appears in ARROW |
|---|---|
| Dose escalation | Phase 1 evaluation of pralsetinib with MTD and RP2D as a registered primary endpoint. |
| DLT | Protocol-defined Grade ≥3 adverse event during Cycle 1 in phase 1 that is not clearly caused by something other than pralsetinib. |
| Binary endpoints | AE/SAE and ORR are registered as binary-type outcomes. |
| Response confirmation | ORR requires confirmed CR or PR over at least two assessments at least 28 days apart with no progression between assessments. |
| Non-randomized design | Observed outcomes are not based on randomized treatment allocation. |
| Denominator interpretation | Safety results are reported with different denominators by dose category. |
| Long follow-up | The AE/SAE endpoint extends up to approximately 6.7 years. |
| Endpoint heterogeneity | Dose selection, safety, and tumor response answer different statistical questions. |
| Statistical reporting discipline |
The trial also illustrates why statistical interpretation should begin with the design. A response percentage can look straightforward, but its meaning depends on whether participants were randomized, what denominator was used, how response was confirmed, and what follow-up was available. Likewise, a safety percentage is inseparable from the exposure and observation period used to define it.
18. A Practical Reading Framework for ARROW
First: identify the design
ARROW is phase 1/2, non-randomized, parallel, and unmasked. That immediately limits the type of comparative causal inference that can be made.
Second: identify the endpoint
MTD/RP2D, AE/SAE, and ORR are different endpoint classes and require different statistical summaries.
Third: inspect the denominator
The overall safety denominator is 590, while the registry-reported dose-specific safety denominators are 37, 540, and 9.
Fourth: separate evidence from inference
The registry states that results are posted, but the registry-reported statistical-analysis field contains no formal analyses.
This framework is broadly applicable to early-phase oncology trials. The statistical sophistication of a study is not determined solely by whether it reports a p-value. In dose-escalation studies, the most consequential statistical decision may instead be the rule that translates observed toxicity into a recommended dose.
19. Statistical Concepts in This Trial
Learn more about the methods used in this trial:
20. Related Statistical Calculators
21. Sources
- ClinicalTrials.gov: NCT03037385 — ARROW. The registry is the source for the trial design, endpoints, time frames, enrollment, status, and the ClinicalTrials.gov record summarized on this page.
- Linked publication: PMID 41886723.
- Linked publication: PMID 40639400.
- Linked publication: PMID 38009200.
- Linked publication: PMID 37934000.
- Linked publication: PMID 37916501.
Continue exploring clinical-trial statistics
Connect the trial's dose-finding, binary-endpoint, safety, and response concepts to deeper statistical tutorials and analysis tools.
22. Record Summary
ARROW is a completed phase 1/2, non-randomized, unmasked, parallel study of pralsetinib in participants with thyroid cancer, non-small cell lung cancer, and other advanced solid tumors, including multiple RET-altered disease populations. Its three registered primary endpoints address dose selection through MTD/RP2D, participant-level adverse events and serious adverse events, and confirmed overall response rate.
The ClinicalTrials.gov record reports 590 participants enrolled and provide serious-adverse-event counts of 26/37 for pralsetinib ≤ 300 mg QD, 381/540 for pralsetinib 400 mg QD, 7/9 for the BID dosing schedule, and 416/590 across all pralsetinib doses. Results are posted for all three primary endpoints, but the registry-reported statistical-analysis field contains no formal analyses. Accordingly, this page does not infer missing ORR estimates, MTD or RP2D values, confidence intervals, p-values, comparative treatment effects, or other unreported statistical quantities.