← Clinical Trials
Scleritis Phase 1/2 Completed NCT01835132

GATSBY: Complete Statistical Analysis of Gevokizumab in Scleritis

An independent statistical review of the GATSBY phase 1/2 trial evaluating gevokizumab for scleritis, with emphasis on its single-group design and the ordinal scleral inflammation endpoint assessed at Week 16.

GATSBY  ·  NCT01835132  ·  National Eye Institute  ·  2013–2014
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

GATSBY was a completed phase 1/2, single-group trial evaluating gevokizumab in participants with scleritis. The registry records 8 participants, 1 arm, no masking, and a treatment-focused primary purpose. Its registered primary endpoint assessed whether scleral inflammation improved by at least 2 grades or reached Grade 0 by the Week 16 visit.

8
Enrollment
Participants
1
Study Arm
Single-group design
16
Primary Time Point
Week 16
0/8
Serious AEs
Affected / at risk
FeatureGATSBY
PhasePhase 1/2
StatusCOMPLETED
ConditionScleritis
Design modelSINGLE_GROUP
AllocationNA
MaskingNONE
Primary purposeTREATMENT
Enrollment8
Number of arms1
InterventionGevokizumab (device)
Lead sponsorNational Eye Institute (NEI)
Sponsor typeNIH
Start date2013-03
Primary completion date2014-07
Results postedYes
ClinicalTrials.govNCT01835132

2. Clinical Question

The registered clinical question can be framed as whether treatment with gevokizumab is associated with improvement in scleral inflammation in the study eye or eyes by the Week 16 visit among participants with scleritis.

Population

Participants with scleritis enrolled in the GATSBY study.

Intervention

Gevokizumab (device).

Comparator

There was no comparator arm. The trial used a single-group design.

Primary question

How many participants had at least a 2-step reduction or reduction to Grade 0 in scleral inflammation on or before the Week 16 visit?

3. Trial Design

01
Enroll 8 participants
02
Baseline Scleral inflammation assessment
03
Treatment Gevokizumab
04
Week 16 Primary endpoint visit
05
Assess Inflammation response
Allocation
The registry records allocation as NA because this was a single-group study rather than a randomized comparison.
Masking
Masking was recorded as NONE.
Design model
SINGLE_GROUP with 1 study arm and an enrollment of 8 participants.
Primary purpose
TREATMENT.
SINGLE STUDY ARM · n = 8

Gevokizumab

  • Intervention: Gevokizumab (device)
  • 8 participants enrolled
  • No separate comparator arm
  • Primary assessment at the Week 16 visit
COMPARATOR

No control group

  • The design is single-group
  • Allocation is recorded as NA
  • Masking is recorded as NONE
  • Between-group treatment effects therefore cannot be estimated from this design

4. Endpoints

EndpointRegistry definitionTime frame
Primary endpoint Number of Participants With at Least a 2-step Reduction or Reduction to Grade 0 in Scleral Inflammation in the Study Eye (or Eyes), According to the National Eye Institute (NEI) Photographic Scleritis Grading System, on or Before the Week 16 Visit. Baseline and Week 16

NEI Photographic Scleritis Grading System

Scleral inflammation was graded following 10% Phenylephrine application using an ordinal scale ranging from Grade 0 through Grade 4+. The registered description defines Grade 0 as no scleral inflammation with complete blanching of vessels. Grade 0.5+ represents minimal or trace inflammation, Grade 1+ mild inflammation, Grade 2+ moderate inflammation, Grade 3+ severe inflammation, and Grade 4+ necrotizing inflammation.

GradeRegistry description
0No scleral inflammation with complete blanching of vessels.
0.5+Minimal/trace inflammation with localized pink appearance of the sclera around minimally dilated deep episcleral vessels.
1+Mild inflammation with diffuse pink appearance of the sclera around mildly dilated deep episcleral vessels.
2+Moderate inflammation with purplish pink appearance of the sclera with tortuous and engorged deep episcleral vessels.
3+Severe inflammation with diffuse significant redness of sclera; the details of superficial and deep episcleral vessels can't be observed.
4+Necrotizing inflammation with diffuse redness of the sclera with scleral thinning and uveal show.

The primary endpoint is therefore a responder definition built from an ordinal clinical grading system. A participant qualifies for the endpoint if the observed change satisfies either part of the registry definition: at least a 2-step reduction, or reduction to Grade 0, on or before Week 16.

5. Results

The ClinicalTrials.gov record indicates that results were posted, but no formal statistical analyses were posted to ClinicalTrials.gov. The registry reports enrollment and safety information, but it does not report a numerical outcome estimate for the registered primary endpoint.

Primary efficacy result: The registry does not report the number or proportion of participants meeting the primary endpoint of at least a 2-step reduction or reduction to Grade 0 in scleral inflammation on or before the Week 16 visit.

What the registry reports

MeasureReported information
Enrollment8 participants
Primary endpoint resultNo numerical outcome estimate reported in the registry record
Formal statistical analysisNo formal statistical analysis posted to ClinicalTrials.gov
Serious adverse eventsGevokizumab: 0/8 affected / at risk

Typical analysis for the primary endpoint

For a single-group endpoint defined as the number of participants meeting a responder criterion, the most direct descriptive analysis would be the responder count and responder proportion among the evaluable participants. With a small sample such as 8 participants, an exact binomial confidence interval would generally be preferable to relying on a large-sample normal approximation.

Single-group responder framework
Responder proportion = number meeting the endpoint / number evaluated

The clinically meaningful quantity is the proportion of participants satisfying the prespecified response definition. An exact confidence interval can communicate the substantial sampling uncertainty that accompanies a small denominator.

Because the registry does not report the primary endpoint outcome, neither a responder proportion nor a confidence interval for that proportion is available from the registry record.

6. Safety Results

The registry reports serious adverse events by arm as 0 of 8 participants in the gevokizumab arm.

Serious adverse events

0/8

Gevokizumab: affected / at risk

Safety measureGevokizumab
Serious adverse events0/8

This is a descriptive safety result rather than a comparative treatment effect because there was no comparator group. The observation of 0 affected participants also should not be interpreted as proof that the intervention has no risk. With only 8 participants, uncommon events may not appear even when they can occur in a larger treated population.

7. Statistical Methodology

Single-group estimation

The central statistical feature of GATSBY is its single-group structure. Every enrolled participant received the same study intervention, so the primary analysis is naturally framed around estimation of the observed response frequency rather than comparison of two randomized treatment groups.

Core estimand
p = P(participant meets the prespecified Week 16 response criterion)

The parameter p represents the underlying probability that a participant satisfies the registered responder definition under the population and treatment conditions represented by the study.

Responder analysis

The primary endpoint converts the ordinal inflammation measurement into a binary clinical response indicator. A participant either meets the specified threshold of improvement or does not. This makes a binomial framework a natural way to summarize the endpoint.

Exact binomial confidence intervals

When the sample size is small, the uncertainty around a response proportion can be large. Exact binomial intervals avoid the poor coverage that can arise from simple normal-approximation intervals when the denominator is small or when the observed proportion is near 0 or 1.

Why exact inference matters here
X ~ Binomial(n, p)

If X is the number of responders among n evaluated participants, the binomial model provides a direct probability model for the responder count.

Ordinal measurement versus binary endpoint

The underlying scleral inflammation scale contains multiple ordered grades, from 0 through 4+. The primary endpoint does not retain every distinction in that ordinal scale. Instead, it defines a clinically meaningful threshold and summarizes whether that threshold was achieved.

Baseline-to-Week-16 assessment

The registry specifies the time frame as Baseline and Week 16. Statistically, the response definition therefore depends on comparing the participant's inflammation grade over the specified assessment period and determining whether the prespecified reduction criterion is met on or before the Week 16 visit.

8. Statistical Methods Explained

Why is this a single-group analysis?

The registry identifies the design model as SINGLE_GROUP and records 1 arm. There is therefore no randomized control group against which the observed response rate can be contrasted.

Why is the primary endpoint binary if the inflammation scale is ordinal?

The underlying measurement is ordinal because its categories have an ordered clinical meaning. The registered primary endpoint then applies a response threshold: at least a 2-step reduction or reduction to Grade 0. That threshold turns the underlying assessment into a yes-or-no responder outcome.

What does a responder proportion tell us?

A responder proportion describes the fraction of evaluated participants who satisfy the predefined clinical response criterion. It does not by itself describe the magnitude of improvement for every participant, the durability of improvement beyond the specified assessment period, or how the result compares with another treatment.

Why use an exact binomial interval with a small sample?

With only 8 enrolled participants, large-sample approximations can be unreliable. Exact binomial methods directly account for the discrete nature of the responder count and are therefore well suited to small single-arm studies.

Why can't a treatment hazard ratio be calculated here?

A hazard ratio is a comparative time-to-event measure that normally contrasts event hazards between at least two groups. GATSBY has one study arm and its registered primary endpoint is a response count at a defined assessment period, so a hazard ratio is not the natural effect measure for the primary endpoint.

What does 0/8 serious adverse events mean?

It means that the registry reports 0 participants affected among 8 participants at risk for the gevokizumab arm. It is a descriptive observation within this small study and does not establish that serious adverse events cannot occur.

9. Interpreting the Primary Endpoint

Clinical Biostats interpretation

The primary endpoint is designed to identify a clinically meaningful improvement in scleral inflammation rather than merely a small numerical change. By requiring at least a 2-step reduction or reduction to Grade 0, the endpoint establishes a threshold that distinguishes responders from nonresponders.

The endpoint does not measure average inflammation across the group, and it does not preserve every level of the underlying ordinal scale in the final responder classification. Two participants with different degrees of improvement can both be counted as responders once they cross the predefined threshold.

Why the sample size matters

Only 8 participants were enrolled. In a sample this small, each individual participant represents a substantial fraction of the observed study population. Consequently, even a seemingly clear responder pattern would have considerable statistical uncertainty when generalized beyond the participants studied.

What a confidence interval would add

A confidence interval around the responder proportion would quantify uncertainty due to the small sample. It would not describe the range of responses experienced by individual participants; instead, it would characterize uncertainty around the estimated population response probability under the statistical model.

Why there is no comparative effect estimate

Because the study contains one treatment arm, the primary endpoint can describe response within the treated group but cannot isolate a between-treatment effect. Without a comparator, observed improvement can be described but cannot be separated statistically from changes that might have occurred in the absence of the intervention using this study alone.

10. What This Design Can and Cannot Establish

What it can estimate

The study can describe the frequency of participants meeting its predefined scleral inflammation response criterion and summarize safety observations within the enrolled group.

What it cannot compare

The single-group design does not provide a randomized comparison of gevokizumab against another intervention or control condition.

What the ordinal scale contributes

The NEI grading system supplies an ordered measure of scleral inflammation that provides the basis for the registered response definition.

Why replication matters

A small single-group study can provide an initial estimate that may inform subsequent research, while larger comparative studies can address questions of relative treatment effect.

11. Analysis Population and Missing Data

The registry identifies an enrollment of 8 participants and provides the primary endpoint time frame as Baseline and Week 16. It does not report a separate analysis-population definition, a missing-data strategy, or an imputation method for the registered primary endpoint.

For a responder endpoint, the handling of participants without a valid Week 16 assessment can materially affect the estimated response proportion. A prespecified analysis would therefore ordinarily define which participants belong in the denominator and how missing or unevaluable assessments are handled.

Statistical caution: with a single-group study of 8 participants, the treatment effect estimate can be sensitive to the treatment of even one participant. The analysis population and missing-data rules should therefore be established before interpreting a responder proportion.

12. Limitations

13. Why This Trial Matters Statistically

GATSBY is a useful teaching example because it illustrates a different statistical problem from the large randomized survival trials that dominate many clinical-trial analyses. Its central task is estimation of a clinically defined response in a very small single-group study.

ConceptHow it appears in GATSBY
Single-group design1 study arm with an enrollment of 8 participants.
Responder endpointResponse is defined by at least a 2-step reduction or reduction to Grade 0.
Ordinal measurementScleral inflammation is graded from 0 through 4+ using the NEI Photographic Scleritis Grading System.
Binary classificationThe primary endpoint converts the ordinal assessment into responder versus nonresponder status.
Exact inferenceAn exact binomial framework is appropriate for estimating a response proportion in a very small single-group study.
Absolute estimationThe primary estimand is naturally expressed as a response probability rather than a hazard ratio.
UncertaintyA small denominator makes confidence intervals especially important when interpreting an observed response rate.
Safety descriptionSerious adverse events are reported as 0/8 for the gevokizumab arm.

14. Trial Timeline

2013-03

Study start

The registry records the GATSBY study start date as 2013-03.

Baseline

Initial inflammation assessment

The registered primary endpoint uses the baseline scleral inflammation grade as part of the response assessment.

Week 16

Primary endpoint assessment

The primary endpoint is assessed on or before the Week 16 visit.

2014-07

Primary completion

The registry records the primary completion date as 2014-07.

15. Statistical Interpretation of the Safety Result

Clinical Biostats interpretation

The reported serious-adverse-event result is 0/8 in the gevokizumab arm. This is an observed count, not a probability that serious adverse events have a true rate of zero in the broader population.

With only 8 participants, the study has limited ability to detect uncommon events. The appropriate interpretation is therefore descriptive: no serious adverse events were reported among the 8 participants at risk in the registry's arm-level result.

16. Primary Endpoint: A Deeper Statistical View

The endpoint is a thresholded change score

The underlying NEI grading system is ordered. A participant's baseline and follow-up grades can therefore be viewed as two positions on an ordinal scale. The registered endpoint then asks whether the change satisfies a specific threshold.

Conceptual responder rule
Responder = [baseline grade − follow-up grade ≥ 2] OR [follow-up grade = 0]

This expression describes the structure of the registered endpoint. The actual registry wording specifies the response criterion as at least a 2-step reduction or reduction to Grade 0 on or before Week 16.

This distinction matters because ordinal clinical measurements contain more information than a binary responder indicator. The responder endpoint is easy to communicate and clinically interpretable, but it intentionally compresses information. A full ordinal analysis could potentially distinguish smaller from larger improvements, whereas the registered endpoint focuses on crossing a prespecified response threshold.

Why the endpoint is not a continuous mean

The grades are ordinal categories rather than measurements for which equal numerical distances necessarily have a quantitative interpretation. Treating the categories as a conventional continuous variable and calculating a mean would impose assumptions that are not part of the registered endpoint definition.

Why baseline matters

The primary endpoint explicitly references reduction from the baseline state. This means the same Week 16 grade can have different responder implications depending on the participant's baseline grade. The statistical analysis therefore needs the paired baseline and follow-up assessments rather than the Week 16 grade alone.

17. Results Interpretation vs Statistical Interpretation

Reported result

The registry reports enrollment of 8 participants and serious adverse events of 0/8 for the gevokizumab arm. It does not report a numerical result for the primary scleral-inflammation response endpoint.

Statistical interpretation

The natural efficacy summary is a single-group responder proportion with uncertainty quantified using an appropriate binomial method. No between-group effect can be estimated from the study design.

18. Sources

Continue through the Clinical Biostats knowledge graph

Explore statistical methods for clinical-trial endpoints, study design, estimation, and interpretation.

19. Record Summary

GATSBY illustrates the statistical structure of a small, single-group clinical trial with an ordinal clinical measurement and a prespecified responder threshold. The study enrolled 8 participants and evaluated gevokizumab for scleritis, with the primary endpoint defined as at least a 2-step reduction or reduction to Grade 0 in scleral inflammation in the study eye or eyes on or before the Week 16 visit.

The key statistical issue is that this endpoint is an estimated response proportion, not a comparative treatment effect. In a study of this size, exact binomial inference is more appropriate than relying on large-sample approximations, and the confidence interval is essential for communicating uncertainty. The absence of a comparator also means that an observed response cannot, from this trial alone, be converted into a randomized estimate of treatment benefit.

Clinical Biostats methodology: For small single-group trials, statistical interpretation should keep three quantities distinct: the observed response count, the uncertainty around the corresponding response probability, and the causal question that the design can support. GATSBY is structured primarily around the first two; it does not contain the randomized comparator needed for a between-treatment causal estimate.