← Clinical Trials
EGFR-Mutation-Positive NSCLC Phase 3 Completed NCT02296125

FLAURA: Complete Statistical Analysis of Osimertinib in EGFR Mutation-Positive Non-Small Cell Lung Cancer

An independent statistical review of the randomized phase 3 FLAURA trial comparing osimertinib 80 mg with standard-of-care EGFR-TKI therapy in patients with locally advanced or metastatic EGFR sensitising mutation-positive non-small cell lung cancer.

Trial period: 2014-12-03 to 2017-06-19  ·  Lead sponsor: AstraZeneca  ·  Enrollment: 674
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record. Numerical estimates on this page are restricted to the ClinicalTrials.gov record.

1. Trial at a Glance

FLAURA was a randomized, parallel, triple-masked phase 3 treatment trial evaluating osimertinib against standard-of-care EGFR-TKI therapy in patients with locally advanced or metastatic EGFR sensitising mutation-positive non-small cell lung cancer. The registry reports 674 participants, two arms, 13 posted outcome measures, and 12 posted statistical analyses.

674
Enrollment
Registry enrollment
2
Arms
Parallel design
0.46
Global PFS HR
95% CI 0.37–0.57
<0.0001
Global PFS P-value
Two-sided
FeatureFLAURA
Trial nameFLAURA
NCT identifierNCT02296125
PhasePhase 3
StatusCompleted
ConditionLocally Advanced or Metastatic EGFR Sensitising Mutation Positive Non Small Cell Lung Cancer
AllocationRandomized
Design modelParallel
MaskingTriple
Primary purposeTreatment
Enrollment674
Lead sponsorAstraZeneca
Sponsor typeIndustry

2. Clinical Question

The central efficacy question was whether treatment with single-agent osimertinib could improve progression-free survival compared with standard-of-care EGFR-TKI therapy in patients with locally advanced or metastatic EGFR sensitising mutation-positive non-small cell lung cancer.

Population

Patients with locally advanced or metastatic EGFR sensitising mutation-positive non-small cell lung cancer.

Intervention

Osimertinib 80 mg, with the registry also listing AZD9291 80 mg/40 mg + placebo among the intervention records.

Comparator

Standard-of-care EGFR-TKI therapy, comprising gefitinib 250 mg or erlotinib 150/100 mg, with corresponding placebo records in the masked regimen.

Primary question

Does osimertinib improve progression-free survival relative to standard-of-care EGFR-TKI therapy?

3. Trial Design

01
Randomize 674 enrolled
02
Parallel arms Two treatment groups
03
Triple masking Registry classification
04
Follow-up PFS and response assessments
05
Survival analysis Time-to-event outcomes
ARM A

Osimertinib group

  • Osimertinib 80 mg is the global-cohort treatment label used in the statistical analyses.
  • The intervention records also identify AZD9291 80 mg/40 mg + placebo and placebo EGFR-TKI components.
  • Global-cohort serious adverse events: 74/279 affected participants among 279 at risk.
  • China-cohort serious adverse events: 25/71 affected participants among 71 at risk.
ARM B

Standard-of-care EGFR-TKI group

  • Standard-of-care EGFR-TKI therapy is the comparator label used in the statistical analyses.
  • The registry intervention records include erlotinib 150/100 mg and gefitinib 250 mg.
  • Global-cohort serious adverse events: 76/277 affected participants among 277 at risk.
  • China-cohort serious adverse events: 12/65 affected participants among 65 at risk.
Registry enrollment caveat. The ClinicalTrials.gov record states that 19 Chinese participants were included in both the global and China cohort, which gives a total of 692 participants instead of a total of 673 participants. This registry caveat is reported here without attempting to reconcile it with the separate registry enrollment value of 674.

Masked treatment structure

The intervention records contain active and placebo components for AZD9291, erlotinib, and gefitinib. This is consistent with a masked comparison in which placebo components can preserve treatment masking. The registry classifies the overall masking as triple; the ClinicalTrials.gov record does not identify the three masked parties, so no further attribution is made here.

4. Trial Timeline and Registry Status

2014-12-03

Trial start

The registry lists 2014-12-03 as the trial start date.

2017-06-19

Primary completion

The registry lists 2017-06-19 as the primary completion date.

Completed

Current registry status

the ClinicalTrials.gov record classifies FLAURA as completed.

5. Primary Endpoints

Registered endpointRegistry time frameTypeDefinition
Median Progression Free Survival (PFS) (Months) At baseline and every 6 weeks for the first 18 months and then every 12 weeks relative to randomisation until progression Time-to-event Progression-free survival was defined as the time from randomization until the date of objective disease progression or death (by any cause in the absence of progression) regardless of whether the participant withdrew from randomized therapy or received another anti-cancer therapy prior to progression and was used to assess the efficacy of single agent osimertinib compared with SoC EGFR-TKI therapy as measured by PFS.
Percentage of Participants in Progression Free Survival at 6, 12, and 18 Months At baseline and every 6 weeks for the first 18 months and then every 12 weeks relative to randomisation until progression Time-to-event Progression-free survival was defined as the time from randomization until the date of objective disease progression or death (by any cause in the absence of progression) regardless of whether the participant withdrew from randomized therapy or received another anti-cancer therapy prior to progression and was used to assess the efficacy of single agent osimertinib compared with SoC EGFR-TKI therapy as measured by PFS.

The two registered primary endpoints describe the same underlying time-to-event construct from different reporting perspectives: the first summarizes PFS through the median, while the second concerns the percentage of participants remaining progression-free at 6, 12, and 18 months.

6. Statistical Methodology

Full analysis set and randomized comparison

The primary PFS analysis was reported for the full analysis set (FAS) and the China-only FAS. The registry definition states that the FAS included all randomized participants prior to the end of global recruitment, while the China-only FAS included all China participants randomized in mainland China. The statistical-analysis records also identify intention-to-treat analysis as a concept used in the analyses.

Analyzing randomized participants according to randomized assignment is important because randomization creates the basis for a treatment comparison that is less vulnerable to baseline confounding. The analysis population therefore differs conceptually from an analysis that selects patients after treatment exposure or according to whether they completed therapy.

Log-rank test

The reported formal method for the primary PFS comparisons was the log-rank test. This test compares the observed pattern of time-to-event outcomes between groups while accounting for the fact that not every participant necessarily experiences progression or death during observed follow-up.

Time-to-event comparison
Observed event experience  ↔  expected event experience under the comparison hypothesis

The log-rank framework uses information from the times at which events occur and the number of participants still at risk at those times. Censoring therefore enters the comparison through the risk sets rather than by simply treating censored observations as if they had experienced the event.

Hazard ratio

The primary PFS effect measure was the hazard ratio (HR). A hazard ratio below 1 indicates a lower estimated instantaneous rate of progression or death in the first group relative to the second group, under the time-to-event model represented by the estimate.

Conceptual interpretation
HR < 1  →  lower estimated instantaneous event rate in the first group

A hazard ratio is not a percentage of patients who benefit and is not the same quantity as a risk ratio or an absolute difference in survival probability.

Logistic regression

Objective response rate and disease control rate were analyzed using logistic regression, with odds ratio as the reported effect measure. This is appropriate for binary outcomes because the model describes the relationship between treatment and the odds of an event or response category rather than directly modeling a time-to-event process.

Linear regression

Duration of response and depth of response were analyzed using linear regression in the posted statistical analyses, with mean difference as the reported effect measure. The registry data identify these analyses as final-value comparisons. The use of a mean difference means the estimate is expressed on the outcome's original scale rather than as a ratio.

7. Primary Results: Progression-Free Survival

Global Cohort

Hazard ratio for progression or death

0.46

95% CI: 0.37–0.57   ·   P < 0.0001

Osimertinib 80 mg versus standard-of-care EGFR-TKI therapy

The formal analysis compared the osimertinib 80 mg global cohort with the standard-of-care EGFR-TKI global cohort. The analysis population was the full analysis set and China-only FAS framework described in the registry; the global-cohort comparison reported above used the FAS. The reported method was the log-rank test, the effect measure was a hazard ratio, the confidence interval was two-sided at 95%, and the hypothesis type was superiority.

Clinical Biostats interpretation

What the estimate means: An HR of 0.46 means the estimated instantaneous rate of progression or death in the osimertinib group was 0.46 times that in the standard-of-care EGFR-TKI group under the reported time-to-event analysis. Expressed as a simple relative-hazard interpretation, this corresponds to an estimated 54% lower instantaneous event rate.

What it does not mean: It does not mean that 54% of participants avoided progression, that every participant experienced the same reduction, or that median PFS was reduced by 54%. A hazard ratio is a relative time-to-event measure, not an absolute probability.

Precision: The 95% CI of 0.37–0.57 describes statistical uncertainty around the estimated hazard ratio. It does not describe the range of outcomes that individual patients might experience.

The p-value: The reported P < 0.0001 addresses evidence against the comparison hypothesis under the specified statistical framework. It does not measure the magnitude of the treatment effect. Effect magnitude is described by the hazard ratio, while its precision is described by the confidence interval.

Cautions: Time-to-event analyses depend on censoring rules and on how the hazard-ratio summary represents the event process over follow-up. A hazard ratio should therefore be interpreted together with the underlying time-to-event estimand and the absolute PFS probabilities when those are available.

China Cohort

Hazard ratio for progression or death

0.56

95% CI: 0.37–0.85   ·   P = 0.0065

Osimertinib 80 mg versus standard-of-care EGFR-TKI therapy

The China-cohort analysis used the same reported log-rank framework and the same two-sided 95% confidence level. The registry specifically notes that the China cohort was not powered for superiority, and its hypothesis type is listed as other / not stated.

Clinical Biostats interpretation

What the estimate means: An HR of 0.56 corresponds to an estimated instantaneous progression-or-death rate of 0.56 times that of the standard-of-care EGFR-TKI group in the China-cohort analysis. As a direct relative-hazard translation, that is approximately a 44% lower estimated instantaneous event rate.

What it does not mean: It is not a 44% absolute improvement in the percentage of participants who remain progression-free, and it does not imply that each participant experienced a 44% reduction in individual risk.

Precision: The 95% CI of 0.37–0.85 is wider than the global-cohort interval of 0.37–0.57. The wider interval reflects greater uncertainty in the China-cohort estimate.

The p-value: The reported P = 0.0065 is a hypothesis-testing quantity, not a measure of effect size. The hazard ratio and its confidence interval provide the information about relative magnitude and precision.

Important design qualification: The registry explicitly states that the China cohort was not powered for superiority. Accordingly, the China analysis should not be read as if it were a separately powered confirmatory superiority trial.

8. Primary Endpoint: Percentage of Participants in Progression-Free Survival at 6, 12, and 18 Months

ClinicalTrials.gov lists Percentage of Participants in Progression Free Survival at 6, 12, and 18 Months as a primary endpoint and indicates that results were posted. However, the ClinicalTrials.gov record does not provide a formal estimate, confidence interval, or p-value for this specific primary endpoint. The page therefore does not supply numerical values that are not present in the trial data.

How this endpoint is normally interpreted: a time-specific PFS percentage is typically obtained from the estimated survival function at the specified time point. For a time-to-event endpoint, this is distinct from the median PFS because it asks for the estimated probability of remaining event-free at a fixed time rather than the time at which the estimated survival probability reaches 0.50.

The registry's posted formal primary analyses instead provide hazard-ratio comparisons for the median PFS outcome measure in both the global and China cohorts. Those analyses are reported above. The ClinicalTrials.gov record does not provide the 6-, 12-, or 18-month PFS percentages themselves, so they are not reconstructed or inferred here.

9. Secondary Endpoint Results

Objective Response Rate

CohortMethodEffect measureEstimate95% CIP-value
Global Logistic regression Odds ratio 1.51 1.03–2.22 0.036
China Logistic regression Odds ratio 1.31 0.61–2.84 0.485

The global-cohort ORR analysis reported an odds ratio of 1.51 with a 95% CI of 1.03–2.22 and P = 0.036. The China-cohort analysis reported an odds ratio of 1.31 with a 95% CI of 0.61–2.84 and P = 0.485. The China cohort was not powered for superiority.

Interpreting an odds ratio

An odds ratio of 1.51 means that the estimated odds of objective response were 1.51 times the corresponding odds in the comparator group in the global-cohort analysis. This is not equivalent to saying that the response probability was 51% higher. Converting an odds ratio into a probability difference requires the underlying response probability or odds in the reference group.

The global 95% CI of 1.03–2.22 indicates uncertainty around the odds-ratio estimate. The China interval of 0.61–2.84 is substantially wider and includes 1, consistent with considerably greater uncertainty in that estimate.

Duration of Response

CohortMethodEffect measureEstimate95% CIP-value
Global Linear regression Mean difference (final values) 2.27 months 1.68–3.08 months <0.0001
China Linear regression Mean difference (final values) 2.48 months 1.21–5.09 months 0.0133

The reported mean difference was 2.27 months in the global cohort and 2.48 months in the China cohort, with both estimates positive for the first-listed group. The corresponding confidence intervals were 1.68–3.08 months and 1.21–5.09 months.

Why the mean difference is different from a hazard ratio

A mean difference is expressed directly in months, so an estimate of 2.27 months represents a difference in the reported final-value means on the duration-of-response scale. It is not a relative effect and should not be interpreted as a percentage increase.

The registry identifies linear regression as the formal method. Because duration of response is inherently related to time-to-event concepts, readers should not automatically assume that a linear-regression estimate has the same interpretation as a survival-model hazard ratio. The statistical estimand matters.

Disease Control Rate

CohortMethodEffect measureEstimate95% CIP-value
Global Logistic regression Odds ratio 2.78 1.25–6.78 0.0110
China Logistic regression Odds ratio 1.67 0.27–12.98 0.5772

The registry notes that an odds ratio greater than 1 favors osimertinib for these disease-control analyses. The global estimate was 2.78, while the China estimate was 1.67. The China confidence interval is particularly wide, extending from 0.27 to 12.98, illustrating substantial uncertainty.

Depth of Response

CohortMethodEffect measureEstimate95% CIP-value
Global Linear regression Mean difference (final values) -6.80 percentage points -11.205 to -2.403 0.0025
China Linear regression Mean difference (final values) -6.59 percentage points -15.246 to 2.072 0.1348

For depth of response, the global-cohort mean difference was -6.80 percentage points, with a 95% CI of -11.205 to -2.403. The China-cohort estimate was -6.59 percentage points, with a 95% CI of -15.246 to 2.072. The ClinicalTrials.gov record does not state the direction convention beyond the numerical mean difference, so the estimates are reported without converting the negative sign into a qualitative label.

Overall Survival

CohortMethodEffect measureEstimate95% CIP-value
Global Log-rank test Hazard ratio 0.799 0.6409–0.9963 0.0462
China Log-rank test Hazard ratio 0.848 0.5568–1.2910 0.4416

The global-cohort overall-survival analysis reported an HR of 0.799, with a 95% CI of 0.6409–0.9963 and P = 0.0462. The China-cohort estimate was 0.848, with a 95% CI of 0.5568–1.2910 and P = 0.4416. The registry again notes that the China cohort was not powered for superiority.

Clinical Biostats interpretation

An HR of 0.799 corresponds to an estimated instantaneous rate of death of 0.799 times that in the comparator group in the global analysis, or approximately a 20.1% lower estimated instantaneous event rate as a simple relative-hazard translation.

The 95% CI of 0.6409–0.9963 is close to 1 at its upper boundary, so the estimate should be read together with its uncertainty rather than as a precise fixed treatment effect. The P-value of 0.0462 describes the statistical evidence under the reported test; it does not quantify the size or clinical importance of the effect.

For the China cohort, the HR of 0.848 has a substantially wider 95% CI of 0.5568–1.2910. This interval includes 1, and the reported P-value is 0.4416. The registry's explicit statement that the China cohort was not powered for superiority is important context when interpreting this estimate.

10. Serious Adverse Events by Arm

The ClinicalTrials.gov record reports serious adverse events by affected participants and participants at risk for both the global and China cohorts.

CohortOsimertinib 80 mgSoC EGFR-TKI
Global cohort 74/279 76/277
China cohort 25/71 12/65

These are reported as affected participants over participants at risk. The ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or p-value for serious adverse events, so no comparative hypothesis test is inferred.

Global cohort

Serious adverse events affected 74 of 279 participants in the osimertinib group and 76 of 277 in the standard-of-care EGFR-TKI group.

China cohort

Serious adverse events affected 25 of 71 participants in the osimertinib group and 12 of 65 in the standard-of-care EGFR-TKI group.

Safety interpretation should remain separate from efficacy interpretation. The existence of a treatment effect estimate for PFS does not establish a corresponding net-benefit conclusion, because efficacy and safety address different outcomes and use different analysis populations and estimands.

11. Statistical Methods Explained

Why was a log-rank test used for PFS?

PFS is a time-to-event endpoint: participants may experience objective disease progression or death at different times, while others may be censored. The log-rank test is designed to compare the event-time distributions between randomized groups while incorporating the timing of events and the number of participants at risk.

What does an HR of 0.46 mean?

The estimate indicates that the estimated instantaneous progression-or-death rate in the osimertinib group was 0.46 times the corresponding rate in the standard-of-care EGFR-TKI group under the reported analysis. The arithmetic translation to a relative hazard reduction is 1 − 0.46 = 0.54, or 54%. This does not mean that 54% of patients avoided progression or that individual risk fell by exactly 54%.

Why does the confidence interval matter?

The point estimate is only one summary of the data. The 95% confidence interval communicates how precisely that effect was estimated under the statistical framework. For global PFS, the interval is 0.37–0.57; for China PFS, it is 0.37–0.85. The wider China interval shows greater uncertainty around that subgroup estimate.

What does an odds ratio of 1.51 mean for objective response?

An OR of 1.51 means the estimated odds of objective response were 1.51 times those in the comparator group in the global cohort. Odds and probabilities are different quantities. Without the comparator response probability, the OR cannot be converted into an absolute response-rate difference.

Why is an overall-survival HR different from a PFS HR?

PFS and OS measure different events. PFS counts objective disease progression or death, whereas the registry-reported OS endpoint measures death from any cause. A treatment can therefore have different hazard ratios for the two endpoints because the event definitions, follow-up processes, censoring, and subsequent treatment pathways differ.

Why should the China analysis be interpreted differently from the global analysis?

The registry explicitly states that the China cohort was not powered for superiority. Its analyses are therefore informative about the observed treatment comparison in that cohort but should not automatically be treated as if the cohort had the same confirmatory statistical role as the global population.

Why does the analysis population matter?

The primary analyses identify the full analysis set and intention-to-treat analysis as relevant concepts. Maintaining participants according to randomized assignment protects the treatment comparison established by randomization. Removing participants after randomization because of treatment discontinuation, for example, can change the population being compared and potentially alter the causal interpretation.

12. Confidence Intervals and P-values

FLAURA provides several useful examples of why effect estimates, confidence intervals, and p-values should be read together rather than substituted for one another.

EndpointEstimate95% CIP-valueWhat each quantity contributes
Global PFS HR 0.46 0.37–0.57 <0.0001 Relative time-to-event effect, precision, and hypothesis-test evidence
China PFS HR 0.56 0.37–0.85 0.0065 Relative time-to-event effect, wider uncertainty, and hypothesis-test evidence
Global ORR OR 1.51 1.03–2.22 0.036 Relative odds of response, precision, and hypothesis-test evidence
Global OS HR 0.799 0.6409–0.9963 0.0462 Relative time-to-event effect, precision, and hypothesis-test evidence
The three quantities answer different questions

Estimate: How large is the observed or estimated effect?

Confidence interval: How much statistical uncertainty surrounds that estimate?

P-value: How compatible are the data with the comparison hypothesis under the specified testing framework?

None of these quantities, by itself, describes the clinical importance of an effect. Clinical interpretation requires the endpoint definition, magnitude, precision, population, follow-up, and relevant design context.

13. Intention-to-Treat Analysis and Censoring

The statistical-analysis records explicitly identify intention-to-treat analysis as a concept for the PFS, response, duration-of-response, disease-control, depth-of-response, and OS analyses. This is especially important for randomized time-to-event comparisons because the treatment contrast is defined by assignment at randomization rather than by treatment received after randomization.

For PFS, the registry definition states that an event is objective disease progression or death in the absence of progression. It also states that the endpoint remains applicable regardless of whether a participant withdrew from randomized therapy or received another anti-cancer therapy prior to progression. This definition prevents treatment discontinuation itself from becoming the PFS event.

Censoring is not the same as no event. In time-to-event analysis, a participant who has not experienced the event by the end of available observation can contribute information until the censoring time. The statistical analysis then uses the observed risk sets and event times rather than simply classifying every participant as either an event or a permanent non-event.

14. Global Versus China Cohort Interpretation

The analyses posted on ClinicalTrials.gov repeatedly report both a global cohort and a China cohort. This creates a useful statistical distinction between the overall randomized comparison and a geographically defined cohort analysis.

EndpointGlobal estimateChina estimateRegistry qualification
PFS HR 0.46 (95% CI 0.37–0.57), P < 0.0001 HR 0.56 (95% CI 0.37–0.85), P = 0.0065 China cohort not powered for superiority
ORR OR 1.51 (95% CI 1.03–2.22), P = 0.036 OR 1.31 (95% CI 0.61–2.84), P = 0.485 China cohort not powered for superiority
Duration of response Mean difference 2.27 (95% CI 1.68–3.08), P < 0.0001 Mean difference 2.48 (95% CI 1.21–5.09), P = 0.0133 China cohort not powered for superiority
DCR OR 2.78 (95% CI 1.25–6.78), P = 0.0110 OR 1.67 (95% CI 0.27–12.98), P = 0.5772 China cohort not powered for superiority
Depth of response Mean difference -6.80 (95% CI -11.205 to -2.403), P = 0.0025 Mean difference -6.59 (95% CI -15.246 to 2.072), P = 0.1348 China cohort not powered for superiority
OS HR 0.799 (95% CI 0.6409–0.9963), P = 0.0462 HR 0.848 (95% CI 0.5568–1.2910), P = 0.4416 China cohort not powered for superiority

The differences in estimates and confidence intervals illustrate why a subgroup result should not be reduced to whether its p-value crosses a threshold. The China cohort generally has wider confidence intervals than the global cohort, and the registry explicitly cautions that it was not powered for superiority. The magnitude and precision of each estimate should therefore be examined directly.

15. Multiplicity and Multiple Endpoints

The registry identifies two primary endpoints and multiple secondary endpoints. The ClinicalTrials.gov record does not provide a detailed multiplicity-adjustment strategy, alpha allocation, or hierarchical testing procedure. It would therefore be inappropriate to infer such a procedure from the reported p-values.

Endpoint roleRegistered / reported outcomeStatistical interpretation
Primary Median Progression Free Survival (PFS) (Months) Formal log-rank analyses with hazard-ratio estimates and 95% CIs were posted.
Primary Percentage of Participants in Progression Free Survival at 6, 12, and 18 Months Results are posted, but the ClinicalTrials.gov record does not provide a formal estimate, CI, or p-value.
Secondary Objective Response Rate Logistic-regression odds ratios were posted for global and China cohorts.
Secondary Duration of Response Linear-regression mean differences were posted for global and China cohorts.
Secondary Disease Control Rate Logistic-regression odds ratios were posted for global and China cohorts.
Secondary Depth of Response Linear-regression mean differences were posted for global and China cohorts.
Secondary Overall Survival - Number of Participants With an Event Log-rank hazard-ratio analyses were posted for global and China cohorts.

Because the ClinicalTrials.gov record does not describe an alpha hierarchy or multiplicity adjustment, individual secondary p-values should be understood in the context of a trial containing multiple endpoints and multiple cohort analyses rather than as automatically independent confirmatory tests.

16. Non-Inferiority, Equivalence, and Bayesian Methods

The statistical analyses posted on ClinicalTrials.gov identify superiority for the global primary PFS comparison and several global secondary analyses. The China PFS, response, duration-of-response, disease-control, depth-of-response, and OS analyses are marked as other / not stated, with the additional note that the China cohort was not powered for superiority.

No non-inferiority margin is reported in the ClinicalTrials.gov record. No equivalence margin is reported. No Bayesian method is identified among the normalized methods. Accordingly, the statistical interpretation on this page does not introduce a non-inferiority or Bayesian framework that is not present in the ClinicalTrials.gov record.

17. What the Hazard Ratio Does — and Does Not — Mean

Global PFS

The reported HR 0.46 is a relative time-to-event measure. Under the reported analysis, the estimated instantaneous rate of progression or death was 0.46 times that of the standard-of-care EGFR-TKI group.

The estimate does not say that the median PFS was 54% longer, that 54% of participants benefited, or that 54% more participants were progression-free at every time point.

Global OS

The reported HR 0.799 means the estimated instantaneous rate of death was 0.799 times that in the comparator group under the reported analysis. A direct arithmetic translation is an estimated 20.1% lower instantaneous event rate.

Again, this is not an absolute survival difference and does not describe the probability that any particular participant will survive.

Confidence intervals

The PFS confidence interval of 0.37–0.57 and OS confidence interval of 0.6409–0.9963 describe uncertainty around their respective estimates. They should not be interpreted as prediction intervals for individual patients.

18. Limitations

19. Why This Trial Matters Statistically

FLAURA is a useful teaching case because the ClinicalTrials.gov record bring together randomized treatment allocation, masking, time-to-event endpoints, hazard ratios, log-rank testing, binary response outcomes, odds ratios, continuous-scale mean differences, intention-to-treat concepts, and separate global and China-cohort analyses.

Statistical conceptHow it appears in FLAURA
RandomizationThe trial is classified as randomized with two parallel arms.
BlindingThe registry classifies masking as triple.
Intention-to-treat analysisIdentified as an analysis concept across the posted efficacy analyses.
Time-to-event endpointsPFS and OS are analyzed using time-to-event methods.
Log-rank testUsed for the reported PFS and OS comparisons.
Hazard ratioUsed for PFS and OS as the relative effect measure.
Confidence intervals95% two-sided intervals accompany the posted effect estimates.
Logistic regressionUsed for ORR and DCR.
Odds ratioUsed as the effect measure for binary response outcomes.
Linear regressionUsed for duration of response and depth of response.
Mean differenceUsed for the reported linear-regression analyses.
Subpopulation analysisGlobal and China cohorts are separately analyzed.
Statistical powerThe registry specifically notes that the China cohort was not powered for superiority.
Safety analysisSerious adverse events are reported separately by global and China cohort.

20. A Statistical Reading of the Complete Evidence Set

The global PFS analysis provides the clearest quantitative example of the trial's primary time-to-event comparison: an HR of 0.46, 95% CI 0.37–0.57, and P < 0.0001. The estimate is substantially below 1, while the confidence interval provides a relatively concentrated range around the estimate.

The China PFS analysis points in the same numerical direction, with an HR of 0.56, but the confidence interval is wider at 0.37–0.85. This is a useful illustration of why an estimate should not be read without its precision and population context. The registry's statement that the China cohort was not powered for superiority adds another layer to the interpretation.

The secondary endpoints show that the statistical signal is not represented by a single universal effect measure. ORR and DCR use odds ratios, duration of response and depth of response use mean differences, and OS uses a hazard ratio. Each effect measure answers a different question and should be interpreted on its own scale.

The overall-survival result also illustrates why PFS and OS should not be conflated. The global OS HR of 0.799 is numerically different from the global PFS HR of 0.46. That difference is not itself evidence of inconsistency: the endpoints have different event definitions and different clinical meanings.

Finally, the China analyses demonstrate why a nonsignificant p-value is not equivalent to proof of no effect. For example, the China ORR estimate is 1.31 with a 95% CI of 0.61–2.84. The interval spans a broad range of plausible odds ratios, so the estimate is substantially uncertain. The same principle applies to the China DCR, depth-of-response, and OS results.

21. Related Tutorials

Learn more about the methods used in this trial:

22. Related Statistical Calculators

23. Sources

Continue through Clinical Biostats

Explore statistical tutorials, calculators, and additional clinical trial analyses.

24. Record Summary

FLAURA provides a compact teaching example of randomized time-to-event analysis with multiple complementary endpoint types. The primary PFS analysis used the log-rank test and reported hazard ratios of 0.46 for the global cohort and 0.56 for the China cohort. The corresponding 95% confidence intervals were 0.37–0.57 and 0.37–0.85, respectively.

The secondary analyses demonstrate why statistical interpretation should remain tied to the outcome scale. ORR and DCR were analyzed with logistic regression and odds ratios; duration of response and depth of response were analyzed with linear regression and mean differences; and OS was analyzed with a log-rank framework and hazard ratios. The global OS estimate was 0.799 with a 95% CI of 0.6409–0.9963, while the China estimate was 0.848 with a 95% CI of 0.5568–1.2910.

The most important statistical lesson is that no single number summarizes the entire trial. Hazard ratios describe relative event rates, odds ratios describe relative odds for binary outcomes, mean differences describe absolute differences on the outcome scale, confidence intervals describe statistical uncertainty, and p-values address hypothesis-testing evidence. These quantities become most informative when interpreted alongside the endpoint definition, analysis population, cohort, and trial design.

Clinical Biostats methodology: This page separates registry-reported numerical results from statistical interpretation. Where the ClinicalTrials.gov record does not provide a numerical result, the page does not reconstruct or infer one.