← Clinical Trials
Diabetes Mellitus Phase 3 Factorial Design NCT00069784

ORIGIN: Complete Statistical Analysis of Initial Insulin Glargine in Diabetes Mellitus

An independent statistical analysis of the randomized phase 3 ORIGIN trial evaluating insulin glargine versus standard care, with cardiovascular, microvascular, mortality, and diabetes-development endpoints analyzed using survival and stratified categorical methods.

ORIGIN Trial  ·  2003-08 to 2011-12  ·  12,537 randomized participants
Scope of this record

This page separates reported trial results from statistical interpretation. The numerical results presented here are restricted to the statistical analyses and trial information from the ClinicalTrials.gov record.

Registry note: This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

ORIGIN was a randomized phase 3 factorial trial enrolling 12,537 participants with Diabetes Mellitus, Non-Insulin-Dependent. The registry reports four arms and evaluates insulin glargine versus standard care using cardiovascular and other clinical endpoints, while the factorial structure also incorporated omega-3 polyunsaturated fatty acids and placebo.

12,537
Randomized
Phase 3 enrollment
4
Arms
Factorial design
1.022
Primary HR
CV death / MI / stroke
0.72
Diabetes OR
95% CI 0.58–0.91
FeatureORIGIN
Trial nameThe ORIGIN Trial (Outcome Reduction With Initial Glargine Intervention)
PhasePhase 3
ConditionDiabetes Mellitus, Non-Insulin-Dependent
DesignRandomized, factorial
MaskingNone
Primary purposeTreatment
Enrollment12,537
Primary endpoints2 binary endpoints registered; both analyzed as time-to-event outcomes
Results postedYes
Statistical analyses posted5
Lead sponsorSanofi
ClinicalTrials.govNCT00069784

2. Clinical Question

The insulin-glargine component of ORIGIN asks whether initial insulin glargine, compared with standard care, changes the time to major cardiovascular events and other clinically relevant outcomes in the enrolled population.

Population

Participants with Diabetes Mellitus, Non-Insulin-Dependent enrolled in the phase 3 ORIGIN trial.

Intervention

Insulin glargine (HOE901), as identified in the registry.

Comparator

Standard Care.

Primary question

Does insulin glargine versus standard care change the occurrence of the two registered cardiovascular composite endpoints?

3. Trial Design

01
Randomize12,537 participants
02
Factorial allocation4 trial arms
03
Insulin comparisonGlargine vs standard care
04
Follow-upUntil study cut-off
05
AnalysisTime-to-event and binary outcomes
Allocation
Randomized. Randomization provides the framework for comparing treatment assignments while reducing systematic differences between groups at baseline.
Design model
Factorial. The registry identifies four arms and interventions involving insulin glargine, omega-3 PUFA, placebo, and a reusable pen device.
Masking
None. The registry classifies the trial as unmasked.
Primary purpose
Treatment. The study evaluates clinical outcomes associated with the randomized intervention strategy.

Understanding the factorial structure

A factorial design evaluates more than one intervention dimension within the same randomized trial. ORIGIN included four arms and lists insulin glargine, omega-3 PUFA, placebo, and a reusable pen device among its interventions. The statistical comparison reported here is specifically insulin glargine versus standard care; the ClinicalTrials.gov record does not provide a separate formal analysis of an interaction between the factorial components.

Why factorial design matters: a factorial trial can use one randomized population to investigate multiple intervention components efficiently. The interpretation of a particular treatment comparison nevertheless depends on the prespecified factorial model and on whether interaction between factors is formally assessed.

4. Endpoints

EndpointRegistry definition / time frameAnalysis reported
Composite of the First Occurrence of Cardiovascular (CV) Death, Nonfatal Myocardial Infarction (MI) or Nonfatal Stroke Number of participants with a first occurrence of one of the above events; from randomization until study cut-off date (median duration of follow-up: 6.2 years). Positively-adjudicated first events were reviewed by an Event Adjudication Committee kept blinded to group assignment. Log-rank test; Cox proportional-hazards model; hazard ratio
Composite of the First Occurrence of Cardiovascular (CV) Death, Nonfatal Myocardial Infarction (MI), Nonfatal Stroke, Revascularization Procedure or Hospitalization for Heart Failure (HF) Number of participants with a first occurrence of one of the above events; from randomization until study cut-off date (median duration of follow-up: 6.2 years). Log-rank test; Cox proportional-hazards model; hazard ratio
Total Mortality (All Causes) From randomization until study cut-off date (median duration of follow-up: 6.2 years). Cox proportional-hazards model; hazard ratio
Composite Diabetic Microvascular Outcome (Kidney or Eye Disease) From randomization until study cut-off date (median duration of follow-up: 6.2 years). Cox proportional-hazards model; hazard ratio
Incidence of Development of Type 2 Diabetes Mellitus in Participants With IGT and/or IFG From randomization until the last follow-up visit or last OGTT (median duration of follow-up: 6.2 years). Cochran-Mantel-Haenszel test; odds ratio

The two registered primary endpoints are composite cardiovascular outcomes. Both are described in the registry as binary primary endpoints, but the posted formal analyses use the timing of the first event and therefore apply time-to-event methods.

5. Statistical Methodology

Intention-to-treat analysis

The primary endpoint analyses were based on the intent-to-treat (ITT) population, defined in the registry as all randomized participants. This means the efficacy comparison is anchored to randomized assignment rather than to whether participants remained compliant with the assigned treatment.

Log-rank test

The primary cardiovascular endpoints were compared using a log-rank test. For time-to-event data, the log-rank test evaluates whether the observed event-time experience differs between randomized groups while incorporating information from participants who are censored before the study endpoint occurs.

Cox proportional-hazards model

The treatment effect for the primary cardiovascular endpoints was reported as a hazard ratio estimated using Cox regression. The registry analysis text identifies stratification by double-blind treatment, baseline diabetes diagnosis, and previous cardiovascular event.

Conceptual hazard-ratio model
HR = estimated hazard in the insulin-glargine group ÷ estimated hazard in the standard-care group

A hazard ratio below 1 indicates a lower estimated instantaneous event rate in the insulin-glargine group under the fitted model; a hazard ratio above 1 indicates a higher estimated instantaneous event rate.

Covariate adjustment

For total mortality and the composite diabetic microvascular outcome, the Cox model included treatment as a factor together with double-blind treatment (omega-3 PUFA, placebo), baseline diabetes diagnosis, and previous cardiovascular event as covariates. This is different from simply comparing crude event proportions: the model estimates the treatment hazard ratio while accounting for the specified covariates.

Cochran-Mantel-Haenszel analysis

The development of type 2 diabetes endpoint was analyzed with a Cochran-Mantel-Haenszel (CMH) test. The odds ratio was stratified by double-blind treatment (omega-3 PUFA or placebo) and previous cardiovascular event (yes or no).

Stratified analysis

Stratification appears repeatedly in the registry's statistical analysis text. For the primary cardiovascular endpoint, the log-rank analysis was stratified by double-blind treatment, baseline diabetes diagnosis, and previous cardiovascular event. For the diabetes-development endpoint, CMH stratification used double-blind treatment and previous cardiovascular event.

Why stratification is useful: stratification allows the comparison to account for prespecified factors that may influence the event process. It is particularly useful when the randomized design and analysis plan identify clinically relevant factors that should be incorporated into the comparison without treating them as treatment effects themselves.

6. Primary Results

Primary Endpoint 1: Cardiovascular Death, Nonfatal MI, or Nonfatal Stroke

Hazard ratio for first occurrence

1.022

95% CI: 0.937–1.114   ·   P = 0.6273

Insulin glargine vs standard care; median duration of follow-up: 6.2 years.

Primary endpointInsulin Glargine vs Standard Care
Effect measureHazard ratio
Hazard ratio1.022
95% CI0.937–1.114
P-value0.6273
MethodLog-rank test; Cox proportional-hazards model
Analysis populationIntent-to-treat; all randomized participants
Time frameFrom randomization until study cut-off date; median duration of follow-up: 6.2 years
Clinical Biostats interpretation

The hazard ratio of 1.022 is slightly above 1. Under the reported Cox model, this corresponds to an estimated hazard that is 1.022 times that of standard care for the first occurrence of the composite endpoint. Put differently, the point estimate is very close to 1.

The estimate does not mean that insulin glargine caused a 2.2% increase in the probability of experiencing the endpoint. A hazard ratio is a relative time-to-event measure, not an absolute risk difference or a percentage of patients affected.

The 95% confidence interval, 0.937–1.114, spans 1. This indicates uncertainty that includes both a lower and higher hazard relative to standard care. The interval describes uncertainty around the estimated treatment effect; it is not a range containing the effects experienced by individual participants.

The P-value of 0.6273 is a measure of compatibility with the statistical testing framework under the null hypothesis. It is not a measure of the magnitude or clinical importance of the hazard ratio. The estimate, confidence interval, endpoint definition, censoring, and model assumptions all contribute to interpretation.

Because this is a Cox analysis, interpretation also depends on the proportional-hazards assumption. A single hazard ratio summarizes the relative event-rate experience through the fitted model; it should not automatically be interpreted as a constant relative risk at every time point.

Primary Endpoint 2: Cardiovascular Death, Nonfatal MI, Nonfatal Stroke, Revascularization, or Hospitalization for HF

Hazard ratio for first occurrence

1.038

95% CI: 0.972–1.109   ·   P = 0.2692

Insulin glargine vs standard care; median duration of follow-up: 6.2 years.

Primary endpointInsulin Glargine vs Standard Care
Effect measureHazard ratio
Hazard ratio1.038
95% CI0.972–1.109
P-value0.2692
MethodLog-rank test; Cox proportional-hazards model
Analysis populationIntent-to-treat; all randomized participants
Time frameFrom randomization until study cut-off date; median duration of follow-up: 6.2 years
Clinical Biostats interpretation

The hazard ratio of 1.038 is close to 1. Under the reported Cox model, the estimated instantaneous event rate for the broader cardiovascular composite was 1.038 times that of standard care.

It does not mean that the probability of the composite endpoint was exactly 3.8% higher. Hazard ratios describe relative event rates in a time-to-event model and cannot be converted directly into an absolute risk difference without additional information.

The 95% confidence interval of 0.972–1.109 includes 1 and is relatively close to the point estimate on both sides. It therefore includes values corresponding to a modestly lower as well as a modestly higher estimated hazard relative to standard care.

The P-value of 0.2692 does not quantify the size of the observed effect. It addresses the statistical evidence against the relevant null hypothesis under the prespecified testing framework. It should be read together with the hazard ratio and confidence interval rather than used as a standalone measure of treatment effect.

The registry specifies a stratified log-rank comparison and a Cox regression model, with treatment as a factor and stratification by double-blind treatment, baseline diabetes diagnosis, and previous cardiovascular event. These design features matter because the reported estimate is model-based rather than a simple ratio of two cumulative event percentages.

7. Secondary Endpoint Results

Total Mortality (All Causes)

Hazard ratio for all-cause mortality

0.983

95% CI: 0.899–1.076   ·   P-value not reported in the ClinicalTrials.gov record

The analysis was based on the ITT population, meaning all randomized participants were included regardless of compliance with the protocol. The Cox model included treatment as a factor, with double-blind treatment, baseline diabetes diagnosis, and previous cardiovascular event as covariates.

Clinical Biostats interpretation

A hazard ratio of 0.983 is close to 1, so the point estimate represents only a small relative difference in the modeled hazard of death between randomized groups. The 95% CI of 0.899–1.076 includes 1, indicating that the uncertainty interval encompasses both modestly lower and modestly higher hazards relative to standard care.

The hazard ratio is not an absolute mortality difference and does not say that 1.7% fewer participants died. The confidence interval concerns the estimated model parameter, not individual patient outcomes. The registry analysis does not report a P-value for this secondary endpoint, so no additional hypothesis-test interpretation is appropriate here.

Composite Diabetic Microvascular Outcome (Kidney or Eye Disease)

Hazard ratio for first microvascular outcome

0.970

95% CI: 0.900–1.047   ·   P-value not reported in the ClinicalTrials.gov record

This secondary time-to-event endpoint was analyzed in the ITT population using a Cox proportional-hazards model. The model included treatment as a factor and adjusted for double-blind treatment, baseline diabetes diagnosis, and previous cardiovascular event.

Clinical Biostats interpretation

The point estimate of 0.970 corresponds to an estimated hazard approximately 0.970 times that of standard care under the reported Cox model. The 95% CI of 0.900–1.047 includes 1, so the interval is compatible with a range of effects around the null value.

Again, this is a relative time-to-event measure, not a 3.0% absolute reduction in kidney or eye disease. The ClinicalTrials.gov record does not report a P-value for this secondary analysis, so the confidence interval and point estimate provide the available statistical description.

Development of Type 2 Diabetes Mellitus in Participants With IGT and/or IFG

Odds ratio for development of type 2 diabetes

0.72

95% CI: 0.58–0.91   ·   P-value not reported in the ClinicalTrials.gov record

FeatureReported analysis
PopulationSubgroup of the ITT population without diabetes at randomization
EndpointIncidence of Development of Type 2 Diabetes Mellitus in Participants With IGT and/or IFG
Effect measureOdds ratio
Estimate0.72
95% CI0.58–0.91
MethodCochran-Mantel-Haenszel test
StratificationDouble-blind treatment and previous cardiovascular event
Time frameFrom randomization until the last follow-up visit or last OGTT; median duration of follow-up: 6.2 years
Clinical Biostats interpretation

An odds ratio of 0.72 means that the estimated odds of developing type 2 diabetes in the insulin-glargine group were 0.72 times the corresponding odds in the standard-care group, using the reported CMH analysis.

This does not mean that the probability of diabetes was exactly 28% lower. Odds and probabilities are related but are not interchangeable, particularly when the outcome is not rare. An odds ratio should therefore not be described as a risk ratio without the additional information needed to make that conversion.

The 95% CI of 0.58–0.91 lies below 1. This gives the reported estimate a confidence interval that does not include the null odds ratio of 1. The interval describes uncertainty around the estimated odds ratio, not the range of individual treatment responses.

The registry-reported analysis does not report a P-value. The appropriate interpretation is therefore based on the reported odds ratio, its confidence interval, the prespecified subgroup population, and the CMH stratification rather than assigning an unreported significance level.

8. Results Summary

EndpointEffect measureEstimate95% CIP-valueMethod
CV death, nonfatal MI, or nonfatal stroke Hazard ratio 1.022 0.937–1.114 0.6273 Log-rank; Cox proportional-hazards
CV death, nonfatal MI, nonfatal stroke, revascularization, or hospitalization for HF Hazard ratio 1.038 0.972–1.109 0.2692 Log-rank; Cox proportional-hazards
Total Mortality (All Causes) Hazard ratio 0.983 0.899–1.076 Not reported Cox proportional-hazards
Composite Diabetic Microvascular Outcome Hazard ratio 0.970 0.900–1.047 Not reported Cox proportional-hazards
Development of Type 2 Diabetes Mellitus in participants with IGT and/or IFG Odds ratio 0.72 0.58–0.91 Not reported Cochran-Mantel-Haenszel
Educational note: the ClinicalTrials.gov record contains effect estimates and confidence intervals but do not contain the underlying participant-level event and censoring times required to reconstruct Kaplan-Meier curves. A valid reconstructed curve should not be fabricated from summary hazard ratios alone.

9. Statistical Methods Explained

Why were log-rank tests used for the primary cardiovascular endpoints?

The primary cardiovascular outcomes were defined by the first occurrence of clinical events after randomization and were followed until a study cut-off. This creates time-to-event data rather than simply a yes/no outcome observed at one fixed time. The log-rank test is designed to compare event-time distributions between randomized groups while incorporating censored observations.

What does a hazard ratio of 1.022 mean?

A hazard ratio of 1.022 is a model-based relative comparison of the estimated instantaneous event rates. It is close to 1, meaning the fitted treatment-effect estimate is close to the null value. It is not a 2.2% difference in cumulative probability and does not indicate that each individual participant had a 2.2% higher risk.

Why is the confidence interval important?

The confidence interval provides information about statistical precision that a point estimate alone cannot provide. For the first primary endpoint, the interval is 0.937–1.114. For the second, it is 0.972–1.109. Both intervals include the null hazard ratio of 1, showing that the reported estimates are compatible with effects on either side of that value under the statistical framework.

Why use a Cox proportional-hazards model as well as a log-rank test?

The two methods serve related but different purposes. The log-rank test provides a formal comparison of the time-to-event experience between groups, while the Cox model provides an estimated hazard ratio and can incorporate specified stratification or covariate information. In ORIGIN, the registry reports both methods for the primary cardiovascular analyses.

Why was a Cochran-Mantel-Haenszel test used for diabetes development?

The diabetes-development endpoint was analyzed as a categorical outcome using the CMH method. CMH analysis provides a stratified comparison of odds while accounting for specified strata. Here, the odds ratio was stratified by double-blind treatment and previous cardiovascular event.

Why does ITT analysis matter?

ITT analysis preserves the treatment comparison created by randomization. In ORIGIN, the primary analyses were based on all randomized participants. This means the estimated treatment effect reflects assignment to insulin glargine versus standard care rather than only the experience of participants who remained fully compliant with their assigned treatment.

10. Factorial Design and Its Statistical Implications

The registry identifies ORIGIN as a factorial randomized trial with four arms and lists insulin glargine, omega-3 PUFA, placebo, and a reusable pen device among its interventions.

Design featureStatistical meaning
Factorial designMore than one intervention component is evaluated within a shared randomized trial structure.
Four armsThe registry reports four treatment arms rather than a simple two-arm trial.
Insulin comparisonThe posted analyses compare Insulin Glargine vs Standard Care.
Additional factorThe reported analyses identify double-blind treatment as omega-3 PUFA or placebo.
InteractionThe statistical analyses posted on ClinicalTrials.gov do not report a formal interaction estimate between the factorial treatment components.

Factorial designs can be statistically efficient because multiple treatment questions can be investigated in one randomized population. The important qualification is that the validity of a main-effect interpretation can depend on the absence, or appropriate handling, of meaningful interaction between factors. Because the registry analyses do not report an interaction result, no conclusion about such an interaction should be inferred here.

Do not confuse factorial design with four independent trials. The four arms arise from one randomized factorial structure. The insulin-glargine comparison therefore needs to be interpreted in the context of the other randomized factor rather than treating each arm as an unrelated experiment.

11. Stratification and Covariate Adjustment

The ORIGIN analyses illustrate two related but distinct approaches to controlling for important variables: stratification and covariate adjustment.

EndpointAdjustment / stratification reported
First primary cardiovascular compositeLog-rank stratified by double-blind treatment, baseline diabetes diagnosis, and previous cardiovascular event.
Second primary cardiovascular compositeLog-rank stratified by double-blind treatment, baseline diabetes diagnosis, and previous cardiovascular event; Cox regression used the corresponding stratification structure.
Total mortalityCox model with treatment plus double-blind treatment, baseline diabetes diagnosis, and previous cardiovascular event as covariates.
Composite diabetic microvascular outcomeCox model with treatment plus double-blind treatment, baseline diabetes diagnosis, and previous cardiovascular event as covariates.
Development of type 2 diabetesCMH analysis stratified by double-blind treatment and previous cardiovascular event.

The distinction is important. A stratified log-rank test compares event-time distributions while accounting for the strata. A Cox model can use stratification to allow different baseline hazards across strata, while covariates can enter the model directly as explanatory variables. These approaches should not be described as interchangeable mathematical operations.

12. Primary Endpoint Construction

Both primary endpoints are composite outcomes based on the first occurrence of specified clinical events. The first primary composite includes cardiovascular death, nonfatal myocardial infarction, or nonfatal stroke. The second expands the composite by adding revascularization and hospitalization for heart failure.

Why use a composite?

A composite endpoint can capture several clinically relevant event types within one prespecified outcome, allowing time to the first qualifying event to become the analysis target.

Why first occurrence matters

The registry specifies that the primary endpoint counts the first occurrence. Repeated events therefore do not contribute in the same way as the first qualifying event.

Why adjudication matters

The first primary endpoint used positively adjudicated events reviewed by an Event Adjudication Committee that was blinded to treatment assignment.

Why composite interpretation requires care

A hazard ratio for a composite describes the combined endpoint. It does not by itself establish that every individual component has the same magnitude or direction of effect.

13. Confidence Intervals and P-values

The two primary analyses provide a useful example of why point estimates, confidence intervals, and P-values answer different statistical questions.

EndpointEstimate95% CIP-value
CV death / nonfatal MI / nonfatal strokeHR 1.0220.937–1.1140.6273
CV death / nonfatal MI / nonfatal stroke / revascularization / hospitalization for HFHR 1.0380.972–1.1090.2692

The point estimate gives the fitted treatment effect. The 95% confidence interval describes uncertainty around that estimate under the statistical model and sampling framework. The P-value summarizes evidence against a null hypothesis under the specified testing framework.

Three quantities, three roles
Effect estimate → magnitude   |   Confidence interval → precision   |   P-value → evidence against the null

No one of these quantities should be used as a substitute for the others. In particular, a P-value is not an effect-size measure, and a confidence interval is not a prediction interval for individual participants.

14. Multiplicity and the Two Primary Endpoints

ORIGIN registered two primary endpoints. The registry-reported analysis notes state that the total required number of first coprimary outcomes, 2200, was based on assumptions that a hazard reduction of 14-16% would be clinically significant, with the overall experiment-wise Type 1 error controlled at 5% and power of 80% for each outcome.

Why two primary endpoints matter

Multiple primary endpoints create a multiplicity issue because repeated confirmatory testing can otherwise increase the probability of a false-positive conclusion.

Prespecified error control

The registry analysis notes explicitly describe control of the overall experiment-wise Type 1 error at 5% with power of 80% for each outcome.

The same analysis notes indicate that approximately 12,500 participants were ultimately estimated to be needed to achieve the targeted number of events within the planned enrollment and treatment periods. The actual enrollment was 12,537.

Interpretation principle: multiplicity is a property of the trial's prespecified testing strategy. A P-value should not be interpreted in isolation from the number and hierarchy of primary hypotheses being tested.

15. Safety Results

The ClinicalTrials.gov record reports serious adverse events by arm for the insulin-glargine comparison.

Safety measureInsulin GlargineStandard Care
Serious adverse events303 / 6231232 / 6273
Serious adverse events: affected participants / participants at risk
Insulin Glargine
303 / 6231
Standard Care
232 / 6273

The reported safety figures are counts of affected participants relative to the number at risk. They are presented descriptively here because the ClinicalTrials.gov record does not provide a formal statistical comparison, confidence interval, or P-value for serious adverse events.

Safety interpretation: the denominators differ between the two groups, so the affected/at-risk counts should not be read as if they were identical-sized samples. More importantly, the ClinicalTrials.gov record does not establish a formal hypothesis test for the difference in serious adverse events.

16. Analysis Populations and Missing Data

The ClinicalTrials.gov record explicitly identifies the intent-to-treat population for the reported efficacy analyses. Primary analyses were based on all randomized participants, and the total-mortality analysis explicitly states that participants were included regardless of compliance with the protocol.

Population / issueWhat the ClinicalTrials.gov record reports
Primary efficacy populationIntent-to-treat; all randomized participants
Total mortalityAll randomized participants regardless of protocol compliance
Microvascular outcomeIntent-to-treat; all randomized participants
Diabetes-development analysisSubgroup of the ITT population without diabetes at randomization
Missing-data / imputation strategyNot specified in the registry-reported statistical-analysis fields

Time-to-event methods naturally accommodate censoring when participants have not experienced the event by their last available follow-up. However, the ClinicalTrials.gov record does not specify an imputation method for missing observations, so no particular missing-data procedure should be attributed to ORIGIN beyond the analysis descriptions provided.

17. Why the Time-to-Event Framework Matters

The cardiovascular endpoints were followed from randomization until the study cut-off, with a median duration of follow-up of 6.2 years. That structure means the analysis is not simply a comparison of the number of participants who eventually experienced an event.

Conceptual survival function
S(t) = P(T > t)

The survival function represents the probability of remaining free of the event beyond time t. Time-to-event methods use both event timing and censoring information rather than reducing every participant to a single fixed-time binary observation.

For a participant who has not experienced the endpoint by the end of available follow-up, the observation is censored rather than treated as an event. This distinction is fundamental to Kaplan-Meier estimation, log-rank testing, and Cox regression.

Important distinction: the registry lists the primary endpoints as binary in its endpoint-type field, but the posted statistical analyses are explicitly time-to-event analyses. The first event and its timing are therefore central to the reported Cox and log-rank results.

18. Statistical Interpretation of the Five Reported Effects

EndpointWhat the estimate representsNull value
CV death / MI / strokeRelative hazard of first composite event under the Cox modelHR = 1
CV death / MI / stroke / revascularization / HF hospitalizationRelative hazard of first broader composite event under the Cox modelHR = 1
Total mortalityRelative hazard of death from any cause under the Cox modelHR = 1
Diabetic microvascular outcomeRelative hazard of first kidney or eye disease outcome under the Cox modelHR = 1
Development of type 2 diabetesRelative odds of the categorical diabetes-development outcome under CMH stratificationOR = 1

The first four estimates are hazard ratios and should therefore be interpreted in a time-to-event framework. The fifth is an odds ratio and belongs to a different statistical framework. Treating all five estimates as though they were the same measure would obscure an important methodological distinction.

19. What the Hazard Ratios Do — and Do Not — Mean

Hazard ratio 1.022

The estimated hazard for the first primary composite was 1.022 for insulin glargine relative to standard care under the reported Cox model. This is close to the null value of 1.

It does not mean that 2.2% more participants experienced the endpoint, nor does it mean that each participant's individual risk increased by 2.2%.

Hazard ratio 1.038

The estimated hazard for the broader cardiovascular composite was 1.038 relative to standard care. This is again a model-based time-to-event measure close to 1.

It does not represent a 3.8% absolute increase in cumulative event probability.

Hazard ratio 0.983

The all-cause mortality estimate of 0.983 is close to the null value. The confidence interval of 0.899–1.076 describes uncertainty around that model estimate.

Hazard ratio 0.970

The microvascular outcome estimate of 0.970 indicates a modeled hazard slightly below 1. The 95% CI of 0.900–1.047 includes the null value.

20. Odds Ratio vs Hazard Ratio

ORIGIN provides a useful teaching example because its statistical analyses include both a hazard ratio and an odds ratio.

FeatureHazard ratioOdds ratio
ORIGIN examplesCV composites, total mortality, microvascular outcomeDevelopment of type 2 diabetes in participants with IGT and/or IFG
Analysis frameworkTime-to-eventCategorical / stratified analysis
Reported methodCox proportional-hazards modelCochran-Mantel-Haenszel test
Null value11
Core interpretationRelative event hazard over follow-up under the fitted modelRelative odds of the outcome under the stratified analysis

An odds ratio of 0.72 should not be casually rewritten as a hazard ratio of 0.72 or a risk ratio of 0.72. The mathematical objects are different, and the appropriate interpretation depends on the outcome definition and analysis design.

21. Limitations and Interpretation Issues

22. Why This Trial Matters Statistically

ORIGIN is a useful teaching case because it combines several core methods in one randomized clinical-trial program: factorial allocation, ITT analysis, stratified log-rank testing, Cox proportional-hazards modeling, covariate adjustment, CMH analysis, hazard ratios, odds ratios, confidence intervals, and composite time-to-event endpoints.

ConceptHow it appears in ORIGIN
Randomization12,537 participants randomized in a phase 3 trial.
Factorial designFour-arm factorial structure involving insulin glargine and other listed interventions.
ITT analysisPrimary analyses based on all randomized participants.
Time-to-event endpointsPrimary cardiovascular outcomes followed from randomization until study cut-off.
Log-rank testUsed for the two primary cardiovascular comparisons.
Cox modelUsed to estimate hazard ratios for primary and secondary time-to-event endpoints.
Hazard ratioReported for both primary cardiovascular endpoints and two secondary time-to-event endpoints.
Stratified analysisUsed in log-rank and CMH analyses and incorporated into the cardiovascular Cox framework.
Covariate adjustmentUsed in the Cox analyses of total mortality and diabetic microvascular outcome.
Odds ratioUsed for development of type 2 diabetes in participants with IGT and/or IFG.
MultiplicityTwo coprimary outcomes with overall experiment-wise Type 1 error controlled at 5%.
Confidence intervalsReported for all five posted statistical analyses.

23. Longitudinal Trial History

2003-08

Trial start

The ORIGIN phase 3 trial began enrollment according to the registry.

2011-12

Primary completion

The registry records primary completion in December 2011.

6.2-year median follow-up

Mature time-to-event analysis

The registered primary cardiovascular endpoints and the reported secondary time-to-event outcomes use follow-up from randomization until study cut-off, with a median duration of follow-up of 6.2 years.

24. A Practical Reading Strategy for the ORIGIN Results

A statistically disciplined reading of ORIGIN starts with the randomized comparison and then moves through the analysis hierarchy.

  1. Identify the estimand: determine exactly which endpoint is being compared and whether it is a time-to-event or categorical outcome.
  2. Identify the population: confirm whether the analysis uses the ITT population or the specified subgroup without diabetes at randomization.
  3. Identify the method: distinguish log-rank/Cox analyses from the CMH analysis.
  4. Read the effect estimate: interpret the HR or OR according to its statistical definition.
  5. Read the confidence interval: assess the precision and whether the interval includes the null value.
  6. Read the P-value only in context: the P-value describes statistical evidence under the relevant hypothesis-testing framework; it does not measure effect size.
  7. Check the design: factorial structure, stratification, composite endpoints, and the ITT framework all affect interpretation.
Clinical Biostats principle: the most informative interpretation is rarely just "significant" or "not significant." The statistical story comes from the endpoint definition, analysis population, model, effect estimate, confidence interval, testing framework, and trial design considered together.

25. Related Tutorials

Learn more about the methods used in this trial:

26. Related Statistical Calculators

27. Sources

Continue through the Clinical Biostats statistical methods

Connect the ORIGIN endpoints to deeper tutorials on survival analysis, factorial trials, stratified methods, confidence intervals, and treatment-effect estimation.

28. Record Summary

ORIGIN provides a rich statistical example of a large randomized phase 3 factorial trial. Its primary cardiovascular endpoints were analyzed using stratified log-rank testing and Cox proportional-hazards models in the ITT population, with hazard ratios of 1.022 and 1.038. Secondary analyses extended the same survival framework to total mortality and diabetic microvascular outcomes, while development of type 2 diabetes in participants with IGT and/or IFG was evaluated with a stratified Cochran-Mantel-Haenszel analysis and an odds ratio of 0.72. The trial therefore illustrates how randomized design, endpoint construction, stratification, time-to-event analysis, covariate adjustment, categorical analysis, confidence intervals, and multiplicity fit together in a clinical-trial statistical analysis.

Clinical Biostats methodology: A trial-results page should not merely repeat the registry. The goal is to reconstruct the statistical story of the trial while clearly separating reported numerical evidence from educational interpretation and avoiding conclusions that are not supported by the registry-reported analysis.