← Clinical Trials
Type 2 Diabetes Mellitus Phase 3 Non-Inferiority NCT00790205

TECOS: Complete Statistical Analysis of Sitagliptin in Type 2 Diabetes Mellitus

An independent statistical review of the randomized, double-blind phase 3 TECOS study comparing sitagliptin with placebo in participants with type 2 diabetes mellitus, focusing on the registered MACE plus endpoint, Cox proportional-hazards methodology, non-inferiority framework, secondary time-to-event analyses, and safety.

TECOS  ·  Phase 3  ·  Completed  ·  Enrollment 14,671
Scope of this record

This page provides an independent statistical analysis and educational interpretation of publicly reported results. ClinicalTrials.gov provides the official trial registry record.

1. Trial at a Glance

TECOS was a randomized, double-blind, parallel phase 3 trial evaluating sitagliptin versus placebo in participants with type 2 diabetes mellitus. The registry reports 14,671 participants, two treatment arms, two registered primary endpoints, and formal Cox proportional-hazards analyses for both primary endpoints.

14,671
Enrollment
Randomized trial
2
Arms
Sitagliptin vs placebo
0.98
Primary HR
Per protocol; 95% CI 0.88–1.09
<0.001
NI P-value
Primary analyses
FeatureTECOS
Trial nameTECOS
Brief titleSitagliptin Cardiovascular Outcomes Study (MK-0431-082)
PhasePhase 3
ConditionType 2 Diabetes Mellitus
DesignRandomized, double-blind, parallel
AllocationRandomized
Primary purposeTreatment
Enrollment14,671
InterventionsSitagliptin and placebo
Primary endpointsTwo registered MACE plus endpoints, analyzed in per-protocol and intent-to-treat populations
Primary analysisCox proportional-hazards model
Effect measureHazard ratio
ClinicalTrials.govNCT00790205
Lead sponsorMerck Sharp & Dohme LLC

2. Clinical Question

The central statistical question was whether the time to a first confirmed cardiovascular event meeting the registered MACE plus definition could be compared between sitagliptin and placebo using a non-inferiority framework.

Population

Participants with type 2 diabetes mellitus enrolled in the randomized phase 3 TECOS study.

Intervention

Sitagliptin.

Comparator

Placebo.

Primary question

Is the hazard of the first confirmed MACE plus event with sitagliptin sufficiently close to placebo to satisfy the prespecified non-inferiority criterion?

3. Trial Design

01
Randomize14,671 participants
02
TreatSitagliptin or placebo
03
FollowUp to 5 years
04
AssessCV time-to-event endpoints
05
AnalyzeCox proportional hazards
ARM 1

Sitagliptin

  • Randomized intervention arm
  • Compared with placebo
  • Primary and secondary endpoints analyzed as time-to-event outcomes
ARM 2

Placebo

  • Randomized comparator arm
  • Compared with sitagliptin
  • Used as the reference group for reported hazard ratios

The registry identifies TECOS as randomized, parallel, and double-masked. This structure is particularly important for interpretation of the treatment comparison: randomization establishes the intended basis for comparing the assigned groups, while double masking reduces the potential for knowledge of assignment to influence treatment or assessment.

4. Endpoints

EndpointPopulationTime frameDefinition
First confirmed MACE plus Per protocol Up to 5 years CV-related death, nonfatal MI, nonfatal stroke, or unstable angina requiring hospitalization.
First confirmed MACE plus Intent to treat Up to 5 years CV-related death, nonfatal MI, nonfatal stroke, or unstable angina requiring hospitalization.

The registry classifies the two registered primary endpoints as binary outcomes, while the posted formal analyses identify them as time-to-event endpoints and use Cox proportional-hazards models. That distinction is statistically important: a participant's follow-up time and censoring status contain information that a simple endpoint proportion does not capture.

Secondary endpoints with formal analyses

Secondary endpointTime frameAnalysis
First confirmed MACEUp to 5 yearsCox proportional-hazards model
All-cause mortalityUp to 5 yearsCox proportional-hazards model
Congestive heart failure requiring hospitalizationUp to 5 yearsCox proportional-hazards model
Initiation of chronic insulin therapyUp to 5 yearsCox proportional-hazards model
Initiation of co-interventional agentUp to 5 yearsCox proportional-hazards model

5. Analysis Populations

The registry provides separate primary analyses for a per-protocol population and an intent-to-treat population. This creates an instructive comparison because non-inferiority trials commonly consider both adherence-sensitive and randomized-population perspectives.

PopulationRegistry definitionPrimary role
Per protocol Included all randomized participants who received study medication except those participants who did not contribute at least 1 day of data to the study analysis due to a major [registry description truncated in the ClinicalTrials.gov record]. Primary MACE plus non-inferiority analysis
Intent to treat Included all randomized participants who received study medication, provided consent, and did not have a major GCP deviation. Primary MACE plus non-inferiority analysis

The distinction matters because the two populations answer somewhat different statistical questions. The ITT analysis retains the randomized treatment assignment for the eligible randomized population, while the per-protocol analysis attempts to focus the comparison on participants satisfying the protocol-defined analysis requirements.

Why both populations matter: In a non-inferiority setting, relying on only one analysis population can be problematic. ITT analyses can be conservative or anti-conservative for non-inferiority depending on the nature of protocol deviations and treatment switching, while per-protocol analyses can lose some of the protection reported by randomization. Concordance between appropriately specified analyses provides a stronger basis for interpreting a non-inferiority result.

6. Statistical Methodology

Cox proportional-hazards model

All 12 posted statistical analyses in the ClinicalTrials.gov record uses a Cox proportional-hazards model. The Cox model relates the hazard of an event to treatment and, where specified, additional covariates or stratification variables.

Cox model
h(t | X) = h0(t) exp(βX)

The hazard ratio associated with treatment is obtained by exponentiating the treatment coefficient. In the simplest two-group comparison, HR = exp(β).

For the MACE plus analyses, the registry states that the model was stratified by region, with treatment group as the explanatory variable. This means the comparison permits the baseline hazard to differ across regions while estimating a common treatment hazard ratio across the strata.

Hazard ratio

The hazard ratio is the primary effect measure reported in TECOS. An HR of 1 corresponds to equal modeled hazards between the treatment groups. An HR below 1 indicates a lower estimated instantaneous event rate in the sitagliptin group relative to placebo under the fitted model, while an HR above 1 indicates a higher estimated instantaneous event rate.

Interpretation
HR = hSitagliptin(t) / hPlacebo(t)

The hazard ratio is not an absolute risk difference, not a probability of experiencing the event, and not a statement that every participant has the same proportional change in risk.

Stratification by region

The registry explicitly states that the primary MACE plus Cox models were stratified by region. Stratification allows the underlying event hazard to differ by region without requiring the region effect itself to be represented by a single proportional-hazards coefficient.

Time-to-event analysis and censoring

A time-to-event analysis uses both whether an event occurred and how long each participant was observed before the event or censoring. This allows participants with different amounts of follow-up to contribute information without treating all participants as though they had identical observation time.

Educational note: a Kaplan-Meier estimator would ordinarily be a natural descriptive tool for a time-to-event endpoint, but the registry-reported TECOS statistical-analysis records specifically identify the formal posted analyses as Cox proportional-hazards models. No Kaplan-Meier estimates or curves are added here because they are not contained in the ClinicalTrials.gov record.

7. Primary Results: MACE Plus — Per Protocol

The first registered primary endpoint was the percentage of participants with a first confirmed cardiovascular event of MACE plus in the per-protocol population, over a time frame of up to 5 years.

Hazard ratio for first confirmed MACE plus

0.98

95% CI: 0.88–1.09   ·   P < 0.001

Non-inferiority analysis; two-sided 95% confidence interval

FeatureReported result
ComparisonSitagliptin vs placebo
EndpointFirst confirmed MACE plus
Analysis populationPer protocol
Time frameUp to 5 years
ModelCox proportional-hazards model
StratificationRegion
Effect measureHazard ratio
Estimate0.98
95% CI0.88–1.09
P-value<0.001
Hypothesis typeNon-inferiority or equivalence
Clinical Biostats interpretation

The estimated hazard ratio of 0.98 means that the fitted Cox model estimated the instantaneous hazard of the first MACE plus event in the sitagliptin group at approximately 98% of the corresponding placebo hazard, under the model's assumptions and over the analyzed follow-up.

It does not mean that 98% of participants experienced an event, that the absolute event probability was 98%, or that every participant had exactly a 2% reduction in risk. The hazard ratio is a relative time-to-event measure.

The 95% CI of 0.88–1.09 describes statistical uncertainty around the estimated hazard ratio. It gives a range of values compatible with the model and data under the stated confidence framework; it is not a range containing the treatment effect for 95% of individual participants.

The P < 0.001 value does not measure the size of the treatment effect. A p-value addresses evidence against a specified statistical hypothesis under the analysis framework. In this case, the registry identifies the hypothesis type as non-inferiority or equivalence and states that the confidence interval was compared with the non-inferiority margin of 1.30.

For a non-inferiority analysis, the key question is therefore not whether the HR is statistically different from 1. Instead, the upper confidence limit must be considered relative to the prespecified margin. Here, the reported upper limit of 1.09 is below the stated margin of 1.30. The analysis therefore supports the registry's non-inferiority framework without requiring the treatment effect itself to be statistically superior to placebo.

8. Primary Results: MACE Plus — Intent to Treat

The second registered primary endpoint used the intent-to-treat population. The endpoint definition and time frame were the same: first confirmed MACE plus over up to 5 years.

Hazard ratio for first confirmed MACE plus

0.98

95% CI: 0.89–1.08   ·   P < 0.001

Non-inferiority analysis; two-sided 95% confidence interval

FeatureReported result
ComparisonSitagliptin vs placebo
EndpointFirst confirmed MACE plus
Analysis populationIntent to treat
Time frameUp to 5 years
ModelCox proportional-hazards model
StratificationRegion
Effect measureHazard ratio
Estimate0.98
95% CI0.89–1.08
P-value<0.001
Hypothesis typeNon-inferiority or equivalence
Clinical Biostats interpretation

The ITT hazard ratio of 0.98 is essentially the same point estimate as the per-protocol analysis. The 95% CI is 0.89–1.08, which is somewhat narrower on the upper side than the per-protocol interval of 0.88–1.09.

Again, the HR should not be interpreted as an absolute risk difference or as the percentage of participants protected from MACE plus. It summarizes the relative event hazard estimated by the Cox model.

The reported P < 0.001 should not be read as evidence that the HR is meaningfully different from 1. Non-inferiority asks a different question: whether the observed uncertainty is sufficiently bounded below the prespecified margin. The registry analysis states that the confidence interval was compared with 1.30.

The upper confidence limit of 1.08 remains below 1.30. Thus, within the registry's stated non-inferiority framework, the ITT analysis provides the same directional conclusion as the per-protocol analysis. The fact that both populations produce closely aligned estimates is statistically informative because it reduces concern that the non-inferiority conclusion is driven solely by the choice of analysis population.

9. Understanding the Non-Inferiority Margin

The ClinicalTrials.gov record identifies 1.30 as the non-inferiority margin for the hazard ratio of sitagliptin relative to placebo. The formal logic is different from a conventional superiority test against HR = 1.

Non-inferiority logic
Upper 95% CI < 1.30  →  compatible with the stated non-inferiority criterion

The margin defines the largest relative hazard considered acceptable for the non-inferiority claim under the prespecified design. It is therefore the comparison boundary for the confidence interval, rather than 1 itself.

Primary analysisHR95% CIUpper CI vs 1.30
Per protocol0.980.88–1.091.09 < 1.30
Intent to treat0.980.89–1.081.08 < 1.30

This is one of the most important statistical features of TECOS. A non-inferiority analysis is not asking whether sitagliptin is superior to placebo. It is asking whether the data exclude a treatment difference sufficiently large to cross the prespecified non-inferiority boundary.

Do not confuse non-inferiority with equivalence. A non-inferiority result does not establish that the two treatments have exactly the same effect. It establishes that the uncertainty around the estimated difference is sufficiently constrained relative to the prespecified non-inferiority margin. The registry labels the hypothesis type as “Non-inferiority or equivalence,” while the specific primary analyses describe comparison with the 1.30 non-inferiority margin.

10. Secondary Results: MACE

The registry reports formal secondary analyses of the first confirmed MACE endpoint in both the per-protocol and intent-to-treat populations.

PopulationHR95% CIP-valueHypothesis
Per protocol0.990.89–1.11<0.001Non-inferiority or equivalence
Intent to treat0.990.89–1.10<0.001Non-inferiority or equivalence

Both analyses used a Cox proportional-hazards model stratified by region, with treatment group as the explanatory variable. The reported point estimate of 0.99 indicates an estimated hazard very close to that of placebo.

Clinical Biostats interpretation

The secondary MACE estimates are close to 1 in both analysis populations. Their confidence intervals quantify the uncertainty around those estimates, while the reported p-values should be interpreted in the context of the stated non-inferiority hypothesis rather than as measures of effect magnitude.

The similarity between the per-protocol HR of 0.99 and the ITT HR of 0.99 is descriptive evidence that the estimated treatment comparison is not highly sensitive to the registry-reported choice of analysis population. It does not, by itself, establish equivalence across every possible analysis population or endpoint.

11. Secondary Results: All-Cause Mortality

PopulationHR95% CIP-valueHypothesis
Per protocol1.060.91–1.240.435Superiority
Intent to treat1.010.90–1.140.875Superiority

The registry states that the all-cause mortality analyses assessed the between-treatment difference using the hazard ratio of sitagliptin to placebo.

Clinical Biostats interpretation

The per-protocol estimate of 1.06 corresponds to an estimated hazard modestly above 1, while the ITT estimate of 1.01 is very close to 1. The confidence intervals for both analyses include 1.

The p-values of 0.435 and 0.875 do not establish that the two treatment groups have identical mortality hazards. A nonsignificant superiority test means that the registry-reported analysis does not provide sufficient evidence for a superiority difference under that test; it does not prove equality.

The wider interpretation should also respect the confidence intervals: 0.91–1.24 and 0.90–1.14 describe the uncertainty around the respective model estimates. They are more informative than the p-values alone about the range of relative hazards compatible with the analyses.

12. Secondary Results: Congestive Heart Failure Requiring Hospitalization

PopulationHR95% CIP-valueModel covariates
Per protocol0.980.81–1.190.858Treatment group; history of CHF at baseline; stratified by region
Intent to treat1.000.83–1.200.983Treatment group; history of CHF at baseline; stratified by region

Unlike the primary MACE plus analyses, the CHF hospitalization models included history of CHF at baseline as an explanatory variable in addition to treatment group, with the model stratified by region.

Clinical Biostats interpretation

The per-protocol HR of 0.98 and ITT HR of 1.00 are close to the null value of 1. Their confidence intervals are also relatively broad compared with the point estimates, illustrating why a point estimate alone should not be treated as the complete result.

The p-values of 0.858 and 0.983 are tests of the specified superiority hypotheses. They do not quantify the probability that the treatment effects are exactly equal, nor do they measure clinical importance.

13. Secondary Results: Initiation of Chronic Insulin Therapy

Per-protocol analysis

HR 0.69

95% CI: 0.61–0.77   ·   P < 0.001

Intent-to-treat analysis

HR 0.70

95% CI: 0.63–0.79   ·   P < 0.001

PopulationHR95% CIP-valueHypothesis
Per protocol0.690.61–0.77<0.001Superiority
Intent to treat0.700.63–0.79<0.001Superiority
Clinical Biostats interpretation

The reported HR of 0.69 in the per-protocol population means that the estimated instantaneous hazard of initiating chronic insulin therapy was approximately 69% of the placebo-group hazard under the fitted Cox model. Equivalently, the point estimate corresponds to an approximately 31% lower estimated hazard relative to placebo.

The ITT HR of 0.70 similarly corresponds to an approximately 30% lower estimated hazard. These are relative hazard interpretations, not absolute reductions in the percentage of participants initiating insulin.

The 95% confidence intervals, 0.61–0.77 and 0.63–0.79, describe the precision of the estimated hazard ratios. The p-values of <0.001 provide evidence against the respective superiority null hypotheses under the registry's specified analyses, but the p-values do not tell us whether the effect is clinically important or how many participants were prevented from initiating insulin.

14. Secondary Results: Initiation of a Co-Interventional Agent

PopulationHR95% CIP-valueHypothesis
Per protocol0.700.65–0.75<0.001Superiority
Intent to treat0.720.68–0.77<0.001Superiority

Both analyses used Cox proportional-hazards models stratified by region, with treatment group as the explanatory variable.

Clinical Biostats interpretation

The per-protocol estimate of 0.70 corresponds to an approximately 30% lower estimated instantaneous hazard of initiation of a co-interventional agent relative to placebo. The ITT estimate of 0.72 corresponds to an approximately 28% lower estimated hazard.

Those percentage statements are simple interpretations of the reported hazard ratios; they should not be converted into absolute percentages of participants experiencing the event. The confidence intervals provide the relevant information about statistical precision, while the p-values indicate evidence under the stated superiority hypotheses.

15. Summary of All Posted Statistical Analyses

The ClinicalTrials.gov record contains 12 formal statistical analyses: two primary MACE plus analyses and ten secondary analyses covering MACE, all-cause mortality, CHF hospitalization, chronic insulin initiation, and co-interventional-agent initiation, each evaluated in per-protocol and intent-to-treat populations where provided.

EndpointPopulationHR95% CIP-valueHypothesis
MACE plusPer protocol0.980.88–1.09<0.001Non-inferiority
MACE plusIntent to treat0.980.89–1.08<0.001Non-inferiority
MACEPer protocol0.990.89–1.11<0.001Non-inferiority
MACEIntent to treat0.990.89–1.10<0.001Non-inferiority
All-cause mortalityPer protocol1.060.91–1.240.435Superiority
All-cause mortalityIntent to treat1.010.90–1.140.875Superiority
CHF hospitalizationPer protocol0.980.81–1.190.858Superiority
CHF hospitalizationIntent to treat1.000.83–1.200.983Superiority
Chronic insulin initiationPer protocol0.690.61–0.77<0.001Superiority
Chronic insulin initiationIntent to treat0.700.63–0.79<0.001Superiority
Co-interventional-agent initiationPer protocol0.700.65–0.75<0.001Superiority
Co-interventional-agent initiationIntent to treat0.720.68–0.77<0.001Superiority

This table is intentionally limited to the 12 analyses reported in the ClinicalTrials.gov record. No additional endpoint estimates, event counts, medians, subgroup results, or analyses are introduced.

16. Statistical Methods Explained

Why was a Cox proportional-hazards model used?

The analyzed outcomes concern the time until an event, with follow-up extending up to 5 years. A Cox model is designed for this structure because it incorporates both event occurrence and follow-up time and can accommodate right-censored observations. It produces a hazard ratio that summarizes the relative event hazard between treatment groups.

What does an HR of 0.98 mean?

An HR of 0.98 means the fitted model estimates the sitagliptin hazard at approximately 98% of the placebo hazard, subject to the model assumptions. It does not mean that 98% of participants avoided an event or that the absolute risk was reduced by 2 percentage points.

Why is the non-inferiority margin 1.30 important?

The margin defines the largest relative hazard compatible with the prespecified non-inferiority criterion. For the TECOS primary analyses, the registry states that the confidence interval was compared with 1.30. Thus, the upper confidence limit—not merely whether the HR differs from 1—is central to the non-inferiority decision.

Why report both per-protocol and ITT analyses?

The two populations emphasize different properties of the randomized comparison. The ITT population preserves treatment assignment for the eligible randomized population, whereas the per-protocol analysis focuses on participants satisfying the protocol-defined analysis requirements. In a non-inferiority trial, agreement between these perspectives is especially useful because protocol deviations can influence the interpretation of non-inferiority.

What does stratifying the Cox model by region accomplish?

Stratification permits the underlying baseline hazard to differ among regions while estimating the treatment hazard ratio across those strata. The registry specifically states that the primary MACE plus models were stratified by region and included treatment group as the explanatory variable.

Why does a p-value below 0.001 not tell us the size of the treatment effect?

A p-value measures the compatibility of the observed data with a specified statistical null hypothesis under the analysis framework. It does not quantify the magnitude or clinical importance of an effect. TECOS illustrates this distinction particularly well: a non-inferiority analysis can have a very small p-value while the key inferential question remains whether the confidence interval stays within the non-inferiority margin.

What assumption is important for interpreting a Cox hazard ratio?

The proportional-hazards model assumes that the relative hazard represented by the model is appropriately described by the proportional-hazards structure. If the treatment hazard ratio changes substantially over time, a single HR can become a less complete description of the treatment-time relationship. The ClinicalTrials.gov record identifies the Cox model but do not provide a formal proportional-hazards diagnostic, so no such diagnostic is inferred here.

17. Confidence Intervals and Precision

The primary estimates are especially instructive when read together with their confidence intervals.

Primary hazard-ratio estimates and 95% confidence intervals
Per protocol
0.98
Intent to treat
0.98

The graphic emphasizes the point estimates only; the exact uncertainty intervals are 0.88–1.09 and 0.89–1.08, respectively.

The two primary point estimates are identical at 0.98, while their confidence intervals differ slightly. This is a useful reminder that the point estimate alone does not determine the strength or precision of statistical evidence.

Point estimate

The HR is the single best estimate produced by the specified Cox model.

Confidence interval

The interval describes statistical uncertainty around the estimated HR under the specified model and confidence framework.

P-value

The p-value evaluates evidence against a specified hypothesis; it is not an effect-size metric.

Non-inferiority margin

The margin of 1.30 provides the prespecified boundary against which the upper confidence limit is compared.

18. Safety

The ClinicalTrials.gov record reports serious adverse events by randomized arm using affected participants over the corresponding at-risk populations.

Safety measureSitagliptinPlacebo
Serious adverse events928/7266909/7274

Sitagliptin

928 participants with serious adverse events among 7,266 at risk.

Placebo

909 participants with serious adverse events among 7,274 at risk.

These safety figures are descriptive counts of affected participants relative to the registry-reported at-risk populations. The ClinicalTrials.gov record does not provide a formal comparative statistical analysis for serious adverse events, so no hazard ratio, risk ratio, confidence interval, or p-value is assigned to these figures.

Safety interpretation: the denominators for the reported serious-adverse-event figures are not identical to the overall enrollment of 14,671. The ClinicalTrials.gov record identifies the specific affected/at-risk pairs, and those figures should be preserved rather than replacing them with an inferred denominator.

19. Multiplicity and Interpretation of Secondary Analyses

The registry supplies two primary analyses and ten secondary statistical analyses. The secondary analyses include both non-inferiority/equivalence hypotheses and superiority hypotheses.

Analysis familyExamplesHypothesis type in the ClinicalTrials.gov record
PrimaryMACE plusNon-inferiority or equivalence
Secondary CV efficacyMACENon-inferiority or equivalence
Secondary mortalityAll-cause mortalitySuperiority
Secondary CHFCHF requiring hospitalizationSuperiority
Secondary treatment-modification outcomesChronic insulin; co-interventional agentSuperiority

Because multiple endpoints and analysis populations are reported, individual p-values should not automatically be interpreted as though they were isolated tests with no multiplicity considerations. The ClinicalTrials.gov record does not provide an alpha-allocation or multiplicity-adjustment scheme for these secondary analyses. Accordingly, this page reports the posted p-values exactly but does not assign them an additional confirmatory interpretation beyond the hypothesis type explicitly provided.

20. Crossover, Missing Data, and Bayesian Methods

Crossover

The registry-reported TECOS trial data do not report a crossover analysis or crossover rate. No crossover adjustment is therefore described here.

Missing data and imputation

The ClinicalTrials.gov record defines the per-protocol and intent-to-treat populations but do not provide a missing-data or imputation method. No specific imputation strategy is inferred.

Bayesian methods

No Bayesian statistical method is reported in the ClinicalTrials.gov record. The posted formal analyses use Cox proportional-hazards models with hazard ratios.

Why this matters: statistical analysis pages should distinguish between methods actually documented in the available registry data and methods that might be reasonable alternatives. A method can be statistically plausible without being part of the reported TECOS analysis.

21. Stratified Analysis

Region is explicitly identified as the stratification variable in the primary MACE plus Cox models and in several secondary analyses. The CHF hospitalization models additionally include history of CHF at baseline as an explanatory variable.

AnalysisStratificationExplanatory variables
MACE plus, per protocolRegionTreatment group
MACE plus, intent to treatRegionTreatment group
MACE, per protocolRegionTreatment group
MACE, intent to treatRegionTreatment group
CHF hospitalization, per protocolRegionTreatment group; history of CHF at baseline
CHF hospitalization, intent to treatRegionTreatment group; history of CHF at baseline
Chronic insulin initiation, per protocolRegionTreatment group
Chronic insulin initiation, intent to treatRegionTreatment group
Co-interventional-agent initiation, per protocolRegionTreatment group
Co-interventional-agent initiation, intent to treatRegionTreatment group

Stratification is not the same as adjustment through an ordinary regression coefficient. In a stratified Cox model, the baseline hazard is allowed to differ by stratum while the treatment effect is estimated across the strata.

22. Clinical Biostats Interpretation of the Primary Analysis

What the result says

The primary MACE plus analyses produced the same hazard-ratio point estimate, 0.98, in both the per-protocol and ITT populations. Their 95% confidence intervals were 0.88–1.09 and 0.89–1.08, respectively. The registry identifies the analyses as non-inferiority analyses and states that the confidence interval was compared with a margin of 1.30.

What the result does not say

The result does not establish that sitagliptin and placebo have exactly identical hazards. It also does not provide an absolute event probability, an absolute risk difference, or a statement about the treatment effect in every individual participant.

Why the confidence interval matters

The confidence interval is central to the non-inferiority interpretation because the question is whether the plausible upper boundary of the relative hazard remains below the prespecified margin. For the two primary analyses, the upper limits are 1.09 and 1.08, respectively.

Why the p-value is secondary to the margin

The p-value of <0.001 is not the reason the non-inferiority criterion is satisfied. Non-inferiority is determined by comparing the confidence interval with the prespecified margin. The p-value and confidence interval answer related but different statistical questions.

23. Limitations

24. Why This Trial Matters Statistically

TECOS is a particularly useful teaching case because the registry-reported analysis combines randomized treatment assignment, double masking, time-to-event methodology, stratified Cox regression, separate ITT and per-protocol analyses, and a clearly stated non-inferiority margin.

ConceptHow it appears in TECOS
RandomizationThe study is registered as randomized with two parallel treatment arms.
BlindingThe registry identifies the study as double-masked.
Time-to-event endpointsThe formal statistical analyses use time-to-event methodology over up to 5 years.
Cox modelAll 12 registry-reported formal statistical analyses use Cox proportional-hazards models.
Hazard ratioThe treatment effect is expressed as a hazard ratio for sitagliptin versus placebo.
StratificationPrimary MACE plus models are stratified by region.
Non-inferiorityThe primary MACE plus analyses compare confidence intervals with a margin of 1.30.
ITT analysisA registered primary endpoint is formally analyzed in an intent-to-treat population.
Per-protocol analysisA second registered primary endpoint is formally analyzed in a per-protocol population.
Confidence intervalsBoth primary analyses report two-sided 95% confidence intervals.
Superiority testingSecondary analyses of mortality, CHF, insulin initiation, and co-interventional-agent initiation are identified as superiority hypotheses.
Safety reportingSerious adverse events are reported as affected participants over at-risk participants by arm.

25. Understanding the Primary Result as a Statistical Workflow

01
DefineMACE plus
02
RandomizeSitagliptin vs placebo
03
FollowUp to 5 years
04
ModelStratified Cox
05
CompareCI vs 1.30

This sequence illustrates an important principle in clinical-trial statistics: the interpretation of a final number depends on the entire analysis specification that generated it. The hazard ratio cannot be separated from the endpoint definition, analysis population, censoring structure, model, stratification, confidence interval, and hypothesis framework.

26. Primary vs Secondary Questions

QuestionEndpointPopulationInference framework
PrimaryMACE plusPer protocolNon-inferiority; HR and 95% CI
PrimaryMACE plusIntent to treatNon-inferiority; HR and 95% CI
SecondaryMACEPer protocol / ITTNon-inferiority or equivalence
SecondaryAll-cause mortalityPer protocol / ITTSuperiority
SecondaryCHF hospitalizationPer protocol / ITTSuperiority
SecondaryChronic insulin initiationPer protocol / ITTSuperiority
SecondaryCo-interventional-agent initiationPer protocol / ITTSuperiority

This hierarchy matters because “statistically significant” is not a universal interpretation. The same p-value notation can occur under very different hypotheses. TECOS therefore provides a useful example of why the endpoint role and hypothesis type should always be displayed alongside the numerical result.

27. Statistical Concepts in This Trial

Learn more about the methods used in this trial:

28. Related Statistical Calculators

29. Sources

Continue with the statistical methods behind TECOS

Explore tutorials and statistical calculators covering survival analysis, hazard ratios, confidence intervals, non-inferiority design, and clinical-trial methodology.

30. Record Summary

TECOS provides a clear example of how a large randomized clinical trial can use time-to-event methodology to address a non-inferiority question. The ClinicalTrials.gov record shows two registered primary MACE plus endpoints, one analyzed in the per-protocol population and one in the intent-to-treat population. Both use stratified Cox proportional-hazards models, both report an HR of 0.98, and both have upper confidence limits below the stated non-inferiority margin of 1.30.

The secondary analyses extend the same Cox framework to MACE, all-cause mortality, CHF hospitalization, chronic insulin initiation, and initiation of a co-interventional agent. The results illustrate why effect estimates, confidence intervals, p-values, analysis populations, and hypothesis types must be interpreted together. In particular, the primary non-inferiority analyses cannot be reduced to the statement that a p-value is below a threshold: the prespecified margin and the confidence interval are central to the inference.

The safety information posted on ClinicalTrials.gov for TECOS is more limited, consisting of serious adverse-event counts and at-risk denominators for each arm. No formal safety comparison is inferred. Likewise, no crossover, imputation, Bayesian method, subgroup result, event count, median time, or additional efficacy estimate is added beyond the ClinicalTrials.gov record.

Clinical Biostats methodology: A trial-results page should separate the reported statistical evidence from educational interpretation. For TECOS, the most important statistical lesson is the distinction between a conventional superiority question and a non-inferiority question: the treatment hazard ratio, its confidence interval, and the prespecified non-inferiority margin must be considered together.