Scientific deep-dive

How to Trust an Observational Study

We have twice warned that a very large hazard ratio from an observational study is telling you about the study. Here is one reporting 0.42 and 0.58 that deserves trust, and the specific techniques that earn it.

By Nora Bissett · Pricing Editor
Editorially reviewed (not clinically reviewed). Not medical advice · How we verify contentLast reviewed
7 min read·1 citation

We have twice warned that a very large hazard ratio from an observational study is usually telling you about the study rather than the drug. Here is one reporting hazard ratios of 0.42 and 0.58 that deserves considerably more trust — and the reason is a set of specific techniques worth being able to recognize.[1]

The technique that earns the trust

Before asking their real question, the researchers used their claims database to emulate two randomized trials whose results were already published — STEP-HFpEF DM and SUMMIT. They rebuilt each trial’s eligibility criteria, exposures and outcomes inside the data, ran the analysis, and compared what came out against what the trials had found.

The emulations agreed with the trials on every prespecified metric.

Only then did they widen the eligibility criteria to the kind of patients seen in ordinary practice — the people trials exclude — and ask the question nobody had randomized evidence for.

That is a positive control for an entire study design. A method that reproduces answers we already know has earned some credit when it reports an answer we do not. It is the same logic as testing a detector on a known signal before trusting it on an unknown one, and almost no observational study in this field does it.

The second technique is negative control outcomes — outcomes the drug has no plausible reason to affect. If the results were being produced by healthier people getting the newer drug, those outcomes would improve too. They did not, which is evidence the effect is not simply a healthier group in disguise.

What they found

The question was heart failure with preserved ejection fraction alongside obesity or type 2 diabetes — a condition where early trials were promising but rested on very few actual events. The outcome was a composite of hospitalization for heart failure or death from any cause, over up to 52 weeks. Sitagliptin, a drug with no expected effect here, served as the comparator.

US claims data, 2018 to 2024, with propensity score weighting.[1]
ComparisonPatientsHazard ratio (95% CI)
Semaglutide vs sitagliptin58,3330.58 (0.51–0.65)
Tirzepatide vs sitagliptin11,2570.42 (0.31–0.57)
Tirzepatide vs semaglutide28,1000.86 (0.70–1.06)

Secondary endpoints, subgroups and sensitivity analyses were consistent, and no substantial safety signal appeared.

The third row is a null, and it matters

Set head to head, tirzepatide showed no meaningfully lower risk than semaglutide — a hazard ratio of 0.86 with an interval running from 0.70 to 1.06, which includes no difference.

That is worth sitting with, because tirzepatide produces substantially more weight loss. If the benefit here ran through weight, the drug that removes more of it should win. It did not, which fits the larger pattern we have covered: the cardiovascular benefit does not track weight loss, and a synthesis of 262 trials similarly found the drug with the mortality evidence sitting fifth on the weight table.

What still limits it

  • Sitagliptin is a placebo proxy, not a placebo. It is chosen because it should not affect these outcomes, and it is still a drug given to people for reasons.
  • Claims data cannot see an echocardiogram. Preserved ejection fraction is inferred from diagnostic codes, so some patients are certainly misclassified.
  • Fifty-two weeks, in a chronic condition.
  • Benchmarking validates the method against two specific trials, in the populations those trials enrolled. It does not guarantee the method behaves as well in the expanded population, which is precisely where no randomized answer exists.

That last caveat is the honest limit of the whole approach, and it does not undo the value of the exercise. A method shown to work where it can be checked is a better instrument than one never checked at all.

The portable version

When you meet a large effect from an observational study, these are the questions that separate the credible from the merely striking.

  • Did they benchmark against a known answer? Emulating an existing trial and reproducing it is the strongest signal available.
  • Are there negative control outcomes? If everything improved, the patients differed, not the treatment.
  • Is the comparator an active drug or nothing? Comparing against another treatment partly controls for who gets prescribed what.
  • Is the effect size physiologically plausible? This remains the first filter — when a hazard ratio is too good.

Frequently Asked Questions

In this analysis of US claims data, semaglutide and tirzepatide were each associated with substantially lower risk of heart failure hospitalization or death compared with sitagliptin — hazard ratios of 0.58 and 0.42 over up to 52 weeks.
Because the researchers first emulated two randomized trials inside the same data and reproduced their published results, then expanded to the wider population. A method that reproduces known answers has earned credit on an unknown one.
No meaningful difference appeared — a hazard ratio of 0.86 with a confidence interval from 0.70 to 1.06, which includes no difference, despite tirzepatide producing more weight loss.
An outcome the treatment has no plausible reason to affect. If it improves too, that suggests the treated group was simply healthier rather than that the drug worked.
Sitagliptin is a stand-in for placebo rather than a placebo, preserved ejection fraction is inferred from diagnostic codes rather than measured, follow-up is 52 weeks, and the benchmarking validates the method in the trial populations rather than the expanded one.

References

  1. 1.Krüger N, Schneeweiss S, Fuse K, et al. Semaglutide and Tirzepatide in Patients With Heart Failure With Preserved Ejection Fraction JAMA. 2025. PMID: 40886075.

Where to get GLP-1 online, safely: sellers our editors have checked

These are telehealth sellers our editors have checked. For each one we hold a price, the form the drug comes in, and the states it reaches.

No insurance needed · vetted by our editors

Some of the links on this page earn us money. If you sign up with a provider after following one, that provider may pay GLP Watchdog a commission. Learn more

8.4

Collective

Flat any-dose pricing, if you can absorb a $199 annual membership on top

8.6

Found

Tirzepatide at $169/month, 37% below the typical price

7.7

HealthRX

Compounded semaglutide at $133/month