Scientific deep-dive

How to Trust an Observational Study

We have twice warned that a very large hazard ratio from an observational study is telling you about the study. Here is one reporting 0.42 and 0.58 that deserves trust — and the specific techniques that earn it.

By Nora Bissett · Pricing Editor
Editorially reviewed & fact-checked against primary sources · How we verify contentLast reviewed
7 min read·1 citations

We have twice warned that a very large hazard ratio from an observational study is usually telling you about the study rather than the drug. Here is one reporting hazard ratios of 0.42 and 0.58 that deserves considerably more trust — and the reason is a set of specific techniques worth being able to recognize.[1]

The technique that earns the trust

Before asking their real question, the researchers used their claims database to emulate two randomized trials whose results were already published — STEP-HFpEF DM and SUMMIT. They rebuilt each trial’s eligibility criteria, exposures and outcomes inside the data, ran the analysis, and compared what came out against what the trials had found.

The emulations agreed with the trials on every prespecified metric.

Only then did they widen the eligibility criteria to the kind of patients seen in ordinary practice — the people trials exclude — and ask the question nobody had randomized evidence for.

That is a positive control for an entire study design. A method that reproduces answers we already know has earned some credit when it reports an answer we do not. It is the same logic as testing a detector on a known signal before trusting it on an unknown one, and almost no observational study in this field does it.

The second technique is negative control outcomes — outcomes the drug has no plausible reason to affect. If the results were being produced by healthier people getting the newer drug, those outcomes would improve too. They did not, which is evidence the effect is not simply a healthier group in disguise.

What they found

The question was heart failure with preserved ejection fraction alongside obesity or type 2 diabetes — a condition where early trials were promising but rested on very few actual events. The outcome was a composite of hospitalization for heart failure or death from any cause, over up to 52 weeks. Sitagliptin, a drug with no expected effect here, served as the comparator.

US claims data, 2018 to 2024, with propensity score weighting.[[cite:1]]
ComparisonPatientsHazard ratio (95% CI)
Semaglutide vs sitagliptin58,3330.58 (0.51–0.65)
Tirzepatide vs sitagliptin11,2570.42 (0.31–0.57)
Tirzepatide vs semaglutide28,1000.86 (0.70–1.06)

Secondary endpoints, subgroups and sensitivity analyses were consistent, and no substantial safety signal appeared.

The third row is a null, and it matters

Set head to head, tirzepatide showed no meaningfully lower risk than semaglutide — a hazard ratio of 0.86 with an interval running from 0.70 to 1.06, which includes no difference.

That is worth sitting with, because tirzepatide produces substantially more weight loss. If the benefit here ran through weight, the drug that removes more of it should win. It did not, which fits the larger pattern we have covered: the cardiovascular benefit does not track weight loss, and a synthesis of 262 trials similarly found the drug with the mortality evidence sitting fifth on the weight table.

What still limits it

  • Sitagliptin is a placebo proxy, not a placebo. It is chosen because it should not affect these outcomes, and it is still a drug given to people for reasons.
  • Claims data cannot see an echocardiogram. Preserved ejection fraction is inferred from diagnostic codes, so some patients are certainly misclassified.
  • Fifty-two weeks, in a chronic condition.
  • Benchmarking validates the method against two specific trials, in the populations those trials enrolled. It does not guarantee the method behaves as well in the expanded population, which is precisely where no randomized answer exists.

That last caveat is the honest limit of the whole approach, and it does not undo the value of the exercise. A method shown to work where it can be checked is a better instrument than one never checked at all.

The portable version

When you meet a large effect from an observational study, these are the questions that separate the credible from the merely striking.

  • Did they benchmark against a known answer? Emulating an existing trial and reproducing it is the strongest signal available.
  • Are there negative control outcomes? If everything improved, the patients differed, not the treatment.
  • Is the comparator an active drug or nothing? Comparing against another treatment partly controls for who gets prescribed what.
  • Is the effect size physiologically plausible? This remains the first filter — when a hazard ratio is too good.

Frequently Asked Questions

References

  1. 1.Krüger N, Schneeweiss S, Fuse K, et al. Semaglutide and Tirzepatide in Patients With Heart Failure With Preserved Ejection Fraction JAMA. 2025. PMID: 40886075.

Where to get GLP-1 online, safely: sellers our editors have checked

These are telehealth sellers our editors have checked. For each one we hold a price, the form the drug comes in, and the states it reaches.

No insurance needed · vetted by our editors

Some of the links on this page earn us money. If you sign up with a provider after following one, that provider may pay GLP Watchdog a commission. Learn more

6.0

SkinnyRx

Starting below a standard dose, with microdose tiers

8.3

SnagRx

Semaglutide at $99/month, 48% under the register median

9.3

Embody

Knowing which pharmacy fills the vial — it names RedRock Pharmacy