Scientific deep-dive

The Statistic That Asks If Groups Differ

A SOUL secondary analysis found oral semaglutide cutting heart-failure events in people who had heart failure and doing nothing in those who did not. The test of whether that split is real came out at P = .06.

By Nora Bissett · Pricing Editor
Editorially reviewed & fact-checked against primary sources · How we verify contentLast reviewed
7 min read·1 citations

A secondary analysis of the SOUL trial found oral semaglutide cutting heart-failure events by 22% in people who already had heart failure, and doing nothing whatsoever in people who did not.[1] That looks like a clean split. The statistic that tests whether the split is real came out at P = .06 — and understanding what that means will change how you read every subgroup table you ever see.

The analysis

SOUL enrolled 9,650 people with type 2 diabetes plus atherosclerotic cardiovascular disease, chronic kidney disease, or both, across 444 centers in 33 countries, and followed them for an average of just under four years. Roughly a quarter — 2,229 people — had a history of heart failure at entry. The analysis asked whether daily oral semaglutide changed a composite of heart-failure hospitalization, urgent heart-failure visit, or cardiovascular death.

Composite heart-failure outcome with oral semaglutide against placebo, by baseline status.[[cite:1]]
GroupHazard ratio (95% CI)Reading
Heart failure at baseline0.78 (0.63–0.96)22% fewer events
No heart failure at baseline1.01 (0.84–1.20)No effect
P for interaction.06The groups are not shown to differ

It is worth knowing what the parent trial found before reading its subgroups. SOUL was an event-driven superiority trial in 9,650 people aged 50 or over with type 2 diabetes and either atherosclerotic cardiovascular disease, chronic kidney disease, or both, followed a median of just over four years.[2] Its primary outcome — cardiovascular death, non-fatal heart attack or non-fatal stroke — occurred in 12.0% on oral semaglutide against 13.8% on placebo, a hazard ratio of 0.86 (95% CI 0.77–0.96, P = 0.006).

Two details belong with that. The confirmatory secondary outcomes, including a five-point kidney composite, did not differ significantly between groups — so the trial’s positive result is the cardiovascular one and not a general sweep. And serious adverse events were slightly less common on the drug, 47.9% against 50.3%, in a population sick enough that roughly half of everyone had one.

What a P for interaction actually asks

Most people read a subgroup table by checking which rows are significant. That answers a different question from the one worth asking. Each row tests whether that group’s result differs from no effect. The interaction test asks whether the groups differ from each other.

Those come apart more often than people expect. One subgroup can land just inside significance and another just outside without the difference between them being distinguishable from chance. It is entirely possible for a table to show “works here, does not work there” while the honest summary is “we cannot tell whether it works differently in these two groups.”

At .06, this interaction test does not clear the conventional threshold. It is also close enough to it that dismissing the split entirely would be its own error — .06 is weak evidence of a difference, not evidence of no difference. The correct posture is genuine uncertainty, which is less satisfying than either headline.

Checking which rows are significant answers a different question from the one you meant to ask.

The split inside the split

Among the people who did have heart failure, the effect divided again along the type of heart failure — and here the contrast is stark.

Composite heart-failure outcome among participants with heart failure at baseline.[[cite:1]]
TypeHazard ratio (95% CI)
Preserved ejection fraction (HFpEF)0.59 (0.39–0.86)
Reduced ejection fraction (HFrEF)0.98 (0.70–1.38)

Essentially the entire benefit sits in preserved ejection fraction — the form where the heart muscle stiffens rather than weakens, which is strongly associated with obesity and for which treatment options have historically been thin.

This subgroup deserves more credence than most. Semaglutide has a dedicated randomized trial in preserved ejection fraction behind it, so this is a pre-existing hypothesis being checked rather than a pattern found by looking. That distinction — a prior expectation versus a discovery in the table — is the main thing separating a credible subgroup from a spurious one. It is still a subgroup.

What did not split

The trial’s main cardiovascular endpoint behaved completely differently. Major adverse cardiovascular events were reduced similarly with and without heart failure history — hazard ratios of 0.83 and 0.86, with an interaction P of .77, which is about as clean a statement of “no difference between groups” as these analyses produce.

That contrast is informative. The drug’s general cardiovascular benefit appears not to care whether you have heart failure. Only the heart-failure-specific endpoint showed any suggestion of differing, and that suggestion did not reach significance.

Serious adverse events were similar in the heart-failure group on drug and on placebo — 53.8% against 57.1% — which is the safety question anyone would ask about giving a new agent to people with established heart failure.

How to read the next subgroup table you see

  • Find the interaction P before reading the rows. Without it, a table of subgroups is a list of separate questions, not a comparison.
  • Ask whether the subgroup was expected. A group singled out because of prior evidence is far more believable than one that emerged from the data.
  • Notice which endpoints did not split. A drug that behaves the same everywhere on its main endpoint, and differently on one specific one, is telling you something narrower than the headline.
  • Treat P = .06 as uncertainty, not as either answer. It is neither a difference established nor a difference ruled out.

The companion problem — a subgroup with the biggest number and the least evidence behind it — is in the biggest number is the weakest.

Frequently Asked Questions

References

  1. 1.Pop-Busui R, Rasmussen S, Deanfield JE, et al. Oral Semaglutide and Heart Failure Outcomes in Persons With Type 2 Diabetes: A Secondary Analysis of the SOUL Randomized Clinical Trial JAMA Internal Medicine. 2026. PMID: 41627802.
  2. 2.McGuire DK, Marx N, Mulvagh SL, et al. Oral Semaglutide and Cardiovascular Outcomes in High-Risk Type 2 Diabetes The New England Journal of Medicine. 2025. PMID: 40162642.

Where to get GLP-1 online, safely: sellers our editors have checked

These are telehealth sellers our editors have checked. For each one we hold a price, the form the drug comes in, and the states it reaches.

No insurance needed · vetted by our editors

Some of the links on this page earn us money. If you sign up with a provider after following one, that provider may pay GLP Watchdog a commission. Learn more

7.0

MyDrHank

An oral route if you will not self-inject

6.0

SkinnyRx

Starting below a standard dose, with microdose tiers

8.3

SnagRx

Semaglutide at $99/month, 48% under the register median