[PSA] STEP is not what most of this community thinks it is
STEP is not what most of this community thinks it is — with the reasoning, because a rule without a reason gets ignored.
On means, which this board treats as targets and which are nothing of the sort.
A reported mean body weight change is the centre of a distribution that in these trials is very wide. Substantial numbers of participants did much better, and substantial numbers did considerably worse while remaining on the drug and in the analysis.
Quoting the mean as an expectation therefore misleads in both directions: it makes ordinary results look like failures and it makes exceptional results look normal. If a paper publishes the distribution — and several do, in the appendix — look at that instead. It is far more informative than the number in the abstract.
Argued for a week about a result and then read the limitations section, which conceded most of my opponent’s point.
How to read one of these papers in fifteen minutes, in the order that actually helps.
Start with the registered protocol and check the primary endpoint against what is reported. Then the methods: who was included, what the comparator was, how long the randomised phase ran. Then the discontinuation numbers, which are a tolerability result and are usually in a supplementary table.
Only then the efficacy figure, and read the interval rather than the point estimate. Finish with the limitations section, which is where the authors say what they actually think.
Fifteen minutes, and you will know more than any thread summarising it.
Please do not ask me what dose you should be on. I genuinely do not know and neither does anyone else here.
best — the order this archive was captured in
Why comparing across trials almost never works, with the specific failure modes.
Different populations: an obesity programme and a diabetes programme enrol different people with different baseline characteristics. Different endpoints: body weight change, glycaemic control and cardiovascular events are not convertible. Different durations: 68 weeks and 72 weeks are not the same, and the curves have not flattened by either.
Different analysis populations: one paper reports intention-to-treat, another emphasises completers. Different support: some trial designs include structured lifestyle contact that no member of this board receives.
Stack those and the "X beats Y" tables that circulate here are comparing five things at once and attributing the difference to the molecule.
Left up and flaired Trial Data. The publication is linked rather than the coverage, which is what we ask for.
open-label extensions are not the same evidence as the randomised phase
Left up and flaired Trial Data.
This is the distinction that would end about half the arguments on this board.
What was the comparator?
phase 2 finds a dose, phase 3 measures the effect
The registered protocol is public. Comparing the registered primary endpoint with the reported one is a two-minute check and it is how outcome switching gets caught.
Yes — the interval is the finding. A point estimate with a wide interval is a hypothesis in a nice font.
Careful with that mean. The distribution around it was wide enough that it describes very few individual participants.
read the endpoint before you read the headline
trial populations get support that nobody on this board gets
read the endpoint before you read the headline
Adding the check nobody runs — the registered protocol is public and takes two minutes to compare.
Adding the check nobody runs — the registered protocol is public and takes two minutes to compare.
nora_oyelaran is right about the programme names. They are different populations with different endpoints.
Went looking for the registered protocol to see whether the endpoint had changed. It had not, which was reassuring and worth checking.
- 1The registered protocol is public. Comparing the registered primary endpoint…8 comments in this branch · started by u/foam_head_fred