What a small effect is, and what the other group got
95 min
Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.
- Read a standardised effect size and say in plain terms what it means, using a conversion this course labels as its own
- Explain why the same treatment can look bigger against a waiting list than against a placebo therapy or another treatment
- Ask of any effect size what the comparison group got, and say why two effects measured against different comparisons cannot be ranked
The next five lessons are about methods people are offered for stress, worry and low mood, and the one after them is about medication. Every one of them comes with a number. This lesson is about how to read that number, because in this subject the number means very little until you know one more thing about it.
This course is education, not care. If you're thinking about suicide or self-harm, or don't feel able to keep yourself safe, contact emergency services (911 in the US and Canada, 999 in the UK, 112 across the EU, 000 in Australia) or a crisis line: call or text 988 in the US and Canada, call Samaritans on 116 123 in the UK and Ireland, or Lifeline on 13 11 14 in Australia. Elsewhere, findahelpline.com lists free, confidential lines by country.
Logic and Argument lesson 9 gave you a question for any number: compared with what? There it was about baselines and about who went missing from a sample. In treatment trials it has a sharper form, because the comparison group isn't a baseline sitting still. It's a condition people are put in, and what they're given, and told, changes how they do. So this course's version is: what did the comparison group get?
The band
This course's research file found these effects for methods offered for depression and anxiety, each with its comparison. They're standardised effect sizes, the Hedges' g that Note-Taking lesson 5 taught: a bigger number is a bigger difference between two groups' averages, measured in standard deviations.
| Method | Effect | Compared with | Source |
|---|---|---|---|
| Unguided self-help CBT, depression | g = 0.45 | not separated in the abstract | 1 |
| Self-guided internet CBT, depression | g = 0.27 | usual care, waiting list or attention control | 2 |
| Smartphone apps, depression | g = 0.28 | control conditions | 3 |
| Smartphone apps, generalised anxiety | g = 0.30 | control conditions | 3 |
| CBT with a therapist, depression | g = 0.79 | control conditions such as care as usual and waitlist | 1 |
CBT is cognitive behavioural therapy, the kind of talking therapy Sleep lesson 6 met as CBT-I. Every row was read at abstract level.
Four of the five sit between about 0.25 and 0.45. That band, roughly 0.2 to 0.6, is where almost every method in this course lands.7 The one number above it is full therapy with a therapist, measured against controls that include waiting lists. Both of those could push it up, and this lesson is about the second; how much of the 0.79 the comparison explains, the abstract doesn't say.
What 0.3 looks like
Sleep lesson 5 used a conversion you can use here too. If scores in both groups are spread in the bell shape statisticians call the normal curve, an effect of g puts the average person in the treated group ahead of a certain share of the comparison group.
| Effect size | Average treated person is ahead of about |
|---|---|
| 0.2 | 58 percent of the comparison group |
| 0.3 | 62 percent |
| 0.5 | 69 percent |
| 0.8 | 79 percent |
That conversion is this course's arithmetic on the normal curve, not a figure from any study, and it only holds if the bell-shape assumption does.7 With no effect at all, the figure would be 50 percent.
One study in the table gives a second way to feel the size. For self-guided internet CBT, the authors report a number needed to treat of 8.2 Logic and Argument lesson 9 taught the idea: roughly, for every eight people given the programme rather than the comparison, one more responded than otherwise would have.
Before you read on. An effect of 0.3 puts the average treated person ahead of about 62 percent of the comparison group. Does that mean the method works for 62 percent of people?
Show the answer
No. It's a statement about two groups' averages, not about how many people it worked for. Some people in the treated group did worse than some in the comparison group, and some in the comparison group got better with nothing at all.
It also says nothing about you. Whether a method helps one person is a question an average can't answer.
What the other group got
Every effect size is a gap between a treated group and a comparison group. So its size depends on two things: how much the treated group improved, and how much the comparison group did.
Trials in this subject use several kinds of comparison group. From easiest to beat to hardest, as this course orders them:7
- A waiting list. People are told they'll get the treatment later, and get nothing meanwhile.
- No treatment. People carry on as they would have, without being told they're waiting for anything.
- Usual care. Whatever care people would ordinarily get, which may include a doctor or medication.
- A placebo therapy, or attention control. Time and contact with a professional, with the treatment's active ingredients left out.
- Another active treatment. A different therapy, or a medicine.
Two groups sign up for a stress course. One is told "you'll start in ten weeks"; the other isn't enrolled in anything and just carries on. Which would you expect to improve more by week ten?
Show the answer
Most people guess no difference: neither group got the course. The research this lesson reads found something else. In one analysis of trials of CBT for depression, the people told to wait did worse than the people who weren't waiting for anything.
Why that might be is a proposal, not a finding, and the next section gives it in the authors' words.
A 2014 network meta-analysis by Furukawa and colleagues took 49 trials of CBT for depression, with 2,730 participants, and compared the comparison conditions with each other. A network meta-analysis can do that because the trials share conditions: one compares CBT with a waiting list, another CBT with no treatment, and the shared arm links them.
Their finding: "the effect size estimates for CBT were substantively different depending on the control condition. The odds ratio of response for NT over WL was statistically significant at 2.9 (95% CI: 1.3-5.7)."4 NT is no treatment, WL is waiting list, and "response" means improving by an amount a trial set in advance.
In plain terms, in this network the odds of responding were about three times higher with no treatment than on a waiting list. Odds aren't the same as chances, so this isn't "three times as many people", and the interval is wide, running from 1.3 to 5.7. That gloss is this course's.7
Before you read on. If people on a waiting list do worse than people given nothing, what happens to a treatment's effect size when it's measured against a waiting list?
Show the answer
It can look bigger. The effect size is the gap between the groups, and a comparison group that does unusually badly widens the gap without the treatment doing anything more.
The paper's title offers a reason, that a waiting list "may be a nocebo condition". This course reads that as: being told to wait may itself hold improvement back. That's the authors' proposal with its "may", in this course's words, and this course read the abstract, not the analysis behind it.47
The authors are careful about their own finding, and the care travels with it: "the quality of evidence, including publication bias, was less than ideal and none of the preplanned sensitivity analyses limiting to high-quality studies could be conducted, while findings of significant differences did not persist in post hoc sensitivity analyses trying to adjust for publication bias."4 In plain terms: they meant to rerun the analysis on only the best-run trials and had too few to do it, and when they tried to correct for unpublished studies, the difference between waiting list and no treatment went away.
So treat it as a warning, not a law: a number measured against a waiting list might be flattered, and you should find out before you compare it with anything else. The authors' conclusion is correspondingly modest: "There may be important differences in control conditions currently used in psychotherapy trials."4
The same pattern turns up in anxiety. A 2014 meta-analysis of twelve trials in generalised anxiety found that "Provision of a psychological placebo was associated with a significantly greater reduction of symptoms than placement on a waiting list", with its own qualifier: "the overall level of evidence was classified as 'moderate', indicating that further research could change the overall results of the meta-analysis."5
One thing to keep separate. A trial's waiting list isn't the "active monitoring" that lesson 1's guideline recommends at step 1. A waiting list is being told you'll get something later, with nothing meanwhile. Active monitoring is education and a clinician checking back. That distinction is this course's; Furukawa's abstract doesn't discuss monitoring.7
One therapy, three comparisons
Cuijpers and colleagues, 2023, is in its authors' words "the largest meta-analysis ever of a specific type of psychotherapy for a mental disorder": 409 trials and 52,702 patients.1 This course read its abstract.
Against control conditions: "CBT had moderate to large effects compared to control conditions such as care as usual and waitlist (g=0.79; 95% CI: 0.70-0.89), which remained similar in sensitivity analyses and were still significant at 6-12 month follow-up."
Against medication: "The effects of CBT did not differ significantly from those of pharmacotherapies at the short term, but were significantly larger at 6-12 month follow-up (g=0.34; 95% CI: 0.09-0.58), although the number of trials was small, and the difference was not significant in all sensitivity analyses."
Against other talking therapies, the difference was 0.06, with an interval from 0 to 0.12.1
Before reading the authors' next clause about that 0.06: using what Sleep lesson 5 taught about intervals, what does an interval running from 0 to 0.12 already tell you?
Show the answer
That the difference sits right at the edge of nothing. An interval that touches zero can't rule out no difference at all, and even its top end is small.
The authors' clause confirms it: CBT "was significantly more effective than other psychotherapies, but the difference was small (g=0.06; 95% CI: 0-0.12) and became non-significant in most sensitivity analyses."1
Same therapy, same meta-analysis, three comparisons. Against controls that include waiting lists and usual care, 0.79. Against medication, no significant difference in the short term and a difference in CBT's favour at six to twelve months, on few trials. Against other talking therapies, 0.06, which didn't survive most of the authors' re-checks.
The trap is setting 0.79 beside a medicine's effect against a placebo pill, which lesson 8 will show is about 0.3. That would compare a therapy against waiting lists and usual care with a pill against a dummy pill, and it would make the therapy look far stronger. The fair comparison in this abstract is the one where the two meet head to head.17
The anxiety literature has the same trap, and there the authors name it themselves. Carl and colleagues, in 79 trials of generalised anxiety, report that "Psychotherapy showed a medium to large effect size (g = 0.76) and medication showed a small effect size (g = 0.38)", and in the same abstract: "Because medication studies had more placebo control conditions than inactive conditions compared to psychotherapy studies, effect sizes between the domains should not be compared directly."6
Nothing in this lesson is a reason to start, stop or switch a treatment. NICE's advice to anyone on an antidepressant who wants to stop is to "talk with the person who prescribed their medication", because "it is usually necessary to reduce the dose in stages over time".8
The app that's "as effective as therapy"
The app and the advert are invented for this lesson.7
An app advertises itself as "clinically proven, as effective as therapy". You find its study: the app against a waiting list, an effect of 0.3.
Run the question on this advert yourself. What would "as effective as therapy" need the comparison group to have got, and what can this study actually support?
Show the answer
Therapy. A claim about matching therapy needs therapy in the comparison, and there wasn't any. So the study can support "better than waiting, in that trial", and its comparison may flatter how much better. It can't support "as good as therapy".
The real evidence on apps, for scale. A 2019 meta-analysis found that "Smartphone interventions significantly outperformed control conditions in improving depressive (g=0.28, n=54) and generalized anxiety (g=0.30, n=39) symptoms", and that these effects were "robust even after adjusting for various possible biasing factors (type of control condition, risk of bias rating)".3 "n" counts trials there, not people. The same abstract found "no significant benefit over control conditions on panic symptoms", and on the direct question: "Smartphone interventions did not differ significantly from active interventions (face-to-face, computerized treatment), although the number of studies was low (n≤13)."3
So apps have real, modest evidence against control conditions that held up to a check on the type of control. The little head-to-head evidence found no significant difference, in too few studies to show the two are equal: "no significant difference" isn't the same as "the same".
The course's question
Four Core courses before this one each added a question to the institute's way of reading a claim. Memory added the sample, Focus and Deep Work the instrument, Note-Taking the setting, and Sleep the clock. This course adds what did the comparison group get? Every method lesson from here gives its number and its comparison together, and lesson 9 joins the question to the rest.
Three things people get wrong
"A bigger effect size means a better treatment." Only if the two were compared with the same thing. CBT's 0.79 against controls and its no-significant-difference against medication in the short term are the same therapy.
"An effect of 0.3 means 30 percent better." It's a difference in standard deviations. On this course's conversion, the average treated person is ahead of about 62 percent of the comparison group.
"If it beat the control group, it'll work for me." An average over a group says nothing certain about one person, and the comparison group sets how big the average looks.
Practice
Take 20 minutes.
Here are six results from this lesson, each with what its comparison group got:
- CBT, depression, 0.79: control conditions including waiting lists and usual care.
- CBT against medication, depression, short term: no significant difference.
- CBT against other talking therapies, depression: 0.06.
- Psychotherapy, generalised anxiety, 0.76: mostly inactive conditions, on the authors' account.
- Medication, generalised anxiety, 0.38: more often placebo, on the authors' account.
- A psychological placebo against a waiting list, generalised anxiety: a significant difference.
Rank them once by size, where there is a size. Then sort them by how hard their comparison was to beat, using the list of five comparison types. Where a comparison is a mix or isn't stated, mark it unrankable; that's a finding too.
Write one sentence on what changed between the two orderings.
Take 20 minutes.
Find one app, programme or course that says it helps with stress, anxiety or mood and claims evidence. This one counts.
- Write down its claim, word for word.
- Find its "science" or "research" page and follow it to the study's summary. Search that for "waitlist", "wait-list", "usual care", "control" or "active control".
- Write one sentence: what the claim would need the comparison group to have got, and whether it did.
If you can't find what the comparison group got, that's the finding.
Connections
Back. Logic and Argument lesson 9 taught "compared with what?" and the number needed to treat. Note-Taking lesson 5 taught Hedges' g; Sleep lesson 5 used the normal-curve conversion and read an interval against zero, and Sleep lesson 6 noticed that the CBT-I trials weren't head to head.
Forward. Lesson 3 is self-help CBT, where what the comparison group got matters a great deal.
Go deeper
- Cuijpers and colleagues, 2023, free at PubMed Central. This course read the abstract. The clearest single place to see one therapy give different numbers against different comparisons.
- Furukawa and colleagues, 2014. Abstract only here. Read the qualifier as carefully as the finding.
Sources
- P. Cuijpers and colleagues, "Cognitive behavior therapy vs. control conditions, other psychotherapies, pharmacotherapies and combined treatment for depression", World Psychiatry 22(1), 2023, pp. 105 to 115, doi 10.1002/wps.21069. Read: the abstract. Supports: the scale, 0.79 against controls and its follow-up, the comparison with medication short and longer term, 0.06 against other therapies, and 0.45 for unguided self-help, whose comparison the abstract doesn't separate.
- E. Karyotaki and colleagues, "Efficacy of Self-guided Internet-Based Cognitive Behavioral Therapy in the Treatment of Depressive Symptoms", JAMA Psychiatry 74(4), 2017, doi 10.1001/jamapsychiatry.2017.0044. Read: the abstract. Supports: g = 0.27, the comparisons, and the number needed to treat of 8.
- J. Linardon and colleagues, "The efficacy of app-supported smartphone interventions for mental health problems", World Psychiatry 18(3), 2019, doi 10.1002/wps.20673. Read: the abstract. Supports: the app figures, their robustness to the type of control, the panic result, the comparison with active interventions, and that "n" counts trials.
- T. A. Furukawa and colleagues, "Waiting list may be a nocebo condition in psychotherapy trials", Acta Psychiatrica Scandinavica 130(3), 2014, doi 10.1111/acps.12275. Read: the abstract. Supports: the 49 trials of CBT for depression, the odds ratio of 2.9 and its interval, the authors' qualifier and conclusion, and the word "nocebo" from the title.
- Z. Zhu and colleagues, "Comparison of psychological placebo and waiting list control conditions in the assessment of cognitive behavioral therapy for the treatment of generalized anxiety disorder", Shanghai Archives of Psychiatry 26(6), 2014, doi 10.11919/j.issn.1002-0829.214173. Read: the abstract. Supports: the placebo-against-waiting-list finding and its qualifier.
- E. Carl and colleagues, "Psychological and pharmacological treatments for generalized anxiety disorder (GAD): a meta-analysis of randomized controlled trials", Cognitive Behaviour Therapy 49(1), 2020, doi 10.1080/16506073.2018.1560358. Read: the abstract. Supports: 0.76 and 0.38, and the authors' warning not to compare them.
- This course's own constructions, labelled where they appear. The normal-curve table is this course's arithmetic. The "band" is this course's summary of the research file. The ordering of comparison types, the plain-language gloss of the odds ratio, the reading of "nocebo", and the distinction between a waiting list and active monitoring are this course's. The app and its advert are invented.
- National Institute for Health and Care Excellence, Depression in adults, NG222, 2022, recommendation 1.4.12. Read in full. Supports: the advice on stopping an antidepressant.
Check your understanding
This lesson has a 6-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.