Antidepressants: one number, two readings, and a hypothesis

115 min

Listen: this lesson as a conversation

Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.

In this lesson you will learn to
  • State the drug-placebo difference that Kirsch's and Cipriani's reviews report, and what each paper concludes from it, in its own words
  • Explain what "the average hides a responder group" claims, and what evidence would test it
  • Describe the serotonin dispute as a disagreement about two claims, cause and association, and explain why neither settles whether a drug helps

Two headlines about antidepressants have run for years, and they seem to contradict each other. One says the pills barely beat a dummy pill. The other says they are an established treatment that works. Then in 2022 a third arrived, saying that depression isn't caused by low serotonin after all, and it's easy to hear that as saying the pills were built on nothing. This lesson reads the papers behind all three, and the surprising part is how much the two sides of the first dispute agree on.

If you're struggling right now

This course is education, not care. If you're thinking about suicide or self-harm, or don't feel able to keep yourself safe, contact emergency services (911 in the US and Canada, 999 in the UK, 112 across the EU, 000 in Australia) or a crisis line: call or text 988 in the US and Canada, call Samaritans on 116 123 in the UK and Ireland, or Lifeline on 13 11 14 in Australia. Elsewhere, findahelpline.com lists free, confidential lines by country.

One thing first, because some readers of this lesson take an antidepressant. Nothing here is advice about yours, in either direction. The lesson describes a dispute among researchers and reports what the guidelines say, and it ends with the one recommendation NICE gives on stopping. Whether a medicine is right for a particular person is a conversation with the person who prescribed it.

The number both sides print

The critique: Kirsch and colleagues, 2008

Kirsch and colleagues wrote the best-known critique of antidepressants (this course read the abstract and the parts of the full text listed in the Sources).1 Their starting point was a worry about what gets published. The abstract opens by saying that earlier meta-analyses had found "only modest benefits over placebo treatment, and when unpublished trial data are included, the benefit falls below accepted criteria for clinical significance."1

So they used the trials submitted to the US Food and Drug Administration, the FDA, for licensing "the four new-generation antidepressants for which full datasets were available", a dataset that in the abstract's words takes in "published and unpublished clinical trials". It came to 35 trials with 5,133 patients.1 Depression was measured on the Hamilton Rating Scale for Depression, which Kirsch abbreviates HRSD. This course hasn't recorded the scale's full range, but for a sense of it, the trials in the larger review below that reported a starting score averaged 25.7 on its 17-item version.2 The result:

"weighted mean improvement was 9.60 points on the HRSD in the drug groups and 7.80 in the placebo groups, yielding a mean drug–placebo difference of 1.80 on HRSD improvement scores. Although the difference between these means easily attained statistical significance (Table 2, Model 3a), it does not meet the three-point drug–placebo criterion for clinical significance used by NICE."1

And as a standardised effect, the kind of number lesson 2 taught you to read: "Thus, the difference between improvement in the drug groups and improvement in the placebo groups was 0.32, which falls below the 0.50 standardized mean difference criterion that NICE suggested."1

Notice the order of those two sentences in the first quotation. Kirsch does not say the drugs failed to beat placebo. He says they beat it, "easily", and then that the margin falls short of a line. The line was NICE's at the time: Kirsch attributes it to NICE's 2004 depression guideline, and this course has not checked whether the current guideline, NG222, uses any such threshold.17

Predict first

The review most often quoted in reply to Kirsch came ten years later and pooled 522 trials. Before you read on: do you expect it to have found an effect much bigger than Kirsch's, about the same, or smaller?

Show the answer

About the same, roughly: 0.30. The next section has it, and why "roughly" matters.

The larger review: Cipriani and colleagues, 2018

Cipriani and colleagues pooled 522 trials with 116,477 participants, covering 21 antidepressants for adults with major depressive disorder, the clinical diagnosis of depression (this course read the abstract and the parts of the full text listed in the Sources).2 It's a network meta-analysis, which uses trials comparing drugs with each other as well as with placebo. Its stated aim was to "compare and rank antidepressants", and every one of the 21 beat placebo.2

For the effect on symptom scores, as opposed to the share of people who responded: "The random-effects summary SMD for all antidepressants was 0·30".2 Its 95% credible interval ran from 0.26 to 0.34. The trials lasted a median of eight weeks.2

So the critique found 0.32 and the larger review found 0.30, on different data (35 FDA trials of four drugs against 522 trials of 21). One caution before you set them side by side. A standardised effect depends on which spread the points are divided by, and the two papers didn't divide by the same one. Kirsch standardised each group's change on its own, "one in which each group's change was represented as a standardized mean difference (d), which divides change by the standard deviation of the change score", and got 1.24 for the drug groups and 0.92 for placebo; the 0.32 is the gap between those two.1 Cipriani's is a standardised difference between the groups, "SMD, Cohen's d".2 So the two figures describe the same size of effect only roughly. The points are the steadier comparison, and a 2022 analysis you'll meet below, with Kirsch among its authors, found 1.82 points for adults, against his 1.80.3

Kirsch and colleagues, 2008 Cipriani and colleagues, 2018
Trials 35, submitted to the FDA 522, from journals, regulators and registers
Drugs 4 21
Compared with placebo placebo, and each other
Average drug-placebo difference 0.32 (1.80 points on the HRSD), computed within each group 0.30 (interval 0.26 to 0.34), between groups

Whatever the measure, neither paper is arguing about whether the average is small. It is. The dispute is about what a small average means. Lesson 2 gave you a way to feel 0.3: on this course's arithmetic on the normal curve, not a figure from either paper, the average person on the drug ends up ahead of about 62 percent of the placebo group, where no effect at all would be 50 percent.7

What the placebo group got

Lesson 2's question, what did the comparison group get, has a clear answer here: a dummy pill, usually in a trial where neither patient nor rater knew who had which. That is the hardest comparison short of another treatment, which is why lesson 2 warned against setting a therapy's 0.79, measured against controls such as waiting lists and usual care, beside a drug's 0.3 against placebo.

And look again at Kirsch's two groups. The drug groups improved by 9.60 points and the placebo groups by 7.80.1 Most of the improvement happened in both. The 2022 analysis says the same in its conclusions: "Patients with depression are likely to improve substantially from acute treatment of their depression with drug or placebo."3 So the 0.3 is the part the trials credit to the drug, over and above everything else that came with being in a trial. Whether that credit is too small or too large is itself argued over, and both arguments are below: Cipriani's team say dropouts can make it an underestimate, and Stone's team say they can't fully exclude "functional unblinding", which would make it an overestimate.23

Two readings of the same 0.3

Both camps accept the number. The question is what it means, and the two readings are best heard in their authors' own words, one after the other.

Too small to matter clinically

Kirsch and colleagues read the 0.32 against a threshold for clinical significance. This course's gloss on the idea is that a difference can be real, in the sense that it isn't chance, and still be too small to matter to a patient; a threshold is an attempt to say where mattering starts.7 On their data the drug-placebo difference fell short of it.

They then asked whether it depended on how depressed people were at the start, and found that it did, but not in the way you might expect. Their conclusion, whole: "Drug-placebo differences in antidepressant efficacy increase as a function of baseline severity, but are relatively small even for severely depressed patients. The relationship between initial severity and antidepressant efficacy is attributable to decreased responsiveness to placebo among very severely depressed patients, rather than to increased responsiveness to medication."1 The gap widened in the most severe cases because the placebo group improved less, not because the drug group improved more. In the abstract's account, the difference reached the conventional criterion only for patients at the upper end of the very severely depressed category.1

From that they drew a conclusion about prescribing: "Given these data, there seems little evidence to support the prescription of antidepressant medication to any but the most severely depressed patients, unless alternative treatments have failed to provide benefit."1 Read the whole sentence. It keeps a place for the drugs in the most severe depression and where other treatments have not helped. And it is a claim about prescribing policy across a population of patients, not about what any one person should do.

Real, reliable and modest

Cipriani and colleagues report the same size of effect as a finding that holds. Every one of 21 drugs beat placebo, across 522 trials and more than a hundred thousand people. Their Discussion: "We found that all antidepressants included in the meta-analysis were more efficacious than placebo in adults with major depressive disorder and the summary effect sizes were mostly modest."2 And they say what the review is for: "These results should serve evidence-based practice and inform patients, physicians, guideline developers, and policy makers on the relative merits of the different antidepressants."2

They also argue that trials like these may understate the drugs. People on the drug in a placebo-controlled trial may drop out early, suspecting they have placebo, and their early scores are carried to the end: "The final result can be an underestimate of the true efficacy of the active drug."2 And they put their figure beside earlier work: "The estimates of treatment effect from our study are in line with previous reviews on the same matter", they write, but "considerably more precise because of our larger quantity of data and resulting statistical power."2

"Modest" is their word, and they don't hide the argument. Their introduction names it: "However, there is a long-lasting debate and concern about their efficacy and effectiveness, because short-term benefits are, on average, modest; and because long-term balance of benefits and harms is often understudied."2 They grade their own evidence too: "46 (9%) of 522 trials were rated as high risk of bias, 380 (73%) trials as moderate, and 96 (18%) as low; and the certainty of evidence was moderate to very low."2 And they say what the review didn't cover, including specific adverse events and withdrawal symptoms.2

In the parts this course read, Cipriani and colleagues don't set their figure against a threshold for clinical significance in either direction. This course hasn't read a primary statement of why a modest, consistent effect like this one matters clinically, and it won't write one for the other side in its own words.27

Where the line comes from

This is the part of the dispute that's easiest to miss. Kirsch's case turns on a threshold for its reading of the average: three points on the HRSD, or 0.50 as a standardised effect. On this course's reading, that threshold isn't something a trial measures. It is a judgement about how big a difference has to be before it matters to a person, made in NICE's 2004 guideline, and a different committee could draw it somewhere else without moving the 0.3.17 His severity finding and his point about placebo response don't depend on where the line sits.

The 2022 analysis makes a similar point about a different line, the convention that a patient has "responded" if their score falls by half: "These threshold definitions, although useful, are arbitrary."3 That doesn't make thresholds useless. It means that "below the threshold" is a finding about the data plus a judgement. This course hasn't read a defender's argument against the 2004 criterion itself. In its sort, the size of the average effect is established, and what it means clinically is contested.7

Sleep lesson 6 split a guideline grade into two parts, the strength of a recommendation and the quality of the evidence behind it, and showed that neither measures how big an effect is. Cipriani's "moderate to very low" is the second kind. It grades how sure anyone can be about the estimate, not how large the benefit is.

Check yourself

Suppose a future committee set the line for a meaningful difference at 0.25 instead of 0.50. What would change about the evidence, and what would change about Kirsch's reading?

Show the answer

Nothing about the evidence. The trials would still show about 0.3.

His sentence that the difference "falls below" the criterion would no longer be true of 0.32, while his severity finding, that the gap grows because placebo groups improve less, would stand as it is. That's the sense in which the threshold is a judgement and not a finding: moving it changes a verdict without changing a single number. It works the other way too. A higher line would widen the same verdict on the same data.

What an average can hide

Both readings so far argue about the average. There is another possibility, and a 2022 paper tested it.

Stone and colleagues analysed individual participants' data from "232 randomized, double blind, placebo controlled trials of drug monotherapy for major depressive disorder submitted by drug developers to the FDA between 1979 and 2016, comprising 73 388 adult and child participants" (this course read the abstract, Introduction, Discussion and Conclusions, and the Results paragraphs on the model, not the Methods).3 Three of the authors were at the FDA, and the paper says it does not represent the agency's views. The last author is Kirsch, and on this course's reading that isn't a change of view about the 2008 data: the paper asks a different question of more data.37

Why individual data? A review that pools trial averages sees only each trial's mean. With every person's score you can see the shape of the spread. The authors' introduction: "Lack of knowledge about the distributions of individual responses has hampered discussions of the clinical significance of mean effects." And the question they asked: "The drug effect might not be a uniform small, and hence clinically unimportant, benefit across patients (ie, a shift in distribution mean without a change in the shape of the distribution). Rather, it could occur as a large, and thus clinically important, difference for a small subpopulation (ie, a difference in response distribution composition)."3

A small invented example, this course's own, shows why that's possible.7 Take 100 people on a drug and 100 on placebo. World one: every person on the drug does 1.8 points better than they would have on placebo. World two: 15 of the 100 do 12 points better and the other 85 do no better at all. Fifteen times twelve is 180, spread over 100 people, so the average advantage is 1.8 points in both worlds. Real response is messier than either.

This is where the two ideas in this lesson meet, on this course's reading.7 NICE's three points was a line for a change a person would notice, and Kirsch applied it to a group average. An average of 1.8 can come from a world in which some people are well past 3 points, or, as in world one, from a world in which nobody is. The average alone can't say which.

Predict first

Stone's team modelled everyone's improvement, on drug and on placebo, to see which world the data looked like. Before you read on: do you expect one smooth spread of improvements, or separate groups?

Show the answer

Separate groups: three of them. The next paragraph has the figures.

Their average first, which you'll recognise: "The random effects mean difference between drug and placebo favored drug (1.75 points, 95% confidence interval 1.63 to 1.86)." For adults alone it was 1.82 points, which they standardise as 0.24.3

Then the shape. In plain terms, they fitted overlapping bell curves to all the improvement scores and let the curves' sizes differ between drug and placebo. The best fit was "a combination of three overlapping normal distributions allowed to vary in relative size between drug and placebo", with mean improvements of 16.0, 8.9 and 1.7 points.3 The authors named the three Large, Non-specific and Minimal. "Those treated with drug were more likely to show a Large response (24.5% v 9.6% with placebo), however, and less likely to have a Minimal response (12.2% v 21.5%)."3 For the broad middle group, the authors suggest "these responses might reflect the diverse interactions of individual characteristics with placebo and other effects not related to drug treatment, such as response to increased clinical contact, spontaneous improvements, and regression toward the mean."3 Regression toward the mean is the tendency of anyone enrolled at a bad moment to be measured nearer their usual level later.

Work the first comparison through. On drug, 24.5% had a Large response; on placebo, 9.6% did. That's 14.9 percentage points, by this course's subtraction, and the authors' conclusion rounds it and states it as a suggestion: "The trimodal response distributions suggests that about 15% of participants have a substantial antidepressant effect beyond a placebo effect in clinical trials".37

Check yourself

Your turn with the second comparison. What share had a Minimal response on drug and on placebo, and what's the difference?

Show the answer

12.2% on drug against 21.5% on placebo, so 9.3 percentage points fewer people had a Minimal response on the drug.

Each arm adds up to 100, and the middle group was about two thirds in both, slightly smaller on the drug: 63.3% against 68.9%.3 So the drug arm's extra 14.9 points at the top came from 9.3 fewer at the bottom and 5.6 fewer in the middle. The authors' own summary: "Thus the observed advantage of antidepressants over placebo is best understood as affecting a minority of patients as either an increase in the likelihood of a Large response or a decrease in the likelihood of a Minimal response."3

What each side can take from it, and what it doesn't settle

This paper is the bridge between the two readings, and it's easy to use it as a trophy for either side.

A critic can point out that the average is small, that it matches the 2008 figure in points, and that most people's improvement fell in the same middle group on drug as on placebo. A defender can point out that, on the model, about one person in seven more landed in the Large group on the drug than on placebo, which is not a trivial thing to happen to someone. Both statements come from the same paper, and its conclusions hold them together: "Although the mean effect of antidepressants is only a small improvement over placebo, the effect of active drug seems to increase the probability that any patient will benefit substantially from treatment by about 15%."3 This is also why "antidepressants barely beat placebo" and "antidepressants work well for some people" aren't contradictory sentences.

The authors are careful about limits. The groups are inferred from the shape of the spread; nobody was observed to be in one, and no individual's result says what they'd have done on the other pill. The trials' patients "are thus likely to have had less clinically complex but more acutely severe depression than is typically seen in the community."3 And: "We cannot fully exclude the possibility that the effects of the drugs are accounted for by functional unblinding", meaning that participants or raters may have been able to tell who was on the drug, which could colour what was reported.3 The authors set out why they think that possibility is unlikely to explain the pattern, and don't claim to have ruled it out.

What would settle the argument is what the paper asks for: "Further research is needed to identify the subset of patients who are likely to require antidepressants for substantial improvement."3 Until someone can predict who will be in the Large group, nobody can say in advance whether a given person is one of them. This course can't tell you either. The paper's conclusion puts the weighing with the person: "Because the benefits and risks might be categorically different (eg, reduced sadness v anorgasmia), weighting should be done at the individual level, jointly by patients and their care providers."3

What the guidelines do with it

NICE's depression guideline, NG222, is where a UK clinician looks, and its positions are reported here as NICE's, from the parts named in the Sources.6

For less severe depression, recommendation 1.5.3: "Do not routinely offer antidepressant medication as first-line treatment for less severe depression. Only offer it if that is the person's informed preference. [2022]"6 In Table 1, the options for less severe depression, SSRIs are ninth of eleven, and as in lessons 1 and 6, the table's heading says the order weighs cost and "consideration of implementation factors" as well as clinical effectiveness, so a place in it isn't an effect size. SSRIs are selective serotonin reuptake inhibitors, the family whose name refers to serotonin.6

For more severe depression, Table 2's first row is "Combination of individual cognitive behavioural therapy (CBT) and an antidepressant", and antidepressant medication alone is the fourth row.6

Check yourself

Is NICE's guideline siding with the critique or with the defence?

Show the answer

Neither, on this course's reading of it. For less severe depression it does not routinely offer the drugs first, which a critic would welcome; for more severe depression it lists a drug with CBT first, which a defender would welcome.

A guideline is a committee's judgement that folds in the evidence, costs, what can be provided and what patients prefer. It is not a referee's decision on the research dispute, and it's shaped by severity, which is also the axis Kirsch's severity analysis ran along.

A hypothesis, at two strengths

The third headline is about how the drugs were supposed to work. For decades a popular explanation of depression has been a shortage of serotonin, a chemical messenger in the brain, sometimes told as a "chemical imbalance" that the pills correct. In 2022 a review said the evidence doesn't support it, and a large group of researchers replied that its conclusion was overstated. Read closely, they agree on less than the headlines suggest in one place and more in another.

The review: Moncrieff and colleagues

Moncrieff and colleagues ran an umbrella review, meaning a review of reviews, across the main areas of serotonin research (this course read the abstract, the Introduction's first two paragraphs and the Discussion).4 Its central finding: "The main areas of serotonin research provide no consistent evidence of there being an association between serotonin and depression, and no support for the hypothesis that depression is caused by lowered serotonin activity or concentrations."4 The Discussion opens the same way: "Our comprehensive review of the major strands of research on serotonin shows there is no convincing evidence that depression is associated with, or caused by, lower serotonin concentrations or activity."4 Notice that it rejects two things, a cause and an association.

The evidence ran across several kinds of study. One kind lowers serotonin on purpose, by depleting tryptophan, an amino acid the body makes serotonin from. In the review's summary such "methods to reduce serotonin availability using tryptophan depletion do not consistently lower mood in volunteers", and its abstract gives the exception: "One meta-analysis of tryptophan depletion studies found no effect in most healthy volunteers (n = 566), but weak evidence of an effect in those with a family history of depression (n = 75)."4 The genetic studies were the largest. The two biggest studies of the gene for the serotonin transporter, a protein involved in the "reuptake" that SSRIs are named for (this course's gloss), with 115,257 and 43,165 people, "revealed no evidence of an association with depression".47 Their conclusion: "We suggest it is time to acknowledge that the serotonin theory of depression is not empirically substantiated."4

Why they think it matters is part of their case. "The general public widely believes that depression has been convincingly demonstrated to be the result of serotonin or other chemical abnormalities", and in their account this belief shapes how people understand their moods, "leading to a pessimistic outlook on the outcome of depression and negative expectancies about the possibility of self-regulation of mood".4 They grade their own sources frankly, AMSTAR-2 being a checklist for rating the quality of a review: "Most of the included studies were rated as low quality on the AMSTAR-2, but the GRADE approach suggested some findings were reasonably robust."4

The response: Jauhar and colleagues

The next year, Jauhar and 34 co-authors replied in the same journal, under the title "A leaky umbrella has little value" (this course read it in full).5 Their charge, from its abstract: "We present reasons for why this conclusion is overstated, including methodological weaknesses in the review process, selective reporting of data, over-simplification, and errors in the interpretation of neuropsychopharmacological findings."5

One of their examples is tryptophan depletion, where they say the review left out the figure for people who were actually depressed and not taking antidepressants. A negative figure here means mood went down: "the effect size for the effects of tryptophan depletion on mood in depressed people not taking antidepressants, from 8 samples, was large (Hedge's g = −1.9 (95% CIs −3.02 to −0.78). Admittedly, a number of these studies came from the same group, and confidence intervals were wide (removal of a potential outlier decreased Hedge's g to −1.06 (95% CIs −1.83 to −0.29))."5 That "Admittedly" is their own concession. Without the outlier the estimate is still large, past the top of lesson 2's table (by the same conversion, about 86 percent), though its interval reaches down into small.7 They also cite evidence on tryptophan in the blood, which they describe as "circulating tryptophan concentrations, which directly influence central serotonin synthesis": "L-tryptophan plasma concentrations show, after adjusting for publication bias, significant decrease in people with MDD (Hedge’s g = −0.45, 95% CIs, −0.66 to −0.23), with a large effect size of g = −0.84 (95% CIs −1.27 to −0.4) in unmedicated people".5 MDD is major depressive disorder.

And they dispute the review's reading of brain imaging of the serotonin transporter: "However, findings in a number of brain regions are consistent, with SERT reductions reported in people with MDD in all reviews."5 The review's abstract had called that evidence "weak and inconsistent" and added that "effects of prior antidepressant use were not reliably excluded."4

Their conclusion, whole: "A more accurate, constructive conclusion would be that acute tryptophan depletion and decreased plasma tryptophan in depression indicate a role for 5-HT in those vulnerable to or suffering from depression, and that molecular imaging suggests the system is perturbed. The proven efficacy of SSRIs in a proportion of people with depression lends credibility to this position."5 5-HT is serotonin's chemical abbreviation.

Both papers declare interests, and a reader should know both sets. Among the review's authors, one co-founded a company to help people stop antidepressants, another receives royalties from a book titled Evidence-biased Antidepressant Prescription, and the first author receives royalties for books about psychiatric drugs and co-chairs the Critical Psychiatry Network.4 Among the response's 35 authors, many declare pharmaceutical honoraria, consultancies or grants, and one is a part-time employee and shareholder of the drug company Lundbeck.5 This course recorded interests for these two papers and for the 2022 analysis, where one author lists consulting roles, and didn't read the statements for Kirsch's 2008 paper or Cipriani's.3

Two claims, not one

Now put the two conclusions side by side and look at the verbs. The review rejects two claims. One is causal: that depression is "caused by lowered serotonin activity or concentrations". The other is an association: it finds "no consistent evidence of there being an association between serotonin and depression".4 The response defends "a role" for serotonin and a system that is "perturbed".5

On the causal claim, the two can both be true, on this course's reading: a system can be involved in an illness without a shortage in it being the cause.7 On the association they flatly disagree. The response's tryptophan and imaging evidence is offered as exactly the link the review says isn't consistently there, and the review reads the same studies as weak, small or confounded by past drug use.45

So the dispute looks like a flat contradiction in the headlines and is a narrower, still real, disagreement in the papers. The response also says the review's methods were flawed and its reporting selective. On the transporter imaging question, what would settle more of it in the response's view is "a conventional umbrella review, where included studies are extracted and meta-analysed".5 This course classifies the question as contested and reaches no verdict on it. What it can say is that neither paper's conclusion claims that depression is a serotonin shortage which the drugs top up, which is the version most readers had heard.457

Check yourself

Three findings. Does each bear on the causal claim, the association claim, or neither? (a) "People with depression have normal serotonin levels on average." (b) "Lowering serotonin worsens mood in some people with a family history of depression." (c) "An SSRI beats placebo in a trial."

Show the answer

(a) Counts against an association, and so against the causal claim, which needs one.

(b) Bears on the association, and fits a role for serotonin in some people. The review itself reports weak evidence of this, and the two sides weigh it differently.

(c) Neither, directly. It's evidence that a drug helps, which is a different question. The response does count SSRIs' efficacy as lending credibility to a role, and the review calls that kind of inference an assumption, which is the next section.

Why whether the drugs help is a different question

The tempting leap from the 2022 headlines is from "the serotonin story is wrong" to "the pills don't work". Those are different sentences, and the two papers show why.

Whether a drug helps is measured in trials that compare it with placebo, which is the 0.3 you've already met. How it helps, if it does, is a separate question, and on this course's reading a treatment can work through a route that has nothing to do with the illness's cause.7 The response makes the point about scope in its opening, objecting to "discussion of antidepressant efficacy in a review that did not present any such data".5 The review's introduction separates the two questions from the other direction: "It is often assumed that the effects of antidepressants demonstrate that depression must be at least partially caused by a brain-based chemical abnormality, and that the apparent efficacy of SSRIs shows that serotonin is implicated."4 They call that an assumption, and they list other explanations that have been put forward for how the drugs have their effects.4

The two papers don't draw the link the same way. The response's conclusion treats the efficacy of SSRIs as lending credibility to a role for serotonin, which is the inference the review calls an assumption. But neither treats a review of causes as a test of whether the drugs help, and neither paper is evidence for starting or stopping a medicine.7

What people get wrong

"Antidepressants are no better than placebo." Both sides measure them as better on average, and Kirsch's own word was that the difference "easily" reached statistical significance. The dispute is over whether the margin is big enough to matter.

"Antidepressants fix a chemical imbalance." Neither paper in the serotonin dispute concludes that. The review found no convincing evidence that depression is associated with, or caused by, lower serotonin; the response defends a role and a perturbed system, not a deficiency the drugs refill.

"If the serotonin theory is wrong, the drugs don't work." How a drug works and whether it works are separate questions, answered by different studies, and the review didn't test the second.

"The improvement someone feels on an antidepressant is all the drug." Most of the improvement in these trials came in both arms. The drug's share is the gap.

"An average of 0.3 means everyone gets a small benefit." An average fits more than one pattern, and Stone's model suggests the advantage is concentrated in fewer people. Whether that model is right is a research question.

Practice

The review and the reply

Take 25 minutes.

Open Moncrieff and colleagues and Jauhar and colleagues side by side. Both are free, and both carry an abstract at the top.

  1. Write each paper's central claim in one sentence, keeping its own key verb or phrase.

  2. Mark the single strongest point in each, in your judgement, and write one line on why it's strong.

  3. Mark one thing each paper concedes or grades down in its own evidence.

  4. For each paper's conclusion, mark which clause the other paper actually contests and which it leaves alone.

Then check your sentences against this lesson's "Two claims, not one". If your summary of either paper uses a word the paper doesn't, change it.

Questions, not decisions

Take 15 minutes.

Write down the questions you'd want answered if you, or someone close to you, were talking with a prescriber about antidepressants. You're writing questions only. Don't answer them, and don't decide anything on the strength of the list.

Some directions to write in, if you're stuck: what the guideline suggests for someone in your situation and why; what the evidence suggests about a medicine for someone in that situation, and what tends to happen without one; what the options are, including ones that aren't a drug; what a person might expect, and by when; what the side effects can be; and, if a person on a medicine ever wanted to stop, how the prescriber would want that done.

Then read your list against this lesson and ask of each question whether it's one a trial can answer, one only a conversation about a particular person can answer, or both. Most of the important ones are in the second group, which is why the list is for taking to someone rather than for answering alone.

If writing the list brings up thoughts of suicide or self-harm, stop and use the callout at the top of this lesson.

Connections

Back. Lesson 2 gave you the tools this lesson ran on: what 0.3 looks like, and what the comparison group got, which here is the hardest comparator short of another treatment. It also warned against setting a therapy measured against waiting lists beside a drug measured against placebo. Lesson 6 met "comparable to antidepressants" and why a comparison across trials misleads. Sleep lesson 6 showed that a grade on the evidence reports certainty rather than size, which is how to read Cipriani's "moderate to very low", and NICE's table order needs the same care about what an order means. Lesson 7's ratio dispute ended with one claim falling and a weaker one left standing; the serotonin dispute has a similar shape on its causal claim, except that in the papers this course read, neither side gives up its central claim.

Forward. Lesson 9 closes the course by joining the comparison question to the claim sort you've used since the first course, and it ends on what this course cannot tell you about yourself.

Go deeper

  • Kirsch and colleagues, 2008, free in PLoS Medicine. This course read the abstract and parts of the Introduction, Methods, Results and Discussion. The page also carries an editors' summary, which is not the authors' text. Read it for the critique in its own words, including the sentence on severe depression.
  • Cipriani and colleagues, 2018, free at PubMed Central. This course read the abstract and parts of the Introduction, Methods, Results and Discussion. Its limitations paragraph is a fair account of what a review this size still can't tell you.
  • Stone and colleagues, 2022, free at PubMed Central. This course read the abstract, Introduction, Discussion and Conclusions, and the Results paragraphs on the model, not the Methods. The best single paper for seeing how both readings can be true at once.
  • Moncrieff and colleagues and Jauhar and colleagues, both free in Molecular Psychiatry. This course read the review's abstract, two paragraphs of its introduction and its Discussion, and the response in full. Read together, as the first exercise asks.
Before you change a treatment

Nothing in this lesson tells anyone to start, stop, reduce or switch a medicine. The research disputes it describes are about averages, mechanisms and thresholds, and none of them can say what's right for one person.

Stopping has its own evidence, which this course hasn't read beyond NICE's summary of it. Recommendation 1.4.14 notes that "withdrawal can sometimes be more difficult, with symptoms lasting longer (in some cases several weeks, and occasionally several months)".6 That's one more reason the conversation below comes first.

NICE's recommendation 1.4.12, in full: "Advise people taking antidepressant medication to talk with the person who prescribed their medication (for example, their primary healthcare or mental health professional) if they want to stop taking it. Explain that it is usually necessary to reduce the dose in stages over time (called 'tapering') but that most people stop antidepressants successfully. [2022]"6

Sources

  1. I. Kirsch, B. J. Deacon, T. B. Huedo-Medina, A. Scoboria, T. J. Moore and B. T. Johnson, "Initial severity and antidepressant benefits: a meta-analysis of data submitted to the Food and Drug Administration", PLoS Medicine 5(2), 2008, e45, doi 10.1371/journal.pmed.0050045, PMC2253608. Read: the abstract in full; from the full text, the Introduction's first paragraph, the Results paragraphs on the overall difference and on baseline severity, and the Discussion's paragraph on NICE's criterion (2026-09-23); the Methods paragraph on the two kinds of analysis (2026-09-24). Supports: the dataset and its scope, the 9.60 and 7.80 improvements, the 1.80 and 0.32 differences and how d was computed (1.24 and 0.92), the NICE criterion and its attribution to the 2004 guideline, the severity finding, and the prescribing sentence.
  2. A. Cipriani, T. A. Furukawa, G. Salanti and colleagues, "Comparative efficacy and acceptability of 21 antidepressant drugs for the acute treatment of adults with major depressive disorder: a systematic review and network meta-analysis", Lancet 391(10128), 2018, pp. 1357 to 1366, doi 10.1016/S0140-6736(17)32802-7, PMC5889788. Read: the abstract in full; from the full text, the Introduction's second paragraph, the Methods on outcomes, timepoint and statistical analysis, the first Results paragraph, the secondary-outcome paragraph, and the Discussion including limitations. Supports: the trial count, participants and drugs, the aim, the 0.30 and its interval, the kind of SMD, the baseline severity, the timepoint, the Discussion's summary, the Interpretation sentence, the dropout argument and the comparison with earlier reviews, the introduction's account of the debate, the risk-of-bias and certainty figures, and the limitation on adverse events and withdrawal.
  3. M. B. Stone, Z. S. Yaseen, B. J. Miller, K. Richardville, S. N. Kalaria and I. Kirsch, "Response to acute monotherapy for major depressive disorder in randomized, placebo controlled trials submitted to the US Food and Drug Administration: individual participant data analysis", BMJ 378, 2022, e067606, doi 10.1136/bmj-2021-067606, PMC9344377. Read: the abstract, Introduction, Discussion and Conclusions, with the author affiliations, disclaimer and competing interests, from the PMC full text (2026-09-24), and the Results paragraphs on the mixture model (2026-09-24); not the Methods. Supports: the dataset, the 1.75 and adult 1.82 and 0.24 figures, the introduction's sentences on individual responses, the three distributions and their shares in each arm, the middle-group sentence, the 15% conclusion, the placebo-arm sentence, the threshold sentence, the limitations, the call for predictors, the sentence on individual weighing, and the declared interests (summarised).
  4. J. Moncrieff, R. E. Cooper, T. Stockmann, S. Amendola, M. P. Hengartner and M. A. Horowitz, "The serotonin theory of depression: a systematic umbrella review of the evidence", Molecular Psychiatry 28(8), 2023 (online 2022), pp. 3243 to 3256, doi 10.1038/s41380-022-01661-0, PMC10618090. Read: the abstract in full; from the full text, the Introduction's first two paragraphs, the Discussion in full, and the competing interests. Supports: the central finding and the Discussion's opening sentence, the tryptophan results including the family-history clause, the genetic result, the SERT-binding summary, the conclusion, the public-belief passage, the quality limitation, the introduction's sentence on what is assumed from the drugs' effects, and the declared interests (summarised, not quoted).
  5. S. Jauhar, D. Arnone, D. S. Baldwin and 32 others, "A leaky umbrella has little value: evidence clearly indicates the serotonin system is implicated in depression", Molecular Psychiatry 28(8), 2023 (online 16 June 2023), pp. 3149 to 3152, doi 10.1038/s41380-023-02095-y, PMC10618084. Read in full, including the competing interests. Supports: the charge in the abstract, the tryptophan-depletion and plasma tryptophan figures with the "Admittedly" sentence, the SERT sentence, the conclusion, the objection to discussing efficacy, the call for a conventional umbrella review, and the declared interests (summarised).
  6. National Institute for Health and Care Excellence, Depression in adults: treatment and management, NICE guideline NG222, 2022, https://www.nice.org.uk/guidance/ng222. Read: recommendation 1.5.3, Table 1 with its heading and row order, Table 2's row order, and recommendations 1.4.12 to 1.4.21. Supports: the first-line recommendation for less severe depression, the SSRI row's place and the heading caveat, Table 2's first and fourth rows, 1.4.14's note on withdrawal, and the closing recommendation on stopping.
  7. This course's own constructions, labelled where they appear. The 62 and 86 percent conversions (lesson 2's arithmetic on the normal curve), the note that NG222 was not checked for any threshold, the observation that Cipriani did not set their figure against a threshold and that the course has read no defender's statement on clinical significance, the reading of the threshold as a judgement, the invented hundred-person example and the point where the threshold and the average meet, the note that Stone's paper is not a change of view by Kirsch, the 14.9-point subtraction, the setting of NG222 against both readings, the gloss on the serotonin transporter, the split of the serotonin dispute into a causal claim and an association claim and its classification as contested, the statement that neither conclusion claims a serotonin shortage, and the point that a treatment can work by a route unrelated to a cause.

Check your understanding

This lesson has a 6-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.