Five meta-analyses, one question

85 min

Listen: this lesson as a conversation

Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.

In this lesson you will learn to
  • Read a table of meta-analytic effect sizes and say what the spread across them means
  • Explain why five syntheses of one literature can give five different answers
  • Say what this course will and will not conclude about note-taking medium, and why

Lessons 3 and 4 were one study and what happened when two teams checked it. This lesson is what happened when five teams pooled the whole literature, and it is the reason this course exists.

Five answers

Five meta-analyses have asked whether handwritten lecture notes beat typed ones for achievement. The most recent of them prints the other four beside its own result, in a table this course read along with the discussion around it.1 The other four were not opened for this course. Their numbers reach you through this paper's table, which is exactly the trace lesson 3 taught you to run.

A positive number means handwriting came out ahead. Hedges' g is a difference measured in standard deviations, so 0.25 means the two groups' averages sat a quarter of a standard deviation apart, and the interval beside it is the range the synthesis cannot rule out.

The synthesis Effect sizes pooled Hedges' g 95 percent interval
Allen and colleagues, 2020 24 +0.250 0.127 to 0.371
The 2024 meta-analysis 49 +0.248 0.181 to 0.315
Lau, 2022 76 +0.144 0.023 to 0.265
Urry and colleagues, 2021 22 +0.040 −0.13 to 0.20
Voyer and colleagues, 2022 73 +0.008 −0.16 to 0.18

The 73 in the last row is the 2024 paper's count for that synthesis; lesson 1's scope table records why a summary of the same paper says 77.1

The table's own note, verbatim: "A positive Hedges' g indicates that handwritten notes were more effective than typed notes on overall achievement."1

Every study pooled in that table is a college student taking lecture notes, tested within the hour or the week. That is lesson 1's scope table, five syntheses deep.1

Five meta-analyses of one question, with their confidence intervals Five horizontal intervals on a scale of Hedges' g from minus nought point two five to plus nought point four five. A positive value favours handwriting. Allen and colleagues 2020 sits at plus nought point two five zero with an interval from nought point one two seven to nought point three seven one. The 2024 meta-analysis sits at plus nought point two four eight, interval nought point one eight one to nought point three one five. Lau 2022 sits at plus nought point one four four, interval nought point nought two three to nought point two six five. Urry and colleagues 2021 sits at plus nought point nought four zero, interval minus nought point one three to nought point two zero. Voyer and colleagues 2022 sits at plus nought point nought nought eight, interval minus nought point one six to nought point one eight. A dashed line marks zero. All five points sit to the right of it, three intervals clear it and two cross it; the two that cross it are drawn in a second colour with unfilled circles. Five syntheses, one question Hedges' g for achievement. Positive favours handwriting no difference Allen and colleagues, 2020 (24) .250 The 2024 meta-analysis (49) .248 Lau, 2022 (76) .144 Urry and colleagues, 2021 (22) .040 Voyer and colleagues, 2022 (73) .008 −0.1 0 0.1 0.2 0.3 Table 6 of the 2024 meta-analysis, drawn

Two facts are in that picture and quoting either one alone misleads.

All five point the same way. Every point estimate is positive. Not one synthesis of this literature has found typing ahead.

Three intervals clear zero. Two do not. In the 2024 paper's own words: "The present study is among three (i.e., Allen et al., 2020; Authors, under review; Lau, 2022) of the five studies included in Table 6 that detected statistically significant achievement advantages stemming from recording handwritten lecture notes."1

"Authors, under review" is the 2024 paper naming itself, which is how a paper cites its own unpublished work, so the three are Allen, Lau, and the paper printing the table. And of the other two, in the same paragraph: "Although the directionality of the overall meta-analytic achievement analyses of the other two studies included in Table 6 (i.e., Urry et al., 2021; Voyer et al., 2022) also favored handwritten notes, those analyses did not achieve statistical significance."1

And read the fourth column against the second. Voyer pooled seventy-three effect sizes and has the widest interval in the table; the 2024 paper pooled forty-nine and has the narrowest. More studies did not buy a tighter answer, because how much the pooled studies disagree with each other matters as much as how many there are. That is the second thing the table teaches and almost nobody looks for it, and reading it that way is this course's own rather than a point any of the five makes.2

Predict first

Before you read on. Five teams pooled overlapping sets of studies about the same question and got pooled effects ranging from +0.008 to +0.250, a factor of thirty. Write down what you think they are disagreeing about.

Show the answer

The guess this course expects most readers to make is that somebody made a mistake, and that isn't what's going on.

They aren't disagreeing about what any individual study found. The studies are published; the numbers are the numbers. Nobody in this table is disputing anybody's data.

They are disagreeing about which studies belong in the pool. A meta-analysis begins with a decision about what counts as an instance of the question, and that decision is made by people. A pool can be drawn tight or wide: one synthesis might take only laboratory work, another might take classroom studies too. This course read one of the five, so it cannot tell you which did what, with one exception, and the exception is the useful one.

Voyer and colleagues, whose pooled effect is the smallest, offer one version of this themselves: distraction may account for the handwriting advantage found in earlier work, because laboratory studies remove it. That reaches this course from a summary of their paper rather than from the paper, which lesson 1's scope table records.2

And that isn't a flaw in the method, and this next part is the course's reading of the table rather than something the table says.2 It's what a meta-analysis is: a single estimate produced by a stated set of choices. Change the choices and you change the estimate, which is why a careful synthesis states its inclusion criteria before its result.2

What the range tells you, then, is how sensitive the answer is to those choices. A question whose pooled effect swings from about zero to about a quarter of a standard deviation depending on who's included is a question whose answer depends on which reader you are.

The moderator that goes both ways

The sharpest version of that is in the 2024 paper itself.

Both it and Lau's 2022 synthesis asked whether reviewing your notes changes the size of the handwriting advantage. They got opposite answers, and the 2024 authors say so, verbatim: "Curiously, Lau (2022) found that the achievement advantage for handwritten notes disappeared after notes were reviewed, whereas our analysis revealed that note review increased the achievement benefits of handwritten notes."1

One team found reviewing erases it. The other found reviewing increases it.

Their own explanation is tentative and worth having, because it shows the mechanism of the disagreement: the two analyses covered different student populations.1 This course did not read Lau, so it cannot say which populations.

This matters more than the headline number, because a moderator analysis splits an already-pooled literature into smaller pieces, so it's the part of a synthesis most sensitive to which studies went in. A moderator finding is a weaker thing than the main effect it qualifies, which is this course's reading and not a sentence any of the five wrote,2 and this pair is the clearest demonstration of it in anything this course read.2

What a reader quoting one row would be entitled to say

Take the top row and the bottom row.

A reader quoting Allen and colleagues alone would say: handwriting produces a small but reliable achievement advantage, pooled across twenty-four effect sizes, with an interval well clear of zero. Every word of that is true.

A reader quoting Voyer and colleagues alone would say: pooled across seventy-three effect sizes, there is no reliable achievement difference between the two media. Every word of that is true too.

Neither reader is misquoting anything. They've each picked a synthesis and reported it correctly, and they will disagree completely.

What neither can say is "the research shows". There's no "the research" here. There are five readings of an overlapping body of studies, and the honest version of either sentence names its source and says there are four others.

Check yourself

So which is it? A course that lays out five answers and refuses to pick one can look like it is dodging the question.

Show the answer

It would be dodging if the question had an answer this course was withholding. It doesn't, and saying why is more useful than a verdict.

What this course can say is the direction. All five point one way. Nobody has found typing ahead. That isn't nothing, and a reader who takes only "the syntheses disagree" out of this lesson has dropped the half that's consistent.

What it can't say is the size, and size is what a decision needs. The largest of the five is a quarter of a standard deviation and the smallest is indistinguishable from zero, and which one describes you depends on whether you are in the population each pooled.

So what the course hands you is the table rather than an instruction, and the section below is the one reader for whom that difference is not a technicality.

One thing this lesson has to add

And there is one thing this lesson has to add that none of the sources does.2 Advice to write by hand is not neutral across readers. For some people a keyboard isn't a preference, and nothing read for this course reports whether any of these studies included them. A course that turned this table into "write by hand" would be giving an instruction to people its evidence never looked at.

So the course's answer is the table, and the decision's yours. Lesson 8 is where you make it, and the project is where you find out what your own notes are actually doing.

The difference everybody can measure, and the one they cannot

One more pair of numbers, and together they are the most useful thing in this lesson.

The volume difference is large and undisputed. Typed notes contain more words, at a pooled effect of about 0.919 in the 2024 analysis, where the sign runs the other way from the table: this one favours typing.1 Urry and colleagues put the same difference the other way round again, at about −0.91 with an interval from −1.18 to −0.65, as the 2024 paper quotes them.1 Two syntheses, two directions of sign, one agreed finding: typing gets more words down, by a lot.

The achievement difference is small and disputed, at most about 0.25, and two of five syntheses cannot distinguish it from zero.

And the sentence that joins them is the 2024 paper's own, verbatim: "there is no evidence across the existing meta-analyses that the advantages in note-taking volume enjoyed by students who type their notes give those students any achievement advantages over students who record handwritten notes, even when those notes are reviewed before testing."1

Read the two effect sizes together. The thing that's easy to measure is large, reliable and agreed. The thing anybody actually cares about is small, contested, and not bought by the first one. That pattern isn't special to note-taking, and lesson 8's sort is where it becomes a question you ask.2

Three things people get wrong about this table

"The meta-analysis settles it." Five of them exist, and naming one is choosing one.

"If the studies disagree, nobody knows anything." All five point the same direction. That's a finding, and it's the half that gets dropped by people who've decided the literature is a mess.

"A bigger effect size means a more important finding." The largest effect here is the number of words in a notebook.

Practice

Find the second one

Take 25 minutes.

Put it down at twenty-five minutes whether or not you found a second one: not finding one inside twenty-five minutes is outcome two below.

Find a claim in any field that is supported by "a meta-analysis", named or not.

Then go looking for a second meta-analysis of the same question. Search in Google Scholar or PubMed, and search the question rather than the claim.

Write down four things.

  1. The claim, and which synthesis it cites, if it names one at all.
  2. Whether a second one exists. Search the question rather than the claim.
  3. If it does: what it found, and whether that is the same answer.
  4. What the first source would have had to say to be reporting the literature rather than a reading of it.

Two outcomes are both results. Either you find a second synthesis, in which case you now know something the person quoting the first one didn't, or you find none, in which case the claim rests on one team's inclusion criteria and nobody has checked them.

One sentence you would be willing to say

Take 15 minutes.

Look at the table again. Write the one sentence about note-taking medium you would be willing to say to a colleague who asked, with nothing left off.

Then check it against four tests.

  1. Does it give a direction? All five syntheses have one.
  2. Does it give a size, and is the size defensible? Look at the spread before you answer.
  3. Does it say who was measured? What this course can vouch for is the 2024 synthesis: college students taking lecture notes, tested within the hour or the week.
  4. Would you be comfortable if the colleague repeated it to somebody else without the qualifications? If not, the qualifications aren't in the sentence yet.

Most first attempts fail test 4, which is the course's own expectation rather than something it has measured,2 and the usual reason is that the qualification was in your head rather than in the sentence.

Connections

Back. Lesson 3 is the study this whole literature was built around, and lesson 4 is the paper whose mini meta-analyses are one of the five rows. Lesson 2's denominator question is why the volume effect is easy to measure and the achievement effect is not. Memory lesson 3 is a pooled figure that halves on one methodological choice, and this lesson is the same lesson with five pools instead of one.

Forward. Lesson 6 is the same question asked with a completely different instrument, and what that instrument can and cannot show. Lesson 8 is where the table becomes a habit rather than a fact.

Go deeper

  • Typed Versus Handwritten Lecture Notes and College Student Achievement: A Meta-Analysis (Educational Psychology Review, 2024). Read in part by this course: the abstract, the introduction, the table this lesson is built on and the discussion around it, and one of its four limitations subsections. Read its Table 6 and the discussion around it, which is the whole of this lesson's evidence and is worth seeing in its own setting.
  • Don't Ditch the Laptop Just Yet (Psychological Science, 2021). Read at abstract level only by this course, and its mini meta-analyses are one of the five rows. Read it if you want the paper behind a row, which this course did not open past the abstract.

Sources

  1. Abraham E. Flanigan, Jordan Wheeler, Tiphaine Colliot, Junrong Lu and Kenneth A. Kiewra, "Typed Versus Handwritten Lecture Notes and College Student Achievement: A Meta-Analysis", Educational Psychology Review 36, 2024, article 78. Read in part: the abstract, the introduction, the Table 6 comparison of meta-analyses, the discussion of note-taking quantity, and the first of the four limitations subsections. The results tables other than Table 6 were not read, and lesson 1's scope table records the read level. Supports: every cell of the table and the chart, which are Table 6 reproduced; the table's own note, quoted; the quoted sentence naming the three syntheses that reached significance; the quoted sentence about the other two; the quoted sentence about Lau's opposite moderator finding; the tentative explanation about student populations; the volume effects of 0.919 and −0.91 with its interval; and the quoted sentence about volume buying no achievement advantage. The four other syntheses in the table were not read: their numbers reach this course through this paper's table, which is what the table is for.
  2. The account of why five syntheses of one literature can differ, and that the disagreement is about inclusion rather than about data, is this course's own explanation of what the table shows, said as such in the predict block. So is the statement that a moderator finding is a weaker thing than the main effect it qualifies, labelled where it appears. So is the note that advice to write by hand is not neutral across readers, and the sentence with it is exact: nothing read for this course reports whether these studies included people for whom a keyboard is not a preference. So is the closing observation that the easy-to-measure difference being the large one is a pattern beyond note-taking. The expectation that most first attempts fail the fourth test is an expectation rather than a measurement, and the exercise says so.

Check your understanding

This lesson has a 6-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.