What happened when people checked
80 min
Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.
- State what the direct replication found, distinguishing the part that replicated from the part that did not
- Explain why "the study failed to replicate" is the wrong description of this result
- Identify, in a described replication, which component of a claim was tested
Lesson 3 left a claim with two halves. Typists transcribe more, and transcribing more costs you. This lesson is what happened when two teams went and checked.
The direct replication
In 2021, Heather Urry and her research-methods class published a direct replication of the first of the three studies.1 This course read its abstract and did not open the paper.
Their abstract, in full, because it is all this course has:1
In this direct replication of Mueller and Oppenheimer's (2014) Study 1, participants watched a lecture while taking notes with a laptop (n = 74) or longhand (n = 68). After a brief distraction and without the opportunity to study, they took a quiz. As in the original study, laptop participants took notes containing more words spoken verbatim by the lecturer and more words overall than did longhand participants. However, laptop participants did not perform better than longhand participants on the quiz. Exploratory meta-analyses of eight similar studies echoed this pattern. In addition, in both the original study and our replication, higher word count was associated with better quiz performance, and higher verbatim overlap was associated with worse quiz performance, but the latter finding was not robust in our replication. Overall, results do not support the idea that longhand note taking improves immediate learning via better encoding of information.
Both replications are the same population as lesson 3's: students, one recorded lecture, a quiz soon afterwards.12
Read the third sentence and then the fourth, because between them they are the lesson.
"As in the original study, laptop participants took notes containing more words spoken verbatim by the lecturer and more words overall than did longhand participants." That is the mechanism, and it replicated.
"However, laptop participants did not perform better than longhand participants on the quiz." That is the outcome, and it did not.
Read that sentence carefully, because it is phrased oddly: the original's claim was that laptops do worse, and this says only that they did not do better. The sentence that closes the gap is the abstract's last one: "results do not support the idea that longhand note taking improves immediate learning via better encoding of information".1
Why "it failed to replicate" is the wrong sentence
You'll hear this study described as a failed replication, and the description throws away the interesting part.
A claim of this shape has two components, and this split is this course's way of putting it, not a division either team makes.3 The mechanism is what the intervention does to the thing in the middle: here, what typing does to the notes. The outcome is what that is supposed to cost or buy: here, what it does to the test score.
A replication can confirm one and not the other, and when it does, the claim is in a more interesting state than either "confirmed" or "refuted".3 The thing typists were supposed to do to their notes, they did, in a replication of a hundred and forty-two people.1 The cost that was supposed to follow from it did not show up.
So the live question moves. It's no longer "do typists transcribe?" It's "does transcribing cost anything?", and that's a narrower and more answerable question than the one everybody was arguing about.
Before the next part. The proposed mechanism was that writing more, and more verbatim, is what costs you. The replication measured word count against quiz score directly. What would you expect to find?
Show the answer
The prediction almost everybody makes is the one the data contradicts, which is why it is worth writing down first.
If writing more is the problem, more words should go with worse scores. That's the mechanism stated as a correlation, and it's checkable in the same dataset.
What they report is the opposite. In their words: "in both the original study and our replication, higher word count was associated with better quiz performance".1
In both. Not only in the replication, where you might explain it away as a fluke, but in the original study whose own authors proposed the mechanism.
The verbatim half did point the way the mechanism predicts, and the same sentence says what happened to it: "higher verbatim overlap was associated with worse quiz performance, but the latter finding was not robust in our replication".1
So one of the two correlations held up and it is the one nobody quotes, and the one that is quoted did not.
One claim, two replications
This case is constructed, and it isn't about note-taking, so that you sort the structure rather than the subject.3
Somebody claims that a new checklist reduces errors in a hospital ward, because it forces staff to state the patient's allergies out loud.
The mechanism is the stating out loud. A replication of it asks: with the checklist, do staff state allergies out loud more often than without it? That is a question about behaviour and you answer it by watching.
The outcome is the errors. A replication of it asks: with the checklist, are there fewer errors? That is a question about results and you answer it by counting.
Before the list. Two of the four combinations are easy to read: both replicate, or neither does. What would it mean if the outcome replicated and the mechanism did not?
Show the answer
It means the thing works and not for the reason given, and it is the case people find hardest to accept.
Nobody is disputing the result in that case. The intervention did what it was supposed to do to the outcome. What failed is the account of how, which is the part everybody quotes.
Write down a claim you believe where that might be true, before you read the list. It is a more uncomfortable exercise than it sounds.
Four things can happen and all four are informative, and this four-way sort is the course's own.3
- Both replicate. The claim stands as stated.
- Neither replicates. The original result is in trouble.
- The mechanism replicates and the outcome does not. The intervention does what it was supposed to do and the benefit doesn't follow. This is the case this lesson is about.
- The outcome replicates and the mechanism does not. The checklist helps and not for the reason given. This course has not read anything that counts how often that happens, and it is the case people find hardest to accept.
Notice what case 3 doesn't license. It does not say the checklist is useless: it says the route from the checklist to the errors is not the route that was described. And it does not say the original researchers were careless. They proposed the most plausible account of a real result, which is what a discussion section is for.
The replication that added a condition
Kayla Morehead, John Dunlosky and Katherine Rawson replicated the original and extended it, and the extension is the part this course finds most useful, which is a judgement rather than anything they claim.2 This course read their abstract and not the paper.
In their words: "We conducted a direct replication of Mueller and Oppenheimer (2014) and extended their work by including groups who took notes using eWriters and who did not take notes."2
Lesson 1 gave you what the no-notes group did, and it is worth having again with its qualification attached. Some trends in their data favoured longhand, they report, and then: "performance did not consistently differ between any groups (experiments 1 and 2), including a group who did not take notes (experiment 2)".2
And their conclusion, in their own words: "Based on the present outcomes and other available evidence, concluding which method is superior for improving the functions of note-taking seems premature."2
"Premature" is a careful word and it is worth reading as one. It does not say the question is unanswerable. It says the evidence in 2019 was not enough to answer it, which is a statement about the state of a literature rather than about the world. Lesson 5 is what that literature looked like once somebody pooled it.
Two teams checked the most famous study in this subject, and neither found the effect everybody repeats clearly: one reports some trends favouring longhand and no consistent difference, the other no longhand advantage at all. Why is this course not simply telling you the original was wrong?
Show the answer
Because that isn't what happened, and getting this right is worth more than the finding.
The original reported something real. One team measured the note-taking difference it reported and found it. Neither abstract gives a size, and this course did not open either paper. Nothing this course read disputes that typists transcribe more.
What did not travel is the step from there to the test score, in an immediate test, with no opportunity to study. That is a specific condition and both replications used it, because a direct replication has to.
So the honest sentence is narrow. Under the original's own conditions, the note-taking difference reproduces and the quiz difference does not. What happens once notes are reviewed is a different question, and the second team's abstract reports group differences decreasing further after students studied their notes, in their second experiment, which is a hint rather than an answer.2
And there is a reason to be careful in the other direction too. Two studies finding no difference isn't the same as evidence of no difference, and lesson 5 has twenty-four studies pooled and a different answer. A reader who leaves this lesson certain that the medium does not matter has overshot exactly as far as the person who arrived certain that it does.
The eighty-eight authors
The replication lists eighty-eight authors. One is the researcher who led it and the rest are her research-methods class.1
That is a real fact about how the work was produced and it's worth a paragraph, because people read it two opposite ways and both are wrong.
It is not a reason to trust it less. A long author list says nothing about a design, a sample or an analysis, and those are the things that bear on a result.
It is not a reason to trust it more either. "Eighty-eight people checked it" is not what an author list means.
What it is, is a piece of information about incentives, and this part is the course's own reading rather than anything reported.3 A class replicating a famous study has nothing invested in the outcome. They did not propose the mechanism and were not going to be embarrassed by either result. The record also shows the study was preregistered, which is a fact rather than a reading.1 The author count tells you nothing; how little anybody stood to lose tells you something.
Three things people get wrong about this
"The laptop study was debunked." Its central note-taking finding was reproduced in a direct replication. What did not reproduce is one step in the account of what that finding costs.
"A failed replication means the original was wrong." It means one result did not appear again under those conditions. Four things can happen, and this was case 3.
"Writing more is the problem." More words went with better scores, in both studies.
Practice
Take 20 minutes.
Find a claim with a mechanism attached. "X works because Y." A management practice, a health recommendation, a study technique, anything where somebody has explained why their thing works.
Write four things.
- The claim, word for word.
- The mechanism: what the thing is supposed to do in the middle.
- The outcome: what that is supposed to buy.
- How you would test each half separately, in one sentence each.
Then one line: which half does the evidence you have been shown actually address?
This is what this course expects rather than something anybody has counted: usually it is the outcome, with the mechanism asserted in the discussion.3 That is not a criticism of anybody: it is what a study is for, and it is why lesson 3 told you to hold a mechanism more loosely than a result.
Take 15 minutes. This needs the trace you did in lesson 3. If you did not do it, pick any claim you have argued about this week.
Read what you wrote, and answer two questions.
- Which half of the claim were you tracing? The mechanism or the outcome.
- Was the source you reached about the half you cared about?
Then one line: if you were to spend another twenty minutes on it, what would you go looking for now?
This is an expectation rather than something anybody has counted: most traces go after the outcome, because that's the half with the number in it, and the mechanism is usually the half doing the persuading.3
Connections
Back. Lesson 3 is the original study and the mechanism it proposed, and this lesson is what happened to each half of it. Memory lesson 2 is what a replication can and cannot establish, and this lesson is that point at a higher resolution: there, a one-subject result reproduced; here, half a claim did.
Forward. Lesson 5 is twenty-four studies of the same question pooled, and five meta-analyses that all point the same way and disagree about whether the effect clears zero. It is also where the word-count difference gets its measured size.
Go deeper
- Don't Ditch the Laptop Just Yet: A Direct Replication of Mueller and Oppenheimer's (2014) Study 1 (Psychological Science, 2021). Read at abstract level only by this course. Read the abstract, which is unusually clear about which parts replicated and which did not, and is the whole of what this lesson rests on.
- How Much Mightier Is the Pen than the Keyboard for Note-Taking? (Educational Psychology Review, 2019). Also read at abstract level only. Read it for the no-notes group and for the word "premature", which is the most careful sentence anybody in this dispute has written down.
Sources
- Heather L. Urry and colleagues, "Don't Ditch the Laptop Just Yet: A Direct Replication of Mueller and Oppenheimer's (2014) Study 1 Plus Mini Meta-Analyses Across Similar Studies", Psychological Science 32(3), 2021, pp. 326 to 339. Read at abstract level only, from the bibliographic record. The paper was not opened, and the body says so. Supports: the abstract as quoted in full, including the 74 and 68 participants, the two correlations and their robustness, and the eight similar studies. The count of eighty-eight authors and the composition of that list come from the same record, which lists them.
- Kayla Morehead, John Dunlosky and Katherine A. Rawson, "How Much Mightier Is the Pen than the Keyboard for Note-Taking? A Replication and Extension of Mueller and Oppenheimer (2014)", Educational Psychology Review 31, 2019, pp. 753 to 780. Read at abstract level only, from the ERIC record, and the body says so. Supports: the three quoted fragments, about the eWriter and no-notes extension, about performance not consistently differing, and about concluding which method is superior being premature. The trends favouring longhand, and the group differences decreasing after students studied their notes, are in the same abstract.
- The split of a claim into a mechanism and an outcome, and the four cases a pair of replications can produce, are this course's own framing, said as such where they appear, and the hospital checklist is constructed. The reading of the eighty-eight authors as a fact about incentives is also the course's own, labelled where it appears; the author count itself is from the record. The observation that most traces go after the outcome is an expectation rather than a measurement, and the exercise says so.
Check your understanding
This lesson has a 6-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.