Retrieval practice: one of the two techniques that hold up
60 min
Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.
- Explain why retrieving information beats restudying it, and what "the testing effect" means
- Design a retrieval routine for a subject you are actually studying
- Distinguish self-testing to check what you know from self-testing to learn, and say what each one changes about what you do next
You've done this. The night before an exam, you reread the chapter, maybe twice. It goes smoothly; by the second pass the sentences feel familiar and you can almost see where each idea sits on the page. Then in the exam you reach for it and there's nothing there. Meanwhile the person next to you, who spent the evening "just doing practice questions" and complained that it felt awful, is writing steadily.
Lesson 1 explained the first half of that: familiarity was being read as knowledge. This lesson is the fix, and it's one of the two best-supported fixes in the whole literature on learning (spacing, in lesson 4, is the other). It's also the one most people already own and use wrongly.
The core idea
Pulling information out of memory strengthens it more than putting it in again. The act of retrieving something you've partly learned, from a question, a blank page, or a flashcard, does more for later memory than another exposure to the same material. Psychologists call this the testing effect, and the technique built on it retrieval practice.
How confident should you be in it? When Dunlosky, Rawson, Marsh, Nathan and Willingham (2013) reviewed ten common study techniques for the Association for Psychological Science, they rated each against how well it held up across learners, materials, and kinds of test. Only two came out "high utility". Practice testing was one; distributed practice, which is lesson 4, was the other. Rereading, the technique 84% of students in one elite-university sample listed and 55% ranked first, was rated low. 1
The founding modern study is worth knowing in numbers, because it tells you exactly when retrieval wins and when it looks like it loses. Roediger and Karpicke (2006), at Washington University in St. Louis, had students read short prose passages. Each student then restudied one passage and took a free-recall test on the other (write down everything you remember, no feedback). After five minutes, the restudied passage was ahead, 81% to 75%. 2
Same students, tested again a week later. Which passage do they remember better now, the one they restudied or the one they were tested on? And by how much?
Show the answer
The tested one, and by a wider margin than restudy had on day one. After two days the tested passage led, 68% to 54%. After a week, 56% to 42%. In a second experiment, students who studied a passage once and then took three recall tests remembered 61% at one week, against 40% for students who studied it four times. The four-time studiers were also the more confident. 2
Read that again with lesson 1 in mind. Restudying boosted performance today and did less for storage than the today number suggested. Testing cost a little today and paid at a week. The students' confidence tracked the today number. If your exam is tomorrow morning, rereading tonight is a defensible choice. This lesson is about what you'll still have in a week.
Does it hold outside the lab?
For a long time the honest answer was "mostly lab studies with word lists and short passages". That's no longer the case.
Yang, Luo, Vadillo, Yu and Shanks (2021) meta-analysed 222 classroom studies covering 48,478 students and found an overall effect of g = 0.499 for classroom quizzing on later performance. 3 Agarwal, Nunes and Blunt (2021) reviewed 50 classroom experiments from 37 studies (5,374 students) in real courses and found 57% of the effects medium or large; their practical conclusion was that nearly every quiz format they saw worked. 4 One caveat they make themselves: only 6% of those experiments came from outside Western, industrialised countries, so the classroom evidence is broad in age and subject and narrow in geography. And one caveat I'll make: everything in this lesson is about low-stakes or no-stakes retrieval with feedback. It says nothing about high-stakes exams as a policy, which is a different argument.
Yang and colleagues also measured what made the effect bigger, and the directions turn into three design rules. The effect was larger when the quiz came with feedback (g = 0.537 with feedback against 0.374 without), larger when the practice quiz matched the format of the final test, and larger the more often quizzing was repeated. 3 So: always check your answers; practise in the form you'll be tested in; and quiz more than once.
Adesope, Trevisan and Sundararajan (2017) pooled practice-testing studies of every kind and found g = 0.61 against all comparison conditions and 0.51 against restudy specifically. 5 Inside that headline there's a puzzle. Multiple-choice practice tests came out at g = 0.70 and short-answer at 0.48. If recognition is weaker than recall, as the next section argues, why would the recognition format win? Two reasons, and both matter for how you build your own questions. First, format match: most final tests in those studies were multiple choice, and Yang's moderator says matching format inflates the effect. Second, a well-built multiple-choice question isn't pure recognition. Little, Bjork, Bjork and Angello (2012) showed that when the wrong options are plausible, answering means retrieving why each alternative is wrong, and students later remembered information about those alternatives too. 6 A multiple-choice item with silly distractors is a recognition test; one with competitive distractors is several retrieval attempts at once. Write the second kind.
Why it works
Nobody fully agrees on the mechanism, and you should know the candidates, because the routine you build depends a little on which one you believe.
The oldest idea is that recall reconstructs the memory rather than reading it off a shelf, and each reconstruction lays down more routes to the same information. Carpenter (2009) sharpened that into the elaborative retrieval account: searching memory from a weak cue activates related information along the way, and that related information becomes part of the trace. The evidence is that the testing effect gets bigger when the cue is weaker, which is what you'd expect if the search is what does the work. 7
Karpicke, Lehman and Aue (2014) offered a different account, episodic context. When you retrieve, you reinstate the context in which you first learned the thing, and then update it with the context you're in now. The memory ends up tied to more, and more distinctive, contexts, so more situations can call it up later. 8
A third account, from Kornell, Bjork and Garcia (2011), doesn't compete with those two so much as explain the strange timing you saw in Roediger and Karpicke's numbers. Their bifurcation model says a test splits your items into two groups: the ones you retrieved get a large boost in strength, and the ones you didn't get nothing. Restudy gives every item a small boost. Right after study, a small boost to everything beats a large boost to some; after a delay, the weakly boosted items have slipped below the threshold and the strongly boosted ones are still there. That's the whole "testing loses at five minutes and wins at a week" pattern in one picture. 9
You don't need to pick a winner, and the field hasn't. What would separate them is whether the effect tracks cue weakness, context change, or the item-level split, and studies that vary one while holding the others fixed are still few. All three agree on the practical point: the effort of pulling the idea out is what does the work, and rereading skips it. The page reconstructs the idea for you and your memory only has to recognise it. This is why lesson 2's "memory is the residue of thought" matters here: when you reread, what you think about is the sentence in front of you. When you retrieve, what you think about is the idea itself, its connections, and the gap where a piece is missing. Retrieval is effortful in exactly the way lesson 1 called a desirable difficulty: it feels worse than rereading because your memory is being made to do something.
Which of the three accounts predicts that a harder cue produces a bigger testing effect, and why?
Show the answer
Carpenter's elaborative retrieval account. A weak cue forces a longer search through memory, and the related information activated during that search gets attached to the trace. A strong cue hands you the answer with little search, so there's less to attach. The episodic context account cares about context change rather than cue strength, and the bifurcation model is about which items get boosted, not about how hard each retrieval was.
One open question you should know about. Van Gog and Sweller (2015), working from the cognitive load theory you met in lesson 2, argued that the testing effect should shrink as the material gets more complex, because a retrieval attempt on material with many interacting parts overloads the working memory the attempt depends on; they reviewed studies on complex materials that found no effect. Karpicke and Aue (2015) replied that those studies had used weak or massed retrieval, and that with proper retrieval the effect shows up on complex material too. Neither side has conceded. What would settle it is a set of studies that vary the complexity of the material directly while holding the retrieval schedule constant and strong, and those are still thin on the ground. 10 What it means for you: if your subject is a web of interdependent ideas rather than a list, expect the effect to be smaller and more dependent on doing retrieval well (from memory, spaced, to a criterion), and watch your own delayed results rather than assuming the lab numbers transfer.
Getting it wrong helps, if you find out
A fear stops many people testing themselves early: "If I try to recall it and get it wrong, I'll learn the wrong answer." Kornell, Hays and Bjork (2009) tested exactly that. They gave people questions they could not possibly answer, let them guess, then showed the answer. In five of six experiments, guessing wrong and then seeing the answer beat spending the same time reading the answer, and the wrong guesses did not stick. 11 Butler and Roediger (2008) found the same pattern for multiple-choice tests: feedback kept the benefit and removed the harm. 12
The condition is the feedback. Roediger and Marsh (2005) showed that a multiple-choice test without feedback, even though it still helped on balance, can teach you some of its own wrong options; a share of the lures got remembered as facts. 13 So the rule isn't "don't guess". It's "guess, then always check". The failed attempt opens a gap; the feedback fills it; the memory that results is stronger than one that never had the gap.
You take a multiple-choice practice test and never check the answers. Which study in this section predicts what happens next, and what does it predict?
Show the answer
Roediger and Marsh 2005. The test still helps overall, but some of the wrong options you picked, or even just read, will come back later as things you believe. The fix costs thirty seconds: check the answers.
How many times, and when
One correct recall isn't much. Rawson and Dunlosky (2011) ran 533 students through schedules of retrieval to find out how much was enough. Their prescription, which they call successive relearning: in the first session, keep testing each item until you've recalled it correctly three times; then, in three later spaced sessions, bring each item back to one correct recall. 14 Lesson 4 is about the spacing of those later sessions. For now, hold on to the shape: retrieve to a criterion, then come back.
Worked example: retrieval versus the "deeper" technique
Here is the study most often cited against the objection you're probably forming: "fine for vocabulary, but I need to understand my subject, not memorise it."
Karpicke and Blunt (2011) took a 276-word science text and gave undergraduates one of four ways to learn it, twenty students in each. One group studied it once. One studied it repeatedly. One group did elaborative concept mapping: drawing the text's ideas as a diagram of nodes and labelled links, with the text and an example map in front of them. Concept mapping isn't a study-guide gimmick; it comes from Novak's tradition of meaningful learning and has a research literature of its own. The last group did retrieval practice: study the text, put it away, write down everything you can recall, study it again, recall again. 15
Before you look at the result, do what they asked the students to do. Rank the four methods (study once, study repeatedly, concept map, retrieval) by how well you think each will hold up on a test a week later.
Show the answer
The students, asked during learning, ranked repeated study best and retrieval worst, with mapping and single study in between. A week later the retrieval group scored 0.67 and the concept-mapping group 0.45, with the two study groups below mapping. If you ranked retrieval last, you're in good company, and wrong in the same way.
The test a week later asked both verbatim questions (facts stated in the text) and inference questions (things you'd have to work out by connecting the facts). Retrieval over mapping was an effect size of d = 1.50, one of the largest you'll see in this field, though it comes from twenty students per group, so hold the exact number loosely. And retrieval won on the inference questions, not just the verbatim ones.
Their second experiment made it personal, and put the size on firmer ground. 120 students each learned two texts, one by concept mapping and one by retrieval, so every student was their own control. 84% did better on the retrieval text. Yet when asked, during learning, which method would work better, 90 of the 120 (three-quarters) expected concept mapping to do as well as or better than retrieval; about half expected it to do better outright. 15 The paper's own table puts the two facts side by side, and the chart below redraws it:
Then the wrinkle inside the wrinkle. For half of those 120 students, the final test was itself a concept map: they had to draw the diagram from memory. The text they had learned by retrieval still came out ahead, d = 1.01. Practising the drawing lost to practising the recall, at drawing. 15 That's not a contradiction of the format-match rule from earlier. Format match helps when both formats are retrieval from memory; it doesn't rescue a format that has the source open in front of it.
Karpicke and Blunt's explanation is this. The mappers had the text in front of them the whole time, so each node could be located on the page rather than reconstructed. The retrieval students had the text taken away and had to rebuild not just the facts but the structure that connected them, because free recall of a passage is impossible without some structure to hang it on. The inference questions rewarded the structure. The concept-map test rewarded it too. On this account, the technique that looked structural was doing less structural work than the one that looked like rote.
The study drew fire, and you should know the objection. Mintzes and colleagues, researchers in the concept-mapping tradition, wrote to Science arguing that the students had been given little practice at concept mapping, so the comparison pitted a novice's map against a familiar act of recall, and that a well-taught map means choosing, ranking and labelling relations, which is not copying. They also argued that a recall test naturally favours a recall method. 16 Karpicke and Blunt replied that the concept-map final test in the second experiment was exactly the test that should have favoured mapping, and retrieval won there too. The critics' point about training is fair, and it's a real limit on how far this one study takes you. The reply is also fair. What both sides would accept is narrower than either headline: a concept map drawn from memory is, by the definition at the top of this lesson, retrieval practice. My own reading is that reconstruction is what does the work; a mapping researcher would add that a well-trained mapper is reconstructing too.
A harder example: testing to check versus testing to learn
Most students already self-test. That's the trap.
Kornell and Bjork (2007) surveyed 472 UCLA introductory psychology students about why they self-test.
Guess the share who said they self-test because testing teaches more than restudying. Then guess the share who said they do it to find out how well they've learned the material.
Show the answer
18% said testing teaches. 68% said they do it to check. 4% said they enjoy it, and 9% said they don't self-test at all. 17 Karpicke, Butler and Roediger (2009) asked 177 undergraduates a different way, to list their strategies: 84% listed rereading, and 11% listed self-testing at all. 18
Why does the reason matter, if the behaviour is the same? Because it usually isn't the same. The survey doesn't say what those students did in the minute after the quiz, so what follows is my reading of how the two motives play out, not something Kornell and Bjork measured. There is measured evidence for the cost, which I'll come to.
Call the first student Tom. Tom is a checker. He looks at the score. If it's high, he stops: the thermometer says done. If it's low, he goes back to the chapter and rereads it, all of it, from the start. Either way the test was a measurement, and the learning was supposed to happen somewhere else, in the rereading. He has just used the strongest technique available as a gauge for one of the weakest.
Call the second student Maya. Maya is a learner. She barely looks at the score. She looks at which questions she missed. Those go on a list. She rereads the sections behind those questions and nothing else. Then she retests the list, and a question stays in that day's session until she has answered it correctly three times. After that it comes back once in each of the later sessions, and a miss at any stage puts it back to three. That's Rawson and Dunlosky's prescription, run informally.
Tom's routine has retrieval in it. Maya's routine is made of retrieval. Roediger and Karpicke's 40% versus 61% is the gap between no retrieval and repeated retrieval; Tom sits somewhere between, and the measured cost of his actual habit is the next study.
Dunlosky and Rawson (2012) had students judge their own answers to concept definitions and stop practising an item once they judged it learned. Students who were overconfident in their judgements stopped early and, on a delayed test, retained less. 19 Tom isn't just wasting the test. His reading of the score is deciding when he quits, and the score he reads is inflated.
There's a second-order cost too. A quiz taken right after rereading measures retrieval strength at its peak, which is the least informative moment (lesson 1 again). A checker who scores 90% the night after studying is measuring fluency and calling it knowledge. A learner who tests two days later, scores 60%, and fixes the 40% has better information and a better memory.
You scored 90% the night after studying. Name two reasons that number tells you less than a 60% scored two days later would.
Show the answer
First, the 90% was taken at peak retrieval strength, right after exposure, so it mostly measures fluency; the 60% was taken after some forgetting, so it measures what's actually stored. Second, the 60% comes with a list of the 40% you missed, which is exactly the information a learner needs to steer the next session. The 90% gives you nine right answers and permission to stop, which Dunlosky and Rawson found is the permission that costs you.
Worked example: one routine, built for a real case
Exercise 2 will ask you to build a routine. Here's one built first, so you can see the shape.
Priya is three weeks from a first-year anatomy test on the bones and muscles of the arm. Day 1, after the lecture, she closes everything and free-recalls onto a blank page for five minutes: every bone, every muscle, its origin and insertion, whatever comes. She opens her notes and marks the gaps. There are eleven. Each gap becomes a question on a card: "Origin and insertion of biceps brachii?", "Which nerve runs through the cubital fossa?" and so on. Still on day 1, she works through the eleven cards until each has been answered from memory three times. Two of them take six or seven tries. The cards live in a shoebox on her desk, and her calendar has three entries: day 6, day 12, day 18.
Day 6 she takes out all eleven cards and answers each until it's right once. Nine go straight through. Two she misses.
What should happen to those two missed cards on day 6, and what does the schedule look like for them from here?
Show the answer
They go back to three. She works each of the two until she has answered it correctly three times in that session, then they rejoin the others for day 12 and day 18, one correct recall each.
Day 12 she misses one card; same treatment. Day 18, all eleven come through first time. The day before the test she does nothing to the cards, because the schedule has already done the work, and she spends the evening on a free recall of the whole topic instead, which is a test of the structure rather than the items.
Notice what's not in the routine: no rereading of the chapter from the start. The only rereading was the sections behind the eleven gaps, once, on day 1.
What people get wrong
"Testing is for measuring, not for learning." This is the checker's belief, and in Kornell and Bjork's survey it was the majority one. A test with feedback is a learning event, and one of the two strongest we have. Treat every quiz on this site, and every question you write for yourself, as study, not as an exam.
"Once I've got it right, I can drop it." Karpicke (2009) gave students control over their own study and watched what they did: most dropped an item from further practice as soon as they had recalled it once, and their later retention suffered for it. 20 One correct recall is the first step of a schedule, not the end of one. Three in the first session, then back to it later.
"I'll get it wrong and learn the wrong thing." Only if you never find out. A wrong attempt that gets corrected is still a retrieval attempt, and the correction is what stays. The rule is not "don't guess" but "always check".
Practice
Take a blank page. Without opening lesson 1 or lesson 2, write ten questions that a fair examiner would ask about them: five on learning versus performance and the fluency illusion, five on working memory, chunking, and cognitive load. Make at least four of them application questions ("A student does X. What happens at a week, and why?") rather than definitions.
Then, still without looking, answer all ten.
Now open the lessons and mark yourself. Every question you got wrong, or couldn't answer, or answered vaguely, goes on a list with the section it came from. Reread only those sections. Tomorrow, retest the list. Do not reread anything that isn't on the list.
Notice two things while you do this. First, writing the questions was itself retrieval, and probably the hardest part. Second, the questions you couldn't write are the ideas you don't yet have a route to; those are the ones to keep.
Pick a subject you are actually studying right now, not this course. Decide, in writing, on the routine you will use for the next two weeks, using Priya's as the template:
- After each study session: close everything and free-recall on a blank page for five minutes. Write what you remember, in whatever order it comes, including diagrams. Then open your notes and mark the gaps.
- Turn each gap into a question, and in that same session keep at it until you've answered it from memory three times. In each later session, once is enough. Anything you miss goes back to three.
- Two days later, and then on a spaced schedule: bring the whole list back. Lesson 4 will tell you how far apart the sessions should be; for now, three sessions across the two weeks.
- Before rereading anything: ask what you'd get wrong if tested on it now. If you don't know, test first.
Write down where the questions will live (a notebook, a card deck, an app) and the dates of the return sessions. A routine you haven't scheduled isn't a routine.
Every quiz question you pass on this site goes into your review bank and comes back later, spaced out, on the Review page. There's also a "Practise 10 random" mode that pulls questions from everything you've done. That gives you the spaced return visits; it doesn't track a three-correct-recalls criterion, so the first-session part of the routine is still yours to run. Use it as a learner, not a checker. When a question comes back and you miss it, reopen that section of the lesson, then let the question return.
Connections
Lesson 1 gave you the diagnosis: performance now is a poor guide to storage, and your sense of knowing tracks fluency. Retrieval practice is the direct treatment. It builds storage strength while feeling worse than rereading, which is precisely why, in both studies above that asked, students preferred the method that worked less.
Lesson 2 explained why reconstruction is the point: memory is the residue of thought, and retrieval makes you think about the idea rather than the page. It also gave you cognitive load theory, which is where the one live objection to this lesson comes from.
Lesson 4 answers the question this lesson leaves open: when to come back. Rawson and Dunlosky's three relearning sessions have to be spaced, and the size of the gap depends on how long you need the memory to last. Lesson 5 will add that the questions you retrieve should be mixed, not blocked. Lesson 6 turns retrieval on understanding hard material through self-explanation, which is retrieval of the "why". The course project asks you to write a self-test on day one and take it on day fourteen; you now know why the test has to be written first.
Go deeper
- Dunlosky, "Strengthening the Student Toolbox", American Educator, Fall 2013: the ten-technique ratings in plain English, free online. Read this before anything else.
- Brown, Roediger and McDaniel, Make It Stick (2014), chapter 2, "To Learn, Retrieve": the narrative version of this lesson, with the classroom studies told as stories; Roediger is an author of the founding study.
- Carpenter, Pan and Butler, "The science of effective learning with spacing and retrieval practice", Nature Reviews Psychology 1, 496–511 (2022): the current review of both this lesson and the next, written for a general scientific reader.
- Agarwal, Nunes and Blunt, "Retrieval Practice Consistently Benefits Student Learning" (2021): the classroom review; read it for the range of formats that worked and the honest list of what the field still hasn't tested. Free to read at the link.
- RetrievalPractice.org: Agarwal's site of short free guides for teachers and tutors setting up quizzing in a class; the main guide is a ten-minute read.
Sources
- Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). "Improving Students' Learning With Effective Learning Techniques." Psychological Science in the Public Interest, 14(1), 4–58. Supports: practice testing and distributed practice as the two high-utility techniques; rereading rated low; the 84% and 55% rereading figures (reporting Karpicke et al.).
- Roediger, H. L., & Karpicke, J. D. (2006). "Test-Enhanced Learning." Psychological Science, 17, 249–255. Supports: Experiment 1 within-subjects; 81% vs 75% at five minutes, 68% vs 54% at two days, 56% vs 42% at one week; 61% vs 40% at one week for study-test-test-test vs study-study-study-study, with higher confidence in the repeated studiers (Experiment 2).
- Yang, C., Luo, L., Vadillo, M. A., Yu, R., & Shanks, D. R. (2021). "Testing (quizzing) boosts classroom learning." Psychological Bulletin, 147, 399–435. Supports: 222 classroom studies, 48,478 students, g = 0.499; larger with feedback (0.537 vs 0.374), with test-format match, and with more repetitions.
- Agarwal, P. K., Nunes, L. D., & Blunt, J. R. (2021). "Retrieval Practice Consistently Benefits Student Learning." Educational Psychology Review, 33, 1409–1453. Supports: 50 classroom experiments from 37 studies, n = 5,374; 57% medium or large; 6% from non-WEIRD countries; nearly every format worked.
- Adesope, O. O., Trevisan, D. A., & Sundararajan, N. (2017). "Rethinking the Use of Tests: A Meta-Analysis of Practice Testing." Review of Educational Research, 87(3), 659–701. Supports: g = 0.61 vs all comparisons, 0.51 vs restudy; multiple-choice 0.70, short-answer 0.48.
- Little, J. L., Bjork, E. L., Bjork, R. A., & Angello, G. (2012). "Multiple-choice tests exonerated, at least of some charges: Fostering test-induced learning and avoiding test-induced forgetting." Psychological Science, 23(11), 1337–1344. Supports: multiple-choice items with competitive alternatives prompt retrieval of why the alternatives are wrong and improve later recall of related information.
- Carpenter, S. K. (2009). "Cue strength as a moderator of the testing effect: The benefits of elaborative retrieval." Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(6), 1563–1569. Supports: the elaborative retrieval account; larger testing effects with weaker cues.
- Karpicke, J. D., Lehman, M., & Aue, W. R. (2014). "Retrieval-based learning: An episodic context account." Psychology of Learning and Motivation, 61, 237–284. Supports: the episodic context account.
- Kornell, N., Bjork, R. A., & Garcia, M. A. (2011). "Why tests appear to prevent forgetting: A distribution-based bifurcation model." Journal of Memory and Language, 65(2), 85–97. Supports: the bifurcation model and its explanation of the short-delay reversal.
- van Gog, T., & Sweller, J. (2015). "Not new, but nearly forgotten: The testing effect decreases or even disappears as the complexity of learning materials increases." Educational Psychology Review, 27, 247–264; and Karpicke, J. D., & Aue, W. R. (2015). "The testing effect is alive and well with complex materials." Educational Psychology Review, 27, 317–326. Supports: the complexity debate as contested, and its cognitive-load rationale.
- Kornell, N., Hays, M. J., & Bjork, R. A. (2009). "Unsuccessful retrieval attempts enhance subsequent learning." Journal of Experimental Psychology: Learning, Memory, and Cognition, 35(4), 989–998. Supports: failed attempts followed by the answer beat reading the answer in five of six experiments; errors did not persist.
- Butler, A. C., & Roediger, H. L. (2008). "Feedback enhances the positive effects and reduces the negative effects of multiple-choice testing." Memory & Cognition, 36(3), 604–616. Supports: feedback as the condition for multiple-choice testing to help.
- Roediger, H. L., & Marsh, E. J. (2005). "The positive and negative consequences of multiple-choice testing." Journal of Experimental Psychology: Learning, Memory, and Cognition, 31(5), 1155–1159. Supports: a net benefit of multiple-choice testing alongside some lures learned as facts when no feedback is given.
- Rawson, K. A., & Dunlosky, J. (2011). "Optimizing schedules of retrieval practice for durable and efficient learning." Journal of Experimental Psychology: General, 140, 283–302. Supports: 533 students; successive relearning prescription of three initial correct recalls followed by relearning to one correct recall in three spaced sessions.
- Karpicke, J. D., & Blunt, J. R. (2011). "Retrieval Practice Produces More Learning than Elaborative Studying with Concept Mapping." Science, 331, 772–775. Supports: Experiment 1, four conditions, 20 per cell, 276-word text, mappers had the text and an example map; one-week test with verbatim and inference questions; retrieval 0.67 vs concept mapping 0.45, d = 1.50; students' predictions during learning ranked repeated study best and retrieval worst. Experiment 2, within-subjects N = 120, 84% better after retrieval; 90 of 120 predicted mapping would be as good as or better (49% better, 26% same); concept-map final test for half the sample, d = 1.01.
- Mintzes, J. J., Canas, A., Coffey, J., Gorman, J., Gurley, L., Hoffman, R., McGuire, S. Y., Miller, N., Moon, B., Trifone, J., & Wandersee, J. H. (2011). "Comment on 'Retrieval Practice Produces More Learning than Elaborative Studying with Concept Mapping'." Science, 334(6055), 453; and Karpicke, J. D., & Blunt, J. R. (2011). "Response to Comment." Science, 334(6055), 453. Supports: the critique (little training in mapping; recall test favours recall) and the reply (retrieval won on the concept-map test).
- Kornell, N., & Bjork, R. A. (2007). "The promise and perils of self-regulated study." Psychonomic Bulletin & Review, 14(2), 219–224. Supports: 472 UCLA students; of the whole sample, 18% self-test because it teaches more, 68% to check, 4% because it is enjoyable, 9% do not self-test.
- Karpicke, J. D., Butler, A. C., & Roediger, H. L. (2009). "Metacognitive strategies in student learning: Do students practise retrieval when they study on their own?" Memory, 17, 471–479. Supports: 177 undergraduates; 84% listed rereading; 11% listed self-testing.
- Dunlosky, J., & Rawson, K. A. (2012). "Overconfidence produces underachievement: Inaccurate self evaluations undermine students' learning and retention." Learning and Instruction, 22, 271–280. Supports: overconfident self-judgements led students to stop practising early and retain less.
- Karpicke, J. D. (2009). "Metacognitive control and strategy selection: Deciding to practice retrieval during learning." Journal of Experimental Psychology: General, 138(4), 469–486. Supports: students given control of their study dropped items after one correct recall, and retained less as a result.
Check your understanding
This lesson has a 5-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.