Learning is not performance
70 min
Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.
- Explain the difference between storage strength and retrieval strength, and why the two can come apart
- Identify the fluency illusion in a described study habit and say which strength it is boosting
- Predict which of two study conditions will win at five minutes and which will win at one week
You already know the feeling. You go over your notes the night before, everything looks familiar, you'd bet money you know it. Then, a week later, someone asks you a question about it and there's nothing there. Most people conclude they have a bad memory. They don't. They have a normal memory and a bad instrument for measuring it: the feeling of knowing.
This lesson gives you a better instrument. It rests on one distinction that everything else in this course depends on: what you can do right now is a poor guide to what you will be able to do next month. Once you can see that gap in your own habits, a lot of study advice that sounds like a matter of taste turns out to be a question you can settle with evidence.
Two strengths, not one
Start with an example. There's a phone number you dialled every day as a child. Try to say it now. For most people nothing comes, or the wrong number comes. But if someone showed it to you once, you'd have it back in seconds, and it would stay back. Now think of a number you read off a screen thirty seconds ago. You can say it right now without trouble, and by tomorrow it will be gone.
Those two numbers are in opposite states, and one word ("remembered") can't describe both. Robert Bjork and Elizabeth Bjork, who have studied memory at UCLA for decades, proposed that every memory has two separate strengths (Bjork & Bjork, 1992; 2011).12
Storage strength is how deeply a memory is entrenched, how well it is woven into what you already know. In the Bjorks' model it does not decrease. Once something is well learned, it stays learned in this sense; what changes is whether you can get at it.
Retrieval strength is how accessible the memory is right now. It rises with recent use and falls with time and with competition from other memories.
The childhood phone number has high storage strength and low retrieval strength today. The number from thirty seconds ago has high retrieval strength and very little storage strength. So the two can come apart in both directions, and that's the whole reason the distinction is worth having.
I should say plainly that this is the Bjorks' model. It is one model among several of how memory works, and researchers argue about the details. The experiments you'll meet in the rest of this lesson are real whichever model you prefer; the two strengths are the clearest way I know to read them.
Here is the problem the model exposes. Every test you give yourself in the moment, every "yes, I know this" as you look over a page, measures retrieval strength. What you actually want for an exam next month, a conversation next year, or a skill for life is storage strength. And you have no direct sense of storage strength at all. You infer it from how easy things feel now.
There's one more rule in the model, and it's the one that does the real work. The Bjorks propose that the gain in storage strength from a study or retrieval event is larger the lower the memory's current retrieval strength is at the time.12 Put crudely: you build the most when the thing is hardest to bring to mind. That single rule predicts that recalling beats rereading (the answer isn't in front of you), that a gap between sessions helps (retrieval strength has dropped), and that mixing topics helps (each one is less accessible when you come back to it). Lessons 3, 4 and 5 take each of those in turn. For now, hold on to the rule.
Your friend says she "has forgotten" all the French she learned at school, because she can't say anything in it any more. In the Bjorks' terms, which strength is she reporting, and what does the model predict about a month of relearning?
Show the answer
She is reporting low retrieval strength: the French isn't accessible now. If she learned it well and used it for years, storage strength is high, and the model says it doesn't decrease. So relearning should be fast, much faster than learning it the first time, because each retrieval event on a high-storage, low-retrieval memory brings it back quickly.
Nicholas Soderstrom and Robert Bjork reviewed nearly a century of research on this gap and concluded that performance during practice is an unreliable index of learning, and that some study conditions move the two in opposite directions: they make you look better today and remember less later (Soderstrom & Bjork, 2015).3 That sentence is the one to test every study session against.
Why familiarity gets read as knowledge
When you reread a chapter, the second pass is easier than the first. The sentences flow, nothing surprises you, your eyes move faster. That ease is called fluency, and it is a real signal: the text is more familiar than it was. But fluency is produced by recognition. The words on the page trigger the memory; you never had to reconstruct the memory yourself.
Reconstructing a memory from nothing, with no page in front of you, is what builds storage strength most. Recognising it while looking at it builds much less. (Not nothing; rereaders in the experiment below still held on to 42% of a passage a week later. But much less.) So rereading raises retrieval strength and fluency while doing little for storage strength. And because fluency is the only signal you get, the chapter feels learned.
This is not a personal failing; the bias is built into the situation. Asher Koriat and Robert Bjork showed the mechanism with word pairs (Koriat & Bjork, 2005).4 Try it yourself.
You study the pair cats–kittens. At test you will see "cats" and have to produce "kittens". Right now, with both words in front of you, how confident are you that you'll manage it?
Show the answer
Most people say very confident. The two words go together obviously. But look at what happens at test. You see "cats" on its own, and "cats" on its own doesn't lead anywhere in particular: dogs, mice, milk, whiskers. In word-association norms, "kittens" brings "cats" to mind 72% of the time, but "cats" brings "kittens" to mind only 2% of the time. The link that made you confident was only visible because the answer was sitting next to the cue. At test, the answer is the thing you don't have.
Koriat and Bjork built two kinds of pairs: ones where the link runs forward from cue to answer, and ones like cats–kittens where it runs backward. For the forward pairs, people's predictions were almost exactly right (they predicted 78% and recalled 79%). For the backward pairs, they predicted 76% and recalled 60%.4 The direction of the link barely changed how confident people felt and changed a lot how much they remembered. Koriat and Bjork called this foresight bias: your judgement of how well you'll remember is made with information in front of you that won't be there at test, and that information can overstate what you'll get back. They stress the bias is selective. For plenty of items people's judgements are well calibrated. It appears when what you're looking at now creates a sense of connection that the test situation won't supply.
Rereading a chapter is that situation on a large scale. Every sentence is a cue with its answer sitting right next to it.
A review by Bjork, John Dunlosky and Nate Kornell summarises the pattern across many studies: judgements of learning track fluency and current performance, so they are wrong in a predictable direction (Bjork, Dunlosky & Kornell, 2013).5 One study they discuss makes the point with numbers. Jeffrey Karpicke and Janell Blunt had 120 students learn a science text two ways, once by drawing a concept map with the text in front of them, and once by closing the text and writing down what they could recall (Karpicke & Blunt, 2011).6 On a test a week later, 84% of the students did better on the material they had recalled. Yet during the learning session, before any test, 90 of the 120 had predicted that concept mapping would work as well as or better than recall (49% said better, 26% said the same).6 The prediction was made while the text was open and the map was taking shape, which is to say while everything felt fluent.
The bias has consequences beyond feeling good. Dunlosky and Katherine Rawson found that students whose judgements of their own learning were overconfident dropped items from practice sooner and recalled less two days later (Dunlosky & Rawson, 2012).7 The illusion flatters you, and it also decides when you stop.
Why does it matter that Karpicke and Blunt's students made their predictions during the learning session rather than after the test?
Show the answer
Because it shows the illusion at the moment it does damage. While you're studying, the feeling of fluency is the only signal you have about how it's going, and that's when you decide whether to keep going, switch methods, or stop. The students weren't being asked to remember an old result; they were reading their own current experience, and it pointed the wrong way.
Worked example 1: a prediction game
Before you read the results, commit to a prediction. Write it down or say it out loud; it matters that you can't quietly revise it afterwards.
Henry Roediger and Jeffrey Karpicke ran a study at Washington University in St. Louis that has become the founding experiment of modern research on testing (Roediger & Karpicke, 2006).8 Each student read two short prose passages. One passage they then restudied: read it again. The other they closed and wrote down everything they could recall, with no feedback. So every student did both things, on different passages, and the comparison is within the same people.
Later, everyone took a final recall test on both passages. Some took it five minutes after the study session, some two days later, some a week later.
Which condition scores higher at five minutes, and which scores higher at one week? Roughly how big is the gap each time?
Show the answer
Most people predict restudy wins at both delays, because the restudied passage was seen twice and the recalled one only once. The actual results are below. If you predicted the crossover, well done; if you didn't, notice that the intuition you just used is the one this lesson is about.
At five minutes, restudy won: 81% recalled, against 75% for the recall condition.8 So far the intuition holds. Reading twice does produce a better score right now, and if you'd asked students how well they knew each passage, the restudied one would have felt better known.
At two days, the order had flipped: the recalled passage came back at 68%, the restudied one at 54%.8 That is a large effect (d = 0.95, in the units researchers use; Jacob Cohen's rule of thumb, still the one most people quote, calls 0.2 small, 0.5 medium and 0.8 large).9
At one week, the recalled passage scored 56% and the restudied one 42%.8 Restudy had lost nearly half of what it had at five minutes; recall had lost a quarter. The one-week effect size was d = 0.83. The chart below draws all six numbers; the crossover between the two lines is the whole lesson in one picture.
Take the two-strengths model and explain the five-minute result and the one-week result with it. Which strength did rereading raise, and which did recalling raise? Say it in two sentences before you look.
Show the answer
Here's mine. Rereading pushed retrieval strength up fast, which is why it won at five minutes, and did little for storage strength, so by a week most of what it built had drained away. Writing the passage from memory was an attempt to retrieve something with lower retrieval strength (the page was closed), which by the Bjorks' rule is exactly when the storage gain is biggest; it scored worse in the moment and was the only condition that built anything durable.
Roediger and Karpicke's second experiment sharpened this. One group studied a passage four times (study, study, study, study). Another studied it once and then recalled it three times (study, test, test, test). Before reading the result, predict the one-week scores for each. Then check: at one week the repeated testers recalled 61%; the repeated studiers recalled 40%.8 And the repeated studiers were more confident they would remember it. The condition that felt best was the one that worked worst, by 21 percentage points.
So: at five minutes, bet on rereading. At a week, bet on retrieval. If you have to guess which one an exam, a job, or your future self will resemble, it is not the five-minute one.
Worked example 2: the illusion survives the evidence
You might hope that experience would fix the illusion. If your own test showed that method B beat method A, surely you'd switch to B. Here is the harder case.
Nate Kornell and Robert Bjork asked people to learn the painting styles of twelve artists, six paintings each (Kornell & Bjork, 2008).10 The task was to look at a painting they had never seen and say which of the twelve artists painted it: not memorising the paintings, but learning to recognise a style.
For some artists, the six paintings were shown one after another in a block. For others, the paintings were spread out among other artists' work, so a painting by one artist was followed by a painting by a different one. The paper calls these massed and spaced; lesson 5 will call the second one interleaving and explain why it helps. For now the point is what the learners believed about it.
Which order do you think produced better identification of new paintings? And which order do you think the participants said had helped them more?
Show the answer
Spaced won on the test. Massed won the vote. The numbers are below.
In the first experiment, 120 people did both orders, with different artists in each, and got feedback after each test answer ("correct", or the right artist's name). Their accuracy on new paintings was 61% for the spaced artists and 35% for the massed ones, a difference of d = 0.99.10 After the test, they were told what "massed" and "spaced" meant and asked which had helped them learn more. Kornell and Bjork report that 78% of participants did better with spaced presentation, and 78% of participants said massing had been as good or better.10
The second experiment removed the feedback during the test entirely and used a different test format, and the pattern held. Of the 72 participants who didn't answer "about the same", 64 said massing had been more effective.10 Across both experiments, 85% did at least as well with spacing and 83% rated massing as equal or better.
Here is the detail that matters for the question we started with. At no point were participants shown their scores by condition. Even in the first experiment, feedback was item by item, and a follow-up with 28 people found they could not even tell which artists had been massed: they picked at chance. So it is not that people saw a scoreboard and ignored it. Their own test performance had just demonstrated the opposite of what they believed, and the belief never registered it. The result is not a quirk of paintings, either. A 2019 meta-analysis of interleaving studies found an effect of g = 0.67 for paintings, one of the larger effects in the literature (Brunmair & Richter, 2019).11
Why does massing win the vote? Because blocked study feels better while you do it. Seeing six paintings by the same artist in a row makes the style feel obvious; you get fluent quickly. Seeing one by that artist, then one by a second, then one by a third is confusing, and confusion feels like not learning. The feeling is about retrieval strength and ease in the moment. The test result is about what stuck. When the two disagreed, nearly nine in ten people believed the feeling.
You will not be able to feel your way out of the fluency illusion, and you won't reliably reason your way out of it either, even with your own results in your hand. The only defence I know of is to decide in advance to trust delayed tests over feelings, and to build your study around things that measure storage strength: recalling from nothing, after a gap.
In the painters studies, what exactly did participants not have, that you might have assumed they had?
Show the answer
A score per condition. They got item-by-item feedback in the first experiment and none in the second, and they were never shown "you scored X on spaced and Y on massed". A follow-up group couldn't even tell which artists had been massed. So the illusion survived their own performance, but it was never tested against an explicit scoreboard, and the lesson shouldn't claim it was.
What people get wrong
"If it feels easy and fluent, I've learned it." Ease measures retrieval strength now. Bjork and Bjork put it this way: learners mistake current retrieval strength for storage strength (Bjork & Bjork, 2011).1 The painters study shows the feeling doesn't correct itself when your own performance says otherwise. Treat ease during study as information about the present, not a promise about the future.
"Rereading and highlighting are studying." They are the most common study methods, which is part of why the illusion is so widespread. When Karpicke, Andrew Butler and Roediger surveyed 177 undergraduates at a selective university, 84% listed rereading among their strategies and 55% ranked it first; only 11% mentioned self-testing (Karpicke, Butler & Roediger, 2009).12 Dunlosky and colleagues, in the most thorough review of study techniques to date, rated rereading and highlighting as low utility. They also report one study in which highlighting hurt performance on questions that required inference, presumably because it pulls attention to isolated phrases rather than to how ideas connect; they note it needs replicating (Dunlosky et al., 2013).13 The two techniques they rated highest, practice testing and spreading study out over time, are the subject of lessons 3 and 4. Rereading does have a role: use it for what you could not retrieve, after you have tried.
"Self-testing tells you what you know. It doesn't teach you anything." This is the belief that keeps students rereading even when they own flashcards. Kornell and Bjork asked 472 UCLA students why they self-test. Only 18% said it was because they learn more that way; 68% said they did it to find out how well they had learned the material, and 9% didn't self-test at all (Kornell & Bjork, 2007).14 Most people, in other words, use the best learning tool they have as a thermometer. Roediger and Karpicke's students who closed the book weren't measuring anything; they were building storage strength, and it showed a week later.
"Cramming works." It does, for tomorrow. Roediger and Karpicke's five-minute result is a real result: if the test is in an hour, rereading will beat self-testing.8 The mistake is expecting cramming to have bought anything beyond the exam. If you want to pass Friday's quiz and never think about the material again, cram. If you want to still have it in a month, you are asking for storage strength, and cramming does not sell it. Know which strength you are buying.
Retrieval was harder than rereading in Roediger and Karpicke's study, and that difficulty was the point. Bjork and Bjork call conditions like this "desirable difficulties". But they are explicit that a difficulty is desirable only if the learner can overcome it (Bjork & Bjork, 2011).1 Trying to recall a passage you never understood in the first place is not desirable; it's just failure. Researchers who work on cognitive load make the same warning from the other side: for a beginner, difficulty that eats up working memory before anything has been understood is a cost, not an investment, and studying a fully worked solution first is often the better move (Sweller, van Merriënboer & Paas, 2019).15 Lesson 2 covers that view and lesson 6 shows how to use it. Later lessons show which difficulties pay and which are simply hard.
Practice
- Write down the three things you most often do when you "study" or "practise". Be honest and concrete: "reread my notes", "watch the video again", "do the end-of-chapter problems with the solutions open", "explain it to my flatmate".
- For each one, decide: does it mainly raise retrieval strength (it makes the material feel accessible right now), storage strength (it forces you to reconstruct the material from nothing, after a gap), or both? The test is simple: could you do this activity with the source material closed? If not, it is almost certainly a retrieval-strength activity.
- Circle the one habit you would need to change first if you wanted next month's memory rather than tonight's confidence.
Now close this page, or scroll to the top so you can't see the text, and answer these five questions on paper without looking back. Don't skip this. It is the first retrieval attempt in the course, and by the argument above it is doing more for you than a fourth reading would.
- Define storage strength and retrieval strength in one sentence each.
- In Roediger and Karpicke's experiment, which condition won at five minutes and which won at one week? Give the four percentages if you can.
- Why does rereading raise fluency without doing much for storage strength?
- In the painters studies, roughly what share of people did better with spacing, and roughly what share said massing was as good or better?
- Give one example from your own life of a memory with high storage strength and low retrieval strength. It counts if you can't produce it now but would be certain of it the moment you saw it, and it would stay with you afterwards.
Then scroll back and check. Mark what you missed. Those items, and only those, are what you should reread.
Connections
Everything else in this course is a way of building storage strength on purpose, and of not being fooled by retrieval strength in the meantime.
- Lesson 2 explains the machinery: why working memory is small, why experts and novices see different things in the same page, and why "understanding" a lecture in the room is another version of today's illusion.
- Lesson 3 is the fix for the problem you just met. Retrieval practice is one of the two best-supported techniques in the literature (spacing is the other), and the five-minute-versus-one-week result is its founding evidence.
- Lesson 4 explains why the gap matters, and why forgetting a little between sessions is the price of remembering a lot. It's the Bjorks' rule again: the storage gain is biggest when retrieval strength has dropped.
- Lesson 5 returns to Kornell and Bjork's painters to explain what the spacing actually does.
The reason this lesson comes first is that it changes what "did that study session work?" means. From here on, the answer is never how it felt. It is what you can do, cold, later.
Go deeper
- Bjork & Bjork, "Making things hard on yourself, but in a good way" (2011): nine pages, free from the Bjork Learning and Forgetting Lab website; the clearest statement of the two-strengths idea by the people who developed it.
- Brown, Roediger & McDaniel, Make It Stick (2014), chapters 1 and 5: "Learning Is Misunderstood" and "Avoid Illusions of Knowing"; the narrative version of this lesson, written by two of the field's central figures.
- Soderstrom & Bjork, "Learning Versus Performance: An Integrative Review" (2015): for readers who want the full evidence that practice performance and long-term learning can move in opposite directions.
- Carpenter, Pan & Butler, "The science of effective learning with spacing and retrieval practice", Nature Reviews Psychology 1, 496–511 (2022): the current review of the two techniques this lesson points towards, with the boundary conditions the older papers didn't have.
- Dunlosky, "Strengthening the Student Toolbox" (American Educator, 2013): free, plain-English ratings of ten study techniques, so you can check your own audit against the evidence.
Sources
- Bjork, R. A. & Bjork, E. L., "Making things hard on yourself, but in a good way: Creating desirable difficulties to enhance learning", in Psychology and the Real World (Worth, 2011), 56–64. Storage strength versus retrieval strength; learners mistake the second for the first; the gain in storage strength from a study or retrieval event is greater when current retrieval strength is lower; difficulties are desirable only if the learner can meet them.
- Bjork, R. A. & Bjork, E. L., "A new theory of disuse and an old theory of stimulus fluctuation", in Healy, A. F., Kosslyn, S. M. & Shiffrin, R. M. (eds), From Learning Processes to Cognitive Processes: Essays in Honor of William K. Estes, vol. 2 (Erlbaum, 1992), 35–67. The original statement of the two-strengths model, including the assumptions that storage strength does not decrease and that increments in storage strength are a decreasing function of current retrieval strength.
- Soderstrom, N. C. & Bjork, R. A., "Learning Versus Performance: An Integrative Review", Perspectives on Psychological Science 10, 176–199 (2015). Performance during practice is an unreliable index of learning; some manipulations move the two in opposite directions.
- Koriat, A. & Bjork, R. A., "Illusions of competence in monitoring one's knowledge during study", Journal of Experimental Psychology: Learning, Memory, and Cognition 31, 187–194 (2005). Foresight bias: judgements of learning are made with the target present, which can overstate later recall. Example pair cats–kittens (kittens elicits cats .72; cats elicits kittens .02 in Palermo & Jenkins 1964 norms). Experiment 2 (N = 20): forward pairs predicted 78.1% and recalled 78.8%; backward pairs predicted 75.7% and recalled 60.3%. The authors stress the bias is selective, not a general feature of judgements of learning.
- Bjork, R. A., Dunlosky, J. & Kornell, N., "Self-Regulated Learning: Beliefs, Techniques, and Illusions", Annual Review of Psychology 64, 417–444 (2013). Judgements of learning track fluency and current performance and are biased in a predictable direction.
- Karpicke, J. D. & Blunt, J. R., "Retrieval Practice Produces More Learning than Elaborative Studying with Concept Mapping", Science 331, 772–775 (2011). Experiment 2, within-subjects, N = 120: 84% of students scored higher after retrieval practice than after concept mapping on a one-week test; during learning, 90 of 120 (75%) predicted that concept mapping would be as good as (26%) or better than (49%) retrieval. Mappers had the text and an example map in front of them.
- Dunlosky, J. & Rawson, K. A., "Overconfidence produces underachievement: Inaccurate self evaluations undermine students' learning and retention", Learning and Instruction 22, 271–280 (2012). Students who were overconfident in judging their own recall dropped items from practice sooner and recalled less on a delayed test.
- Roediger, H. L. & Karpicke, J. D., "Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention", Psychological Science 17, 249–255 (2006). Washington University in St. Louis. Experiment 1, within-subjects (each student restudied one passage and recalled the other): 5 minutes, restudy 81% vs test 75% (d = 0.52); 2 days, 68% vs 54% (d = 0.95); 1 week, 56% vs 42% (d = 0.83). Experiment 2: study–test–test–test 61% vs study–study–study–study 40% at one week, with repeated studiers more confident.
- Cohen, J., Statistical Power Analysis for the Behavioral Sciences, 2nd ed. (Erlbaum, 1988). The conventional benchmarks for d: 0.2 small, 0.5 medium, 0.8 large.
- Kornell, N. & Bjork, R. A., "Learning Concepts and Categories: Is Spacing the 'Enemy of Induction'?", Psychological Science 19, 585–592 (2008). Twelve painters, six paintings each. Experiment 1a (N = 120, within-participant, item feedback during the test): spaced .61 vs massed .35, d = 0.99; 78% did better with spacing, 78% rated massing as good or better. Experiment 1b (N = 72, between-participants): .59 vs .36, d = 1.28; no preference question. Experiment 2 (N = 80, recognition test, no feedback): of 72 who did not say "about the same", 64 said massing was more effective. Combined 1a and 2: 85% did at least as well spaced, 83% rated massing equal or better. Participants were never shown scores by condition; a 28-person follow-up could not identify which artists had been massed (M = .55, chance).
- Brunmair, M. & Richter, T., "Similarity matters: A meta-analysis of interleaved learning and its moderators", Psychological Bulletin 145, 1029–1052 (2019). Overall g = 0.42 across 59 studies; paintings g = 0.67.
- Karpicke, J. D., Butler, A. C. & Roediger, H. L., "Metacognitive strategies in student learning: Do students practise retrieval when they study on their own?", Memory 17, 471–479 (2009). 177 undergraduates: 84% list rereading, 55% rank it first, 11% self-test.
- Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J. & Willingham, D. T., "Improving Students' Learning With Effective Learning Techniques", Psychological Science in the Public Interest 14, 4–58 (2013). Rereading and highlighting rated low utility; practice testing and distributed practice rated high. The inference finding for highlighting rests on one study (Peterson 1992, Reading Research and Instruction 31, 49–56), which the authors say needs replication.
- Kornell, N. & Bjork, R. A., "The promise and perils of self-regulated study", Psychonomic Bulletin & Review 14, 219–224 (2007). 472 UCLA introductory psychology students: 18% self-test because they learn more that way, 68% to find out how well they have learned, 4% because it is enjoyable, 9% do not self-test.
- Sweller, J., van Merriënboer, J. J. G. & Paas, F., "Cognitive Architecture and Instructional Design: 20 Years Later", Educational Psychology Review 31, 261–292 (2019). Cognitive load theory: working memory is limited when handling new material; instruction for novices should cut extraneous load and use worked examples before problems, fading them as expertise grows.
Check your understanding
This lesson has a 5-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.