How memory actually works

70 min

Listen: this lesson as a conversation

Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.

In this lesson you will learn to
  • Explain why the limits of working memory make a novice and an expert see different things in the same material
  • Apply the idea of chunking to a subject you are learning, by listing the chunks on a page and marking which ones you already own
  • Identify extraneous cognitive load in a piece of instruction and rewrite it to remove that load

Pick a paragraph you found impossible recently. A clause in a lease. A page of a statistics textbook. A block of code someone else wrote. A verse in a commentary that assumed you knew three Greek words. You read it, you reached the end, and you had nothing. Then someone who knows the field glanced at it and said "oh, that's just saying X".

That person doesn't have a better brain than you in any way that matters. They have the same tiny working memory you do. What they have that you don't is a set of pre-built bundles in long-term memory. This lesson is about those bundles: what they are, why they make experts see differently, and what it takes to build one (the next lessons turn that into method). It also explains the most common way instruction and study go wrong, which is overloading the tiny memory before anything reaches the big one.

Two memories, one bottleneck

The textbook picture of memory, and the one cognitive load theory is built on, has two stores.1 (Some models treat working memory as the active part of long-term memory rather than a separate box. For what follows, the difference doesn't matter.)

Working memory is where you think. It holds whatever you're attending to right now, and it's small and brief. In a classic 1959 experiment, people heard three letters and were then stopped from rehearsing them. Within about twenty seconds, most had lost them.15 And there are only a handful of places to put things.

Long-term memory is where everything you know lives, and for practical purposes it has no capacity limit. Nobody has ever run out.

For the kind of learning this course is about, everything has to pass through the small store to reach the large one. That's the bottleneck, and most of what goes wrong in studying goes wrong there. The diagram below is the whole architecture this lesson runs on; keep it in view as you read.

The bottleneck: new material passes through a small working memory to reach an unlimited long-term memory Three boxes stacked top to bottom. New material flows down into working memory, a small box labelled about four chunks of new material, held for seconds. An arrow labelled thinking about meaning carries material down from working memory into long-term memory, a large box labelled everything you know, no practical limit. A return arrow up the right-hand side shows stored chunks coming back to working memory as single units, without spending a slot on each part. Two memories, one bottleneck New material Working memory about 4 chunks of new material, for seconds thinking about meaning Long-term memory everything you know, no practical limit stored chunks return A stored chunk comes back as a single unit, without spending a slot on each of its parts. The two-store picture cognitive load theory is built on (Sweller; capacity estimate from Cowan 2001). Sizes not to scale.

How small is small? You've probably heard the number seven. George Miller's 1956 paper, one of the most cited in psychology and free to read, put the limit at "seven, plus or minus two".2

Predict first

Miller's people were allowed to rehearse and regroup. If you stop them doing that, does the number go up, go down, or stay the same?

Show the answer

Down. Nelson Cowan reviewed the evidence in 2001 and argued that when you prevent rehearsing and regrouping, the limit is three to five, about four on average.3 Four is what you get when you meet material for the first time and have no tricks available. Miller's seven is what you get once people have already found ways to group.

So Miller's number is famous and, for learning, misleading. Four is the number to keep in mind.

Baddeley is the researcher whose 1974 model made "working memory" the standard term, and this five-minute interview is him on exactly the relation this section describes: the small active store, the large permanent one, and why what you already know changes what the small one can hold.

Two honesties about that number. First, "slots" is a metaphor, and a contested one. Memory researchers still argue over whether working memory has a fixed number of discrete places or a single resource that spreads thinner as you add items. For a learner, the two pictures give the same advice, so this lesson uses slots. Second, the limit is a limit on new material. Once something is stored in long-term memory, working memory can pull it in without spending a slot in the same way. The four-item ceiling is a ceiling on what you have never seen before.1

But four what? Not words. Not facts. Chunks.

What a chunk is

A chunk is anything your long-term memory already treats as a single unit. The letters F, B and I are three chunks to a child learning to read and one chunk to you. The phrase "compound interest" is two words and, if you took the finance course, one idea. A working-memory slot holds one chunk, and a chunk can be as big as your knowledge lets it be.

This is why the same four slots make novices and experts see different things. The classic demonstration is chess. It starts with Adriaan de Groot, who found in the 1940s that strong players could reproduce a position from a real game after a glance, and weak players couldn't.4 In 1973 William Chase and Herbert Simon took three people, a master, a club player and a beginner, and showed each of them positions from real games for five seconds at a time. Then they asked them to rebuild the board from memory. On the first attempt the master placed about four-fifths of the pieces correctly, the club player about half, the beginner a third.4

Then came the part that matters. Chase and Simon scattered the same pieces on the board at random and ran the test again.

Predict first

What happened to the master's advantage on the random boards?

Show the answer

It disappeared. In their data there was no relation at all between skill and recall for random positions. (A later review of thirteen such studies found a sliver of it survives, roughly one extra piece per 400 rating points, against about five for real positions, so "largely vanished" is the fair summary.)5 The master didn't have a better memory. He had thousands of stored patterns (a castled king, a fianchettoed bishop, a pawn chain), so a real position was a handful of chunks to him and a random one was as many separate pieces as it was to anyone else.

How many patterns? Simon and Kevin Gilmartin estimated in 1973 that a master carries somewhere between 10,000 and 100,000 of them. The figure that stuck is 50,000, and when Fernand Gobet and Simon re-tested it in the 1990s it held up, with later models using even larger stores.5 Later theories argue the stored structures are bigger and more flexible than simple chunks (Gobet's "templates"; Ericsson and Kintsch's "long-term working memory"). The lesson's point survives all of them: the advantage lives in long-term memory.16

So expertise, at the level of memory, is having bigger chunks. Not more slots. Bigger contents per slot.

Check yourself

A chess master recalls a real position far better than a beginner. Is that because she has more working-memory slots, or because each slot holds more?

Show the answer

Each slot holds more. Her working memory is the same size as yours, about four chunks. What differs is that a real position matches patterns she already has stored in long-term memory, so it takes a few chunks instead of twenty-odd separate pieces. The random-board result is the proof: take away the patterns and her advantage goes with them.

If you've done "Learning How to Learn"

Barbara Oakley's course (with Terrence Sejnowski) draws "chunking" from the same research this lesson uses, but she uses the word for two things at once: the stored unit, and the whole process of building one through focused practice until it comes as one piece. This lesson keeps to the first, narrow sense, Chase and Simon's stored pattern that working memory can treat as one item, because the cognitive-load rules below depend on counting units. Her process sense returns in the later lessons on practice, where it belongs.6

Memory is the residue of thought

If chunks are built in long-term memory, how do things get there? Not, for this kind of learning, by exposure. Daniel Willingham's summary of decades of memory research is one sentence: "memory is the residue of thought".7 You remember what you thought about.

The experiment behind the sentence is Craik and Tulving's from 1975. They showed people a list of words and asked a different question about each one: is it in capitals? does it rhyme with train? would it fit in this sentence? Then came a surprise memory test. The words people had judged for meaning were recognised several times as often as the words they had judged for typeface.11

You might suspect that's just time on task: thinking about meaning takes longer, so the word got more exposure. Craik and Tulving suspected it too, so in a later experiment in the same paper they used a harder non-meaning question, about the word's pattern of vowels and consonants, which took twice as long as the meaning question.

Predict first

Which words were remembered better: the ones judged for meaning, or the ones judged (for twice as long) for vowel pattern?

Show the answer

Meaning still won.11 Same words, same exposure, more time on the non-meaning question, and the meaning-judged words were still remembered better. What differed was what the person was thinking about while looking at them.

The consequence is uncomfortable. You can read a page while thinking about the sound of the words, or that you're on page 40 of 300, or lunch, and you'll remember those things. The meaning leaves no trace because you never processed the meaning. This is why the previous lesson's rereading habit fails: you can reread while thinking about almost nothing, so it deposits almost nothing, while feeling fluent the whole time.

It also gives you a design rule for study, which a 1977 follow-up sharpened: memory is best when the thinking you did at study matches the thinking the test will ask for.11 Whatever you want to remember, arrange to think about it, in the form you want it back. If you want to remember why something is true, spend your thinking on the why. If you want to remember how to do it, spend your thinking on doing it.

Check yourself

Two students spend the same twenty minutes on the same page. Why might one remember far more than the other a week later?

Show the answer

Because what they thought about differed. Memory is the residue of thought: the trace you keep is of whatever you processed, and exposure on its own leaves almost nothing. If one student spent the twenty minutes on the meaning and the other on the typeface, the page number, or a podcast, the first will remember the page and the second will remember the podcast.

Cognitive load: two kinds, one to cut

John Sweller's cognitive load theory turns these limits into rules for instruction and study.1 The load on working memory during learning comes from two sources.

Intrinsic load comes from the material itself. Specifically it comes from what Sweller calls element interactivity: how many elements you have to hold at once because they only make sense together. Learning ten Spanish nouns is low interactivity; each word can be learned alone. Learning to conjugate a verb in a sentence is high interactivity; person, tense, ending and word order all interact, and you can't understand one without the others. Intrinsic load isn't the enemy. It's the content. But it can be managed, by building up the sub-parts first so they become chunks, and by sequencing so that only a few new interacting elements arrive at once.

Extraneous load comes from how the material is presented, and it's pure waste. A diagram on one page and its labels on another, so you shuttle between them. A paragraph that explains a term three sentences after using it. Decorative animation, background music, a lecturer's tangent. All of it occupies slots that aren't then available for the content. The first two are Sweller's split-attention and redundancy effects. In one early experiment, trainees learning to read electrical wiring diagrams learned more from a version with the explanatory text written onto the diagram than from the same text printed beside it, and adding a second, redundant description made things worse rather than better.112 The decoration is Richard Mayer's coherence effect: in his experiments, adding background music and sound effects to a short science lesson lowered both recall and transfer.12 Integrate the parts a learner has to combine, remove what they don't need, and learning improves without changing what is taught.

Here's what that looks like on the page. A pharmacology paragraph, before:

Dose adjustment depends on t½, which for renally excreted drugs varies inversely with GFR (see Table 3, p. 212). t½ is the time for plasma concentration to fall by half. GFR (glomerular filtration rate) is the volume of plasma the kidneys filter per minute.

Count the extraneous load. "t½" is used before it's defined. "GFR" is used, then expanded, then defined: three separate items to hold until they merge. And the table the sentence depends on is forty pages away. After:

Half-life is the time for a drug's blood level to fall by half. For drugs the kidneys remove, half-life lengthens as the kidneys' filtration rate (GFR, the volume filtered per minute) falls, so the dose must fall with it. [Table 3 here.]

Nothing has been removed from what is taught. The elements arrive in the order they're needed, and the table is where the eye already is. That's the test for extraneous load: if the rewrite changes nothing about the content and the passage gets easier, the load was extraneous. If you had to cut content to make it fit, the load was intrinsic, and the fix is to build chunks, not to edit.

There used to be a third category, germane load, for the effort of actually building chunks. Critics pointed out that no experiment could measure it separately from the other two, so almost any result could be explained after the fact by moving load between categories. In 2010 Sweller redefined it as the share of working memory you devote to intrinsic load, rather than a separate kind. That answers part of the objection and simplifies the rule: cut extraneous load so that the whole of your four slots goes to the intrinsic load of the thing you're learning.1

Check yourself

A tutorial is hard to follow. How do you tell whether the difficulty is intrinsic or extraneous?

Show the answer

Rewrite it and see what you had to change. If you can make it easier without touching the content (define terms before using them, put labels on the diagram, drop the music), the load was extraneous. If the only way to make it fit was to cut content, the load was intrinsic, and the fix isn't editing; it's building the sub-parts into chunks first.

The mechanism: why prior knowledge does the work

Put the three ideas together and a mechanism appears.

Material that overflows working memory can't be assembled into a chunk, because assembling a chunk means holding all its parts together at once. So if a paragraph contains twelve interacting new elements, no amount of rereading will store it. There's no moment when the whole thing is in mind.

What rescues you is what you already know. If eight of those twelve elements are already chunks, the paragraph needs four slots, it fits, and now it can be built into a new, larger chunk that next week will take one slot. This is the engine of expertise, and it runs on prior knowledge. Willingham draws two conclusions from it.7

First, knowledge of a domain comes before skill in that domain. You can't reason critically about a text or a case or a dataset whose elements are all new to you, because reasoning needs working memory and yours is full just holding the elements. That's the cognitive claim, and it's well supported. How far it should dictate the order of teaching, facts first or knowledge built up through inquiry, is a live and partly political argument in education (the "knowledge-rich" camp against the inquiry and constructivist camps), and this lesson takes no side on it. The narrower point stands whichever route you take: the reasoning runs on stored chunks, so if you don't have them yet, building them is the first job.

Second, abstractions are understood through what you already know. A new idea gets in by attaching to an old one. This is why good explanations reach for analogies and concrete cases, and why an explanation that is perfectly precise but connects to nothing you own tends to bounce off.

So when something feels impossible, the problem is usually not that it's hard in the abstract. It's that you're missing the chunks it assumes. Find and build those; don't stare harder.

Check yourself

A paragraph has twelve new elements that only make sense together. Why won't rereading it ten times store it?

Show the answer

Because a chunk is built by holding all its parts in working memory at once, and twelve is more than four. On every reread the paragraph overflows, so there's never a moment when the whole thing is in mind to be assembled. The fix is to shrink the count: build some of the twelve into chunks separately, until what's left fits.

Worked example: one clause, two readers

Here's a sentence from a commercial contract. Read it once at normal speed.

Notwithstanding the foregoing, the indemnifying party shall have no obligation under this Section to the extent the claim arises from the indemnified party's gross negligence or wilful misconduct.

Now count what someone meeting this for the first time has to hold in working memory. I'll start. "Notwithstanding the foregoing": an unfamiliar phrase, so its two words plus a guess at what it does. "Indemnifying party" and "indemnified party": two near-identical terms that must be kept apart, each needing its own definition looked up and held.

Check yourself

Stop there and do the rest yourself. For each remaining phrase ("no obligation under this Section", "to the extent", "gross negligence", "wilful misconduct"), write down what a first-time reader has to hold, then total the count.

Show the answer

Here's my count. "No obligation under this Section": which Section, what obligation? "To the extent": a proportional qualifier that changes the whole meaning. "Gross negligence" and "wilful misconduct": two legal standards, each a chunk you don't have yet. That's nine or ten interacting elements, depending how you count. Working memory holds four. The sentence doesn't fit, so it can't be chunked, so you reach the full stop with nothing.

A second-year law student, marking the same list, would put "indemnifying" and "indemnified" down as own, "gross negligence" as half (she recognises it but couldn't state the test), and "notwithstanding the foregoing" as new. Three items over the limit, and now she knows which three to build.

The lawyer reads the same sentence as three chunks. Carve-out ("notwithstanding the foregoing... to the extent"): this sentence is an exception to the rule above it, sized to the fault. Indemnity direction (indemnifying versus indemnified): a pattern she has seen hundreds of times, so which is which is automatic. Fault standard (gross negligence or wilful misconduct): the standard high bar, as opposed to ordinary negligence. She knows immediately that this is the version that favours the indemnified party, because the exception only bites when that party was seriously at fault; the other side's lawyer would have pushed for plain "negligence". Three chunks, one slot to spare, and she has read it in the time it took you to find "notwithstanding".

Same sentence, same four slots. Her advantage is entirely in long-term memory. And notice what that says about learning to read contracts. You don't get there by reading more of them faster. You get there by building the three chunks (carve-out, indemnity direction, fault standards) one at a time, until a sentence like this fits.

Swap the domain and the structure is identical. Take one line of Python:

top = sorted(rows, key=lambda r: r[1])[-3:]

To a beginner that's eight or nine separate things: sorted makes a new list; key= says what to sort by; lambda is a nameless function; r is each row; r[1] is its second column; [-3:] is a slice; the minus counts from the end; the whole thing is assigned to top. To a working programmer it's two: "sort by the second column" and "take the last three". A verse of Romans is a wall of clauses to a first reader and one argument with a hinge to someone who knows Paul's long chained sentences.

A harder example: the lecture that made sense

Now the wrinkle, because the mechanism has a trap in it.

You sit in a lecture on, say, how interest rates affect exchange rates. It makes complete sense. The lecturer builds it step by step, every step follows, you nod. The next evening you sit down with the problem set and can't begin. You're baffled and a little alarmed: you understood it.

What happened is that in the lecture, the lecturer's working memory was doing the assembling. She held the chunks and put them together in the right order, out loud. Yours only had to follow one step at a time, which four slots can do. Following felt like understanding, and in the moment it was a kind of understanding. But following doesn't build a chunk. Alone, you have to hold all the steps together yourself to see how they connect, and you have twelve elements and four slots. The chain that was assembled for you can't be reassembled by you.

This is the previous lesson's fluency illusion, seen from the inside of the machine. Retrieval strength was high in the room because the cues were all present. Storage strength was never built because you never did the assembling.

It also explains something about notes. Notes that copy the slides don't help much, because copying needs almost no thought about the meaning (residue of thought again) and records the surface rather than the links. Notes that reconstruct do help. Close the slides and write, in your own words, why step three follows from step two, and writing that sentence forces you to hold two and three together, which is the act of chunking. The evidence here needs stating carefully. The one popular claim the note-taking literature doesn't support is that handwriting beats typing: Mueller and Oppenheimer's 2014 result failed to replicate in two later studies.14 What it does support is Chi and colleagues' 1989 finding: the students who learned physics from worked examples were the ones who explained each step to themselves as they went, and the ones who did badly reread and copied.14 The medium isn't the point. The thinking you do is the memory you get.

Check yourself

Everything in the lecture made sense, and the next day you can't start the problems. What did the lecture give you, and what didn't it give you?

Show the answer

It gave you the experience of following: each step fit in your four slots as the lecturer supplied the links, so retrieval was easy while the cues were in front of you. It didn't give you the chunk, because you never held the steps together yourself. That assembling is the work that builds storage strength, and nobody can do it for you.

What people get wrong

"Smart people have bigger working memories." People do differ in working memory capacity. The differences are real, and they matter: capacity is one of the better predictors of reasoning and reading comprehension.13 But in slot terms the spread is roughly three to five, and decades of training studies show you can improve on the trained task and barely anywhere else. Two large meta-analyses of working-memory training found no reliable transfer to reasoning, reading or arithmetic.13 Chunk size, by contrast, differs between novice and expert by orders of magnitude and is entirely trainable. When someone reads a page in their field effortlessly that defeats you, the difference is mostly their chunks. The lever you can pull is knowledge.

"I can multitask." Working memory is the thing you're trying to multitask with, and it has four slots. A review by May and Elder (2018) of media multitasking, meaning phones, laptops and messaging during study or class, found negative associations with recall, comprehension and grades in nearly every study it covered, in lectures and in private study alike.8 Most of that evidence is correlational, and the exceptions are instructive: messaging about the lecture hurt far less than messaging about something else, and students who could re-read at their own pace recovered comprehension, at the cost of time. But experiments that assign students to multitask during a lecture find the same direction. In one, students told to do unrelated tasks on a laptop scored lower on a comprehension test afterwards, and so did the students sitting in view of their screens.8 The mechanism is the one this lesson has been building: a notification you attend to takes a slot, and the chunk you were assembling has one fewer part in mind.

"Novices and experts think in all the same ways." This appears, in those words, in the misconceptions table of The Science of Learning, a research summary written for teacher-training programmes with Willingham as its scientific adviser.9 It's the one most likely to hurt you, because it suggests you should imitate what experts do. Experts skim, skip steps, and think in abstractions, because they have chunks for the steps and the abstractions attach to their knowledge. A novice who skims, skips and reaches for the abstraction first is running expert habits on novice memory, and overflows. Willingham is blunt: novices can't shortcut to expert thinking; the chunks come first.7

"Chunking means memorising the terms." A list of names is not a chunk. Priya (in the quiz) could recite "nucleophile, electrophile, leaving group, transition state" and still have four separate items, because a chunk is the relations between parts held as one thing: that a good leaving group makes the reaction faster, that hindrance around the carbon slows it. Memorising the vocabulary is a fair first step, since it turns each term into a single item instead of a definition you have to hold. But the chunk is built when the parts are held together and the connection is thought about, which is why the exercise below asks you to explain, not to list.

Practice

Do it now

Take one page, or one screen, of something you're actually learning right now: another course, or your own work. If you have nothing else to hand, use lesson 1 of this course (storage strength, retrieval strength, the fluency illusion, the five-minute versus one-week result); you'll find out quickly which of those you own.

  1. List the chunks. Go through the page and write down every element a first-time reader would have to hold: the things it assumes you already know and the things it's introducing. Every term, formula, step, name, or move. Don't summarise; itemise.
  2. Mark each one as own (you could explain it to someone without looking), half (you recognise it but couldn't explain it), or new.
  3. Count the new and half ones in the hardest paragraph. If it's more than four, you now know why that paragraph felt impossible, and you know what to do: build those chunks separately, one at a time, before rereading the paragraph. To build one: look it up, write its parts and how they relate, close the source, and explain it in one sentence from memory. That act of holding the parts together is the chunking. Do this for the two most-used. If the count is four or fewer, the paragraph fits, and the problem wasn't load; it was that you read it without thinking about it. The fix is step 4, or a why-question at every sentence.
  4. Cut extraneous load. Find one explanation on the page that confused you. Rewrite it so that every term is defined before it's used, the diagram (if any) has its labels on it, and anything decorative is gone. Compare it with the original. Apply the test from above: if nothing about the content changed and it got easier, the load was extraneous. If you had to drop content to make it fit, it was intrinsic, and go back to step 3.

Keep the chunk list. The new and half items on it are the first things you'll self-test in lesson 3.

Connections

The previous lesson established that what you can do today is a poor guide to what you'll retain, and that fluency fools you. This lesson gives the machinery underneath: fluency is working memory following along, retention is chunks in long-term memory, and they're built by different activities.

Lesson 3 follows directly. If memory is the residue of thought, the question is which kind of thinking leaves the most residue. The answer, which has the strongest evidence in the whole field, is reconstructing from memory rather than recognising on the page.

Lesson 6 is where cognitive load theory pays off in method. Worked examples work for novices precisely because they hand you the chunks instead of making you find them with a full working memory, and they stop working once you have the chunks, which is the expertise-reversal effect.

And from here on you have a diagnostic for any dense page: if a paragraph won't go in, count its chunks before you blame yourself.

Go deeper

  • *Willingham, Why Don't Students Like School?, chapters 2 to 4.* The plainest account in print of why reasoning in a domain depends on knowledge of that domain, why memory is the residue of thought, and why abstractions are hard; each chapter ends with what to do about it.
  • Sweller, van Merriënboer & Paas, "Cognitive Architecture and Instructional Design: 20 Years Later" (2019). The authors' own summary of cognitive load theory, its effects, and what changed; readable without a psychology background.
  • Chase & Simon, "Perception in chess" (1973). The original chunking experiment. Twenty-seven pages, but the random-position section alone is a lesson in experimental design.
  • Deans for Impact, The Science of Learning. A free PDF; its one-page misconceptions table is worth pinning above your desk.

Sources

  1. Sweller, J., van Merriënboer, J. J. G. & Paas, F. (2019). "Cognitive Architecture and Instructional Design: 20 Years Later." Educational Psychology Review 31, 261–292. Sweller, J. (2010). "Element interactivity and intrinsic, extraneous, and germane cognitive load." Educational Psychology Review 22, 123–138. Also Sweller, Ayres & Kalyuga, Cognitive Load Theory (Springer, 2011). Supports: the two-store architecture with a capacity- and duration-limited working memory for novel information and an effectively unlimited long-term memory; intrinsic load as element interactivity; extraneous load from presentation; the split-attention and redundancy effects; the 2010 redefinition of germane load as working-memory resources devoted to intrinsic load, and the measurement criticism that preceded it.
  2. Miller, G. A. (1956). "The magical number seven, plus or minus two: Some limits on our capacity for processing information." Psychological Review 63(2), 81–97. Supports: the original seven-plus-or-minus-two estimate.
  3. Cowan, N. (2001). "The magical number 4 in short-term memory: A reconsideration of mental storage capacity." Behavioral and Brain Sciences 24(1), 87–114. Supports: the revised estimate of three to five chunks, averaging about four, when rehearsal and grouping are prevented.
  4. Chase, W. G. & Simon, H. A. (1973). "Perception in chess." Cognitive Psychology 4, 55–81. de Groot, A. D. (1965). Thought and Choice in Chess (Mouton; Dutch original 1946). Supports: de Groot's original finding of skilled recall for game positions; Chase & Simon's three-subject, five-second design; first-trial recall of about 81%, 49% and 33%; no relation between skill and recall for random positions.
  5. Simon, H. A. & Gilmartin, K. (1973). "A simulation of memory for chess positions." Cognitive Psychology 5, 29–46. Gobet, F. & Simon, H. A. (1996). "Recall of rapidly presented random chess positions is a function of skill." Psychonomic Bulletin & Review 3, 159–163. Gobet, F. & Simon, H. A. (1998). "Expert chess memory: revisiting the chunking hypothesis." Memory 6, 225–255. Supports: the 10,000–100,000 range and the 50,000 midpoint; the small residual skill effect on random positions (about one piece per 400 rating points); the 1990s re-tests upholding the estimate, with later models using larger stores.
  6. Oakley, B. (2014). A Mind for Numbers, chapter 4 ("Chunking and Avoiding Illusions of Competence"); and the UC San Diego / Coursera course "Learning How to Learn" (Oakley & Sejnowski). Supports: Oakley's dual use of "chunking" for the stored unit and the process of building it.
  7. Willingham, D. T. (2009). Why Don't Students Like School? (Jossey-Bass; 2nd ed. 2021). Chapter numbers follow the first edition: chapter 2 (factual knowledge precedes skill), chapter 3 ("memory is the residue of thought"), chapter 4 (abstractions are understood via what you already know), chapter 6 (novices cannot shortcut to expert thinking). Supports: all claims attributed to Willingham above.
  8. May, K. E. & Elder, A. D. (2018). "Efficient, helpful, or distracting? A literature review of media multitasking in relation to academic performance." International Journal of Educational Technology in Higher Education 15, article 13. Sana, F., Weston, T. & Cepeda, N. J. (2013). "Laptop multitasking hinders classroom learning for both users and nearby peers." Computers & Education 62, 24–31. Supports: negative associations between media multitasking and recall, comprehension and grades, with the on-task and self-paced caveats; the experimental result for multitaskers and their neighbours.
  9. Deans for Impact (2015; 2nd ed. 2026). The Science of Learning. First edition developed with Daniel Willingham and Paul Bruno; second edition with Veronica Yan. Supports: the misconceptions table, including "Novices and experts think in all the same ways."
  10. Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J. & Willingham, D. T. (2013). "Improving Students' Learning With Effective Learning Techniques." Psychological Science in the Public Interest 14(1), 4–58. Supports: the low-utility rating for highlighting and summarisation cited in the quiz.
  11. Craik, F. I. M. & Tulving, E. (1975). "Depth of processing and the retention of words in episodic memory." Journal of Experimental Psychology: General 104(3), 268–294. Morris, C. D., Bransford, J. D. & Franks, J. J. (1977). "Levels of processing versus transfer appropriate processing." Journal of Verbal Learning and Verbal Behavior 16, 519–533. Supports: the levels-of-processing result (meaning judgements remembered far better than typeface judgements, not explained by time on task); the transfer-appropriate-processing rule that study thinking should match the form of the test.
  12. Chandler, P. & Sweller, J. (1991). "Cognitive load theory and the format of instruction." Cognition and Instruction 8(4), 293–332. Moreno, R. & Mayer, R. E. (2000). "A coherence effect in multimedia learning: The case for minimizing irrelevant sounds in the design of multimedia instructional messages." Journal of Educational Psychology 92(1), 117–125. Supports: the integrated-format (split-attention) and redundancy results with electrical-diagram materials; the coherence effect of background music and sounds on recall and transfer.
  13. Shipstead, Z., Harrison, T. L. & Engle, R. W. (2016). "Working memory capacity and fluid intelligence: Maintenance and disengagement." Perspectives on Psychological Science 11(6), 771–799. Melby-Lervåg, M. & Hulme, C. (2013). "Is working memory training effective? A meta-analytic review." Developmental Psychology 49(2), 270–291. Melby-Lervåg, M., Redick, T. S. & Hulme, C. (2016). "Working memory training does not improve performance on measures of intelligence or other measures of 'far transfer': Evidence from a meta-analytic review." Perspectives on Psychological Science 11(4), 512–534. Simons, D. J. et al. (2016). "Do 'brain-training' programs work?" Psychological Science in the Public Interest 17(3), 103–186. Supports: working-memory capacity as a strong correlate of reasoning; near-transfer-only results of working-memory training.
  14. Chi, M. T. H., Bassok, M., Lewis, M. W., Reimann, P. & Glaser, R. (1989). "Self-explanations: How students study and use examples in learning to solve problems." Cognitive Science 13, 145–182. Mueller, P. A. & Oppenheimer, D. M. (2014). "The pen is mightier than the keyboard." Psychological Science 25(6), 1159–1168; non-replications in Morehead, K., Dunlosky, J. & Rawson, K. A. (2019), Educational Psychology Review 31, 753–780, and Urry, H. L. et al. (2021), Psychological Science 32(3), 326–339. Supports: self-explainers versus rereaders and copiers; the failure of the handwriting-beats-typing result to replicate.
  15. Peterson, L. R. & Peterson, M. J. (1959). "Short-term retention of individual verbal items." Journal of Experimental Psychology 58(3), 193–198. Supports: loss of unrehearsed items within about twenty seconds.
  16. Gobet, F. & Simon, H. A. (1996). "Templates in chess memory: A mechanism for recalling several boards." Cognitive Psychology 31, 1–40. Ericsson, K. A. & Kintsch, W. (1995). "Long-term working memory." Psychological Review 102(2), 211–245. Supports: the template and long-term-working-memory accounts as successors to simple chunk theory.

Check your understanding

This lesson has a 6-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.