Spacing: when to come back

65 min

Listen: this lesson as a conversation

Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.

In this lesson you will learn to
  • Calculate a sensible review gap from how long you need to remember something
  • Explain why the same study time spread over gaps beats the same time in one block
  • Set up a spaced review schedule for a real goal, including one measured in years

You reviewed something yesterday and it felt solid. You'll review it again tomorrow. That plan feels responsible, and it's one of the least efficient ways to spend study time the research has measured.

The question this lesson answers isn't whether to review (you should) but when. Come back too soon and the review is easy, pleasant, and nearly worthless. Come back at the right moment and the same ten minutes does several times the work. There's a rule of thumb for finding that moment, there's a study with over a thousand participants behind it, and by the end of this lesson you'll have used it on something you actually need to remember.

The core idea

Take a fixed amount of study time, say one hour on a set of material. Spend it in one block and you've got massed practice. Split it into two or more sessions with a gap between them and you've got spaced (or distributed) practice. At a delayed test, spaced wins, and it wins at the same total time.

This is one of the oldest findings in experimental psychology and one of the best replicated. Cepeda, Pashler, Vul, Wixted and Rohrer (2006) reviewed 184 articles containing 317 experiments and 839 separate assessments of the effect in verbal recall tasks. Spacing reliably beat massing, and the review found something more useful than the headline: the best gap isn't a fixed number. It grows with how long you need to remember the material.1 Dunlosky and colleagues (2013), rating ten study techniques for their monograph, put distributed practice in the top tier alongside practice testing, the only two techniques to earn a "high utility" rating.2

One distinction is worth having before we go further. "Spacing" is really two findings. The first is that any gap beats no gap. The second, sometimes called the lag effect, is that the size of the gap matters too. Most of this lesson is about the second, because that's the one you have to make a decision about.

How big should the gap be?

The gap-sizing numbers come from Cepeda, Vul, Rohrer, Wixted and Pashler (2008). They ran more than 1,350 participants through a study with two learning sessions, a gap between them of anywhere up to 3.5 months, and a final test up to a year later. Then they plotted the best gap against the test delay.3

Predict first

Before you see the numbers: for a test one week away, the best gap turned out to be one day. For a test a year away, should the best gap be a bigger or a smaller share of the horizon than that?

Show the answer

Smaller. The best gap for a one-year test was about three weeks, which is 6% of the horizon, against one day out of seven, which is 14%. The best gap grows as the horizon grows, but nowhere near in proportion. The authors called the shape a "temporal ridgeline".

Here are the four horizons they tested and the gap that produced the best recall at each.

Time until the test Best gap observed Gap as a share of the horizon
7 days 1 day about 14%
35 days 11 days about 31%
70 days 21 days about 30%
350 days 21 days about 6%

The fitted curve put the one-year optimum at 23 days, or 7% of the horizon.3 The same four rows as a picture:

The best first gap grows with the horizon, but not in proportion A bar chart with four bars. For a test 7 days away the best gap was 1 day, about 14 percent of the horizon. For 35 days, 11 days, about 31 percent. For 70 days, 21 days, about 30 percent. For 350 days, 21 days, about 6 percent. The bars rise steeply and then level off. Best gap between the two study sessions 1 day 11 days 21 days 21 days 7 days 35 days 70 days 350 days (14%) (31%) (30%) (6%) Time until the test, with the gap as a share of it Cepeda and colleagues (2008), best gap tested per horizon.

Two things to take from the table. The best gap gets longer as the horizon lengthens, from a day to three weeks. And the share isn't constant: it's about a seventh for a week, closer to a third for a month or two, and back down to a sixteenth for a year. There is no single percentage that fits every row.

That's why the working rule is a rule of thumb and not a formula. Carpenter, Cepeda, Rohrer, Kang and Pashler (2012), reviewing this and related work for teachers, suggested a gap of roughly 10–20% of the time until you need the material.4 Against the table, that rule runs a little short for the middle horizons (it would give 3 to 7 days for a 35-day test, where 11 was best). That doesn't much matter, for a reason that's the most useful single fact in the study.

The curve is lopsided. In Cepeda's data, a gap that was too short cost far more recall than a gap that was too long by the same amount.3 Miss on the short side and you've done a fluent, low-value review. Miss on the long side and you've done a harder review that still mostly works. So when you're torn between two gaps, take the longer one.

Check yourself

You have a test in 30 days. The rule of thumb says 3 to 6 days; Cepeda's nearest data point (35 days) says 11. You're deciding between 4 days and 10. Which way should you err, and why?

Show the answer

Ten. Both are inside the evidence, but the cost of undershooting was much larger than the cost of overshooting, so the safer mistake is the longer gap. Four days is not wrong; it's just the more expensive way to be slightly off.

Three honest caveats. The ridgeline study used two study sessions and one test; real study involves many reviews, so the ratio is a guide for the first gap rather than a formula for all of them. Nearly all of this evidence is from verbal material: word pairs, facts, vocabulary, trivia. It's a reasonable bet that the pattern holds for other kinds of learning, but the percentages haven't been mapped for, say, a motor skill or a maths procedure. And most of it was collected in the lab or online with volunteers, not in classrooms with a syllabus and a term to get through; the classroom studies that exist point the same way, but there are fewer of them and the effects are noisier.

Spacing plus retrieval

The spacing effect was originally studied with restudy: read the material, wait, read it again. That works. Spaced rereading beats massed rereading, and most of the experiments in Cepeda's 2006 review were exactly that.1 But you already know from lesson 3 that retrieval beats rereading at every gap. Put the two together and you get spaced retrieval practice, which Latimier, Peyre and Ramus (2021) examined across 29 studies. Spaced retrieval beat massed retrieval with an effect size of g = 0.74, large by the field's conventions.5 So the ordering is: spaced retrieval, then spaced rereading, then massed anything. Every review in the schedules below starts with the book closed for that reason.

The same review compared expanding schedules (gaps that grow) with uniform ones (gaps that stay the same) and found them about equal. We'll come back to that under misconceptions, because a lot of people believe otherwise.

Kang (2016) is the accessible policy summary of all this. If you want one short paper to put in front of a teacher or a manager, that's the one.6

The mechanism: forgetting is the price

Why should a gap help? Nothing happens during the gap except forgetting, and forgetting sounds like the enemy.

Seven minutes of Bjork on why forgetting is not the enemy of learning but part of its machinery, which is exactly the claim this section has to earn. Watch it after reading the section and notice that his account is the storage-and-retrieval model from lesson 1 doing the work.

Here's the account that fits the most data. Recall lesson 1's distinction between storage strength (how entrenched a memory is) and retrieval strength (how accessible it is right now). Immediately after study, retrieval strength is high. Review then, and the memory is simply recognised: "yes, I know this." Recognition is easy, fluent, and does almost nothing to storage strength. It's the reread-and-nod experience that feels like learning and isn't.

Wait a few days and retrieval strength drops. Now the review is different. You have to reconstruct the memory: search for it, rebuild it from cues. That reconstruction is retrieval practice, the act lesson 3 showed to be the strongest tool in the literature. So a spaced review isn't "the same review, later." It's a harder review, and the difficulty is the kind Robert Bjork (1994) called desirable: effortful reconstruction that raises storage strength.7 Researchers call this the study-phase retrieval account, and it's the one this lesson leans on.

It isn't the only account, and I'd be misleading you if I said the question was closed. Three rivals are still in play. Encoding variability says two sessions in different contexts (different room, mood, time of day) attach more cues to the memory, so there are more routes back to it. Deficient processing says you simply pay less attention to something that feels already known, so an immediate review is half-attended. And consolidation says the brain does work on a memory between sessions, some of it during sleep, so the second session builds on a better foundation. Cepeda's 2006 review looked at all of these against the data and concluded that no single theory accounts for everything.1 They aren't mutually exclusive, and the practical advice is the same whichever wins. But if a friend explains spacing to you in terms of sleep or context, they're not wrong; they're backing a different horse in an unfinished race.

Check yourself

A friend reviews a chapter the same evening she learned it and gets everything right. On the study-phase retrieval account, how much has that review done for her?

Show the answer

Not much. Retrieval strength was still high, so nothing had to be reconstructed; the review was recognition, which barely moves storage strength. On the deficient-processing account the answer is similar, for a slightly different reason: the material felt known, so she processed it shallowly. Either way, the evening review was pleasant and cheap and mostly wasted.

The retrieval account also explains why the best gap scales with the horizon. Too short, and nothing has faded, so there's nothing to reconstruct. Too long, and too much has faded: you fail to retrieve and have to relearn from scratch, which wastes the original session. The sweet spot is where the memory is fragile but recoverable. A one-week test needs a modest memory, so a short gap gets you there. A one-year test needs a memory that survives a long stretch, and the wider gaps that build it would be too wide for the one-week case.

Why cramming feels right and works badly

Cramming maximises retrieval strength for tomorrow. Every pass is fluent, the material feels owned, and the exam the next morning may go fine. But nothing forced reconstruction, so storage strength barely moved, and the material fades faster afterwards than spaced material does. Cramming isn't stupid; it's optimising for the wrong test.

Worked example: an exam in 30 days

You have an exam in 30 days and a chapter to learn today.

Step 1: the horizon. Thirty days from first study to the moment you need it.

Step 2: the first gap. The rule of thumb gives 10–20% of 30 days, which is 3 to 6 days. The nearest row in Cepeda's table is the 35-day horizon, where 11 days was best. So anything from about 4 to 10 days is inside the evidence, and the lopsided curve says lean long. Call it 6 days.

Step 3: the criterion for each session. Rawson and Dunlosky (2011) gave 533 students a practical rule they called successive relearning: in the first session, keep going until you've recalled each item correctly three times; in each later session, keep going until you've recalled each item correctly once.8 That turns "review" from a vague intention into a stopping rule. Some days that takes five minutes; on a bad day it takes twenty.

Predict first

Step 4 is yours before you read on. With a 6-day gap and a 30-day horizon, what dates go on the calendar, and where does the last review fall?

Show the answer

Days 6, 12, 18 and 24, with a final review on day 28 or so, two days before the exam. Five reviews. If you put the final one on the day before the exam instead, nothing is wrong with that; the point is that the last gap is short because the horizon is now short.

Day What you do
0 Learn the chapter. Write 10–15 questions on it before you close the book; you'll use these at every review. Keep going until you've answered each one correctly three times.
6 First review: answer your questions from memory. Expect to miss some; that's the point. Relearn the misses until every item comes back once.
12 Second review, same criterion. Fewer misses.
18 Third review.
24 Fourth review.
28 Final review, two days before the exam.
30 Exam.

Now the time. Say the first session took 90 minutes and each review takes about 20, so the whole plan is a little over three hours. Compare the schedule most people run: the same 90 minutes today, nothing for 27 days, then the same 100 minutes across the last two evenings. The total time is the same; only where the second half of it sits has changed. The spaced version has forced you to rebuild the material five times; the crammed version rebuilds it once, on the night before, when it's too late to fix what you can't recall.

A variation. If you'd rather expand the gaps (reviews on days 2, 5, 10, 17 and 27), you can. Latimier's review says you'll do about as well, and there's a small hint in the data that expansion pays slightly more when there are many reviews.5 What you shouldn't do is compress the front end into daily reviews "to make sure it's in." A one-day gap was optimal in Cepeda's study for a one-week horizon and for nothing longer.

When you have many chapters, the same rule applies to each, and the reviews interleave naturally: by week three you're reviewing chapter 1 for the third time on the same day you review chapter 4 for the first. That's the schedule a spaced-repetition app builds for you. Here you've built it by hand and know why it looks the way it does.

A harder example: a language for life

Now suppose the goal isn't an exam but a language you want to keep for the rest of your life. The horizon isn't 30 days; it's decades. Cepeda's data stop at one year, so extending the rule past that is an extrapolation, and I want to be plain that it is one. But two findings from Harry Bahrick's long-term studies give you something to extrapolate with.

Predict first

Bahrick and his family relearned foreign vocabulary over nine years on two schedules: thirteen sessions 56 days apart, or twenty-six sessions 14 days apart. Which schedule do you think produced better retention years later?

Show the answer

Neither, really: they came out comparable. Half the sessions, four times the gap, about the same memory. The wider schedule was a little slower to get the words in at first, and that's the only cost the study found.

The nine-year study. Bahrick, Bahrick, Bahrick and Bahrick (1993) tracked foreign-language vocabulary relearning over nine years and compared schedules directly. Thirteen relearning sessions 56 days apart produced retention comparable to twenty-six sessions 14 days apart, with the wider gaps slowing the initial acquisition slightly.9 Read that again slowly: half as many sessions, each four times further apart, and about the same result. There's a catch you should know about before you build your life around it. The study had four participants, and they were the four authors: Harry Bahrick, his wife and their two children. It's a remarkable piece of self-experimentation and it fits everything else in the literature, but it's four people, not four hundred.

Permastore. Bahrick (1984) tested 733 people on Spanish they had learned at school, some of them fifty years earlier. This was a cross-sectional study: different people at different distances from their last class, not the same people followed for fifty years. Retention fell over the first 3–6 years after study stopped, then stayed roughly unchanged for up to 30 years before a final decline in old age.10 He called the stable remainder "permastore". What matters for planning is the shape: the loss is front-loaded. What you still have six years out, you'll probably have at thirty-six.

There's a second finding in that paper that cuts against the easy story. How much reached permastore was predicted mainly by how much you'd learned in the first place; rehearsal after leaving school had little measurable effect.10 So the claim "relearn it in the first few years and it will reach permastore" is not something Bahrick showed. It's an inference, and a reasonable one: relearning in the steep years raises the amount you have, and the amount you have is what the plateau preserves. But hold it as an inference, not a finding.

Put the two studies together and you get a plan that looks nothing like the exam plan:

  1. Years 1 and 2: build with wide gaps. After initial learning, schedule relearning sessions at gaps of one to two months, not one to two weeks. Each session is retrieval: produce the vocabulary and structures from memory, relearn the misses. Bahrick's 56-day schedule is a sensible template.
  2. Years 3 to 6: keep going. This is the rest of the steep part of the curve. Sessions can thin out, but they shouldn't stop, because whatever you hold at the end of this stretch is roughly what you'll hold for decades.
  3. After year 6: use it. What survived is stable without a schedule; Bahrick found that rehearsal did little to what was already in permastore. Ordinary use will do.
Check yourself

Why is the deliberate schedule concentrated in the first six years and not spread evenly across the decades?

Show the answer

Because that's where the loss happens. Bahrick's curve falls over the first 3–6 years and then holds for up to thirty. A relearning session in year two is fighting a decline that's actually under way; a session in year twenty is maintaining something the data say was going to stay anyway.

The wrinkle, then, is that the exam learner and the lifelong learner are using the same rule and getting opposite-looking advice: review every six days, or review every eight weeks. Neither is "the right gap." The horizon sets the gap.

What "I've forgotten it all" usually means

Adults who say they've lost their school French usually haven't lost what reached permastore; they've lost retrieval strength for it. Bahrick's curves are why. If you learned it properly, more is there than it feels like.

What people get wrong

"Cramming is efficient." It's efficient at one thing: passing a test within a day of studying. Every experiment in Cepeda's review that held total time constant found the massed learner behind at a delayed test.1 And it doesn't feel that way from the inside. Bjork, Dunlosky and Kornell (2013) summarise a line of studies in which 90% of students remembered more after spaced study, yet 72% said massed study had worked better for them.11 If you need the material past tomorrow, cramming costs more than it saves, because you'll have to learn it again.

"I forgot it, so the studying was wasted." This is the mistake that makes people abandon spacing. You review after six days, miss a chunk of the items, and conclude the first session failed. It didn't. The forgetting is what made this review count. A review where you miss nothing has taught you very little; a review where you miss and reconstruct has done the work. The measure of a session isn't how it felt on the day but what survives to the delayed test, which is the lesson-1 principle again in a new setting. Keep going.

"The gaps have to expand or it doesn't work." Expanding schedules (1 day, 3 days, 7, 14, 30) are the default in most apps, and most study guides present expansion as the way it must be done. To be fair to them, the research record before 2021 was mixed, with some early studies favouring expansion. Then Latimier, Peyre and Ramus (2021) pooled the comparisons and found expanding and uniform schedules about equal: the difference was g = 0.034, not distinguishable from zero, with a hint that expansion does relatively better when there are many reviews.5 Expanding schedules are convenient, because they let you review new material often and old material rarely, and the site's Review page uses one for that reason. But if a fixed gap suits your calendar, use it. The effect comes from the presence of gaps and from retrieving at each review, not from the shape of the sequence.

"Spacing means spreading out the reading." Half right. Spaced rereading does beat massed rereading; that's what most of the classic experiments measured. But spaced retrieval beats both, and Latimier's g = 0.74 is for retrieval.5 If you're going to space something, space the thing that works best. Every review in your schedule begins with the book closed.

Practice

Do it now: one thing you must know in three months

Pick something specific you need to know in about 90 days: a set of terms for a course, the arguments of a book you're reading for a group, verb conjugations, the sections of a standard you'll be examined on.

  1. Write the date you need it.
  2. Apply the rule: 10–20% of 90 days is 9 to 18 days, and Cepeda's 70-day row says 21. Choose a first gap, and if you're torn, choose the longer one.
  3. Write the review dates, uniform or expanding, up to a final review two or three days before the deadline. Five to ten reviews is typical.
  4. Write down, now, the ten questions you will answer from memory at each review, and the stopping rule: three correct recalls of each today, one at every review after.
  5. Put the first review date somewhere you will actually see it.

Checkable output: a dated list of reviews and a question set. If you have neither, you have a plan to make a plan, which is not the same thing.

Do it now: read the site's schedule

Open the Review page on this site and look at how it schedules the quiz questions you have already passed. The scheme is a version of the SM-2 algorithm many flashcard apps use. A question you get right comes back after 1 day, then 3 days, then the previous interval multiplied by an "ease" factor that starts at 2.5 and creeps up with each success (so roughly 8 days, then 22, then two months). A question you miss goes back to a 1-day interval and its ease drops.

Answer from this lesson, in writing: why do the gaps lengthen? Why does a miss reset the gap? Which rows of Cepeda's table does the 1, 3, 8, 22 sequence resemble?

Check yourself

What should your three answers to the Review-page exercise look like?

Show the answer

The gaps lengthen because each success is evidence the memory is stronger, and a stronger memory takes longer to reach fragile-but-recoverable, so the next gap can be wider; put another way, the horizon you can trust it over has grown, and the gap grows with the horizon. A miss resets the gap because a failed retrieval means the memory has fallen past recoverable; the item is back to roughly newly learned, and a newly learned item needs a short gap. The 1, 3, 8, 22 sequence is a stack of Cepeda's first gaps for lengthening horizons: 1 day matches the one-week row, and 8 to 22 days spans the 11- and 21-day optima for the 35- and 70-day rows. Note that this is an expanding schedule, and that Latimier's evidence says a uniform one would do about as well. The site expands because it's convenient, not because it's necessary.

Connections

Lesson 1 gave you the storage-versus-retrieval distinction and the finding that people prefer the study method that teaches them less; this lesson is why the second follows from the first. Lesson 3 gave you the tool, retrieval; this lesson says when to use it. Together they're the two techniques Dunlosky rated high utility, and combining them is the single most reliable change you can make to how you study.

The site's Review page is this lesson in software: each quiz question you pass enters a schedule with lengthening gaps and reset-on-miss, and the page tells you when things are due. You now know the reasoning behind every number in it.

Lesson 5 adds the third technique, interleaving, which mixes what you practise the way spacing spreads when, and shows where it helps and where it doesn't. Lesson 8 builds all of this into a week you can actually keep.

Go deeper

Sources

  1. Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). "Distributed practice in verbal recall tasks: A review and quantitative synthesis." Psychological Bulletin, 132, 354–380. Supports: spacing beats massing across 184 articles, 317 experiments, 839 assessments; the optimal gap grows with the retention interval; most experiments used spaced restudy; no single theory of the mechanism fits all the data.
  2. Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). "Improving Students' Learning With Effective Learning Techniques." Psychological Science in the Public Interest, 14(1), 4–58. Supports: distributed practice and practice testing rated high utility; the other eight techniques moderate or low.
  3. Cepeda, N. J., Vul, E., Rohrer, D., Wixted, J. T., & Pashler, H. (2008). "Spacing Effects in Learning: A Temporal Ridgeline of Optimal Retention." Psychological Science, 19, 1095–1102. Supports: 1,350+ participants, two learning sessions with gaps to 3.5 months, tests to one year; observed optimal gaps of 1, 11, 21 and 21 days for retention intervals of 7, 35, 70 and 350 days; fitted one-year optimum 23 days (7%); a too-short gap cost far more than a too-long one.
  4. Carpenter, S. K., Cepeda, N. J., Rohrer, D., Kang, S. H. K., & Pashler, H. (2012). "Using Spacing to Enhance Diverse Forms of Learning: Review of Recent Research and Implications for Instruction." Educational Psychology Review, 24, 369–378. Supports: the rule of thumb of a gap roughly 10–20% of the time until the material is needed.
  5. Latimier, A., Peyre, H., & Ramus, F. (2021). "A Meta-Analytic Review of the Benefit of Spacing out Retrieval Practice Episodes on Retention." Educational Psychology Review, 33, 959–987. Supports: spaced versus massed retrieval practice g = 0.74 across 29 studies; expanding versus uniform schedules g = 0.034, not significant, with expansion relatively better when there are more reviews.
  6. Kang, S. H. K. (2016). "Spaced Repetition Promotes Efficient and Effective Learning: Policy Implications for Instruction." Policy Insights from the Behavioral and Brain Sciences, 3, 12–19. Supports: the accessible summary of the spacing evidence for practitioners.
  7. Bjork, R. A. (1994). "Memory and metamemory considerations in the training of human beings." In J. Metcalfe & A. Shimamura (Eds.), Metacognition: Knowing about Knowing (pp. 185–205). MIT Press. Supports: the term "desirable difficulties" for conditions that make practice harder and learning more durable.
  8. Rawson, K. A., & Dunlosky, J. (2011). "Optimizing schedules of retrieval practice for durable and efficient learning: How much is enough?" Journal of Experimental Psychology: General, 140, 283–302. Supports: 533 students; the successive relearning criterion of three correct recalls in the first session and one correct recall in each of three later spaced sessions.
  9. Bahrick, H. P., Bahrick, L. E., Bahrick, A. S., & Bahrick, P. E. (1993). "Maintenance of Foreign Language Vocabulary and the Spacing Effect." Psychological Science, 4, 316–321. Supports: nine-year study with four participants (the authors); 13 relearning sessions at 56-day intervals gave retention comparable to 26 sessions at 14-day intervals; wider gaps slowed acquisition slightly.
  10. Bahrick, H. P. (1984). "Semantic memory content in permastore: Fifty years of memory for Spanish learned in school." Journal of Experimental Psychology: General, 113, 1–29. Supports: 733 people, cross-sectional design; retention declines for 3–6 years then stays roughly unchanged for up to 30 years before a final decline ("permastore"); permastore content predicted mainly by original level of training, with little effect of later rehearsal.
  11. Bjork, R. A., Dunlosky, J., & Kornell, N. (2013). "Self-Regulated Learning: Beliefs, Techniques, and Illusions." Annual Review of Psychology, 64, 417–444. Supports: 90% of students remembered more after spaced study while 72% judged massed study to have worked better.

Check your understanding

This lesson has a 6-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.