Why it always takes longer

85 min

Listen: this lesson as a conversation

Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.

In this lesson you will learn to
  • State the measured size of the planning fallacy in at least one study, with what it was measured on and when
  • Explain why being instructed to be pessimistic moves the prediction without improving it, and why being instructed to remember does almost nothing at all
  • Measure the ratio between your own predicted and actual times on real tasks, and say why the ratio transfers where the difference does not

Lesson 1 found that your picture of a week is unreliable. This lesson is the same failure at the scale of a single job, where it has been measured much more precisely and where the fix that everybody reaches for turns out not to work.

The finding has a name, the planning fallacy, which Buehler and colleagues credit to Kahneman and Tversky in 1979 and put in quotation marks in their own title.2 Buehler, Griffin and Ross, 1994, in the Journal of Personality and Social Psychology.1 It's the best-evidenced thing in this course, and it is worth knowing its scope before its numbers: every subject in the studies below was a Canadian university student, and the paper was published in 1994.

The thesis study

Thirty-seven psychology students in the final semester of the Honors Thesis course at the University of Waterloo were telephoned and asked to predict when they would submit. Not once, but three times over, in three different ways:

  • their best estimate of the date
  • the date if "everything went as well as it possibly could"
  • the date if "everything went as poorly as it possibly could"

The course coordinator recorded when each thesis actually arrived. The means below are from 33 of them.1

Predict first

Before you look at the table: write down two numbers. First, how many days past their best estimate you think the average student took. Second, what share of them you think met their own worst-case date, the one they gave assuming everything went as badly as it possibly could.

Show the answer

Most people who do this get the first number roughly right, in the sense of guessing that the work ran over, and get the second one badly wrong, usually somewhere between 80 and 95 percent.

The reasoning behind the high guess is sound as far as it goes: a worst case is supposed to be the end of the range. If you set out to describe a disaster and then no disaster happens, you should finish early.

Hold your two numbers and read on.

Best estimate "Everything went as well as it possibly could" "Everything went as poorly as it possibly could"
Predicted days 33.9 27.4 48.6
Actual days 55.5 55.5 55.5
Finished within the prediction 29.7% 10.8% 48.7%

Read the last column on its own, because it is where the lesson is.

These students named a date on the assumption that everything that could go wrong would, and the average of those dates was 48.6 days. The work took 55.5. Fewer than half of them finished by their own worst case. The paper states it plainly: "Interestingly, fewer than half of the respondents (48.7%) finished by the time they had predicted assuming that 'everything went as poorly as it possibly could.'"1

So the worst case a person can imagine isn't the worst case. It is a slightly slower version of the run that goes well.

What this does to the obvious fix

You are probably already thinking it, and it is the first thing anybody thinks: fine, so I'll pad my estimates.

The same study tested the nearest thing to it. Reading the pessimistic date as a padded estimate produced under instruction is this course's own reading of that condition rather than a claim the authors make,4 and it is worth saying so, because a worst-case scenario and an arbitrary multiplier are not obviously the same operation. Here is what the authors found: "Although the instructions to make a pessimistic prediction decreased the optimistic bias in prediction, it did not increase the accuracy of respondents' forecasts."1

The absolute error was 23.2 days under the pessimistic instruction against 22.6 days under the instruction to be accurate. Six tenths of a day apart, on a task that ran to fifty-five.

That distinction is the one to carry out of this section, and it recurs twice more in this course. Bias is which side of the truth you land on. Accuracy is how far from it you are. Padding moves you across the truth without moving you closer to it, because it adds no information about this particular job. You are still guessing; you have just changed the direction you guess in.

It isn't only theses

Study 2 took 104 undergraduates and asked for predictions on two projects each, one academic and one not, that they intended to finish in the coming week.1

Academic Nonacademic
Predicted days 5.8 5.0
Actual days 10.7 9.2
Finished within the prediction 37.1% 42.5%

And the figure that closes off the easy explanation: these people were not hedging or being modest. "[S]ubjects reported feeling 74.1% certain that they would meet their forecasts for academic projects and 69.9% certain for nonacademic tasks."1

Three quarters certain, and right about a third of the time.

The ratio, which is the part you can use

Something worth noticing across those studies, and it's this course's own arithmetic on the published means rather than a finding the authors report.4

Divide the actual by the predicted:

  • Thesis study, best estimate: 55.5 over 33.9, about 1.6
  • Study 2, academic projects: 10.7 over 5.8, about 1.8
  • Study 2, nonacademic projects: 9.2 over 5.0, about 1.8

Different tasks, different lengths, ratios in the same neighbourhood. Now do the same with the differences: 21.6 days, 4.9 days, 4.2 days. Those have nothing in common and nothing to say to each other.

This is why the exercise at the end of this lesson asks for a ratio. A shortfall of forty minutes on a forty-minute job tells you nothing about a job that will take a fortnight. "About double" does.

Worked: five tasks and one number

Rosa's week is constructed, and so is every figure in it.4 She is a district nurse on a fixed rota, and these are five things she did around it.

Task Predicted Actual Ratio
Ring the bank about the direct debit 10 min 47 min 4.7
Weekly shop 60 min 75 min 1.25
Write up two case notes 40 min 55 min 1.38
Change the tyre on the car 45 min 40 min 0.89
Fill in the insurance renewal form 20 min 35 min 1.75

Line the five ratios up in order: 0.89, 1.25, 1.38, 1.75, 4.7. The middle one is 1.38, and that is her median. She would have got 1.99 by averaging, and the average is the wrong tool here, because one task is doing all the work in it.

That one task is the finding, not the noise. The bank call ran nearly five times over, and it is the only one of the five where the length was not hers to decide. She could not have known it would be forty-seven minutes. What she could have known, from the last three times she rang them, is that it is the kind of task where somebody else sets the clock.

So her sheet gives her two things rather than one. A working number of about 1.4 for tasks she controls, which is close enough to the studies above to be worth trusting on herself. And a named category, anything that involves waiting on an institution, where the ratio is not a ratio at all and the only honest plan is to give the whole morning to it.

Notice what a median did that an average would not: it kept the outlier visible as a separate fact instead of smearing it across the other four.

Now run the same division on the Study 4 table further down this lesson, which is a computer assignment rather than a thesis: 6.8 over 5.5 is about 1.2. So the neighbourhood is loose, and it is loosest on the shortest task.

That is the first of two warnings, and the three figures above would have hidden it. Three ratios from two studies of students in one country do not give you a constant, and your own number is the one that matters anyway. The second warning is that a ratio is only stable within a kind of work, which the exercise is designed to show you.

Check yourself

Someone objects that all of this is about students and essays, and that a person doing paid trade work quotes jobs accurately because they get paid by the quote. Is that a good objection?

Show the answer

It is a good objection and it should not be waved away, but notice what it's actually claiming.

It says the effect is smaller, or absent, where there's a feedback loop with money attached to it. That is plausible and this course has no measurement of it either way, because the literature it read studied students and office workers.3 Lesson 1 named that gap and this is one of the places it bites.

What the objection does not do is rescue the general case. The person quoting jobs for money has a narrow band of work they have quoted hundreds of times, and their accuracy is a fact about that band. Ask the same person how long the extension on their own house will take.

And there's a finding coming in the next section that applies directly to their situation, which is that people hit dates set from outside far more reliably than dates they set themselves. A quoted price with a penalty clause is an external deadline wearing a different hat.

The finding that inverts the moral

Inside Study 2 there is a smaller analysis that is, practically speaking, the most useful thing in the paper.1

Predict first

Sixty-two of those students had an external deadline on their academic project, a date somebody else set. Two numbers before you read on: what share of them do you think met the deadline, and what share met their own earlier prediction?

Show the answer

The two numbers are 80.6 percent and 38.7 percent, and the gap between them is the finding.

Most people guess the two close together, and usually both low, because the section you have just read is about people who cannot forecast. The reason the guess goes wrong is that forecasting and finishing turn out to be different skills, and this study measures them separately.

Sixty-two of those students had an external deadline on their academic project. A date somebody else set.

  • 80.6 percent of them finished in time to meet it.
  • The projects were due in 12.9 days on average, and were reported finished in 11.0.
  • They had predicted 5.9 days.
  • Only 38.7 percent finished within their own prediction.

Look at what that describes. These were not people who couldn't finish things. Four out of five delivered against the date they'd been given. Their problem was entirely in the forecast, which sat "well in advance of both the deadline and the reported completion time".1

The authors put the conclusion carefully: "subjects underestimated the importance of deadlines in determining when they would finish their projects."1 The correlations make it concrete. Their predictions correlated only weakly with their deadlines, r = .23. Their actual completion times correlated strongly with them, r = .82.

In plain terms: the deadline was running their schedule and their forecast didn't know. They were predicting from the work, and the work was not what governed when they finished.

You should take two things from this and they pull in different directions, which is why the lesson gives you both rather than choosing.

One. If you want to know when something will be done, the date somebody else set is better evidence than your own estimate. That is uncomfortable and it is what the correlations say.

Two. Only a certain kind of task has one. A deadline is not available for the things nobody is waiting for, and those are exactly the things that slide for years. Lesson 4 is about what to do there, and lesson 6 is about deciding whether they should be on the list at all.

Why it happens, in one set of figures

Study 4, the one in the next section, also asked its 123 subjects to list what they were thinking about while they predicted, and the answer is the whole mechanism.1

  • 93.5 percent reported considering future plans and scenarios for the task, mostly about how they'd successfully complete it.
  • 9.8 percent mentioned any potential impediment.
  • 8.9 percent reported thinking about their own past experiences.

Read those three together and the picture is not one of people ignoring their history. It is of people for whom the history never came up.

Forecasting a task means imagining how it will go, step after step, and an imagined run is made of steps that work. That is what makes it imaginable. The specific thing that will actually delay you, the document you can't find, the reply that takes four days, the morning you're too tired, has no place in the story because you don't know which one it will be. You know there will be one. You cannot picture it.

So you do not produce a bad forecast through carelessness. You produce it by doing the thing forecasting is, and your history isn't in the room while you do it.

And remembering doesn't fix it

If the history is absent, the obvious repair is to fetch it. Study 4, the source of the three figures above, tested exactly that on 123 students doing a computer assignment.1

Check yourself

One group described their past experience with similar assignments just before predicting, and was told to keep those experiences in mind. Given everything in this lesson so far, what do you expect that did to their predictions?

Show the answer

The honest answer is that this lesson has set you up to expect very little, and very little is what happened: 29.3 percent finishing on time became 38.1 percent.

What is worth sitting with is why that is surprising anyway. The other two corrections in this lesson failed for reasons you can see. Pessimism adds no information about this job. Imagining the run that goes well is what planning is. But fetching the history directly ought to work, because the history is precisely what the forecast is missing, and these subjects fetched it accurately and out loud.

One group described their past experience with similar assignments just before predicting, and was told to keep those experiences in mind.

Control Told to recall
Predicted days 5.5 5.3
Actual days 6.8 6.3
Finished within the prediction 29.3% 38.1%

Almost nothing. And the authors' description of why is worth reading in full, because it names something most people will recognise:

"The absence of an effect in the recall condition is rather remarkable. In this condition, subjects first described their past performance with projects similar to the computer assignment and acknowledged that they typically finish only 1 day before deadlines. Following a suggestion to 'keep in mind previous experiences with assignments,' they then predicted when they would finish the computer assignment. Despite this seemingly powerful manipulation, subjects continued to make overly optimistic forecasts. Apparently, subjects were able to acknowledge their past experiences but disassociate those episodes from their present predictions."1

They said out loud that they always finish the day before it's due. Then they predicted that this time they'd be early.

Remembering is not connecting, and the gap between those two is where this whole subject lives. There was a third group in that study who were made to do the connecting, and what happened to them is lesson 4's subject. Their share finishing on time roughly doubled, though their forecasts were no more accurate than anybody else's, which is a distinction this lesson has already made you careful about. Lesson 3 stops here on purpose: you should feel the size of the problem before you meet what can be done about it.

Four things people believe about their own estimates

"I just need to pad them." Measured, in the study this lesson opens with. Padding cut the bias and left the error where it was.1 The pessimistic estimate was still a week short of the truth.

"I should remember how long it took last time." Measured, and it moved the on-time share from 29.3 to 38.1 percent while the students in question had just recited their own history accurately.1

"Other people manage this." The paper's first hypothesis was that "[p]eople underestimate their own but not others' completion times",1 and the pattern it reports is that the bias belongs to predicting for yourself rather than to predicting in general. This course has read the abstract and Studies 1, 2 and 4, not the observer study, so treat this as what the paper reports rather than as something the course has checked line by line. Either way, the folk version has it backwards. You are not unusually bad at this, and on what the paper reports the asymmetry is not between you and other people but between predicting for yourself and predicting for somebody else. Which is a reason to ask somebody, and lesson 4 gives you a way of doing the same job on your own.

"Deadlines are the problem." In Study 2, 80.6 percent of the students with deadlines met them.1 Deadlines were the thing that worked. What failed was the forecast.

A word about what this lesson is not saying

Nothing here says you should work to deadlines, or that external pressure is good for people, or that you ought to be doing more. Lesson 1 named the question of whether a person should aim to do more or to do less and said the evidence does not settle it. Lesson 6 is where that question is faced, because it is the lesson that would otherwise decide it silently, and lesson 8 asks what the whole evidence base is worth.

What this lesson claims is narrower and, hopefully, more useful: your estimate of how long a specific job will take is unreliable in a known direction and by a measurable amount, and the two corrections people reach for first have both been tested and are weak. What you do with a better estimate is a separate question, and it stays yours.

Practice

Predict five, then measure five

Take 15 minutes to set up, then a few seconds per task across the week.

Pick five real tasks you'll do in the next seven days. They must be things with a clear start and a clear end, and at least two should be small: answer an email properly, change a tyre, do the weekly shop, write a page, ring the bank about the thing you've been avoiding.

For each one, before you start:

  1. Write the task and your predicted time, in minutes. Commit to a number rather than a range.
  2. Write one sentence about how you pictured it going. One line. This is the thought listing from the study above, run on yourself, and it is the part people skip.

Then do the task and write down what it actually took.

At the end of the week, for each task, work out the ratio: actual divided by predicted. Not the difference. Then answer three questions in writing:

  • What is your median ratio across the five, which is the middle one when you line all five up in order?
  • Which task was furthest out, and was it the largest task or the least familiar one?
  • Look back at your five sentences about how you pictured it going. How many of them mention anything going wrong? Study 4's subjects mentioned an impediment 9.8 percent of the time, and your own five is a small enough sample that you can just count.

Keep the sheet. Lesson 4 uses your median ratio, and lesson 7 uses it again when you build a week.

Find the tasks nobody is waiting for

Take 20 minutes.

Write down three things you intend to do that nobody is waiting for. No client, no boss, no form with a date on it. The dentist, the will, the qualification, the loft, the letter.

For each one, write two lines:

  1. How long have you intended to do it? A date or a season is fine. "Since we moved" counts.
  2. What is your best estimate of how long the task itself would take, once started?

The pair of numbers is the exercise, and for most people reading this the second is small while the first runs into years. That gap is not a forecasting problem and no better estimate will close it, which is the honest thing to say now rather than to imply otherwise for four more lessons.

Keep the list. Lesson 4 puts a cue on one of them, lesson 6 asks whether all three belong on the list at all, and one right answer there is to cross one off and stop carrying it.

Connections

Back. Lesson 1 established that your impression of your own time is unreliable and asked you to measure rather than to introspect; this lesson is that argument at the scale of one task, with much better evidence behind it. Lesson 2's habit of writing the prediction down before the record is what makes the exercise above work at all. From earlier Core courses: Logic and Argument lesson 6 gave you the habit of asking how much a piece of evidence should move you, which is what the correlations of .23 and .82 in the deadline analysis are doing. Using AI Effectively taught you to ask what a measurement was taken on before you carry it anywhere, which is why the scope of these studies is in the second paragraph rather than in a footnote.

Forward. Lesson 4 is the two corrections with measured effects behind them, the third condition of the study this lesson stopped short of and the one planning technique in this subject with a meta-analysis, and it is careful about what each does and does not improve. Lesson 5 is why the day itself doesn't hold still. Lesson 7 builds a week using your own ratio. Lesson 8 returns to what the whole evidence base is worth.

Go deeper

  • Exploring the "Planning Fallacy" (Journal of Personality and Social Psychology, 1994). Hosted openly at MIT. This course has read the abstract, Studies 1 and 2 in full with their tables, and Study 4 with its table and discussion. If you read one thing, read the tables: they carry the argument on their own, which is rare. Reading Well lesson 8 gives you the questions to put to a figure, and they work unchanged on these tables.
  • The American Time Use Survey appears here for a third time and a third reason, after lessons 1 and 2. This one: it is a case of a large organisation deciding that asking people how long things take is not good enough, and building an instrument instead. That is the institutional version of the move this lesson asks you to make on yourself. United States only, and this course has read the article that uses its data rather than the survey's own documentation.

Sources

  1. Roger Buehler, Dale Griffin and Michael Ross, "Exploring the 'Planning Fallacy': Why People Underestimate Their Task Completion Times", Journal of Personality and Social Psychology 67(3), 1994, pages 366 to 381. Read in substantial part: the abstract, Studies 1 and 2 in full with their tables, and Study 4 with its table and discussion. Supports: Study 1's design, its three predicted means and its actual mean, and the 48.7 percent figure with its quoted sentence; the quoted finding that the pessimistic instruction decreased the bias without increasing accuracy, and the 23.2 against 22.6 absolute errors; Study 2's means and completion shares and the quoted confidence figures; the deadline sub-analysis including 80.6 percent, 12.9 and 11.0 days, 5.9 predicted, 38.7 percent, the quoted sentence about underestimating the importance of deadlines, and the correlations of .23 and .82; the thought-listing figures of 93.5, 9.8 and 8.9 percent; Study 4's control and recall columns and the quoted passage about the absence of an effect in the recall condition; and the abstract's first hypothesis about own against others' completion times. Scope: university students in Canada, published 1994. Study 1 is 37 Honours Thesis students at the University of Waterloo, means from 33; Study 2 is 104 undergraduates; Study 4 is 123 students.
  2. The term is credited by Buehler and colleagues to Daniel Kahneman and Amos Tversky, "Intuitive Prediction: Biases and Corrective Procedures", 1979. This course has not opened that paper and the attribution above is the only claim it makes about it.
  3. Brigitte J. C. Claessens, Wendelien van Eerde, Christel G. Rutte and Robert A. Roe, "A review of the time management literature", Personnel Review 36(2), 2007. Read in part: the abstract and pages 255 to 257. Supports the standing point, made in lesson 1 and relied on here, that the evidence base for this subject is mostly students and office workers.
  4. Three things in this lesson are the course's own and are labelled as such in the body where the reader meets them. First, the ratios of actual to predicted time are this course's arithmetic on the published means in source 1: the authors report means and completion shares, they do not report ratios, and nothing in the paper establishes that a ratio is stable for an individual. Second, reading Study 1's pessimistic condition as a test of padding an estimate is this course's reading of that condition and not a claim the authors make. Third, Rosa and her five tasks are constructed, along with every figure in them; no source read for this course records one person's predictions against their own times.

Check your understanding

This lesson has a 6-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.