What an interruption actually costs

80 min

Listen: this lesson as a conversation

Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.

In this lesson you will learn to
  • State what the CHI 2008 experiment measured and found, with its sample, its conditions and its numbers
  • Explain why a measure that looks at time and errors alone can miss the cost of an interruption
  • Predict, for a described case, what an interruption would and would not show up in

Everybody knows interruptions are expensive. This lesson is about an experiment that measured the expense and found it somewhere nobody looks.

What Time Management already gave you

Time Management lesson 5 covers the shape of an interrupted working day: how long people work before switching, what a field study saw when it followed people around an office, and the trace of the most quoted figure in the subject back to a source nobody can produce.

This lesson doesn't repeat any of that. It takes a different paper, by one of the same authors, and asks a narrower question: not how often you are interrupted, and not how long it takes to get back, but what the interruption does to the work and to you.

One thing does overlap and it is worth saying exactly what. Time Management lesson 5 gives this experiment's sample, its design and its abstract: forty-eight mostly German students answering simulated emails, and the conclusion that interrupted tasks were finished faster and paid for in stress, frustration, time pressure and effort. If you have met that, you have met who was in it and what they concluded, and not the study. This course has read the paper in full, and what follows is the condition-by-condition times, the workload ratings one measure at a time, the question the experiment was actually designed to answer, and the one thing that changed about the work itself. The headline is the part that travels; the rest is where the teaching is.

The experiment

Mark, Gudith and Klocke ran it and published it at a computing conference in 2008.1 Six pages, freely available, and the experiment this course quotes most.

The people. Forty-eight subjects, 81 percent of them German university students, mean age 26.

The task. Each played a human-resources manager who had just come back from holiday, answering twelve emails from staff, with a fact sheet to answer from. They were instructed and given incentives, in the paper's words, "to answer all emails in their inbox as quickly, correctly and politely as possible".1

The three conditions, counterbalanced so everybody did all three.

  • No interruption. The baseline.
  • Interruption on the same topic as the task. A "supervisor", who was the experimenter in another room, telephoned or messaged with a question related to the emails.
  • Interruption on a different topic. Same thing, unrelated question.

What they measured. Time to complete the task, with time spent on the interruptions subtracted. Errors, counted as spelling mistakes and typos. A politeness measure. And a modified NASA Task Load Index, on which participants rated mental workload, stress, frustration, time pressure and effort, each on a twenty-point scale.

Predict first

Before the result. Interrupted people had to stop, answer a question, and come back. Write down what you expect happened to the time they took, and to the errors they made.

Show the answer

This course expects most readers to predict the same two things, which is an expectation rather than a measurement, and one of the two is right.

The errors did not differ between conditions. That part is the less surprising of the two: the emails were short, the fact sheet was in front of them, and there's no obvious reason an interruption would make somebody misspell a name.

The time is the one to sit with. If you wrote "longer", you've got the folk model, and so did the researchers: they set out to measure what they called a disruption cost.

Interrupted tasks were finished faster. Not by a trivial amount, and not in one condition only. That is the result, and the rest of the lesson is about what it means and what it hides.

The result

In the authors' own words, from the abstract: "We found that context does not make a difference but surprisingly, people completed interrupted tasks in less time with no difference in quality."1

The numbers, with standard deviations, as the paper reports them.1

Condition Time to perform the task Average errors Words per email
No interruption 22.77 minutes (7.60) 1.94 (.91) 31.49 (8.1)
Same-topic interruption 20.31 minutes (5.94) 1.93 (.88) 29.17 (7.02)
Different-topic interruption 20.60 minutes (4.93) 1.84 (.92) 30.16 (7.18)

The uninterrupted condition was the slowest. Read that twice, because everything else in the lesson depends on it.

And the two kinds of interruption did not differ from each other. That was the question the study set out to answer, and the answer was no: whether the interruption was about your task or about something else didn't change the time. It is the first thing the abstract reports.

So what did it cost?

The workload ratings, on a scale of one to twenty.1

Measure No interruption Same topic Different topic
Mental workload 10.02 10.83 11.50
Stress 6.92 9.46 9.13
Frustration 4.73 6.63 6.48
Time pressure 11.02 12.69 12.17
Effort 9.50 11.04 11.52

Every one of those is higher in both interruption conditions, and all five differ at conventional levels of significance, four of them at the stricter one.[3] The authors put it plainly in their discussion: after only twenty minutes of interrupted performance, people reported significantly higher stress, frustration, workload, effort and pressure.

The authors' own statement of it, from the abstract: "Our data suggests that people compensate for interruptions by working faster, but this comes at a price: experiencing more stress, higher frustration, time pressure and effort."1

And their closing sentence, which is the one to remember: "So interrupted work may be done faster, but at a price."1

The mechanism, and one thing it changed about the work

Compensation is the proposed mechanism and it is the authors'. People know they have lost time, they speed up, and the speeding up is what costs them.

There's a second piece of evidence for it in the table above, and it's easy to miss. The emails were longest in the uninterrupted condition, at 31.49 words on average against 29.17 and 30.16. People wrote less. The paper offers this as part of the interpretation: some of the speed came from producing slightly less.

One more finding, worth a paragraph and not more. A regression found that two personality measures, openness to experience and need for personal structure, both predicted how quickly somebody finished an interrupted task, together accounting for about 14 percent of the variance in completion time.1 That is a real effect and a small one, and the reason it gets a paragraph rather than a section is that a course built on it would be telling you your personality decides how interruption treats you, on 48 people in one session.

So the work wasn't identical, even though the error count and the politeness rating were. That's a specific and checkable thing, and it is the seam this lesson is really about: two measures said nothing had changed, and a third said something had.

Check yourself

A manager reads this and concludes that interruptions are fine, since the work gets done faster and the quality holds. What is wrong with that, and what is right about it?

Show the answer

Something's right about it, which is why this is worth doing carefully.

What is right. On the two measures most workplaces actually use, output and error rate, this study found no damage and a small gain. If your entire account of the cost of interruption was "things take longer and go wrong", this experiment doesn't support you, and pretending otherwise would be the thing this course exists to teach against.

What's wrong is the inference from two measures to no cost. The study measured five more things, and all of them moved. Stress went from 6.92 to about 9.3. Effort went from 9.50 to about 11.3. A manager who sees the output and concludes the interruptions are free has drawn a conclusion the measurements do not support, because the cost landed in the column she isn't looking at.

And there's a third thing, which this course flags as its own reading rather than a finding.[4] The study ran for about ninety minutes. Nothing in it tells you what happens to somebody compensating like that for a week, or a year, and the authors say in their discussion that they don't know whether people would cope over time or whether these measures would only increase. That's a question rather than a result, and it is the honest way to hold it.

Worked: the same finding, seen from a dashboard

This case is constructed.[4] It's built to show the seam above and isn't taken from the paper.

Dana leads a support team of nine. Her dashboard gives her two numbers per person per week: tickets closed, and tickets reopened by the customer. It has done for two years.

Her team moved to a new chat tool in March. Since then, interruptions have roughly doubled: everybody can reach everybody, instantly, and they do.

The dashboard says nothing has happened. Tickets closed is very slightly up. Reopens are flat. On the evidence Dana has, the new tool has been fine or mildly positive.

What this experiment suggests she is not seeing, and the word is suggests, because her team is not 48 German students answering twelve emails in a lab.[4]

  • The tickets may be getting shorter. The one thing the experiment found had changed about the output was its length, and a shorter reply is the kind of thing a reopen count catches only sometimes.
  • The people may be paying for the speed. Stress, frustration and effort all rose in the study, and none of them appears on her dashboard.
  • She has no baseline for either. The measures that moved in the experiment are self-reports, and nobody was collecting them in March.

What she should do about it isn't something this experiment can tell her, and a lesson that finished by prescribing a policy would be overreaching. What it can tell her is what to measure next, which is the useful thing: ask the team the five workload questions, now, and again in a month, and she will have the column she is missing.

Three things people get wrong about this

"Interruptions make you slower." Not in this experiment, and it's the clearest result in it. The recovery-time question, which is a different question, belongs to Time Management lesson 5.

"Interruptions about the same thing are less disruptive." This is the belief the study was designed to test, and it found no difference between same-topic and different-topic interruptions on any measure it reported.

"So interruptions are harmless." Five measures of workload moved and every one of them moved significantly. What the study undermines is one particular account of the harm, not the harm itself.

Practice

Measure the column nobody measures

Take 20 minutes across one working day, then 10 at the end.

Pick two work blocks of about an hour each, on the same day if you can. One where you expect interruptions, and one where you expect few.

In each, keep one count: how many times somebody or something took your attention away from the task. A tally on paper. Don't try to change anything.

At the end of each block, rate four things, one to twenty, exactly as the study did: stress, frustration, time pressure, effort. Write them down before you start the next block, because rating them afterwards from memory is a different measurement.

Then, at the end of the day, three lines.

  1. The two tallies, side by side.
  2. The two sets of four ratings, side by side.
  3. One sentence on whether your output differed, and how you would even know.

One day proves nothing about you, which is Time Management lesson 2's point and it holds here. What the exercise gives you is the experience of collecting the column a dashboard doesn't have.

Find the missing measure

Take 20 minutes.

Find a real measure that somebody applies to work: yours, your team's, a school's, a service's. A target, a metric, a report, a league table.

Write down three things.

  1. What it counts. Exactly, in the terms the measure itself uses.
  2. What would have to be true for the count to stay flat while something got worse. This is the seam. In the experiment it was the time and the error count staying put while the stress rating moved.
  3. What you would measure to catch it, and what it would cost to collect.

If you can't find a way for the count to stay flat while things worsen, say so, and say why. Some measures are hard to game and hard to mislead, and finding one is a real result.

Connections

Back. Lesson 1's instrument question is what makes this lesson intelligible: the time and the error count came from a task, and the stress and effort ratings came from a self-report, and this experiment is a case where the two do not contradict each other but measure different halves of the same event. Time Management lesson 5 has the field study, the work segment and the twenty-three-minute trace, none of which is repeated here. Time Management lesson 6's habit of naming what a decision costs is what the checkpoint asks of Dana.

Forward. Lesson 3 is the interruption that comes from inside, where there's nobody to blame and no tally to keep. Lesson 6 builds its working hour partly on this result. And lesson 7's sorting habit applies to every claim about interruption you will meet after this one.

Go deeper

Sources

  1. Gloria Mark, Daniela Gudith and Ulrich Klocke, "The Cost of Interrupted Work: More Speed and Stress", Proceedings of CHI 2008, pages 107 to
    1. Read in full. Supports: the sample of 48 subjects, 81 percent German university students with a mean age of 26; the task and its quoted instruction; the three counterbalanced conditions and the supervisor delivering the interruptions; the quoted abstract sentence about interrupted tasks taking less time; the quoted sentence about compensation and its price; the quoted closing sentence; and every figure in both tables, which are Tables 1 and 3 of the paper.
  2. This lesson introduces no other evidence. The recovery-time literature and the field study of fragmented work belong to Time Management lesson 5 and are cited there.
  3. The statement that the authors report stress, frustration and effort as differing at conventional levels of significance is from the paper's results section, which gives F statistics and p values for each measure against the baseline. This course has read those and reports the direction and the significance rather than reprinting the statistics, because the raw F values would tell a general reader nothing the table does not.
  4. Dana and her team are constructed, and so is every detail attached to them; the lesson says so where she appears, and says which of the three things she may not be seeing are extrapolations from a lab study rather than claims about her. The reading that a ninety-minute session cannot tell you about a year of compensating is also this course's own, marked inline in the checkpoint, and it is supported by the authors' own statement in their discussion that they do not know how people would cope over time.

Check your understanding

This lesson has a 6-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.