Four instruments and one word

80 min

Listen: this lesson as a conversation

Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.

In this lesson you will learn to
  • Name the four instruments sleep research uses and say what each one can and cannot measure
  • Read a sensitivity and specificity pair together and say what the two numbers mean about a device
  • State what this course read, at what level, and what it therefore cannot say

You have done this thing every night of your life and you have never once measured it. That's unusual. Most things a course can teach you about are things you could in principle watch yourself doing, and sleep is the one activity where the participant isn't, by definition, available to observe it.

So everything anybody knows about sleep came through an instrument. There are four of them in common use, they don't measure the same thing, and the word "sleep" gets used for all four outputs without anybody saying which one produced the number.

This lesson is that list, and one pair of numbers that shows why it matters.

If you're reading this because something is wrong with your sleep, go to the section near the end on what this course leaves out, and read that first.

The four instruments

Polysomnography. Electrodes on the scalp and face, plus breathing and heart traces, recorded overnight in a laboratory and scored by a technician in thirty-second blocks. It is what "the gold standard" means in this subject, and every other instrument in the list is validated against it. It also costs a night with wires glued to your head. The descriptions of the four instruments here are general background rather than anything this course's sources describe, which is why no lesson prints a figure about any of them.6

Actigraphy. A research device worn at the wrist that records movement and infers sleep from its absence. It doesn't know anything about your brain and it knows everything about whether your arm moved.

A consumer device. A watch, a ring, or a sensor beside the bed, recording movement and heart rate and feeding both into a model the manufacturer has not published. The model is the part to notice, because two devices with identical sensors can disagree about the same night when the software disagrees.

A sleep diary. A person writing down when they think they fell asleep and when they woke. It's the only instrument in the list that measures what somebody experienced, which isn't a weakness so much as a different subject.

Four instruments, four different quantities, one word. Grouping them as four is this course's own organisation: the sources treat them separately and none of them prints this list.6

What happens when you check the devices against the electrodes

In 2021 a team put thirty-four healthy young adults in a sleep laboratory for three consecutive nights, wired for polysomnography, and measured them at the same time with a research actigraph and a subset of seven consumer devices: four worn at the wrist, and three sitting beside or under the bed.1 This course read the abstract of that paper and did not open it, which is the level everything below is held at.

The design, in the authors' own words: "34 healthy young adults (22 women; 28.1 ± 3.9 years, mean ± SD) were tested on three consecutive nights (including a disrupted sleep condition) in a sleep laboratory with PSG, along with actigraphy (Philips Respironics Actiwatch 2) and a subset of consumer sleep-tracking devices."1

And the result that matters, verbatim: "Overall, epoch-by-epoch sensitivity was high (all ≥0.93), specificity was low-to-medium (0.18–0.54), sleep stage comparisons were mixed, and devices tended to perform worse on nights with poorer/disrupted sleep."1

An epoch is one of those thirty-second blocks, so "epoch-by-epoch" means the two records were compared block by block across the whole night rather than as two totals.

Two numbers in one sentence, and they have to be read together.

Predict first

Before the two numbers. Every device in this study caught 93 per cent or more of the minutes you were genuinely asleep. What would you expect the same devices to do with the minutes you were genuinely awake?

Show the answer

Worse, and the size of the gap is the point.

The two jobs are not symmetrical. Lying still with your eyes shut and asleep, and lying still with your eyes shut and awake, look almost identical to anything reading movement, and fairly similar to anything reading heart rate. A device catching your sleep is catching the easy case.

So the number to look for is the one for wake, and it is the one that almost never appears in an advertisement. Read on for what it was.

Sensitivity here is the share of genuinely asleep minutes the device called asleep. At 0.93 and above, every device caught almost all of your sleep.

Specificity is the share of genuinely awake minutes the device called awake. At 0.18 to 0.54, the devices caught between a fifth and a half of the time you'd spent awake in bed, and scored the rest as sleep.

And the same abstract says something that cuts the other way, which this lesson would be dishonest to leave out. Most of the consumer devices "performed as well as or better than actigraphy on sleep/wake performance measures", and the conclusion calls their performance "promising".1 The research instrument is living with the same problem. So the specificity figure isn't a complaint about cheap watches, it's a property of measuring sleep by movement at all, and that makes the point structural rather than a grumble.

One night, three versions of it

Say you were in bed for eight hours and genuinely asleep for seven of them, so you spent an hour awake: settling at the start, a stretch in the middle, twenty minutes before the alarm.

At a specificity of 0.54, the device finds about thirty-two of those sixty minutes and calls the other twenty-eight sleep. At 0.18, it finds about eleven and calls forty-nine of them sleep.

The same night, as it happened and as two devices would report it Three horizontal bars, each eight hours of time in bed. The first is what happened: seven hours asleep and one hour awake. The second is a device at specificity 0.54, which finds 32 of the 60 awake minutes and reports the other 28 as sleep. The third is a device at specificity 0.18, which finds 11 and reports 49 as sleep. The awake portion shrinks from 60 minutes to 32 to 11 while the night itself does not change. One night, three versions of it Eight hours in bed. The end nearest 8 h is time awake What happened: 7 hours asleep, 1 hour awake 60 min At specificity 0.54, it finds 32 of the 60 32 min At specificity 0.18, it finds 11 of the 60 11 min 0 h 4 h 8 h The course's own arithmetic on the quoted range

The arithmetic is this course's own and the simplification in it is deliberate: the picture holds sensitivity at a perfect 1.00 so that you can see the specificity effect on its own.6 In reality a sensitivity of 0.93 also sends some genuine sleep the other way, which pushes the reported total back down, and on a night with little time awake in bed it can outweigh the overcount. Which effect wins depends on how long you lay awake, which the abstract doesn't let you work out.

Notice which way the error leans. A device with those two numbers tends to report more sleep than you got on a night with a lot of lying awake, which is the night you'd most want it to get right. If your watch says you slept well and you feel as though you didn't, the instrument's weakness points exactly that way, and that isn't reassurance. This reading is the course's own.6

Predict first

Before you read on. Why would a device with a specificity of 0.20 still be advertised as "over 90 percent accurate", without anybody lying?

Show the answer

Because of how the night is made up, and once you see it you cannot unsee it.

Take the night above. Four hundred and twenty minutes asleep, sixty awake. A device that simply called every single minute "asleep", with no sensor at all, would be right 420 times out of 480. That's 87.5 percent accurate and it knows nothing.

So a single accuracy figure on a mostly-sleep night is close to meaningless, which is why the paper reports the pair instead of one number. Sensitivity grades the device on the easy majority. Specificity grades it on the minority you actually care about, and that is where the numbers fall apart.

This isn't specific to sleep. Any test for something rare or lopsided has the same shape, and Logic and Argument taught it as base rates. What this lesson adds is that the lopsidedness here is the night itself.

The wrinkle, which is the useful part

The last clause of that quoted sentence is the one to keep: devices "tended to perform worse on nights with poorer/disrupted sleep".1

The people most likely to buy a sleep tracker are the people whose nights are disrupted, and those are the nights the devices read worst. That sentence is this course's own, and the first half of it is a guess about who buys trackers rather than anything measured.6 It's the reason this lesson comes first.

Check yourself

Your diary says you fell asleep at half past eleven. Your watch says ten o'clock. Ninety minutes of your evening is in dispute. Which one is wrong?

Show the answer

Neither, and the question has the wrong shape, which is the whole of this lesson in one example.

Your diary records when you experienced falling asleep, which is the moment you stopped noticing you were awake and is therefore reported from the wrong side of the event. Your watch records when a model decided your movement and heart rate looked like sleep. Those are two different quantities about the same night, and there's no reason for them to agree.

What the gap tells you is worth having. A large one would fit a long quiet settling period, when you were lying still and awake, which is exactly the stretch the specificity figure says a device gets wrong. How often that's the explanation is not something this course has read anything about. A reader who throws out the night because the numbers disagree has discarded the most informative night of the week.

And notice what neither instrument gives you. Polysomnography would've called it by the electrodes, and that number would be a third thing again. None of the three is your sleep; each one is a measurement of it, and this course names the instrument every time it prints a number.

What this course read, and how far in

Eight of the sources behind this course produce a figure that appears in a lesson. Here is what each one measured, on whom, and how far this course got into it.

The source What it measured On whom Read at
The device comparison1 Seven consumer devices and an actigraph against polysomnography 34 healthy young adults, 3 nights Abstract verbatim in full. The paper was not opened
The restriction experiment2 Cognitive performance and sleep physiology at 4, 6 and 8 hours in bed 48 healthy adults aged 21 to 38, across a 14-night experiment and a 3-night one Abstract verbatim in full. The paper was not opened
The light experiment3 How far one hour of bright light moves the body clock, by circadian phase 36 participants in a laboratory Abstract verbatim in full, plus one sentence of the introduction. The methods and results were not read
The consensus statement4 Nothing. It is a panel's recommendation about a literature A panel of experts; this course did not read how the panel worked Read in part: the recommendation and four statements
The mortality meta-analysis5 All-cause mortality against sleep duration 1,382,999 people in 27 cohort samples Read in part, in fragments, from a proof copy
The insomnia guideline[7] Nothing. It is a college's recommendation about trials Adults with chronic insomnia disorder Read in part: both recommendations, the opening paragraph and one box
The chronotype paper[8] A short questionnaire against the longer standard questionnaire it shortens Participants in a validation study Read in part: the abstract and the measure
The sleep and memory meta-analysis[9] Episodic memory after sleep against after waking 823 effect sizes from 271 samples Read in part, from the authors' accepted manuscript

Read the last column before you read anything else. Three abstracts were read in full and five sources were read in part. Six other sources are named in the lessons that use them, each with its own level stated there, and this course opened the whole text of nothing at all.

Read the third column too. Every experiment in this course happened in a laboratory, to healthy adults, over days. The two sources that cover years and millions of people are not experiments. That split is the subject of lesson 3 and it is the single most important thing in the course.

What this course leaves out on purpose

Sleep disorders. Apnoea, narcolepsy, restless legs and the parasomnias are a clinical literature this course has not searched, and it isn't going to improvise about them. If you snore loudly and somebody has seen you stop breathing, if you have fallen asleep at the wheel, if severe sleeplessness has come on suddenly, or if sleeping more does not fix your sleepiness, those are reasons to see a doctor, and no lesson here is a substitute for that.

Children and adolescents. A separate literature with live policy arguments attached. Everything here was measured on adults.

Screens, lamps and ordinary indoor light. Lesson 4 has a controlled measurement of what an hour of very bright light does to the body clock, and nothing whatever about a phone. The course says so once, there, and then doesn't mention screens again.

Dreams, naps and supplements other than melatonin, which appears only because a named guideline names it.

And medical advice of any kind. Lesson 6 reports what three guidelines recommend, with the strength each attaches, and stops at the point where a description would become an instruction.

Three things people get wrong about measuring sleep

"My tracker knows how much I slept." It tends to count quiet wakefulness as sleep, and it did worse on disrupted nights, which are the nights you most want to measure.1

"The sleep stages on my watch are measured." The comparison study calls the stage results "mixed" and "inconsistent" in its own words.1

"A laboratory number is the truth about my sleep." It's the truth about one night with electrodes glued to your head in an unfamiliar room, which is a real measurement of a slightly unusual night.

Practice

Say what would have to exist

Take twenty minutes.

Pick last night. Write down, before you look at anything, when you think you fell asleep, when you woke for the last time, and how long you think you were actually asleep in between.

Then three things.

  1. What instrument would have had to be running to check each of your three numbers.
  2. Which of the three you would trust least, and why.
  3. What you would accept as evidence that you were wrong about the one you trust most.

The third line is the one that makes this worth doing. Most people can say which number feels shakiest and very few can say what would settle it, and an answer you can't imagine checking is lesson 8's subject arriving early.

The course project starts here too. It compares two measurements of the same nights across a week, and its first section is a prediction that cannot be written later: how long you sleep on a work night and on a free night, the midpoint of your sleep on a free night, and how confident you are out of ten, on a scale where ten is certain. Write those four things down now, before lesson 2, and do not edit them afterwards.

Find the instrument

Take fifteen minutes.

Find a claim about sleep with a number in it. A headline, an app's home screen, a podcast, a product. Write it down in the words you met it in.

Then answer one question: which of the four instruments produced that number?

There are only four outcomes, and all four of them are results.

  1. The claim names it. Rare, and it tells you the writer has thought about this.
  2. The claim doesn't name it but the study behind it does, and you can find it in a few minutes.
  3. The trail stops at a summary that doesn't say, which is the commonest outcome.
  4. There is no measurement behind it at all. Some claims about sleep are advice with a number attached for confidence.

Write one line saying which you got. You'll do a fuller version of this in lesson 8, on a claim you care about, and this is the short form, to get the habit in.

Connections

Back. Focus and Deep Work lesson 1 is the instrument question in general, and this lesson is it applied to the one activity you cannot watch yourself doing. Memory lesson 7 is the sample question, which the third column of the table above is an instance of. Logic and Argument is where base rates were taught, which is why a mostly-sleep night makes an accuracy figure useless.

Forward. Lesson 2 is the one experiment in this course that takes sleep away on purpose, and it used polysomnography overnight and performance tasks its abstract never names. Lesson 3 is the same subject measured by asking a million people a question. Lesson 4 is the fifth instrument, a hormone assay, and the quantity it measures isn't duration at all.

Go deeper

  • Performance of seven consumer sleep-tracking devices compared with polysomnography, Chinoy and colleagues, 2021. This course read the abstract and nothing else, and the abstract is the part that carries the two numbers this lesson is built on. If you own a tracker, the device list is worth a look to see whether yours is in it.
  • Your own device's support pages. Not a source and not read for this course, but worth ten minutes: find out whether the manufacturer says anywhere what their sleep staging is validated against, and see whether you can find it.

Sources

  1. Evan D. Chinoy and colleagues, "Performance of seven consumer sleep-tracking devices compared with polysomnography", Sleep 44(5), 2021, article zsaa291. Read at abstract level: the abstract verbatim and in full, from the journal's own page. The paper was not opened, which the body says where the figures appear. Supports: the 34 participants and three nights, the device list, the sensitivity and specificity ranges, the inconsistent staging, and the worse performance on disrupted nights.
  2. Hans P. A. Van Dongen and colleagues, "The Cumulative Cost of Additional Wakefulness", SLEEP 26(2), 2003, pp. 117 to 126. Read at abstract level. Named in the scope table only; lesson 2 is where it is used.
  3. Melissa A. St Hilaire and colleagues, "Human phase response curve to a 1 h pulse of bright white light", The Journal of Physiology 590(13), 2012, pp. 3035 to 3045. Read at abstract level. Named in the scope table only; lesson 4 uses it.
  4. Nathaniel F. Watson and colleagues, "Recommended Amount of Sleep for a Healthy Adult", Sleep 38(6), 2015, pp. 843 to 844. Read in part. Named in the scope table only; lesson 3 uses it.
  5. Francesco P. Cappuccio and colleagues, "Sleep duration and all-cause mortality", SLEEP 33(5), 2010, pp. 585 to 592. Read in part, in fragments. Named in the scope table only; lesson 3 uses it.
  6. The course's own constructions, each labelled where it appears in the body. The one-night arithmetic under the chart is this course's, worked from the quoted specificity range, and it holds sensitivity at 1.00 to isolate one effect, which the body says at the chart. The observation that the people likeliest to buy a tracker are the people whose nights the devices read worst is this course's reading of the paper's last clause, not a statement the paper makes, and the body says so. The four-instrument list is this course's own organisation of material the sources treat separately.
  7. Amir Qaseem and colleagues, "Management of Chronic Insomnia Disorder in Adults", Annals of Internal Medicine 165, 2016, pp. 125 to 133. Read in part. Named in the scope table only; lesson 6 uses it.
  8. Neda Ghotbi and colleagues, "The µMCTQ: An Ultra-Short Version of the Munich ChronoType Questionnaire", Journal of Biological Rhythms 35(1), 2020, pp. 98 to 110. Read in part. Named in the scope table only; lesson 4 uses it.
  9. Sabrina Berres and Edgar Erdfelder, "The sleep benefit in episodic memory", Psychological Bulletin 147(12), 2021, pp. 1309 to 1353. Read in part, from the authors' accepted manuscript on PsyArXiv: the abstract, the design definitions and results, and the Discussion's paragraphs on size and design. Lesson 5 uses it.

Check your understanding

This lesson has a 6-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.