Bad reasoning in the news and everyday life

120 min

Listen: this lesson as a conversation

Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.

In this lesson you will learn to
  • Evaluate a source with the four rules for sources (cite, informed, impartial, cross-check) and explain what to look for when qualified experts disagree
  • Convert a relative-risk headline into absolute terms and natural frequencies, and find the missing base rate, denominator or axis baseline in a claim or a chart
  • Identify survivorship bias in a described sample and explain what Wald's analysis of returning aircraft actually did
  • Apply a repeatable checklist to a news argument and identify motivated reasoning in your own reading of it

On 26 October 2015 the International Agency for Research on Cancer, the World Health Organization's cancer agency, put out a press release. It classified processed meat as "carcinogenic to humans (Group 1)", put red meat one rung down at "probably carcinogenic to humans (Group 2A)", and said that "each 50 gram portion of processed meat eaten daily increases the risk of colorectal cancer by 18%."[3] The release said one more thing in the same breath, quoting the head of its monographs programme: for an individual "the risk of developing colorectal cancer because of their consumption of processed meat remains small, but this risk increases with the amount of meat consumed."[3] Of those four statements, one travelled.

Notice what that pair of sentences tells you about the label. A classification that puts something in the top group and calls the individual risk small cannot be a statement about how much harm the thing does. It is a statement about how sure the evidence is that the thing causes cancer at all. Tobacco is in the same group, and nobody at the agency was saying a rasher is a cigarette.

You met a version of that argument this week. Maybe it was about a food, a drug, a policy, a chart, or a study that "proves" something. Its premises had been squeezed into a headline, its source was standing in for the warrant, and you probably weren't shown the study. This lesson is about checking it anyway. You already own every tool you need, because the previous eight lessons built them; what's new is where to point them, and a short list of questions you can run in your head in the time it takes to read a headline twice.

Start with that 18%. It's a relative risk: a comparison of two rates, given as the ratio between them and nothing else. The number you'd actually want is the absolute risk, which is how many people in a group this happens to. Before you go on, answer this.

Predict first

The 18% is a relative risk. Roughly how many more cases of bowel cancer, per thousand people, are there among those who eat the most processed meat than among those who eat the least?

Show the answer

Cancer Research UK worked it out the same day. About 61 people in every 1,000 in the UK develop bowel cancer. Among 1,000 people who eat the least processed meat, about 56 do. Among 1,000 who eat the most, about 66 do.[4] Ten more per thousand. That's a real risk, and it isn't the risk from smoking; the same article puts about 3 in every 100 UK cancers down to red and processed meat, against about 19 in every 100 for smoking.[4] If you guessed a much larger number, notice that the headline gave you nothing to guess from. That was the problem.

Bowel cancer is what British sources call the disease the release calls colorectal cancer. Same disease, two vocabularies, and this lesson uses whichever word its source used.

News is an argument with the premises squeezed out

A news story isn't a different kind of thing from the arguments in lessons 1 to 8. It's the same thing, compressed. "Bacon causes cancer" is a conclusion. The premises are a study, a statistic, and a source, and the source is standing in for the warrant: the unstated general step, from lesson 7, that licenses you to move from "IARC says" to "it's true". Often you never see the study. You see the source's name and a number, and you decide in a second.

Check yourself

Before you read on, write out the six-step check from lesson 1, in order, from memory. You're about to hang five new questions on it, and they only fit if the frame is already there.

Show the answer

Find the conclusion. Find the premises. Supply what's missing. Test the link. Test the premises. And only then consult your opinion of the conclusion. If you couldn't reach all six, that's the thing to go back to; everything in this lesson hangs off them.

So the check still applies, and its six steps still run in that order. What this lesson adds is five questions that fit the compressed form. Three are for the source: who says, how do they know, who disagrees. Two are for the number: share of what, compared with what.

Three questions for the source

Here is the argument the headlines were actually running, written out the way lesson 1 taught.

1. IARC is an expert body in cancer research.
2. Whether processed meat causes cancer is a claim in cancer research.
3. IARC asserts that it does.
------------------------------------------------------------------------
C: Processed meat causes cancer.

That argument is fine. It's an appeal to authority of the good kind, and by the end of this section you'll be able to say why. What went wrong happened after it: the conclusion got repeated with a size attached that nobody had claimed. Hold that, because it's the shape of most news failures. The source is sound and the sentence about the source isn't.

Anthony Weston's Rulebook for Arguments gives four rules for arguments from authority, which are the first four of his chapter on the subject; the fifth is about the internet. They're the best short statement of how to treat a source that I know.[1]

Rule 13, cite your sources: a precise claim needs a citation, and the reason is that a reader has to be able to go and check. Rule 14, seek informed sources: "Sources must be qualified to make the statements they make." His example is that "for the best information about global climate change, go to climatologists, not politicians", and he warns that "experts on one subject are not necessarily informed about every subject on which they offer opinions." Rule 15, seek impartial sources: "People who have the most at stake in a dispute are usually not the best sources of information about the issues involved." He adds the part people miss, which is that honesty isn't enough, because "the truth as one honestly sees it can still be biased. We tend to see what we expect to see." Rule 16, cross-check sources: "Consult and compare a variety of sources to see if other, equally good authorities agree."[1]

Those four rules are the three source questions. Who says (rule 13, and is the source named at all)? How do they know (rule 14, and are they informed on this)? Who disagrees (rules 15 and 16, and what would an interested source have reason to leave out)? Notice what Weston says an interested source does. It doesn't usually lie. It notices, remembers and passes on what fits what it already expects, which is his point about seeing what we expect to see.[1] So the question to ask of a manufacturer's press release, a campaign's chart, or a charity's appeal is rarely "is this false?" It's "what is this not telling me?"

Check yourself

A pharmaceutical company's press release reports its own trial, and every number in it is true. Rule 15 says an interested source is usually not the best source. Which question does that licence you to ask, and which one does it not?

Show the answer

It licenses "what is this not telling me?" and not "is this false?". Weston's reason is that an interested source doesn't usually lie; it notices and passes on what fits what it expects. So the thing to go looking for is the missing half: the trials that didn't report, the harms given in a different unit from the benefits, the comparison group nobody mentions. Treating a verified number as suspect because of who published it isn't the impartial rule. It's a different mistake, and the last question in this lesson's quiz turns on it.

Check yourself

You met the fuller version of "how do they know" in lesson 8: Walton's six critical questions for an appeal to expert opinion. Write down as many of the six as you can reach before you open this.

Show the answer
  1. Expertise. Is the source credible as an expert at all? 2. Field. Is the claim inside their field? 3. Opinion. What did they actually assert, as against what the story says they asserted? 4. Trustworthiness. Are they personally reliable? 5. Consistency. Is the claim consistent with what other experts say? 6. Backup. Is it based on evidence?[2]

In the news, the field question and the opinion question do most of the work. A physicist on diet, a surgeon on economics: the qualification is real and it's a qualification in something else. Expertise doesn't automatically carry across, which is Weston's point in rule 14. And run the IARC argument above through all six. It passes every one, which is the uncomfortable part: an appeal to authority can pass all six critical questions and still be repeated as something the authority never said. That's question 3 doing its work, and question 3 is the one headlines break.

One source is missing from all of this, and it's the one you trust most: your own experience. The critical-reasoning syllabus this course consulted at California State University Northridge puts media, experts and personal experience side by side as the three sources of belief.[17] Personal experience is the one nobody cross-checks. It's a sample of one, drawn by a process connected to your own life, which is the survivorship problem two sections from now with you as the filter. Run the same three questions on it: who says, how do they know, who disagrees.

When the experts really disagree

Sometimes you run all six questions on both sides and both pass. Two qualified, honest, informed people, in the same field, looking at the same evidence, and disagreeing. Weston's advice for that case is blunt: "reserve judgment yourself. Don't jump in with two feet where truly informed people tread with care."[1]

Reserving judgement isn't the same as shrugging, and it helps to know why two honest experts end up apart, because the reason tells you which question to reach for. They may be looking at different evidence, because one has seen a study the other hasn't. They may weigh the same evidence differently, because they disagree about how far a study of that kind should move anyone, which is lesson 6's likelihood ratio and is a real disagreement about method. They may agree on all of it and disagree about what should be done, which is a value question wearing an empirical coat. Or the disagreement may be manufactured, with one funded outlier standing in for a field.

So, three things to look for. First, is there a consensus statement, a review, or a position from a body that represents the field rather than a person in it? Second, how big is each camp? A field split down the middle is a different situation from a field with a few outliers, and Weston notes that "on some topics the appearance of controversy may be created even when there is virtually no disagreement among qualified authorities."[1] Third, and this is lesson 6's question in new clothes, what evidence would settle it, and has anyone gathered it? A disagreement where both sides can say what would change their minds is a live scientific question. A disagreement where neither can is something else.

That's the editorial standard of this university turned into a reading habit. Sort the claim: is it established across the field, contested inside it, or a question of value that no study could settle? News arguments blur the three, and a good deal of your work is un-blurring them.

Check yourself

Two qualified experts, both inside their field, disagree in public about whether a common supplement helps memory. Name three things you would look for before deciding what to think.

Show the answer

A consensus statement or systematic review from the field, not from either person. The size of each camp: is this a split field or one outlier against the rest? And what evidence would settle it, whether it exists, and whether either expert has said what would change their mind. If the third question has no answer, you may be looking at a value disagreement dressed as an empirical one.

Share of what: the first number question

Now the other half of the checklist. A news number usually arrives as a rate or a ratio, and both hide something.

In October 1995 the UK Committee on Safety of Medicines warned that third-generation oral contraceptive pills raised the risk of blood clots "twofold, that is, by 100%". It went out in "Dear Doctor" letters to 190,000 general practitioners, pharmacists and directors of public health, which in the UK means family doctors and the officials who run local public health.[5] Gigerenzer and his colleagues, who tell the story in their 2007 review of health statistics, ask the only question that matters.

Predict first

"Twofold" is a relative risk. How many extra cases of thrombosis, per 7,000 women, was the committee warning about?

Show the answer

One. The studies behind the warning had shown that, "of every 7,000 women who took the earlier, second-generation oral contraceptive pills, about 1 had a thrombosis; this number increased to 2 among women who took third-generation pills."[5] The absolute increase was 1 in 7,000. The relative increase was indeed 100%. Both statements are true. Only one of them tells a woman what she needs to know.

The consequences are documented. Women stopped taking the pill. Gigerenzer and colleagues report an estimate of "13,000 additional abortions in the following year in England and Wales", with one extra birth for every extra abortion and about 800 additional conceptions among girls under 16.[5][6] Their conclusion: "Had the committee and the media reported the absolute risks, few women would have panicked and stopped taking the pill."[5]

Two things this section isn't doing. It isn't saying the warning was wrong: the studies did find a difference between the two generations of pill, and 1995 is where these figures come from and where they stay. And it isn't advice about any medicine. As lesson 6 said once for the whole course, this is education and not personal medical advice; nothing here tells you what to take, or which screening to attend. What it tells you is how to read the number when somebody else reports it to you.

A relative risk is a ratio of two rates, and a ratio throws the base away. That's the whole trick, and it isn't even a trick; it's arithmetic. Divide 2 in 7,000 by 1 in 7,000 and the 7,000 cancels. Cancelling is what ratios are for, which is why a relative risk is so easy to put in a headline: one number, no denominator. But the denominator is where the answer was living. "Doubles the risk" of something that happens to 1 in 7,000 people means 1 extra case in 7,000. "Doubles the risk" of something that happens to 1 in 10 means 1 extra case in 10. The word "doubles" can't tell you which, so a headline that gives you only the ratio has left out the base rate, which is the share of people this happens to before anything else is taken into account, and you're the one who has to put it back.

You put it back by counting, which is lesson 6's method: take a round number of people, apply the base rate to it, apply the change, and compare the two counts. Gigerenzer and Hoffrage found that switching a Bayesian problem from probabilities to counts raised correct answers from 16% and 28% to 46% and 50%.[16] It works on a risk headline for the same reason: counting forces the denominator back onto the page.

Now do the bacon story yourself. Cancer Research UK's figure for the people who eat the least processed meat is 56 in 1,000. The press release says 18% higher for each 50 gram portion eaten daily.

Predict first

Take 1,000 people who eat the most. Work out roughly how many of them develop bowel cancer, then say what the honest one-sentence version of the headline would have been.

Show the answer

Raise 56 by 18% and you get about 66, which is the figure Cancer Research UK published the same day.[3][4] Ten more in a thousand. So the honest sentence is "ten more people in every thousand, at the top of the range", not "bacon causes cancer". Notice the assumption you had to make to get there: the 18% is per 50 gram portion eaten daily, so applying it once assumes the heaviest eaters are about one daily portion above the lightest. Cancer Research UK's 66 came from its own data rather than from that multiplication, and the two agree. And notice which way the conversion runs. From 56 and 66 you can always recover the ratio; from the ratio alone you can never recover 56 and 66. That's why the ratio is the number that gets printed and the counts are the numbers you have to go and find.

Bowel cancer per 1,000 people. Among 1,000 people who eat the least processed meat, about 56 develop it. Among 1,000 who eat the most, about 66 do. Each outlined bar is 1,000 people, so both filled portions are small and the difference between them is ten people. Eat the least 56 Eat the most 66 the 10 extra

Each outlined bar is 1,000 people, which is the denominator the headline dropped; the dark block is the count that develops bowel cancer, and the sliver on the second bar is the ten extra. The sliver is almost invisible at this width, and that is the honest picture of an 18% relative rise on a base of 56 in 1,000. Figures from Cancer Research UK, 26 October 2015.[4]

Two more things the story needed. The Group 1 label grades the strength of the evidence, not the size of the risk, which you can read off the release itself: the same document that assigns the top category also says the individual risk "remains small".[3] And that second sentence was in the source. The version that reached most readers didn't carry it, and Weston's rule 15 names what that is: not a lie, an omission.

Mismatched framing

Gigerenzer and colleagues report an analysis of the BMJ, JAMA and The Lancet from 2004 to 2006: of the studies that reported both the benefits and the harms of a treatment, one in three used different metrics for the two, mostly relative risks for the benefits and absolute frequencies for the harms.[5] A 40% reduction in one line, 1 in 1,000 in the next. When you see two numbers in different units in the same story, convert both to counts per thousand before you compare them.

About five minutes, on the same relative-versus-absolute distinction this section teaches, from the researcher whose review the pill scare comes from. Watch it after you've done the bacon conversion yourself, so you have your own method to set against his. We have verified the channel, the title and the length, and nobody working on this course has watched it end to end, so if it says something different from this lesson, back the lesson's sources.
Check yourself

A headline says a new screening test "reduces deaths from the disease by 25%". Gigerenzer and colleagues open their review with a case exactly like it. What one question do you ask, and what does the answer look like?

Show the answer

Share of what? A 25% relative reduction means nothing until you know how many die without screening. Their own example: the statement that mammography screening reduces the risk of dying from breast cancer by 25% "in fact means that 1 less woman out of 1,000 will die of the disease."[5] One woman in a thousand is a real benefit and a small number, and both facts belong in the story. The review gives the reduction and not the base, and you can recover the base from what it does give, which is a piece of arithmetic worth doing once: if a quarter of the deaths is one in a thousand, the deaths were four in a thousand. That last step is this lesson's, not the review's.

Compared with what: the second number question

The second number question is about the denominator's other job: which people, or which cases, went into it. The best-documented example I know is Abraham Wald's work on aircraft, and you have probably heard a version of it. This section tells the version in the primary account, which is Mangel and Samaniego's 1984 exposition in the Journal of the American Statistical Association, together with the reprint of Wald's own wartime memoranda.[7][9]

The problem is one of decision under weight. A commander has to decide how much armour to put where, and armour costs range and payload. Mangel and Samaniego state the obstacle in one sentence: "The operational commander does not know the distribution of hits on an aircraft that did not return. This is the basic difficulty in making a decision."[7]

Predict first

Returning planes show many hits on the fuselage and few on the engines. You have a limited weight of armour. Where does the pattern on the survivors tell you to put it, and why?

Show the answer

On the engines. The planes are the sample; the population includes the ones that didn't come back. A part that shows few hits on survivors, when it ought to have been hit as often as any other part of its size, is a part where a hit tends to stop the plane returning. The damage you see is the damage a plane can survive. That's the intuition. What follows is what Wald actually did with it, which is more careful than the intuition and worth knowing.

Wald wrote the method up in 1943 as a series of eight memoranda for the Statistical Research Group and the National Defense Research Committee, over a hundred pages in all. The Center for Naval Analyses reprinted them in 1980.[9] Mangel and Samaniego redo his calculations on one hypothetical data set so you can follow the logic. Suppose 400 aircraft fly a mission. Of those returning, 320 have no hits, 32 have one hit, 20 have two, 4 have three, 2 have four and 2 have five.[7] Add those up and 380 came back, so 20 were lost, and the returning planes carry 102 hits between them.

The first move is an assumption you should notice, because the whole method rests on it. Wald also assumed that "an unhit aircraft always returns."[7] Every one of the 20 missing planes was hit. From the pattern of hit counts on survivors, and assuming for the simple version that each hit has the same chance of bringing a plane down, Wald's basic equation gives the probability that a plane survives one hit. For these data it's 0.851: about 15 in 100 hits are fatal.[7]

Now the part the retellings keep, worked the way the paper works it. Divide the plane into four parts: engines, fuselage, fuel system, and everything else. Mangel and Samaniego take the engines to be 26.9% of the plane's surface area, the fuselage 34.6%, the fuel system 15.4% and the rest 23.1%. If bullets landed at random, that's the share of hits each part should get. Set that against the share of the 102 hits each part actually carried on the returning planes: 18.6% on the engines, 38.2% on the fuselage, 17.6% on the fuel system, 25.5% elsewhere.[7] (In counts, that's about 19, 39, 18 and 26.)

For each of four parts of the aircraft, the share of surface area and the share of hits observed on returning planes. Engines 26.9% of area against 18.6% of hits; fuselage 34.6% against 38.2%; fuel system 15.4% against 17.6%; everything else 23.1% against 25.5%. The engines are the only part where the observed share falls below the expected share. share of surface share of hits seen Engines 26.9% 18.6% Fuselage 34.6% 38.2% Fuel system 15.4% 17.6% Everything else 23.1% 25.5%

Read each pair against itself, not against the other pairs. Three of the four parts took a slightly larger share of the hits than their share of the surface. The engines, ringed in red, are the one part where the second bar is shorter than the first: 26.9% of the surface, 18.6% of the hits that came home. Figures from Mangel and Samaniego's worked example, which uses a hypothetical data set.[7]

Under Wald's assumptions, the reason for that gap is that planes hit in the engines tended not to come home. The formula is a ratio: the probability of surviving a hit to a part is the overall survival probability, times the part's share of observed hits, divided by the part's share of area. For the engines, 0.851 times 0.186 divided by 0.269 gives 0.588. A hit on an engine is survived about 59% of the time, against 85% for a hit anywhere.[7]

Predict first

Your turn, on the fuselage: 0.851 times 0.382, divided by 0.346. Work it out, then say what the number means and why it comes out so much higher than the engines' 0.588.

Show the answer

About 0.94, and Mangel and Samaniego's table gives 0.940 for the fuselage, 0.973 for the fuel system, 0.939 for everything else, and 0.588 for the engines. They conclude: "For these data, the most vulnerable portion of the aircraft is the engine area."[7] It means a plane that takes a hit on the fuselage comes home about 94 times in 100. It's high because the fuselage collected a larger share of the hits on survivors, 38.2%, than its share of the surface, 34.6%, and that surplus is the fingerprint of a part you can be hit on and still fly. The engines' deficit is the same fingerprint reversed. So the four numbers aren't four facts about armour; they're one comparison, expected against observed, run four times. If your answer came out above 1, you divided by the hit share instead of the area share; the area share is always the denominator, because it's what you expected.

Wald's own memorandum, on its own example, says such a table "can be used as guides for locating protective armor and can be used to make a prediction of the estimated loss of a future mission."[9] So the armour does go where the survivors weren't hit. Not because Wald had a hunch, but because the observed share of hits, set against the expected share, measures how lethal a hit to each part is, and that comparison is only possible once you have written down the sample you can see and the population you can't.

Outline of an aircraft seen from above, with red dots marking hits spread over the wings, tail and fuselage and few elsewhere
An illustration of the idea, not Wald's data: a hypothetical damage pattern on returning aircraft, with hits clustered where a plane can be hit and still fly home. Wald's memoranda worked from counts of hits per part set against each part's share of surface area, as in the example above. Image by Martin Grandjean (vector), McGeddon (picture), US Air Force (hit plot concept), Wikimedia Commons, CC BY-SA 4.0.

This belongs in a lesson about the news, and not because of the clever statistician. Survivorship is a sampling error, the one from lesson 5 where the sample fails to represent the population. The returning planes are a sample drawn by a process that depends on the very thing you want to measure. So are the businesses in a book about successful businesses, the graduates a university puts on its website, the people who answer a survey about the survey's own subject, the studies that got published because they found something. Whenever the process that put a case in front of you is connected to the outcome you're judging, ask what the process filtered out. That's "compared with what?" in its most general form: compared with the cases you didn't see.

There's a second lesson here, and it's about sources. The popular version of this story, the one with the generals wanting to armour the bullet holes and Wald correcting them, is a retelling. The memoranda contain a method and worked examples; they don't contain that scene. Even the question of whether the method was used in the war has two answers in the same journal issue. The paper's summary says the work "was used in World War II and in the wars in Korea and Vietnam."[7] In their rejoinder to the discussion, the same authors are more careful: "We do not know whether it was used during World War II, although it was produced early enough in the war to have been available", and they then document its use by the Operations Evaluation Group on the A-4 aircraft during the Vietnam War and at Wright Patterson on the B-52.[8] When a story about reasoning has been polished into an anecdote, go to the primary account and see what it actually supports. The primary account here is more interesting than the anecdote, and it's also less certain.

Check yourself

A magazine profiles ten people who dropped out of university and founded successful companies, and concludes that university is overrated for founders. In Wald's terms, which planes are missing, and what would you need to see?

Show the answer

The dropouts who founded companies that failed, and the graduates who founded companies at all. The ten profiles are the planes that came back. To judge whether dropping out helps, you'd need the rate of success among dropouts who tried and among graduates who tried, which is the population the magazine's sample was filtered from. Ten survivors can't give you a rate.

Charts: when the picture is the argument

A chart is a compressed argument too, and it has its own ways of removing the base. One useful set of examples comes from the Economist's own data team, because these are the outlet's own charts and the outlet published the corrections itself. In March 2019 Sarah Leo, of that team, went back through the paper's archive and published a set of its own charts that had gone wrong, sorted into three kinds: misleading, confusing, and failing to make a point.[10]

Two of the misleading ones matter here. One was a line chart from the paper's Espresso app, tracking how people felt week by week about the result of the UK's 2016 referendum on leaving the European Union. Drawn on a scale that was too sensitive, it made respondents look as if they "had a rather erratic view of the referendum result".[10] The other was a bar chart of the average number of Facebook likes on posts by pages of the political left, drawn with the vertical axis truncated so that it didn't start at zero. Leo's verdict was that the chart "not only downplays the number of Mr Corbyn's likes but also exaggerates those on other posts", Jeremy Corbyn then being the leader of the UK Labour Party.[10] Both of these happen to be about British politics, because that's what the paper was charting that month, and both were caught and published by the outlet that drew them. The pattern has nothing to do with the subject: an axis that doesn't start at zero exaggerates every difference on it, whatever it's measuring.

Predict first

A chart shows two bars, this year and last, and this year's bar is twice the height. The printed numbers are 1,040 and 1,120. Before you read on: what has to be true of the vertical axis, and what would the chart look like if it started at zero? (Invented numbers, real chart trick.)

Show the answer

The axis has been cut, so it starts somewhere near 1,000 rather than at zero, and the bars show the part above the cut rather than the quantities themselves. Started at zero, the two bars are almost the same height, which is what an 8% rise looks like. Nothing here is a lie: the numbers are printed and the picture is drawn from them. What's gone is the base the bars are measured from, which is the same thing "doubles the risk" removes.

Both are the same error as "doubles the risk". A truncated axis drops the zero, which is the base the bars are measured from; a hypersensitive scale magnifies the change and hides how small it is against the whole. A third way to move a picture without moving the numbers is to change what the axis measures: on a logarithmic scale each step up is a multiplication rather than an addition, which is honest and useful for quantities that grow by factors, and misleading if a reader takes the steps for equal amounts.

So ask the two number questions of any chart. Share of what: where does the axis start, and what is the full range? Compared with what: is the change drawn against its base, or against itself? A chart that doesn't let you answer those isn't necessarily dishonest, and it is unfinished, because the reader can't do the check the chart invites. The Economist's own examples are the evidence that a bad chart is usually a busy afternoon rather than a plan. And notice who caught these: the outlet, on itself. A source that publishes its own errors has handed you evidence about how it works, and evidence is what you were looking for. It's narrow evidence, about this team's charts rather than about the paper's politics or its reporting, and what you do with it is your call. But a source that never publishes a correction has given you nothing to weigh.

A good place to stop

That's the five questions and the four cases they were built on: a press release, a pill warning, a set of returning aircraft, and a newspaper's own charts. If you're reading in one sitting and want a break, take it here. What's left is the same five questions run on two politicians' numbers, and then on you, and the second of those lands better on a fresh head. Before you start again, see how many of the five you can name without looking.

A matched pair: the same argument, two sides, one verdict

One pattern turns up in political news of every stripe, so here are two documented instances of it, chosen because the checklist reaches the same verdict on each. The pattern is: a good number happened while I was in charge, so I caused it. Lesson 5 called its bad version post hoc; the checklist calls it "compared with what?"

On 3 February 2024 President Biden posted on X: "The last guy had the worst jobs record since the Great Depression", with a chart titled "Jobs Created by President" showing a monthly average loss of 57,000 jobs under President Trump against a large monthly gain under Biden. FactCheck.org checked it on 9 February. The numbers were real. But the chart's baseline for Trump included April 2020, when pandemic shutdowns removed 20.5 million jobs in a month, and its figure for Biden included the recovery of those same jobs. The verdict: the chart also "leaves the misleading impression that presidents are responsible for all the job creation, or loss, during their time in office. But there are many economic factors outside the control of a president (see: COVID-19)."[11] The honest comparison it supplies is the post-recovery rate: "Since then, the job growth under Biden has been an average of 282,000 per month", which it notes is "still 100,000 more than the pre-pandemic average under Trump."[11] A real gain, a smaller one, and one the chart could have shown.

Predict first

You have just seen a fact-checker confirm Biden's two figures and reject his inference. The next case is a White House release under Trump saying the stock market "rebounded strongly under President Trump's leadership", and the rise it rests on is real too. Before you read the verdict, write down what you expect the checker to say, and how confident you are.

Show the answer

The same verdict, for the same reason: the number stands and the comparison is missing. What's worth catching is your own confidence a moment ago. If you were surer about one of the two before you had read either check, that wasn't the evidence talking, because you had the same evidence about both. Hold on to whichever way it went; a later section is about it.

On 16 February 2026 a White House press release said the stock market had "rebounded strongly under President Trump's leadership". FactCheck.org checked it on 19 February. Again the number was real: the S&P 500 had risen 14.5% from the close on 17 January 2025 to the close on 18 February 2026. And again the comparison was missing: "The stock market performed well in Biden's final two years in office", with the S&P 500 "rising over 20% each of those years", which the checkers note was "better than the 13% gain Trump saw in his first year."[12] The market rose under both presidents; it rose faster in the years before this one.

One is a personal post and the other an institutional release. It makes no difference to the check: in both cases the claim comes from the side whose record is at stake, and that's what rule 15 is about. Now run the checklist on both, in its own order.

Conclusion and premises first, written out. Here is the first:

1. The monthly average under Trump was a loss of 57,000 jobs.
2. The monthly average under Biden was a large gain.
3. [Whatever happens to jobs on a president's watch is caused by that
   president.]
------------------------------------------------------------------------
C: The last guy had the worst jobs record since the Great Depression.

And the second:

1. The S&P 500 rose 14.5% between the close on 17 January 2025 and the
   close on 18 February 2026.
2. [Whatever happens to the market on a president's watch is caused by
   that president.]
------------------------------------------------------------------------
C: The stock market has rebounded strongly under President Trump's
   leadership.

Step 3 of the check is to supply what's missing, and in both cases the missing step is the same sentence with a different noun in it. Written down, it's not a claim either side would sign, which is why the same lines work with the names swapped. Step 4, the link: it fails there, on the bracketed premise, and "compared with what?" is the number question that finds it, because the comparison neither claim supplies is the same series before the president arrived. Step 5, the premises: share of what is fine, both figures are honest as stated, and government statistics and market indices are sound ways to know. Who says: the person whose record is at stake, which is rule 15 exactly, so the question is what's left out rather than whether the number is false. Same pattern, opposite sides, same verdict: the number stands, the argument doesn't.

Who disagrees? In both cases FactCheck.org, which confirmed the number and rejected the inference. Now run rule 15 on the checker too, because the rule applies to everyone. What is a fact-checker's stake? Its standing rests on being seen to treat both parties alike, which gives it a reason to be even-handed and also a reason to want to look it. So don't take "they have no agenda" on trust; look for the checkable version. Does it apply the same rule to both sides? It says so in its own words, about one president: "Opinions also differ on how much credit or blame a president should get for what happens while he is in office", and about the other: "We also make no judgement on how much credit or blame the president should receive."[13] That's a thing you can go and verify. "No stake" is not.

One more thing, and it's the kind of thing this lesson exists to make you notice. Look at what each correction leaves behind. In the first case the honest number still beats the predecessor's; in the second the honest number is smaller than the predecessor's. Both leftovers happen to favour the same man. That's what these two checks found, and it isn't what the section is about: what the section is about is that both arguments failed the same test for the same reason. Whether either leftover number means anything is a separate question, and it would take two more checks to answer. If you caught yourself reading the leftovers as a scoreboard, the next section has arrived early.

You may have noticed which of the two you wanted to defend.

Motivated reasoning: belief bias at scale

In lesson 2 you met belief bias. Evans, Barston and Pollard gave people syllogisms and found that they accepted invalid arguments with believable conclusions and rejected valid ones with unbelievable conclusions; the effect "was more marked on invalid than on valid syllogisms", and some of the reasons people gave afterwards were "rationalizations for prejudiced decisions", prejudiced there in its older sense of pre-judged.[14] That was in a lab, with syllogisms about cigarettes and vitamin tablets. The news is the same experiment with your own commitments as the material.

Predict first

Kahan, Peters, Dawson and Slovic gave people a table of results, the kind where you have to compare ratios rather than eyeball the biggest number, and asked what it showed. For half the sample the table was about a skin-rash treatment; for the other half the identical numbers were about a gun-control ban. Among the people who were best at arithmetic, what happened on the gun version?

Show the answer

They got further apart, not closer. On the rash version the most numerate "did substantially better" than the rest; on the gun version responses "became politically polarized, and even less accurate", and the polarisation "did not abate among subjects highest in numeracy; instead, it increased."[15] Most people guess that skill helps a little or makes no difference. The finding is that skill made it worse.

That study ran in 2017. The problem it set "turned on their ability to draw valid causal inferences from empirical data", and the authors' reading is that skilled reasoners "use their quantitative-reasoning capacity selectively to conform their interpretation of the data to the result most consistent with their political outlooks."[15] The design was symmetric: the data were arranged to favour each side for half the participants, so the polarisation isn't an artefact of which answer the numbers happened to support. This is one study, and what this course has read of it is its abstract, so treat the direction as better supported than any single number in it. And note what direction it points. This is not a finding about one political camp. It's a finding about people who are good at reasoning, which is what this course is trying to make you.

The mechanism is lesson 2's, with the volume turned up. Your mind checks a conclusion against what you already believe, fast, and checks the link from premises to conclusion, slowly. On a topic where you have no stake, the slow check runs. On a topic where you have a conclusion already, the fast check fires first, and when it says "yes" the slow check often never starts, and when it says "no" the slow check runs at full power, looking for the flaw it's sure must be there. So you scrutinise the link on the other side's arguments and the conclusion on your own. Both feel like thinking. Only one of them is checking.

Check yourself

You found a fallacy in an argument you disliked in about ten seconds. What should that speed tell you?

Show the answer

That the fast conclusion-check fired and the slow link-check may not have run. Ten seconds is long enough to notice that you disagree; it's rarely long enough to reconstruct an argument, supply its missing premise, and test the link. The fallacy you found may be real. The speed is a reason to run the check again, and then to run it on the last argument you agreed with just as quickly.

The defence is procedural, because the bias can't be felt from the inside. Two rules. First, step 6 of the check stays at step 6: test the link before you consult your opinion of the conclusion, and if you notice you consulted it first, start again. Second, when you catch yourself agreeing quickly, run the checklist on that argument as if the other side had written it. Not to change your mind. To find out whether the argument you accepted is one you'd have accepted from anyone.

What a claim that survives looks like

Everything checked so far has failed, which is a distorted picture of what checking is for. So run the same five questions on the other article from 26 October 2015, the one that gave you the numbers you've been using. Who says: Cancer Research UK, named, with a named author. How do they know: from the IARC release and from UK incidence figures, both cited. Is it their field: yes, and the claim is inside it. Who disagrees, and what has this source reason to leave out: it's a charity with a stake in how cancer is talked about, which is a real interest and the right question to ask, and what it did with that interest was print the absolute numbers on the day. Share of what: 56 and 66 in 1,000, given. Compared with what: about 3 in 100 UK cancers against about 19 in 100 from smoking, given.[4]

That's a pass, and a pass has a shape. It doesn't mean the claim is certain; it means the argument earns the confidence it asks for, and you can hold it at that strength and pass it on with its base rate attached. Most of what you read will land somewhere between this article and the coverage it was correcting. The checklist has three outcomes, not one: the argument holds, the argument fails, or you can't tell yet and you've written down the one thing that would settle it. A reader who only ever reaches the second isn't a careful reader; they've swapped one reflex for another, and the new one is cheaper because it never has to be right about anything.

The checklist

Check yourself

Close the page. Write out the six-step check in order, then the three questions for the source and the two for the number, and against each of the five write the lesson it came from.

Show the answer

Find the conclusion; find the premises; supply what's missing; test the link; test the premises; and only then consult your opinion of the conclusion. For the source: who says (lesson 1's premises, and Weston's rule 13); how do they know, and is it their field (lesson 8's critical questions); who disagrees, and what has this source reason to leave out (Weston's rules 15 and 16). For the number: share of what (lesson 6's base rates); compared with what (lesson 5's samples). If you couldn't reach the six steps, that's the thing to reread; everything else in this lesson hangs off them.

The whole check, in one place. It's the six-step check from lesson 1 with five questions added for the compressed form, and one step at the end that says what the answer is worth.

  1. What is the conclusion? (What is this trying to get me to accept?)
  2. What are the premises? (The study, the statistic, the quote, the source.)
  3. What is missing? (Write the unstated step down. In the news it's often "the source is right" or "what happened after X was caused by X".)
  4. Does the link hold? (Valid, or strong? If a number is doing the work, go to the number questions.)
  5. Are the premises acceptable? (Go to the source questions.)
  6. Only now: what do I think of the conclusion, and did I think it before step 1?
  7. What has this argument earned? Say it in a sentence: it holds and I can repeat it with its numbers; it fails, and here is the step it fails at; or it isn't settled, and here is the one thing that would settle it.

For the source: who says? How do they know, and is it their field? Who disagrees, and what has this source reason to leave out?

For the number: share of what? And compared with what, in both of its senses: compared with which cases, meaning who or what got left out of the denominator, and compared with what baseline, meaning what the number was doing before the thing you are crediting.

It's short enough to run on a headline in the time it takes to read the headline twice, and it's worth ten minutes on anything you are about to pass on.

What people get wrong

"The media lies." As a general claim it can't be checked, which is the first thing wrong with it, and it's a hasty generalisation from the cases you noticed, which is lesson 5's reason. It's also useless, because it gives you no way to tell a good story from a bad one. The rules in this lesson are source by source and claim by claim. The Economist publishing its own mistaken charts, and Cancer Research UK putting the absolute numbers out on the day of the IARC release, are both media too.

"Experts disagree, so nobody knows." Sometimes true. Usually it means one of the questions hasn't been asked: is it their field, how big are the camps, what would settle it. And sometimes the disagreement is manufactured, which is what cross-checking reveals.

"A doubling of risk is a lot." It depends entirely on the base. 1 in 7,000 to 2 in 7,000 is a doubling. So is 1 in 10 to 2 in 10. The word tells you nothing until you have the count.

"I'm not biased; they are." Kahan's result is that the most numerate people on both sides were the most polarised. If you're good at this, you're better equipped to fool yourself, and the only defence is the procedure.

"Wald was clever." He was. But the lesson is about samples, not about Wald, and the reason to learn it is that the next time you're shown a set of survivors, of any kind, the question to ask is what filtered them.

Practice

Four items, then your own. Items 1 to 3 are documented and cited; anything invented in this lesson says so where it appears. For each of the first three, write the verdict, the one comparison or sentence that would change it, and the one question you would put to the writer. You have seen these cases; what you have not done is the last two steps, and those are the ones you'll need on a story nobody has worked for you.

  1. The bacon story, from the side the body didn't check. The lesson ran the source questions on IARC. Run them on the coverage instead: who wrote the headline, how did they know, is it their field, and what did they have reason to leave out?[3][4] Then write the sentence the story could have carried that would have stopped the bacon-and-cigarettes comparison.
  2. A follow-up article says "the WHO has declared red and processed meat carcinogenic". IARC put red meat in Group 2A, "probably carcinogenic to humans", and processed meat in Group 1.[3] Which of Walton's six questions does that sentence fail, and what would the honest sentence say?
  3. The credit-taking pair: Biden's jobs chart of 3 February 2024 and the White House statement of 16 February 2026.[11][12] Write the missing premise in the arguer's own words, name the one comparison that would settle it, and then check yourself: did you reach the same verdict on both? If not, which of the two did you want to be right, and how long did each one take you?
  4. A charity reports that of the 200 people who completed its job-training programme last year, 140 were in work six months later, and concludes that the programme works. (Invented figures.) Two numbers are missing. Name them, say which one makes this a survivorship problem rather than a sampling problem in general, and write the comparison you'd need before believing the conclusion.

Then the conversion. A newspaper reports that a drug "cuts the risk of hip fracture by 40%". The trial's own figures, which the article doesn't print, are that 5 women in every 1,000 broke a hip over three years without the drug and 3 in every 1,000 with it. (The numbers are invented for the exercise; the shape is not.) Write the two natural frequencies, then the relative risk reduction, then the absolute one, then the number of women who would have to take the drug for three years for one fracture to be prevented. Say which of those four numbers a newspaper prints, and which one a woman deciding whether to take it needs.

Check yourself

Model answers for items 1 to 4 and the conversion. Open this only when all five are written out.

Show the answer
  1. Who says: a newspaper desk working from a press release. How do they know: from the release, and behind it a monograph the desk almost certainly did not read. Is it their field: IARC's is; the headline writer's is not, and the claim that got printed is not the claim IARC made. Who disagrees, and what has this source reason to leave out: nobody disagrees with the 18%, and what was left out was the release's own sentence that the individual risk "remains small", plus the fact that the classification grades the strength of the evidence and not the size of the risk. Verdict: the premises are acceptable and the link breaks at step 3, on an unstated premise the coverage never wrote down, that the same group means the same danger. The question to the writer: what does Group 1 mean, and what is the absolute risk? Cancer Research UK answered both on the day.

  2. IARC put red meat in Group 2A, "probably carcinogenic to humans", which is a weaker verdict resting on weaker evidence. A sentence that says "the WHO has declared red and processed meat carcinogenic" welds a Group 2A finding to a Group 1 one and reports the pair at the strength of the stronger. That's Walton's third question, what did they actually assert, failing in a headline. The honest sentence names the two grades separately and says what each one measures.

  3. Who says: the person whose record is at stake, so ask what is left out rather than whether the number is false. How do they know: government statistics and market indices, which are fine. Who disagrees: a checker with no visible stake in either party's fortunes, which is a claim you check by looking at whether it treats both sides alike. Share of what: honest as stated. Compared with what: the same series before the president arrived, which is the comparison neither claim supplied. Missing premise, in the arguer's words: "the figures for the months I was in office are a measure of what I did." The one change that would flip the verdict: a comparison series, or any evidence that a named policy moved the number. The question to the writer: what would this chart look like for a president who did nothing?

  4. Missing: how many started and did not complete, and what share of comparable people who never enrolled were in work six months later. The first is the survivorship half; the 200 completers are the planes that came back, and dropping out may be connected to the same things that keep a person out of work. The second is "compared with what". Wald's move was to set the share observed against the share expected; here the share observed is 70% in work, and until you have a figure for people who didn't do the programme there is nothing to set it against.

The conversion. Without the drug, 5 women in 1,000 break a hip in three years; with it, 3 in 1,000. Relative risk reduction: 2 out of 5, which is 40%. Absolute risk reduction: 2 in 1,000. Number needed to treat: 1,000 divided by 2, so 500 women take the drug for three years for one fracture prevented. A newspaper prints the 40%. The woman needs the 2 in 1,000 and the 500, and the honest story carries all three. If you wrote "40% of 5" and stopped, you did the arithmetic and not the reasoning: a percentage of a rate is still a rate, and a decision gets made on the count.

On paper, with the argument you agreed with fastest

Open the folder of arguments you've been keeping since lesson 1. Pick the one you agreed with most on first reading. Run the checklist on it as if the other side had written it: find the conclusion, the premises, the missing step; test the link; run the three source questions and the two number questions; and only then consult your opinion. Write the verdict in a sentence, and under it write whether it changed, and whether you noticed any point where you wanted to stop checking, which is the line to write most carefully.

Now close the page and write down, from memory: the six steps of the check, in order; the three questions for the source, and which of Weston's rules each one carries; the two questions for the number, and both senses of the second; what a relative risk hides and the one move that puts it back; the three things to look for when two qualified experts disagree; what the returning aircraft are a sample of, and what Wald set against what to get his answer; the two ways a chart can drop a base; what the speed of your own agreement tells you; and the three outcomes the checklist can reach. Then check. Reread only what you missed.

Connections

Everything in this lesson was built earlier. The missing premise (lesson 1) is what "compared with what?" finds in the credit-taking pair. Belief bias (lesson 2) is motivated reasoning at the scale of one syllogism. Sampling and representativeness (lesson 5) are what survivorship breaks. Base rates and natural frequencies (lesson 6) are what "share of what?" restores. The six critical questions (lesson 8) are "how do they know?" spelled out. Lesson 10 turns all of it on your own writing: you will build an argument that a reader running this checklist cannot stop.

Two other Foval courses will own parts of this properly when they are published. How We Know will cover evidence hierarchies and how to weigh expert disagreement in depth. Statistics for Citizens will cover sampling, risk and the design of studies. This lesson taught each only as far as it bears on checking an argument you meet in the news.

Go deeper

  • Weston, A Rulebook for Arguments, 5th ed. (Hackett, 2017), chapter IV, "Arguments from Authority": six pages, and the best short statement of how to treat a source, with rule 17 on the web.
  • Gigerenzer, Gaissmaier, Kurz-Milcke, Schwartz and Woloshin, "Helping Doctors and Patients Make Sense of Health Statistics", Psychological Science in the Public Interest 8(2), 2007: the pill scare, mismatched framing, and the mammography example this lesson uses, and a great deal more on how health numbers are reported; free full text, and readable by anyone.
  • Mangel and Samaniego, "Abraham Wald's Work on Aircraft Survivability", JASA 79(386), 1984, with the discussion and rejoinder in the same issue: the primary account, and worth reading for the method's assumptions, which the anecdote leaves out.
  • Leo, "Mistakes, we've drawn a few", The Economist on Medium, March 2019: a journalist correcting her own outlet's charts, sorted into misleading, confusing, and failing to make a point.
  • Kahan, Peters, Dawson and Slovic, "Motivated numeracy and enlightened self-government", Behavioural Public Policy 1(1), 2017: the study behind the motivated-numeracy result. The abstract is free and states the finding; the full text sits behind a paywall.

Sources

Seventeen entries. Three carry no link: two because no free public address is on file (Walton; Gigerenzer and Hoffrage 1995), and one because the only copy on file is a complete in-print book on a school district's file server, which this project does not link to. Where a free document exists, it is linked in the body and in Go deeper.

  1. Weston, A., A Rulebook for Arguments, 5th ed. (Hackett, 2017), chapter IV, rules 13 to 17, pp. 25 to 30. Read from the PDF of the book; not linked, because the copy on file is a complete in-print book posted without a visible licence. Rule 13 "Cite your sources"; rule 14 "Seek informed sources" ("Sources must be qualified to make the statements they make"; "For the best information about global climate change, go to climatologists, not politicians"; "experts on one subject are not necessarily informed about every subject on which they offer opinions"); rule 15 "Seek impartial sources" ("People who have the most at stake in a dispute are usually not the best sources of information about the issues involved"; "The truth as one honestly sees it can still be biased. We tend to see what we expect to see"); rule 16 "Cross-check sources" ("Consult and compare a variety of sources to see if other, equally good authorities agree"; when experts disagree, "reserve judgment yourself. Don't jump in with two feet where truly informed people tread with care"; "on some topics the appearance of controversy may be created even when there is virtually no disagreement among qualified authorities"). The course research file gives the 5th edition as 2017 in the entry read from the PDF and as February 2018 elsewhere; the year is on the list of items owed to a session with network access.
  2. Walton, D., Appeal to Expert Opinion (Penn State, 1997). This book has not been read for this course. What is used here is the expert-opinion scheme and its six critical questions (expertise, field, opinion, trustworthiness, consistency, backup evidence) as this course's research file records them, and nothing else. Taught in full in lesson 8.
  3. International Agency for Research on Cancer, Press Release No. 240, 26 October 2015. Read from the PDF at https://www.iarc.who.int/wp-content/uploads/2018/07/pr240_E.pdf. Processed meat "carcinogenic to humans (Group 1)"; red meat "probably carcinogenic to humans (Group 2A)"; "each 50 gram portion of processed meat eaten daily increases the risk of colorectal cancer by 18%"; Kurt Straif: "For an individual, the risk of developing colorectal cancer because of their consumption of processed meat remains small, but this risk increases with the amount of meat consumed."
  4. Dunlop, C., "Processed meat and cancer – what you need to know", Cancer Research UK news, 26 October 2015, https://news.cancerresearchuk.org/2015/10/26/processed-meat-and-cancer-what-you-need-to-know/. Out of every 1,000 people in the UK about 61 develop bowel cancer; about 56 per 1,000 among those who eat the least processed meat; 66 per 1,000 among those who eat the most; around 3 in every 100 UK cancers (about 8,800 a year) attributed to red and processed meat against 64,500 cases a year (19%) attributed to smoking. The en dash in the article title is the title's own.
  5. Gigerenzer, G., Gaissmaier, W., Kurz-Milcke, E., Schwartz, L. M. and Woloshin, S., "Helping Doctors and Patients Make Sense of Health Statistics", Psychological Science in the Public Interest 8(2), 53 to 96 (2007). Read from the full text at https://kops.uni-konstanz.de/server/api/core/bitstreams/faa74b7c-5f6e-4abe-b6f9-d980100c9c1f/content. The 1995 pill scare (p. 54): the Committee on Safety of Medicines warning of a risk raised "twofold, that is, by 100%"; "Dear Doctor" letters to 190,000 practitioners, pharmacists and directors of public health; "of every 7,000 women who took the earlier, second-generation oral contraceptive pills, about 1 had a thrombosis; this number increased to 2 among women who took third-generation pills"; "an estimated 13,000 additional abortions (!) in the following year in England and Wales" (citing Furedi 1999), one extra birth per extra abortion, about 800 additional conceptions among girls under 16; "Had the committee and the media reported the absolute risks, few women would have panicked and stopped taking the pill." The exclamation mark in the abortions phrase is the review's own and is dropped in the lesson's quotation. Mismatched framing: 1 in 3 studies in BMJ, JAMA and The Lancet 2004 to 2006 reporting both benefits and harms used different metrics, mostly relative risks for benefits and absolute frequencies for harms (citing Sedrakyan and Shih, 2007); the lesson paraphrases this rather than quoting, because the research file records it as a paraphrase. Summary: a 25% reduction in breast-cancer mortality from mammography screening "in fact means that 1 less woman out of 1,000 will die of the disease"; the base rate of 4 in 1,000 is not in the review and is the lesson's own arithmetic from that sentence.
  6. Furedi, A., "The public health implications of the 1995 'pill scare'", Human Reproduction Update 5(6), 621 to 626 (1999), https://academic.oup.com/humupd/article/5/6/621/745751. Abstract read: oral contraceptive use among under-16s fell from 40% to 27%; estimated NHS cost of about £46 million for abortion provision. The 13,000 figure is quoted in this lesson from [5], who cite this paper; the full text was paywalled.
  7. Mangel, M. and Samaniego, F. J., "Abraham Wald's Work on Aircraft Survivability", Journal of the American Statistical Association 79(386), 259 to 267 (1984), https://people.ucsc.edu/~msmangel/Wald.pdf. Read from a scan of the paper. The operational problem and "the basic difficulty"; eight SRG memoranda for the National Defense Research Committee, over 100 pages, reprinted by the Center for Naval Analyses in 1980; the assumption that "an unhit aircraft always returns"; the hypothetical data (400 aircraft; returning with 0 to 5 hits: 320, 32, 20, 4, 2, 2; 102 hits in all; area fractions engines .269, fuselage .346, fuel system .154, others .231; observed hit fractions .186, .382, .176, .255); overall single-hit survival probability .851; Table 4 single-hit survival by part: engines .588, fuselage .940, fuel system .973, others .939; "For these data, the most vulnerable portion of the aircraft is the engine area." Summary states the work "was used in World War II and in the wars in Korea and Vietnam." The hit counts of 19, 39, 18 and 26 are this lesson's arithmetic from the recorded fractions and the recorded total of 102, not figures read from the paper.
  8. Mangel, M. and Samaniego, F. J., "Abraham Wald's Work on Aircraft Survivability: Rejoinder", Journal of the American Statistical Association 79(386), 270 to 271 (1984), https://jhanley.biostat.mcgill.ca/bios601/CandH-ch0102/WaldAircraft.pdf. "We do not know whether it was used during World War II, although it was produced early enough in the war to have been available"; use by the Operations Evaluation Group on the A-4 in Vietnam, and at Wright Patterson on the B-52.
  9. Wald, A., A Reprint of "A Method of Estimating Plane Vulnerability Based on Damage of Survivors", Center for Naval Analyses Research Contribution 432 (July 1980), in the same PDF as [8]. Eight memoranda written in 1943 for the Statistical Research Group and the National Defense Research Committee; the vulnerability table "can be used as guides for locating protective armor and can be used to make a prediction of the estimated loss of a future mission."
  10. Leo, S., "Mistakes, we've drawn a few: Learning from our errors in data visualisation", The Economist on Medium, March 2019, https://medium.economist.com/mistakes-weve-drawn-a-few-8cdd8a42d386. The original blocks automated fetching, so the article was read in full at a verbatim repost. Charts sorted as misleading, confusing, or failing to make a point. The truncated-axis bar chart of average Facebook likes on posts by pages of the political left, which "not only downplays the number of Mr Corbyn's likes but also exaggerates those on other posts"; the Espresso line chart of weekly attitudes to the Brexit referendum result, whose scale made respondents look as if they "had a rather erratic view of the referendum result". The research file notes that exact chart titles and original print dates were not confirmed, and none is given here.
  11. FactCheck.org, "Biden's Job Growth Chart Ignores Impact of Pandemic", 9 February 2024, https://www.factcheck.org/2024/02/bidens-job-growth-chart-ignores-impact-of-pandemic/. Biden's 3 February post on X ("The last guy had the worst jobs record since the Great Depression") with a chart titled "Jobs Created by President" showing a monthly average loss of 57,000 jobs under Trump; 20.5 million jobs lost in April 2020; nearly 14.8 million jobs added under Biden, 5.4 million above the pre-pandemic peak; the chart "leaves the misleading impression that presidents are responsible for all the job creation, or loss, during their time in office. But there are many economic factors outside the control of a president (see: COVID-19)"; post-recovery average of 282,000 a month, "still 100,000 more than the pre-pandemic average under Trump."
  12. FactCheck.org, "A Pre-SOTU Guide to Trump's Economic Claims", 19 February 2026, https://www.factcheck.org/2026/02/a-pre-sotu-guide-to-trumps-economic-claims/. A 16 February White House press release saying the stock market has "rebounded strongly under President Trump's leadership"; the S&P 500 up 14.5% between the close on 17 January 2025 and the close on 18 February 2026; "The stock market performed well in Biden's final two years in office", with the S&P 500 "rising over 20% each of those years", "better than the 13% gain Trump saw in his first year."
  13. FactCheck.org, "Biden's Final Numbers", 9 October 2025 (updated 5 March 2026), https://www.factcheck.org/2025/10/bidens-final-numbers/: "Opinions also differ on how much credit or blame a president should get for what happens while he is in office." The same series' "Trump's Numbers, July 2026 Update", 28 July 2026, https://www.factcheck.org/2026/07/trumps-numbers-july-2026-update/: "We also make no judgement on how much credit or blame the president should receive." Authors are not recorded in this course's research file and are not given here.
  14. Evans, J. St. B. T., Barston, J. L. and Pollard, P., "On the conflict between logic and belief in syllogistic reasoning", Memory & Cognition 11(3), 295 to 306 (1983). Read from the free full text at https://core.ac.uk/download/81101245.pdf. Belief bias "was more marked on invalid than on valid syllogisms"; some verbal protocols were "rationalizations for prejudiced decisions". The syllogism materials about cigarettes and vitamin tablets are recorded in the course research file from Evans's 2003 review. Taught in full in lesson 2.
  15. Kahan, D. M., Peters, E., Dawson, E. C. and Slovic, P., "Motivated numeracy and enlightened self-government", Behavioural Public Policy 1(1), 54 to 86 (2017), https://doi.org/10.1017/bpp.2016.2. Abstract read from the publisher: a problem that "turned on their ability to draw valid causal inferences from empirical data"; skin-rash treatment versus gun-control ban framings of the same data; the most numerate "did substantially better" on the rash version; on the gun version responses "became politically polarized – and even less accurate", and polarisation "did not abate among subjects highest in numeracy; instead, it increased"; numerate subjects "use their quantitative-reasoning capacity selectively to conform their interpretation of the data to the result most consistent with their political outlooks." The lesson renders the dash in the polarisation quotation as a comma. Sample size and cell percentages are not quoted here because the full text was not read; the symmetric design is stated as the abstract's logic and the replication literature describe it.
  16. Gigerenzer, G. and Hoffrage, U., "How to improve Bayesian reasoning without instruction: Frequency formats", Psychological Review 102(4), 684 to 704 (1995). Study 1: Bayesian answers 16% and 28% in the two probability formats, 46% and 50% in the two frequency formats. As recorded in the course research file and taught in lesson 6; no free address is on file.
  17. California State University Northridge, PHIL 200 Critical Reasoning (Sun, online, Fall 2011). Syllabus PDF read in full. Objective 2 of seven: "evaluate sources of belief (media, experts, personal experience)", which is where this lesson's third source comes from.

Check your understanding

This lesson has a 6-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.