Reading a claim about AI

95 min

Listen: this lesson as a conversation

Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.

In this lesson you will learn to
  • Sort a claim about these systems into measurable, unmeasurable or false, and say what would settle it if it is the first
  • Identify what is missing from a capability claim that carries no model name and no date, and say why each omission matters
  • Explain why your own impression of whether one of these systems helped you is not evidence, and say what would be

In its first lesson this course told you it would not promise to make you faster, because two randomised trials pointed opposite ways. Everything since has been about what these systems do and how to work them. This last lesson turns the same habits outward, onto the claims other people make about them, including the claims you will meet next week.

It's also the lesson to revise when the world moves. Everything in this course that can go stale is here on purpose, so that a reader in three years can read lessons 2 to 10 and know that the argument still holds.

Three questions for any claim

Is it measurable, unmeasurable, or false?

This is Digital Literacy's move, from its lesson on security tools, adapted to a new subject. That course taught it on VPN marketing, and sorted into measurable, unmeasurable, and false or empty. The adaptation here splits that third pile: a claim that says nothing at all goes under unmeasurable, and false is kept for a claim that contradicts something established. The edges are worth having separately, for the reason the next three paragraphs give.

Measurable means somebody could run something that would come out one way or the other. "Reduces drafting time by a third on standard contracts" is measurable: you could measure it.

Unmeasurable means nothing in the sentence could fail. No result would embarrass it. "Understands your documents" names no task, no comparison and no threshold, so no result contradicts it. That's not the same as false, and the two get treated as one more often than any other pair here.

False means it contradicts something established, and you can say what it contradicts. These are rarer than the other two, because a claim that contradicts something established is a claim somebody can be held to, and marketing tends to avoid those.

What is missing? For a capability claim, the two things that are almost always absent are the model and the date. This course has carried both in every sentence that needed them, and the reason applies to other people's writing as much as to this one: a sentence about what "AI" can do, with no model and no date, is a sentence about nothing in particular, and it can't be checked, revisited, or fairly disputed.

Is it a benchmark result or a demonstration? A benchmark is a set of questions somebody chose, scored the same way for everybody. A demonstration is one thing working once, in front of you. Both are worth something and neither is the other, and writing about this subject slides between them constantly.

Four claims, sorted

"Scored at the top of a professional examination." Measurable, and this kind of claim is usually measured before it is made, which is what makes it the hardest of the four to read. Suppose it checks out exactly as stated. What it still does not tell you is whether the system can do that profession's work, because the examination is a set of questions somebody chose and the work is not. Notice the two-step: the claim can be perfectly true, and the inference almost everybody draws from it is a different claim that nobody has tested.

"Cuts the time your team spends on reporting by up to 60%." Measurable, and unsettled, and the weasel is "up to", which makes the sentence compatible with a 0% reduction. What would settle it: matched teams, a real reporting cycle, the time measured rather than reported. What you have instead is a number with no study attached.

"Understands your business." Unmeasurable. Nothing in it could come out the other way, so there is nothing to check.

"Our system does not hallucinate." This is the interesting one, and it needs its own section.

The one that is not a lie and is not true

Three commercial legal research products were sold to lawyers on the claim that retrieval had dealt with the problem. Their providers claimed that retrieval "eliminat[es]" or "avoid[s]" hallucinations and guarantees "hallucination-free" citations.1

Tested in 2024 in a preregistered evaluation, they hallucinated "between 17% and 33% of the time".1

And yet the vendors weren't simply lying, which is what makes this worth a section rather than a sentence. Retrieval really did reduce the rate against a general chatbot; the study says so in the same breath.1 The direction of their claim was right and they weren't inventing the mechanism.

What was wrong was the scope. "Eliminating" is a claim too broad to be true of any system of that kind, and it was written in the present tense about a product you could buy. That's a different failure from a lie and it's much harder to see, because every individual thing in the sentence is defensible and the sentence as a whole isn't.

Digital Literacy found exactly this in a different market. When somebody evaluated sixteen consumer VPN products, twelve of them made claims that were inaccurate or too broad, and most of those twelve were not lying either. They were making claims too broad to be true of any product of that kind. Two industries, one failure, and the reason to notice the pattern is that it will be the shape of the next claim you meet as well.

Predict first

You see: "Our model now matches expert physicians on diagnostic reasoning." Sort it, and name what is missing before you read on.

Show the answer

Measurable, and probably already measured, since that phrasing usually points at a benchmark.

Missing, in order of how much it matters. Which model, because the claim is about one and will be read as being about all of them. When, because both the model and the benchmark move. Which physicians and on what, since "expert" is doing enormous work and diagnostic reasoning on written cases is not diagnosis. And what "matches" means, which is a threshold somebody chose.

Then the step that is not about the claim at all: even taken at its best, this is a result on a set of written cases with known answers, and a morning in a clinic is not that. Lesson 6 gives you the sorting to say why. A written case with a known answer is constrained by what is in front of the system. A clinic is full of facts that have to match something outside the room.

Notice that you can do all of that without knowing any medicine. That is what makes this a transferable skill rather than an expert one.

Why your own year of using it is not evidence

The hardest part of this lesson is the part about you.

You have now spent this course forming impressions. You have run exercises, noticed things, built a sense of what your system is good at. That sense feels like data and it isn't.

Lesson 1 gave you the measurement and it lands differently now. Sixteen experienced developers, working in repositories they knew well, predicted the tools would make them 24% faster. They were 19% slower. Afterwards, having been slowed, they still estimated a 20% speed-up.2

Twice wrong, in the same direction, by people with every reason to know.

The mechanism is this course's own account rather than anybody's finding, and it is offered because you can test it against your own week. The waiting is short and the output arrives complete. That is the part you experience as the cost. The reading, the checking, the correcting, the second attempt and the bit you ended up writing yourself all happen afterwards and are filed under other headings. The felt duration of using one of these systems is the wait, and the wait is genuinely fast.

So what would be evidence? The same thing it always is. Time something. Do a task the usual way and one like it the other way, and write down the minutes including the checking. A sample of two isn't a study, and it's enormously better than an impression, because it counts the parts an impression drops.

Check yourself

You timed two tasks last month, and the one you did with a system was quicker. A colleague says that proves nothing, because two tasks is not a study. Is she right, and what have you got?

Show the answer

She's right that it isn't a study, and she's wrong if she means you have nothing.

What a study buys is the claim that the result holds generally, across people and tasks you did not run. Two tasks buy you none of that, and a difference that size could easily be which task it was, what mood you were in, or how long the phone stayed quiet.

What your two tasks do buy is the thing your impression was missing. You counted the checking. You counted the second attempt and the bit you wrote yourself. That is the part the felt duration drops, and it is the part the developers in the METR trial were wrong about while believing they were right.

So the honest statement is narrow and it's real: on these two tasks, done this way, the clock said this. Say it that way and you have an observation. Say "it makes me faster" and you have gone back to the thing that was measured and found wanting.

The upgrade, if the answer matters, is to keep doing it. Ten pairs over a month still isn't a study, and it's a great deal harder to fool than a feeling.

Five things people believe about claims

"The benchmark score means it can do the job." A benchmark is a set of questions somebody chose, all of them constrained by what is in the question. Lesson 6's sorting is what tells you which parts of a real job are like that and which are not.

"It worked in the demonstration, so it will work for me." A demonstration is a single run of a chosen case by somebody who chose it. Lesson 3 explains why a single run tells you little even when nobody is choosing.

"The newest model is better at everything than the last one." Better on average on chosen measures, which is what "better" usually means in these announcements. Whether it is better on your task is a thing to find out, and the frontier from lesson 6 is the reason it may not be.

"I have used it a lot, so I know what it is good at." The METR gap.2 Use is exposure rather than measurement, and it's the belief this lesson exists to disturb.

"The research is out of date, so there is no point reading it." Capability figures do go stale fast, and that much is true. Findings about people don't. What happened to 758 consultants, sixteen developers and nearly a thousand high-school mathematics students are dated facts that stay true as history, and they are what this whole course is built on.

Practice

Sort one you actually met

Take 20 minutes.

  1. Find a claim about these systems from the last week. A product page, a news article, a colleague, a post, a slide in a meeting. Write it down word for word.

  2. Sort it. Measurable, unmeasurable, or false.

  3. If measurable: what exactly would settle it? Write the study you would want, in two sentences. Who, doing what, compared with what.

  4. Name what is missing. Model, date, task, comparison, threshold. "Not stated" is the usual answer, and that's the finding.

  5. Then the question the whole course has been building to: what would have to be true for this to apply to your work? Use lesson 6's sorting to answer it.

You did a cold version of this in lesson 1. Dig that out and compare the two. The difference is what eleven lessons bought you.

Write the baseline again, from memory

Take 25 minutes. This is the last exercise in the course and it has a particular shape, so do it in this order.

  1. Do not look at your old list. From memory, write two columns again: three tasks you would hand to one of these systems, three you would not, with a reason beside each.

  2. Now, still without looking, mark each reason with which lesson it came from.

  3. Only now get out the original from lesson 1, and the revised copy you made in lesson 6.

  4. Put all three side by side and answer three questions in writing. Which entries moved, and when? Which reasons changed from a feeling to a mechanism? And is there anything in the original list you would now defend that you could not defend then?

The order matters and it isn't a formality. Retrieving before you look is what makes the comparison teach you anything, which is How to Learn Anything's argument used on this course's own material, at the end of it.

Connections

Back. This lesson is the whole course turned outward. Lesson 1's studies come back as claims to be read rather than results to be learned. Lesson 3's single draw is why a demonstration proves little. Lesson 6's sorting is what a benchmark result has to be run through before it says anything about a job. Lesson 7's measurement of the legal tools is the worked case in the middle of this lesson. And Digital Literacy's sorting of security claims is the move itself, borrowed and pointed at a new market.

Forward. Out of the course. What you have is not a list of what these systems can do; a list like that would have been out of date by the time you finished reading it. What you have instead is a mechanism, a sorting, a checking procedure and a way of reading a claim. Each was chosen because it survives the models changing, and the test of whether that was the right choice is whether this course is still useful to you in three years.

Go deeper

  • Becker, Rush, Barnes and Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity" (2025). This course has read the abstract and not the paper, so the recommendation is the paper rather than any part of it. The abstract alone carries the three numbers this lesson turns on, and they are the most useful three numbers in it.
  • Magesh and colleagues on the legal research tools (Journal of Empirical Legal Studies, 2025; lesson 7's Go deeper has the link). Open access. Read the introduction, which sets the vendors' claims out in their own words before testing any of them. That order is the whole method of this lesson, done by people who had to be fair about it in print.
  • The Prompt Report (2024). Fifty-eight named techniques for text and forty more besides. Worth ten minutes as a picture of what a young field's vocabulary looks like before anybody knows which parts of it matter.

Sources

  1. Varun Magesh, Faiz Surani, Matthew Dahl, Mirac Suzgun, Christopher D. Manning and Daniel E. Ho, "Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools", Journal of Empirical Legal Studies, 2025. Read in part. Supports: the quoted vendor claims of "eliminating" or "avoid[ing]" hallucinations and "hallucination-free" citations; the quoted finding of between 17% and 33%; and the finding that hallucinations were reduced relative to a general-purpose chatbot. Tools were tested in 2024.
  2. Joel Becker, Nate Rush, Elizabeth Barnes and David Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity", arXiv 2507.09089, July 2025. Abstract read; the paper was not opened. Supports: sixteen experienced developers, the 24% forecast beforehand, the measured 19% slowdown, and the 20% speed-up still estimated afterwards.
  3. Sander Schulhoff and colleagues, "The Prompt Report", arXiv 2406.06608, 2024. Abstract read. Supports: fifty-eight text-based techniques and forty more for other kinds of output.
  4. The account of why using one of these systems feels fast is this course's own, and no source states it. It is offered as a mechanism the reader can test against their own week rather than as a finding, and the body says so.

Check your understanding

This lesson has a 6-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.