When not to use one
100 min
Two hosts talk the lesson through. The voices are synthetic; the script was written from this lesson and checked against it, and asserts nothing the lesson does not.
- State what the evidence shows about using one of these systems while learning something, giving all three of the study's figures and the scope they carry
- Identify a task where the doing was the point, and say specifically what handing it over would cost you
- Decide, for one task of your own, to stop or to redesign how you use it, with the reason written down and a replacement named
A course that teaches you to use something has an obvious temptation, which is never to say when to stop. This lesson exists because the best-evidenced study in this course's research is about exactly that, and it would be a dishonest course that put it in a callout.
What happened to a thousand students
In 2025 a team published a field experiment in PNAS with nearly a thousand high school mathematics students.1 During practice sessions the students were in one of three conditions: two of them with access to GPT-4, and one with no access at all.
GPT Base was a standard chat interface.
GPT Tutor was the same model with prompts designed to safeguard learning, giving teacher-written hints rather than answers.
And a group with no access, which is the comparison that makes the rest of it mean anything. Without it there is nothing to be 17% below.
Then the results, and you need all of them.
During practice, both helped, and the guarded one helped far more. Grades improved "48% ... for GPT Base and 127% for GPT Tutor".1
When access was taken away, the unguarded group was worse off than students who had never had it at all, by "17% ... in grades for GPT Base".1 The authors' own summary of why: "Without guardrails, students attempt to use GPT-4 as a 'crutch' during practice problem sessions, and subsequently perform worse on their own."1
And the guarded group did not suffer that. The negative effects were "largely mitigated by the safeguards in GPT Tutor".1
| Condition | During practice | After access was removed |
|---|---|---|
| No access | the baseline | the baseline |
| GPT Base | grades up 48% | 17% below the no-access group |
| GPT Tutor | grades up 127% | the harm "largely mitigated" |
The 48 and the 17 travel together, always, because either one on its own is being used to make an argument the study doesn't support. And the 127 travels with them, because it is what turns this from a verdict on the technology into a question about how it is used.
The scope travels too. High school mathematics, practice problems, GPT-4, published 2025. It is one study in one subject with one age group, and it is the best-evidenced thing this course found on the question.
Before you read on: why would using a tool during practice leave somebody worse off afterwards than never having used it? Not the same as, worse than.
Show the answer
The mechanism is one you already have, from Term 1 of the Core.
How to Learn Anything taught that reconstructing something from nothing is what builds the kind of memory that lasts, and that a study method which feels fluent is often the one teaching least. Rereading a clear answer raises fluency and does little for what you will still have next month. Asking for the answer removes the reconstructing, and it removes it at exactly the moment when the struggle was doing the work.
So the worse-off part isn't mysterious. The unguarded group spent their practice sessions reading answers rather than producing them, and got grades in those sessions that reflected the tool rather than themselves. They had the same hours and none of the retrieval, and the feeling of following a clear worked answer is very close to the feeling of understanding one.
That is the fluency illusion with a much better interface.
The mechanism, in one sentence, because it should not live only behind that button. Practice works by making you produce the answer, an answer you read instead is an answer you didn't produce, and the grade you get in that session is a measurement of the tool rather than of you.
The generalisation, and it is this course's rather than the study's
The study is about maths homework, and most readers here aren't doing maths homework. So what follows is reasoning from it rather than a finding, and the label matters because the reasoning is doing a lot of work.3
The cost appears when the doing was the point.
Sometimes the output is the point: a letter you need, a summary somebody will read, a first draft you will rewrite anyway. Hand those over, check them, and get on with your day.
Sometimes the doing is the point. You are learning something. You are building a judgement you will need later, in a room, without the system. You are writing something whose value is that you thought it through rather than that it exists.
And sometimes it's neither, and the honest answer is that nobody can tell you. Whether to trade the skill for the hour is a question about what you want your life to look like, and this course describes the trade and declines to make it. That is a value question in the institute's standards, and it is treated as one.
Two people doing the same thing differently
"When the doing is the point" is easy to nod at and hard to apply, so work it on one person. The case is constructed, and what matters in it is the difference between the two approaches rather than any detail of the trade.
A trainee claims handler, first year, learning to spot the problems in a claim file. She has a system open. Two ways to use it.
One. She reads the file, writes down what she thinks is wrong, and then asks: "Here is the file and here is what I flagged. What have I missed, and why would somebody flag it?"
Her list said three things: the incident date is a Sunday and the policy is a commercial one, the repair quote has no VAT line, and the claimant's address on the form does not match the address on the policy. What comes back adds two she did not have. The claim was notified eleven weeks after the incident, and the policy wording she was sent has a notification window in it. And the quote is from a firm that shares a surname with the claimant.
She reads why each of those matters. The next file she does, she catches the late notification herself and she misses the surname again.
Two. She pastes the file in and asks what is wrong with it. Back come five things, the same five. She reads them, agrees, and sends the file on. The next file, she pastes in.
Same tool, same subject, same fifteen minutes. The first builds the thing she came to the firm to acquire, which is a trained eye: three flags, then five, then the two she missed becoming one. The second produces the same five flags every week and builds nothing, and it works right up to the morning she is in front of somebody with a question the list did not cover.
Notice what the first one is. It is the guarded condition, run by her rather than by a designer. She attempted first and asked for what she missed, which is hints after effort rather than answers instead of it.
The paralegal's second approach produces better flagged titles this week than her first approach does. Is that a reason to prefer it?
Show the answer
For this week's files, yes, and that's exactly what makes the trade hard rather than obvious.
The study found the same shape: the unguarded group's practice grades went up 48%. The work in front of them got better. What went down was what they could do without it.
So the useful question is not which produces better output now, but whether you need the capacity later, and how much later, and what it costs you not to have it. For a first-year paralegal who intends to be a conveyancer for thirty years, the answer is not close. For somebody covering a colleague's caseload for three weeks, it might be.
Writing down which of those you are is the whole exercise at the end of this lesson.
The case where it does not matter, which the course has to say
A man in his last month at a charity he is leaving has to write to forty regular donors explaining that the minibus appeal has been withdrawn and their standing orders will be cancelled. He has never written that letter and will never write it again.
He hands the facts over, gets a draft, reads it, changes two sentences that sounded like a press release, and sends it.
He has lost nothing. Ask the question this lesson asks: if he could not use a system for this for six months, what would he be worse at? Writing withdrawal letters to donors. He is leaving in three weeks and will not write one again, so the answer is nothing he needs. No judgement was going to be built, because none is needed again. The doing wasn't the point; the letter was.
Now change one detail and the answer changes with it. If he were staying, and the charity ran three appeals a year, the same letter would be the third or fourth of its kind he had written, and learning what tone keeps a donor is exactly the judgement a fundraiser is paid for. Same task, same tool, opposite verdict, and what moved it was one fact about his next twelve months.
A course that can't say that plainly has lost you for the cases where the cost is real, because you'll correctly notice that its rule is being applied without its condition, and you'll stop trusting the condition.
The same goes in the other direction, and the standards body says it in the same section as the warning everybody quotes. Alongside over-reliance, which it calls "excessive deference to automated systems", NIST names the opposite failure: "human experts may be unnecessarily 'averse' to GAI systems, and thus deprive themselves or others of GAI's beneficial uses".2
This course is against both. Declining to use one of these systems isn't a mistake, and it isn't a virtue either. It's a decision with a cost on each side, and the point of this lesson is to make the cost on the less obvious side visible enough to weigh.
Five beliefs worth testing
"Using it to do the work and using it to learn are the same thing." The study is the answer, and the paralegal above is the same answer without a study. The output improves in both cases. What differs is what you can do next week.
"If I read the answer carefully, I have learned it." Reading feels like learning, and the feeling tracks how easy something was to follow rather than how much of it you will retain. Term 1 taught this as the fluency illusion; here it arrives with a better interface.
"The study shows these tools are bad for education." The same paper's guarded tutor produced the biggest gain of the three conditions.1 The finding is about design, not about whether. What a school should do about that is a further question, about policy and about what teachers can supervise, and this lesson doesn't answer it either way.
"Refusing to use one is just being difficult." NIST names unnecessary aversion as a real cost in the same breath as over-reliance, and the word doing the work there is "unnecessarily".2 Somebody who does not want to use one is not thereby being averse in NIST's sense, and the reason can be this lesson's trade, or it can be a reason this course has no standing to weigh: what the work is for, what they want to be good at, what they are content to be part of. This lesson supplies one consideration. It doesn't supply the licence.
"I can tell whether I have learned something." The least comfortable one, and the reason the exercises below ask you to test rather than to reflect. The study measured grades rather than what the students believed, so this next part is inference: their practice grades went up while what they could do alone went down, which means the most visible signal they had was pointing the wrong way.
Practice
Take 20 minutes.
List five or six candidate tasks. Ones you currently hand over, if you do. Ones you have been asked or expected to hand over. And ones you have considered and decided against, which belong on the list for the same reason the others do.
Beside each, answer one question: if I could not use a system for this for six months, what would I be worse at? Not what would take longer. What would I be worse at.
Pick the one where the answer is most uncomfortable.
Act on it. If it is a task you hand over now, stop handing that one over: write the reason in a sentence, and write what you will do instead, including how much longer it will take. If it is a task you have already declined, the step is the mirror image: write down what declining is costing you, in hours or in output, and write down what you would need to see to change your mind. Both are the same exercise, which is putting a number against a decision you have been making by feel.
Put a date three weeks from now in your calendar to read what you wrote and say whether you kept to it.
If nothing on your list produces an uncomfortable answer, write that down too, with the list, and say why. That is a legitimate result and it is worth having on paper rather than assumed. So is a list on which every entry is a task you have declined.
Take 25 minutes. This is the guarded-tutor condition applied to yourself, and it matters more than stopping, because you won't stop most things.
Take one task you use a system for and intend to keep using it for. If there is no such task, take one you are expected to use a system for, or one you would use a system for if you were going to use one at all; the redesign works the same way on a hypothetical, and it is the cheapest way to find out whether your objection is to the tool or to a particular shape of using it.
Write down which part of it builds the judgement. Usually one specific step: the noticing, the choosing, the first attempt, the diagnosis.
Redesign the request so that you do that part and the system does the rest. The usual shape is: attempt first, then ask what you missed and why, rather than asking for the answer.
Predict what it will cost you in time, then do it that way three times and record the real number.
Then the judgement, in writing: is the extra time worth what it buys, for this task, for you? A no is a legitimate answer if you've actually measured it, and so is deciding that the redesigned shape is the only one you want to use.
Keep this. It is section six of the course project, which asks for one task stopped and one redesigned.
Connections
Back. Lesson 1's evidence and this lesson's are the two halves of one decision: what handing a task over gains you, and what it costs you. Lesson 6 told you which tasks these systems are reliable on, which is a different question from which tasks you should hand over. Lesson 8 priced the checking, and that price belongs in this calculation too. And Term 1's How to Learn Anything supplies the mechanism, in its lessons on the fluency illusion and on retrieval practice, which this lesson uses and does not re-teach.
Forward. Lesson 10 is about what you hand over of a different kind, which is your data rather than your practice. Lesson 11 closes the course by turning its habits on claims about these systems, and the habit it teaches is the one this lesson just used on a study: ask what was measured, on whom, and what travels with the figure.
Go deeper
- Bastani and colleagues, "Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics" (PNAS, 2025). Every figure in this lesson comes from the abstract, which is on the first page. This course read the abstract and did not open the body, so the recommendation is the paper rather than any part of it, and the first page is where its figures come from.
- NIST AI 600-1, section 2.7 on human-AI configuration. Short, and the only official document this course has found that names both failures in the same section.
- How to Learn Anything, lessons 1 and 3, on the Foval Core. If the mechanism here felt thin, that is where it is taught properly, and this lesson deliberately does not repeat it.
Sources
- Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı and Rei Mariman, "Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics", PNAS, 2025. Abstract read in full and verbatim; the body was not read. Supports: nearly a thousand high school mathematics students; the two conditions, GPT Base and GPT Tutor; the quoted figures of 48% and 127% improvement during practice and 17% reduction for GPT Base once access was removed; the quoted "crutch" explanation; and the statement that the negative effects were largely mitigated by the safeguards. The scope is high school mathematics practice problems with GPT-4, and every use of these figures carries it.
- NIST AI 600-1, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, July 2024, section 2.7. Read in part. Supports: the quoted description of automation bias as excessive deference to automated systems, and the quoted sentence about experts being unnecessarily averse and depriving themselves of beneficial uses.
- The generalisation from maths homework to work in general is this course's reasoning and is not in source 1. The body says so where it is introduced. The study supports the mechanism and one subject; the claim that the cost appears wherever the doing was the point is an extension, and it is the kind of extension a reader should be able to test against their own week.
Check your understanding
This lesson has a 6-question quiz. Pass it and the questions come back on a schedule in Review, so what you learned stays learned. Your progress is saved in your browser; no account needed.