A couple of weeks ago I handed in a master’s assignment that had passed a check nobody ran. The references lined up, and the Artificial intelligence (AI) I had used to pull the sources had reported that they checked out. Nothing about it felt unfinished, and that feeling turned out to be the whole problem. I will come back to it, because it is the same problem I want you to be able to see in your learners before it costs them something.
If you are arriving cold, this is chapter three of a field guide for people who teach, on bringing AI into your teaching without faking the learning. Chapter one argued that this is a teaching question rather than a technology one, and chapter two built a working model of the machine. You do not need either to follow this one.
Why good work is not proof of learning
When you want to know whether someone learned something, you look at two things. You look at the work they hand you, and you look at the person. The essay is coherent, the analysis is correct, the code runs. The learner is nodding, and the end-of-session sheet says the day was useful. For most of the history of teaching, that was as good as evidence gets.
I want to be fair to that instinct, because it was never lazy. A fluent, correct piece of work used to be strong evidence of understanding, because producing it required the understanding. You usually could not write a clear paragraph about a chemical process without holding the process in your head long enough to order the sentences. The artifact and the learning were welded together, so grading one told you about the other. A learner’s confidence was worth something too, because it used to be earned the slow way, by trying and failing until the thing held.
Then the weld broke. A learner can now produce the artifact without ever holding the understanding, because something else held it for them, and the confidence arrives anyway, on schedule, because watching a good answer come together leaves you as sure you understand it as producing one would. Both readouts still light up, and neither one is wired to the learning any more.
The room I teach in most is at an international training facility, and a single class runs from finishers with decades on the spray gun to people who have never held one. They go back to their production shops the same week. When what they felt they learned in my room does not survive contact with a real job, I hear about it within days, from my boss, and sometimes from his. I have a fast, unforgiving readout on the gap between the nod at the end of the day and what transferred, and the nod is the least reliable instrument in the building.
What the research says about feeling like you learned
This gap is older than the tools, and the cleanest demonstration I know of has no AI in it. In 2019 a team at Harvard took physics students and assigned them at random to two versions of the same class, with identical content and handouts. Two instructors taught it and swapped methods between the two topics, so that differences between the teachers cancelled out. One version was a polished lecture. In the other, the students worked the problems themselves in small groups while the instructors circulated, and the explanation came afterward. The active group learned more on a test at the end of that same class. They also rated their own learning lower, and the paper’s own title says it plainly: measuring actual learning versus feeling of learning.
In that experiment the feeling of learning pointed the other way from the learning itself, which is the opposite of what your instinct predicts. It does not always. A 2024 replication with medical students, by Boedeker and colleagues in Medical Science Educator, found the active group both learned more and felt they had. What you cannot do, and what most of us do all day, is assume the feeling tracks the learning.
In 2025 a different team ran something close to that experiment with AI in the room, on nearly a thousand high school students in Turkey doing maths practice. Some had no AI, some had an ordinary chat tool that would answer whatever they asked, and while they had it their practice scores went up substantially. A third group had a version that gave hints and never answers, and more on that in a later chapter. Then the tool was taken away for an exam. The students who had used the ordinary chat scored about seventeen percent below classmates who never had access, and the hint group came out even with them. Nothing in their practice scores showed it until the tool was gone. The paper is in the Proceedings of the National Academy of Sciences, and it is the same gap with the machine in it.
| Study | Who | What they felt | What they learned |
|---|---|---|---|
| Harvard physics, 2019 (PNAS) | Undergraduates, same content, lecture versus working the problems in groups | The active group rated its own learning lower | The active group scored higher on the same-day test |
| Medical students, 2024 (Boedeker and colleagues, Medical Science Educator) | Replication of the same design | The active group felt it had learned more | The active group learned more |
| Maths practice in Turkey, 2025 (PNAS) | Nearly a thousand high school students, with and without an AI chat tool | Practice scores rose substantially while the tool was on | About seventeen percent below no-AI classmates once the tool was gone; the hints-only group came out even |
I have a name for this. Years ago I trained for a pilot’s licence, and part of instrument training is the day your instructor covers the windows so you are flying blind, on the gauges alone. With nothing outside to look at, an acceleration feels exactly like a climb, and your inner ear keeps insisting on the climb no matter what the gauges say. I was certain we were climbing. Bill, the navy pilot next to me, knew better, and the instruments agreed with Bill. Pilots who trust the feeling push the nose down to correct a climb that is not happening, which is how level flight turns into a dive. The fix is to read the instrument and believe it over your own body.
The learner’s feeling of having understood is that inner ear. Sincere and strong, and reading something other than what it claims to report. Your job is to stop flying on it and find the gauge.
Why AI makes the gap so easy to fall into
To see why the tool makes this worse, you need one idea from chapter two and one idea from how memory works. From chapter two: the machine is a confidently-wrong pattern machine. It predicts the shape of a good answer and hands it over fully formed, fluent from the first word, with the same steady voice whether it is right or inventing. It produces the finished thing, with the learner watching.
From memory: the part that does most of the building is the effortful part, pulling a half-formed idea out of your own head and forcing it into words or steps. It is the struggle you feel when you try to explain something you only sort of know, and it is the work. Recognising that an explanation makes sense is far easier than producing one, and, this is the trap, it feels much better. So when the machine does the producing and leaves the learner the recognising, the learner gets all of the good feeling and almost none of the construction, and the feeling of learning arrives on time while the learning does not.
The learner’s feeling of having understood is that inner ear. Sincere and strong, and reading something other than what it claims to report.
Which brings me back to my assignment. The course I am in warned us that generative AI “invents scholars, institutional affiliations, and citations with considerable confidence.” So I set myself a rule: nothing goes in unless it was pulled from an actual fetch of the source, and I would re-verify anything load-bearing myself. I used Claude to fetch the primary sources and draft a first pass on a profile of Rian Rietveld, who led the WordPress accessibility team until she stepped down in 2018 and has spent the last couple of years building a public accessibility knowledge base for people who use and build WordPress. The model came back with references and a report that the sources checked out.
They mostly did. One did not. A 2026 post about her documentation project had been credited to the organisation that hosts it, when Rietveld had written it herself, in the third person, about her own work. A small byline miscredit. You would never notice it unless you looked at the page with your own eyes, which is how I found it, and I fixed the citation before the assignment went in. Here is the honest part. Nothing felt wrong. I had a finished draft, matching references, and the model’s word that the check was done, and I felt done. What saved me was a rule I had set in advance, that my feeling did not count as the check. I keep a public page on how I use and disclose these tools, and this is why. The appearance of a completed check is not a completed check, and the feeling of one is worth even less.
In chapter one I told you about a cup of jelly that was supposed to be liquid. Same trap. The difference this time is that the gap only showed up because I went looking before I felt any need to, and your learners will not do that on their own, because nothing in the experience tells them to. That is the part you have to design.
The one move: put a cold moment in the lesson
Here is the move, reason first. The feeling of learning is produced by the experience of watching an answer arrive. The fact of learning is only visible when the learner has to produce without help. An AI-assisted task only contains the first one unless you add the second, so add it. Somewhere in every task where the tool is allowed, put a moment where the tool is closed and the learner has to produce something the task depended on. Then treat what happens in that moment as the readout, and treat the nod and the polished artifact as readings to compare against it, never as the verdict.
This is not a ban, and it is not a return to closed-book everything. The tool stays in the room, and the cold moment can be two minutes long. What it changes is where you look for the evidence, and in my experience, once learners know the cold moment is coming, the way they use the tool during the warm part starts to change on its own, because they know they will have to hold something afterward. For now the rung is just this: install the gauge, and read it.
Try this before the next chapter
This one is aimed at your learners, on work you have already marked. Pick one piece of recent learner work where AI was allowed and where you gave a good mark, because the good marks are exactly where the gap hides. Sit down with the learner, or set it up in your next session, and put the work away where neither of you can see it.
Ask for two things. Have them talk you through one specific decision in the work, why that step, in their own words with nothing open. Then ask them to change one thing: how would this look if the audience changed. Do not grade the answers. Some learners will walk through both without a pause, and you have just confirmed your mark. Some will explain the decision fluently and then stall completely on the change, and that stall is the strongest signal you will get that the artifact ran ahead of the understanding, so check it a second way before you act on it. Write down which learners landed where, because that list is a map of where the feeling of learning and the fact of learning came apart in your own classroom.
The appearance of a completed check is not a completed check, and the feeling of one is worth even less.
Common questions about the feeling of learning
If a learner’s AI-assisted work is correct, does it matter whether they understood it? It matters the moment they have to do the next thing without the tool. Correct work tells you the task got done. It no longer tells you who did the understanding.
Is asking learners to rate their own understanding useless, then? It is one reading, and it earns its place once there is a second reading beside it. Collect the rating, then put a cold moment next to it and read the two together.
Does this mean I should keep AI out of assessed work? Keeping it out avoids the question, and chapter one is about what that dodge costs. The tool can stay as long as the task also holds a moment where the learner produces without it. If you want help bringing this into your own teaching, that is the learning work I do.
Next chapter: learn it yourself first
So that is the problem the rest of the book is built to solve, and the instrument you can trust is the cold moment.
The catch applies to you before it applies to them. You cannot design a cold moment for a tool you have only read about, and you cannot teach someone to learn with AI honestly if you have not done it yourself, jelly and miscredited bylines included. So in two weeks, in chapter four, the educator does the reps first: learn something in your own subject with the tool, and put your own feeling of understanding under the same test. Bring the list of where your learners stalled. In chapter four you make the same kind of list about yourself, and we put the two side by side.
If you want a shorter weekly version of this thinking between chapters, I write one at the newsletter.

Leave a reply