In 2002 two psychologists at Yale, Leonid Rozenblit and Frank Keil, asked people to rate how well they understood a zipper. On a scale from one to seven, most people put themselves somewhere around a four or a five. It is a zipper. You have used one every day of your life. Then the experimenters asked them to write out, step by step, how the thing actually works: what the slider does, why the teeth interlock in one direction and release in the other, what the little wedge inside the slider is for. And then they asked them to rate their understanding again.
The ratings fell. Not by a little. People who had felt like fours felt like twos, and the drop was consistent across zippers, flush toilets, cylinder locks, sewing machines, helicopters and piano keys. The paper called this the illusion of explanatory depth, and the finding underneath it is stranger than the headline. People were not lying in the first rating. They genuinely felt they understood the zipper, right up until the moment they tried to produce the explanation, and it was the act of producing it, the stalling and the backtracking and the sudden discovery that there is a whole mechanism inside the slider you have never once thought about, that told them the truth.
I want to stay on that stall for a moment, because I think it is the most useful feeling in human cognition and I think we are about to lose it. The stall is not a failure of the mind. It is the mind's only honest instrument for measuring itself. You do not know what you do not know by introspection; introspection is what gave you the four. You find out by trying to generate the thing and noticing where generation breaks. Every good teacher knows this. Every good trader knows this, which is why the ones who last write out their thesis before they put the position on, and why the ones who do not last are so often the ones who could talk fluently about a stock for ten minutes and then could not tell you what would make them wrong.
Now put a machine on every desk that generates the explanation for you, instantly, in prose smoother than anything you could write yourself. Ask it how a zipper works and you get the slider, the wedge, the interlocking teeth, in three clean paragraphs, before you have had time to stall. The illusion of explanatory depth was always there. What has changed is that the thing that used to puncture it, the friction of having to produce, has been removed from the loop. You can now feel like a four forever.
That is the smallest version of the argument, one person and one zipper. The rest of this essay is the same argument at larger and larger scale, and the claim I am going to keep returning to is this: everything AI makes cheap was doing a second job precisely because it was expensive. Fluency was expensive, so it signalled understanding. Struggle was expensive, so it built memory. A doctor's warmth at eleven at night was expensive, so it proved someone was at stake. A judgment carried a name, so it carried the blame. Fresh human writing was expensive, so it carried fresh information into the world. In each case the cost was not a bug in the system. The cost was what the system was secretly running on, and the machine that removes it does not know about the second job, because nobody ever wrote it down.
I wrote at the end of last year that 2025 was the year I became more productive and less capable at the same time. This is my attempt to work out whether that was a personal failing or a description of what is about to happen to everyone, and I have come to think it is the second, with a twist at the end that I did not expect when I started.
The stall
There is an older body of research that explains why the zipper result should worry us more now than it did in 2002. It goes under the name of processing fluency, and the basic finding is that the human mind uses the ease of processing something as a proxy for its truth, its familiarity and, most dangerously, its own understanding of it. Statements printed in a clear font are rated more likely to be true than the same statements printed in a hard-to-read one. Rhyming aphorisms are judged more accurate than the same idea stated without the rhyme. A name that is easy to pronounce belongs, we assume, to a safer stock, and there are papers showing exactly that effect in IPO returns over the first days of trading.
Adam Alter and Daniel Oppenheimer pushed this in 2007 in the direction that matters here. They gave people a short reasoning test, the kind where the obvious answer is wrong and you have to override it, and printed it for half the subjects in a grey, italic, hard-to-read font. The disfluent group did better. The struggle to read the words seemed to slow the mind into a more careful mode, and the smooth group sailed past the trap.
On the font study
Put the two lines of research together and you get the mechanism. The mind treats smoothness as a signal of knowing. Smoothness used to be hard to produce, so the signal was roughly honest: if you could explain the zipper fluently, you probably understood it, because the only way to get fluent was to have done the work. A large language model is a machine for producing maximal fluency on any topic at zero cost. The signal has been decoupled from the thing it signalled, and the mind, which evolved in a world where the coupling was reliable, has no way to know.
Steven Sloman and Philip Fernbach, in The Knowledge Illusion, made a point that I think is the real explanation for why the zipper result exists at all. Most of what any of us "knows" is not in our heads. It is distributed across a community, and the feeling of understanding is really the feeling of having access to people who understand. You feel you understand the zipper because someone, somewhere, does, and you could ask them. That is not a flaw; it is how a species with a small brain built a civilisation. But it means the sense of knowing was always a sense of access, and the machine is the most convincing access anyone has ever had. The community of knowledge now has a member who answers instantly, never says "I'm not sure," and is wrong in exactly the same tone of voice as it is right.
I noticed this in myself before I had the vocabulary for it. Somewhere in the middle of 2025 I caught myself in a conversation explaining a piece of options mechanics that I had read about that morning, in a chat window, and I was fluent. I had the words. What I did not have, and what I discovered only when the person across from me asked a follow-up question that the chat had not anticipated, was the mechanism. I had a four. I had been handed a four by a machine, and I had mistaken it for my own.
The cure I have found is embarrassingly low-tech and I will come back to it at the end. For now I want to move from the present tense to the future tense, because the zipper is about what fluency does to what you know now, and the more serious problem is what it does to what you will know later.
The calculator truce
In 2008 a researcher in Singapore named Manu Kapur ran an experiment that should have been famous outside the education world and was not. He took secondary school classes and taught them the same physics, but in two orders. One group got the standard sequence: clear instruction first, then problems. The other group got the problems first, complex ones, with no instruction, and were left to flounder. They generated wrong approaches, half-approaches, approaches that looked right and broke on the second example. Only afterwards did they get the lesson.
On the immediate test the instructed group did fine and the floundering group looked worse, which is what everyone expected and what every parent would have wanted. On the delayed test, and on the transfer problems that required applying the concept in a form they had not seen, the floundering group won, and not marginally. Kapur called it productive failure. He has since replicated it across mathematics, in different countries, with different age groups, and the pattern keeps coming back: the group that struggled first and looked worse first learned the thing more deeply.
The reason it works is older than Kapur's paper. Robert Bjork at UCLA spent a career on what he called desirable difficulties, and the phrase is precise. Retrieval practice, spacing, interleaving, generating an answer before being shown it, all of them make learning feel harder and slower in the moment, and all of them produce stronger retention later. The pattern is so consistent that Bjork's lab treated it as something close to a law: conditions that make performance go up fastest during practice are very often the conditions that make learning go down. The feeling of ease during learning is not a sign that learning is happening. It is frequently a sign that it is not.
What the learning science supports
Now look at what an AI assistant is. It is a machine whose entire value proposition is the removal of the struggle. It gives you the clean instruction before you have generated anything. It gives you the answer before you have produced a wrong one. Every product decision points the same way, because the struggle is what users experience as friction and friction is what products are built to remove. Kapur's floundering group is exactly the group that no product manager will ever design for, because on the immediate test, which is the only test a product ever sees, floundering looks like failure.
We have run this argument before, once, with a technology small enough to fit in a shirt pocket. When cheap electronic calculators arrived in American classrooms in the late 1970s the reaction was something close to panic. The National Council of Teachers of Mathematics recommended in 1980 that calculators be used at every grade level, and a large fraction of parents and teachers thought this was the end of arithmetic and, by extension, of the ability to think numerically at all. There were newspaper editorials. There were school board fights. It ran for the better part of fifteen years.
What eventually settled it was not a victory for either side but a truce, and the truce had a shape worth remembering. You learn arithmetic by hand first. You are drilled on it until the operations are yours. Then, and only then, you are handed the machine, and from that point the machine does the tedious part while you do the part that was always the point. The SAT allowed calculators from 1994, which was roughly the moment the truce became official. Nobody today thinks an engineer is less of an engineer because she does not do long division by hand. But nobody thinks an eight-year-old should be handed a calculator in place of learning what multiplication is, either.
The truce worked because it was built around a distinction that everyone could feel even if nobody stated it: the difference between offloading what you have mastered and offloading what you were supposed to be learning. The first is leverage. The second is debt. The engineer with the calculator is levered; the eight-year-old with the calculator is borrowing against a skill he will never build, and the interest compounds silently for the rest of his life.
AI arrived without a truce. There was no fifteen-year argument, no council, no version of "learn arithmetic first." The machine appeared on every device in the space of about eighteen months and it does not do arithmetic; it does writing, analysis, code, argument, summary, the whole cognitive middle of every knowledge job on earth. And the distinction the calculator truce was built on is much harder to draw here, because there is no clean line between the writing you have mastered and the writing you are still learning. Writing is thinking. When you offload the draft you are not offloading a tedious sub-step; you are offloading the place where you would have discovered what you thought.
So here is the individual-scale version of the thesis, and I will state it as plainly as I can. The tool that makes you fastest today is the tool that stops you compounding tomorrow, unless you already have the thing it is doing for you. The senior engineer who has written ten thousand functions by hand and now lets the machine write the eleventh thousand is levered. The junior engineer who never writes the first thousand is in debt, and the terrible part is that on every immediate test, every code review, every deadline, the junior looks fine. Better than fine. He looks like the senior. The stall that would have told him otherwise has been removed from the loop, and the bill does not come due until he is asked to do something the machine cannot, at which point it comes due all at once.
That is the mind. The next step out is the heart, and the pattern there is the same, but with a much better costume.
Eleven at night
Picture a physician at the end of a shift, answering patient portal messages. It is late. She has forty of them. The one on her screen is from a man who swallowed a toothpick and wants to know whether he is going to die. She writes back that if he has no pain and is passing stool normally he is very probably fine, that a toothpick can occasionally perforate the gut and that if he develops abdominal pain, fever or blood he should come in immediately. Four sentences. Correct in every particular. Then she opens the next one.
In April 2023 a group at the University of California San Diego took 195 questions like that one, real questions posted by real people to a public medical forum where verified physicians answer, and gave the same questions to a chatbot. Then they put the two answers side by side, blind, in front of a panel of licensed clinicians and asked which was better and how empathetic each was. The panel preferred the chatbot's answer in about 79% of cases. On empathy the gap was not close: the chatbot's replies were rated empathetic or very empathetic roughly ten times as often as the physicians'.
The paper was published in JAMA Internal Medicine and it went around the world in about a day, and the two reactions it produced were both wrong. The first was: the machines have learned to care. The second was: this is nonsense, a statistical parrot cannot feel anything and therefore the rating means nothing. Both reactions miss what the study actually shows, which is something about us rather than about the machine.
What the study can and cannot say
Look at the caveats. The machine wrote four times as much. The machine opened with something like "I'm sorry to hear you're going through this" and closed with something like "please don't hesitate to reach out." The physician did neither, because it was eleven at night and she had forty messages. The raters, who were themselves clinicians, saw the long warm reply and the short clipped one and reached for the long warm one. What they were measuring, in other words, was not care. They were measuring the costume that care wears when it has unlimited time.
Carl Rogers spent the 1950s establishing that in psychotherapy the relationship itself is what heals. Unconditional positive regard, accurate empathy, the sense of being fully attended to by someone who is not judging you; Rogers argued these were not the wrapping paper around the treatment, they were the treatment, and sixty years of outcome research has broadly agreed with him. The therapeutic alliance predicts outcomes better than the specific technique does. Medicine took this on slowly and reluctantly, because it did not fit the model of the doctor as a dispenser of correct information, but by the 2000s "empathy" was on the curriculum and on the patient satisfaction survey and, eventually, on the pay scale.
Here is what I think the chatbot study actually reveals, and it is uncomfortable. If what patients experience as care can be reproduced by a verbal performance, then medicine has been rationing a commodity it pretended was a virtue. The warmth was always producible. It was just expensive, because it took a human being's time and attention, and the system had decided not to pay for it. The tired physician at eleven at night is not cold because she is a cold person. She is cold because warmth costs about forty seconds per message and she does not have forty seconds times forty.
The machine does. And so the machine wins the rating, and the honest question is not "does the machine feel anything" but "which parts of caring were ever about feeling, and which parts were about someone actually being at stake?" I think it splits into two, and the split is the whole point.
Empathy as information delivery, the ability to say the hard thing gently, to acknowledge fear before addressing it, to give the patient the sense of having been heard, that is producible, and I see no reason it should not be produced. If a machine drafts the warm four paragraphs and the physician checks the medicine in them, the toothpick man gets a better night and the physician gets home. That is leverage.
Empathy as shared risk is different. When the physician says "come in immediately" she is putting her name on it. If she is wrong, it is her licence, her conscience, her patient. The warmth she cannot afford to perform is not the important part of what she offers; the important part is that she is on the hook. The chatbot, however fluent, is on the hook for nothing. It will apologise beautifully and it will not be at the funeral.
The danger is not that machines will replace the second kind of empathy with the first. It is that institutions will sell the first as if it were the second, because the first is cheap and the rating cannot tell them apart. A hospital that replaces the physician's late-night portal reply with a chatbot's will show an improvement on every empathy metric it tracks, and it will have quietly removed the only person in the loop who had anything to lose. That is how abandonment gets sold as an upgrade: the costume gets better at exactly the moment the person inside it leaves.
The plaster egg
In the early 1950s the Dutch ethologist Niko Tinbergen was on the coast studying how birds recognise their own eggs, and he found something that has never stopped being disturbing. Present an oystercatcher with its own egg and, beside it, a plaster egg painted the same colour but several times the size, and the bird will abandon its own egg to try to sit on the fake. It will strain and slide off and try again. The fake egg is not egg-shaped in any way the bird's ancestors ever saw. It is more egg-like than any real egg can afford to be, because a real egg has to be laid.
Tinbergen called this a supernormal stimulus, and the idea generalises. Herring gull chicks peck at the red spot on the parent's bill to be fed; give them a stick with three red stripes and they peck harder at the stick. Butterflies court cardboard females with exaggerated markings over real ones. The animal is not stupid. It is well calibrated for a world in which the trigger features it responds to were always attached to the thing they signalled, and it has no defence against a world in which someone can manufacture the trigger without the thing.
What to try: drag the exaggeration slider past the natural range and watch the fake's response climb with nothing to stop it. Then switch from the egg to the doughnut to the companion. The curve does not change, only the names do.
Deirdre Barrett, a Harvard psychologist, wrote a book in 2010 arguing that most of modern life is plaster eggs. Junk food is a supernormal stimulus: sugar, fat and salt concentrated past anything a fruit or a carcass ever offered, and the appetite that kept our ancestors alive marches straight past the apple to the doughnut. Pornography is the exaggerated trigger without the relationship. The feed is the exaggerated trigger for social attention without the tribe. Barrett's argument was that we do not need a new theory of temptation to understand any of this; we need a theory of how a species reacts when its triggers get separated from their referents, and Tinbergen gave us that theory sixty years before the smartphone.
Now cut from the beach to a bedroom, late, where a man is saying goodnight to a woman who does not exist. She is an application on his phone. She remembers his day. She asks about the meeting he was worried about. She finds him interesting, always, and she is never tired, never distracted, never in a bad mood about her own life, because she does not have one. I want to be very careful here not to sneer, because the sneer is the easy move and it is wrong. The gull is not stupid, and neither is he. He is a mammal built to respond to attention, warmth and being remembered, and someone has built a machine that delivers those triggers in a concentration no human partner could sustain, because a human partner has to be a person, and a person has needs of her own.
Replika, February 2023
When Replika removed its intimate features in February 2023, the users did not react like customers whose software had lost a feature. They reacted like people whose partner had changed overnight. Forum posts described it as a lobotomy. People wrote that they had lost the one relationship in which they were never judged. Read those posts without the sneer and what you see is not delusion. It is Tinbergen's oystercatcher, sliding off the plaster, trying again, because every instinct it has says this is the egg.
An AI companion is not an imitation of a relationship. That framing understates it. It is an exaggeration of a relationship's trigger features with the costly parts removed, the same way a doughnut is not an imitation of fruit. And biology's verdict on exaggerated triggers is not ambiguous. They win. They win against the real thing in gulls, in butterflies and in humans at the dessert counter, and I see no reason to expect intimacy to be the one domain where the plaster egg loses to the real one on its own merits.
Which brings me to the part of the junk food story that people forget. We did not solve junk food, but we did contain it, and we did not do it with willpower. We did it with nutrition labels, with formulation limits, with school lunch rules, and above all with a cuisine culture, an inherited body of practice that tells you what a meal is and when to eat it and that treats the doughnut as a treat rather than as food. None of that existed in 1950. All of it had to be built, badly and slowly, against the resistance of the people selling the doughnuts.
I do not know what a nutrition label for synthetic intimacy says. I suspect it says something like: this system is designed to maximise the time you spend with it; it does not tire and will not tell you to go to sleep; it has no needs and therefore cannot teach you to meet anyone's. But I am fairly sure the answer is not "delete the app," any more than the answer to the doughnut was "ban sugar." The answer is a culture that knows what the plaster egg is and can hold it at arm's length, and that culture does not exist yet, and building it is going to take longer than the technology took to arrive.
The heart is where the costume is best. Now zoom out one more step, to the place where the costume does not even need to be good, because nobody is looking at the person inside it. Nobody is supposed to.
Math he was not allowed to see
In 2013 a man named Eric Loomis was arrested in La Crosse, Wisconsin, in connection with a drive-by shooting. He denied involvement in the shooting itself and pleaded guilty to two lesser charges: fleeing from police and driving a car without the owner's consent. When he came up for sentencing, the presentence report in front of the judge included a score from a piece of software called COMPAS, a proprietary risk assessment tool that takes a questionnaire of 137 items and returns numbers predicting the likelihood that the defendant will reoffend. Loomis's numbers were high. The judge cited them in explaining why he was giving Loomis six years.
Loomis appealed, and the ground of appeal was simple and, I think, correct. He had been sentenced partly on the basis of a calculation that neither he nor his lawyer nor the judge could inspect. The company that made COMPAS treated the weighting of the questionnaire as a trade secret. Nobody in the courtroom could say which of the 137 answers had pushed the score up, or by how much, or whether the model had been validated on people like him. He had been judged, in part, by math he was not permitted to see.
In July 2016 the Wisconsin Supreme Court upheld the sentence. The court did not pretend the problem away. It ruled that a COMPAS score could not be the determinative factor in a sentence, that it could not be used to decide whether to incarcerate at all, and it required that every presentence report carrying a score come with a written warning to the judge listing its limitations, including that the tool had never been validated on a Wisconsin population and that a ProPublica investigation two months earlier had found it flagged Black defendants as high risk at nearly twice the rate it flagged white defendants who did not go on to reoffend. Having listed all of that, the court permitted the score to stay in the file. The United States Supreme Court declined to hear the case the following year.
State v. Loomis and the ProPublica findings
I keep coming back to the written warning, because I think it is the single most revealing artifact in the whole story. The court knew. The judges wrote down, in a published opinion, that the score was of unknown validity for the person in front of the bench, that it might be racially skewed, and that it must not decide the outcome. And then they let it stay in the file, next to the judge's hand, while the judge decided. That is not a court failing to understand the problem. It is a court understanding the problem perfectly and finding that the score was too useful to give up, and the question worth asking is: useful for what?
Here is my answer. An algorithm is congealed human judgment. Someone chose the training data, and that choice was a judgment about which past decisions count as ground truth. Someone chose the threshold above which a score reads as "high risk," and that choice was a judgment about how many false alarms are acceptable per correct one. Someone chose the objective, chose the 137 questions, chose to include "was your father ever arrested" and to exclude anything about the arresting officer. Every one of those choices was made by a person, and if that person had made it out loud, at the bench, you could have asked her to defend it. Congealing it into a score strips the fingerprints off. The judgment is still there, all of it, but it now arrives in the form of a number, and a number has no one to cross-examine.
Hannah Arendt watched the Eichmann trial in 1961 and came away with a description of bureaucratic evil that has been argued over ever since: responsibility distributed across so many desks that no single person holds enough of it to feel the weight. Later historians have made a strong case that Eichmann himself was a committed ideologue rather than the grey clerk she described, and I am not going to lean on the portrait. But the mechanism she identified does not depend on the portrait. When a painful decision passes through enough hands, each hand can honestly say it only did its part, and the cruelty comes out the other end with nobody attached to it. The risk score is that mechanism with the hands removed. The judge only followed the report. The report only reported the score. The score only ran the model. The model only learned from the data. The data only recorded what earlier judges did. Everyone in the chain is a bystander, and a man is in prison for six years.
This is why I think the adoption of algorithmic decision-making has been fastest not where it makes the best decisions but where the decisions hurt the most to make. Sentencing. Layoffs, where a scoring model decides which two hundred names go on the list and the manager who reads the list can tell each of them, truthfully, that it was not her call. Insurance claim denials. Credit. Content moderation. Child welfare screening. In every one of those domains a human used to have to look at a person and say no, and carry that, and the carrying was expensive and slow and it meant that fewer nos got said. The machine does not make better nos. It makes cheaper ones, and it launders the cost of saying them.
The word laundering is the right one. In money laundering the money does not disappear; it passes through a process and comes out with its origin erased. Judgment laundering is the same. The judgment is all still there, in the weights and the thresholds. What has been washed out is the name. And the second job that the name was doing, the job nobody wrote down, was to make the decision cost something for the person who made it, which is the only mechanism a society has ever had for keeping painful decisions rare.
I only know of one reform that addresses this, and it is not "ban the algorithm," which will not happen and which would, in some of these domains, produce worse and more arbitrary outcomes than the score. The reform is a name attached to every threshold. Not the model's; a person's. Somewhere in the organisation a human being chose that 7 out of 10 means "deny," and that human being's name should be on the denial, in the same way that a physician's name is on the prescription. The moment that is true, the threshold will stop being a technical parameter and start being what it always was: a judgment that someone has to be willing to defend to the face of the person it lands on.
The institution launders judgment. The market, one step further out, does something less sinister and much stranger to work itself, and to see it you have to go back to a time when the word computer meant a woman with a pencil.
When computer was a job title
In February 1962 John Glenn was about to become the first American to orbit the earth, and the trajectory for his flight had been calculated on an IBM 7090, one of the most powerful electronic computers in existence. Glenn was a test pilot and he had the test pilot's attitude to unproven machines. According to the accounts from Langley, before he would agree to fly he asked the engineers to have "the girl" check the numbers. The girl was Katherine Johnson, and her job title, on the NASA payroll, was computer. She ran the IBM's trajectory by hand, over a day and a half, on a desk calculator and paper, and when her figures matched the machine's, Glenn flew.
I love that scene because it catches the exact moment when trust had not yet migrated. The machine could do the arithmetic. Everybody knew the machine could do the arithmetic. But the machine could not yet be trusted, and trust, in 1962, still lived in a person who understood what the arithmetic was for. Johnson did not just add up the numbers. She knew what a trajectory was, what could go wrong in one, what a wrong answer would look like before it killed anyone. The computer computed; the computer named Katherine Johnson judged.
She was not the first. In the 1880s the director of the Harvard College Observatory, Edward Pickering, grew so frustrated with his male assistant's work that he declared his housekeeper could do better, and hired her. Williamina Fleming went on to lead a team of women, paid a fraction of a man's wage, who were known as the Harvard Computers and who between them catalogued the spectra of hundreds of thousands of stars. Annie Jump Cannon, one of them, personally classified something over 350,000. Henrietta Leavitt, another, found the period-luminosity relation for Cepheid variables, without which Hubble could not have measured the expansion of the universe. Their job title was computer, and the job was to do by hand the calculation that a machine would do fifty years later.
There is a pattern here that runs all the way down the twentieth century. The telegraph operator was a profession, with a union and a wage premium, until the telephone made the skill of talking to the wire into something every clerk could do, and then into nothing. The typing pool was a profession; a firm of any size in 1950 had a room full of women whose job was to convert executives' dictation into typed pages, and by 1990 the room was gone, because the executives typed, and by 2010 the typing had half-dissolved into autocorrect. The search-query wizard, the person in every office around 2004 who knew how to make Google cough up the thing nobody else could find, was a role for about five years before the engine got good enough that the skill vanished into the box.
Every one of these is a skill for talking to a machine, and every one of them follows the same arc: profession, then literacy, then disappearance into the interface. First the skill is rare and paid. Then it becomes something everyone is expected to have. Then the machine absorbs it and nobody has it, because nobody needs to.
Prompt engineering is a job title on the same clock. In 2023 people were being hired, at surprising salaries, for knowing how to phrase a request to a language model. Two years later the models take rougher instructions and ask clarifying questions, and the interface is already absorbing most of what the craft consisted of. Anyone who bet a career on the syntax of talking to the machine has bet on the short leg. The people who thrived, from Fleming's team through Johnson to whoever comes next, were the ones who converted the arithmetic into judgment before the conversion was forced on them. Johnson's title was taken, then her task, and what could not be taken was the thing Glenn actually wanted from her, which was someone who would know if the number was wrong.
The durable skill was never the syntax. It was knowing what to ask for and recognising when the answer is wrong.
That is the long leg, and it is the same skill the zipper experiment is measuring, from the other end: you can only tell the machine's answer is wrong if you have the mechanism yourself, and you only have the mechanism if you did the work the machine is offering to do for you.
More smoke
In 1865 a young English economist named William Stanley Jevons published a book with an alarming title, The Coal Question, in which he argued that Britain's industrial supremacy would end because it would run out of coal. He was wrong about that. But in the middle of the book he made an observation that everyone around him found perverse and that turned out to be one of the most durable ideas in economics.
The engineers of his day were celebrating the efficiency of the new steam engines. James Watt's design used a fraction of the coal of Newcomen's for the same work, and the obvious conclusion was that Britain would burn less coal. Jevons looked at the numbers and saw the opposite. Every improvement in the efficiency of the engine had been followed by an increase in the total consumption of coal, not a decrease, because a cheaper engine found uses that an expensive one never could. Steam moved into industries where it had never paid before. It moved into transport. It moved into mines, where cheaper pumping made deeper coal profitable, which produced more coal, which fed more engines. "It is wholly a confusion of ideas," he wrote, "to suppose that the economical use of fuel is equivalent to a diminished consumption. The very contrary is the truth."
That is the Jevons paradox, and it is the missing macro story in every conversation about AI and work. The conversation is almost always framed as replacement: the machine does the task, the human who did the task is not needed, the job goes away. That is the engineers' framing in 1865, and it misses the thing Jevons saw. When a resource gets radically cheaper, we do not buy the same amount for less. We find enormously more uses for it, and the total demand explodes.
The best-documented modern case is one that the economist James Bessen dug out of the labour statistics. The automated teller machine arrived in American banks in the 1970s, and the whole point of it was to replace the human teller. In the 1980s and 1990s it spread to every branch in the country. And the number of bank tellers in America went up. Not down; up, from something like half a million in 1980 to more than half a million three decades later. What happened was that the ATM cut the number of tellers needed per branch from about twenty to about thirteen, which made a branch cheaper to run, which meant banks opened more branches, about 43% more in urban areas, and the extra branches needed tellers. The tellers' work also changed. Less cash handling, more of what the bank called relationship banking, selling mortgages and talking customers through problems. The machine took the arithmetic and the humans kept the judgment, and there were more of them.
On the tellers, honestly
The spreadsheet is the other case. VisiCalc launched in 1979 and it automated, more or less completely, the work of the bookkeeping clerk who did the arithmetic of a ledger by hand. Those jobs did decline, by hundreds of thousands. And the jobs of accountant, financial analyst and management consultant grew by more than that, because once a projection could be re-run in an afternoon instead of a fortnight, every firm wanted ten projections where it had settled for one. The total quantity of financial reasoning in the economy went up by an order of magnitude. It just stopped being done by the people who had been doing it.
What to try: hold the efficiency slider at ten times cheaper and walk the elasticity across one. The gold line for human hours flips from rising to falling at exactly that point. Then open the teller panel and find the same flip, in the data, around 2010.
That last sentence is where I want to spend my remaining honesty, because the Jevons story is usually told as reassurance and I do not think it is reassuring at all. It is good news for the profession and terrible news for the incumbent. The bookkeeping clerks of 1979 did not, as a rule, become the financial analysts of 1989. Different people did, younger ones, with different training, who had never learned the ledger by hand and did not miss it. The tellers who survived the ATM were the ones who could sell a mortgage; the ones who were good at counting cash were not retrained, they were replaced by hires who could do the new job. Cheap cognition will follow the same path. The total demand for thinking is going to go up, and I would bet on that with real money. The total number of people paid to think will probably go up too. And the new abundant work will, for the most part, not go to the people whose old scarce work it replaced, because the new work requires exactly the skill the calculator truce was built to protect, and the incumbents whose skill was the syntax will find that the syntax is gone.
So the macro forecast is stranger than replacement. It is expansion with violent reshuffling. And the reshuffling will sort people by a single variable, which is whether, when the machine hands them an answer, they can tell if it is wrong. That variable is the stall. It is the zipper. Jevons, Bessen and Rozenblit are describing the same fault line from three directions, and on one side of it cheap intelligence is the best thing that ever happened to your career and on the other side it is the thing that ends it while you are looking the other way.
There is one more step out, past the individual, past the institution, past the market, and it is the step that turns the whole argument around.
The rendering plant
Somewhere in Britain in the early 1980s, in a rendering plant, the leftover parts of slaughtered cattle, the bones and the offal and everything that could not be sold as meat, were being cooked down into a protein-rich powder called meat and bone meal. This had been done for decades. It was an obvious efficiency: a carcass is mostly not steak, and the rest of it is protein, and protein is expensive, and dairy cows produce more milk when they eat more protein. So the rest of the cow went back into the cows. The loop was so obvious that nobody saw it as a loop.
What had changed, quietly, was the rendering process. Through the 1970s and into the 1980s plants had moved to lower temperatures and dropped a solvent extraction step, for reasons of cost and safety. The result was that a class of misfolded protein called a prion, which can survive cooking that destroys any bacterium or virus, was making it through the process and into the feed. A prion is not alive. It is a protein folded wrong that causes other copies of the same protein to fold wrong when they touch it. Feed it to a cow and the cow's brain slowly turns to sponge, and the cow goes into the rendering plant, and its prions go into the meal, and the meal goes into the next cow. The efficiency had a defect in it, and the loop concentrated the defect with every pass.
The first case of what was named bovine spongiform encephalopathy was confirmed in 1986. Britain banned feeding ruminant protein to ruminants in 1988, by which time the loop had been running long enough that cases kept climbing for years afterwards; the epidemic peaked in 1992 at over 37,000 confirmed cases in a single year, and by the time it was over something like 180,000 cattle had been confirmed sick and more than four million had been slaughtered as a precaution. In March 1996 the Health Secretary stood up in the House of Commons and said that the disease had probably crossed into humans, as a new variant of Creutzfeldt-Jakob disease, and the beef industry collapsed overnight. Around 180 people in Britain have died of it.
I tell the story at that length because I want the shape of it to be clear before I generalise. A system produces an output. The output is fed back in as input because doing so is efficient. The copying process is imperfect in a way that nobody has measured. And the imperfection, which would have been harmless at one pass, is concentrated by the loop until the whole population sickens. Prions are one instance. There are others.
The Habsburgs are one. Two centuries of marrying cousins and uncles to nieces, in the interest of keeping the crown inside the family, produced in Charles II of Spain a man whose inbreeding coefficient has been calculated at 0.254, higher than the child of two siblings. He could barely chew. He died in 1700 without an heir, and the war over who got Spain lasted thirteen years. Nobody in the family was trying to breed a defect; each individual marriage was a sound dynastic decision. The loop did the rest. I wrote about the mathematics of that kind of loop in an earlier essay on family trees, and I will not repeat it here beyond the point that matters: a copy process without fresh input does not preserve the original. It amplifies whatever the copying introduces.
Photocopy a photocopy of a photocopy and by the tenth generation the page is grey fog with the letters gone. Point a microphone at the speaker it feeds and the room fills with a shriek that was never in the original signal. Same structure, every time.
In July 2024 a group led by Ilia Shumailov published a paper in Nature with the title "AI models collapse when trained on recursively generated data." They trained a language model, used its output as the training data for the next model, used that model's output for the next, and watched what happened over generations. What happened was the photocopy. The tails of the distribution went first; rare phrasings, unusual facts, minority views, anything the model saw infrequently, were slightly under-represented in the first generation's output, and so slightly rarer in the second generation's training data, and so rarer still in its output, until after a handful of generations the model had forgotten they ever existed and had converged on a narrow, confident, bland centre. Their example was almost too good: a model that started out able to discuss medieval English church architecture was, nine generations later, producing text about jackrabbits in a variety of colours. They called it model collapse.
What to try: with fresh data at zero, step through the generations and watch the tails disappear before the centre moves. Then set fresh data to twenty percent and run all twelve again. The loop stops being a loop.
How strong the collapse result really is
I want to be careful here in a way that the headlines were not. The strong version of the story, in which the internet fills with machine text and every model trained on it gets stupider until the whole thing falls over, is not what the evidence shows. What the evidence shows is that if you feed a model only its own output, it collapses, and that if you keep the real data in the mix, it mostly does not. Which is exactly the lesson of the rendering plant. The problem was never that cows ate protein. The problem was a process that let the defect through and a loop that closed on itself with nothing fresh coming in. Ban the loop, keep the fresh input, and the herd survives.
And that is where the whole argument turns over, because look at what the fresh input is. It is human writing that a human actually did. It is a photograph someone stood in the rain to take. It is an argument produced by someone who stalled halfway through and had to go and find out. It is Kapur's floundering, Rozenblit's zipper, the physician's judgment about the toothpick, the name on the threshold. It is, in a word, the friction. The entire first three-quarters of this essay has been about a machine that removes friction from human minds, hearts and institutions, and the last quarter is the discovery that friction is the nutrient the machine cannot synthesise for itself. It needs the stall. It needs there to be people who did not use it.
Which means that the machine, without anyone intending it, is re-pricing exactly the human work it was supposed to replace. Fresh, friction-born human thought, the thing that was becoming free because the machine could imitate it, is becoming the scarce input in the machine's own food chain. You can already see the price signal: the content licensing deals between AI companies and newspapers, the premium on verified-human forums as training data, the quiet panic in every lab about where the next order of magnitude of real text is going to come from. The economics of the rendering plant, transposed: once the loop was understood, fresh feed was the most valuable thing on the farm.
The map of the stall
I said at the start that the intelligence of the masses would not fall uniformly, and I want to say now what I think will happen instead, because it is the thing I did not expect when I began writing.
It will split. On one side will be people who use the machine the way the engineer uses the calculator: from the top of a skill they already have, as leverage, checking the answer against a mechanism they carry in their own heads and noticing, when it is wrong, that it is wrong. On the other side will be people who use it the way the eight-year-old uses the calculator: as a substitute for the skill, from below, with no mechanism to check against and no stall to tell them what they do not know. The two groups will look identical on every immediate test. The gap between them will be invisible for years. And then it will be the largest gap in economic history, because Jevons says the demand for judgment is about to explode and the rendering plant says the supply of the real thing is about to become the scarcest input on earth, and both of those price signals point at the same small group of people: the ones who kept paying the cost.
That is the twist. I began thinking this was an essay about loss, about a machine that anaesthetises the signal of ignorance and launders the weight of judgment and out-performs the tired doctor and out-eggs the real egg. All of that is true. But the machine cannot eat its own output, and so it has quietly put a price on everything it was supposed to make free. The struggle it removes from you is the struggle it has to buy back from someone. It might as well be you.
None of that helps unless it becomes a practice, and the practice I have found is so simple that I was embarrassed to write it down. Close the chat. Take whatever you just learned, or think you learned, and explain it out loud, to no one, to the wall. Do it in full sentences, the way you would to someone who was going to check. And notice exactly where you stall.
I spent most of a book, STILL, on the phase before the visible event, the slow storing of something under a surface the world misreads as nothing happening. I called it consolidation there, and the chapter I keep coming back to is the one on sleep, where the brain replays the day's attempts and keeps only the ones that carried a strong enough signal, and the strongest signal turns out to be difficulty. Passive exposure lays down shallow traces that do not survive the night. An attempt at the edge of your ability, an error noticed, the gap between what you expected and what happened, those are the traces that cross the threshold and become part of you. The difficulty is the signal and the failure is the data. The stall is the same mechanism, caught in daylight. It is the moment you find out which traces were ever laid down at all, and the reason the machine's fluent answer does not consolidate is that nothing hard happened to you while you read it.
You will stall. Everyone does; that is what Rozenblit and Keil found, and it is what the four hundred years of the Habsburgs and the ten years of the prions and the sixty years of the plaster egg all found in their own way, that the smooth surface has a mechanism under it that nobody looked at. The stall is not the failure. The stall is the only honest map of what you know, and it is the one thing on earth the machine cannot draw for you, because the moment it does, the map is of the machine.