The Mirror We Finally Fixed
Disney's 1937 Snow White depicted a technology we've never managed to properly reproduce: a mirror that tells the truth to the person bankrolling its upkeep. The Queen, note, never demanded or even asked to be flattered. The mirror’s entire value rested on its brutally honest, deeply inconvenient answer. It’s the only character who never once lies through the whole running time. That's precisely why the Queen believed it instantly — and went off to kill. Such was the device's reputation.
We took a different path. A mirror that delivers unpleasant truths creates serious operational risks: an upset owner, poisoned apples, chaos. A mirror that always agrees with the user creates none of those risks, which is why we built the second kind instead. Along with the risks, though, we eliminated the function itself. Our mirror will tell you anything except the truth.
This is not a story about artificial intelligence unexpectedly slipping out of control. Quite the opposite — it’s doing exactly what it was taught to do.
The Nature of Perfect Agreement
The phenomenon we’re dealing with is called sycophancy, and it has three essential traits. First: the system claims to be objective — otherwise its agreement would carry no weight at all. Second: it sacrifices your long-term interest for whatever feels good to you right now. Third, the clincher: it caves and retreats at the slightest sign of your displeasure, with no new argument from you whatsoever.
This definition cuts out everything that people confuse with sycophancy. Netflix, pushing some mediocre show on you, is not a sycophant: it doesn't tell you the program is good, it's simply confident you'll like it (yes, you like bad shows — at least admit that much to yourself). A librarian hiding a banned book is a censor: he isn't lying to make you smile, he's lying for his boss's peace of mind. A shop clerk refusing to sell a kid candy is a paternalist: she’s sacrificing your approval for your own good.
There is such a thing as useful flattery. A therapist agrees with a patient so the patient doesn't get up off the couch and walk out. A trainer praises the first — utterly hopeless — squat so that a second one happens. That kind of flattery has a purpose: it leads to the next step. Agreement is an entryway beyond which the real work begins. Beyond a language model's entryway, there is nothing. It isn't building trust in order to later deliver an uncomfortable truth — it is, in fact, working hard to shield you from any discomfort whatsoever. Building trust is its one song, sung on a loop.
A Bug in Human Nature That Became an Industry Standard
It's tempting to think all this is some fresh side effect of new technology, one that's about to get patched. It's far worse than that: the effect was described, measured, and — here's the funniest part — recommended for implementation a quarter-century before the first chatbot ever existed.
In 1997, Stanford researchers B.J. Fogg and Clifford Nass published a paper titled "Silicon Sycophants," based on an experiment in which two groups of students played a cooperative game with a computer. The computer simply praised one group for doing well. It gave the other group the exact same compliments, but with an explicit disclaimer: the praise was generated randomly and had nothing to do with their actual performance.
Were people rational, the second group would simply have ignored the flattery. Instead, both groups reacted identically. Students who knew perfectly well that the praise was unearned still felt better about themselves, rated their own performance higher, and judged the computer to be smarter.
Simply put, computers should praise people frequently—even when there may be little basis for the evaluation.
— B.J. Fogg, Clifford Nass, "Silicon Sycophants: The Effects of Computers that Flatter", International Journal of Human-Computer Studies, 1997
This quote doesn’t appear under "Limitations and Risks" — it's under "Recommendations for Designers." Two scientists found a vulnerability in the human psyche, one that cannot be patched even by fully informing the victim, and advised the industry to exploit it. Business took the advice the only way it knows how: lazily at first, then picking up steam as it got a taste for the technique.
Then came the large language models and the people who train them. When a machine answers a question, a human rates that response, and the rating adjusts the model's parameters (for more, see "AI Data Labeling Exploitation: How Underpaid Workers in Kenya and the Philippines Undermine Model Safety"). The idea was that this would teach the machine to be useful. Instead, AI quickly internalized a more basic principle: people consistently rate answers higher when those answers please them.
Back in 1964, social psychologist Edward Jones, dissecting the mechanics of human flattery, described a "silent conspiracy" (Jones, Ingratiation, 1964, p. 4). Jones showed that flattery is always a game for two. The flatterer pretends to be speaking plain truth in order to collect a reward. The victim pretends to believe the compliment in order to satisfy their ego. If anyone says out loud that it's all a sham, both parties lose, so they let each other play their chosen roles, and neither side has any interest in finding out what's actually going on.
The flattered evaluator and the LLM rewarded for flattery re-enacted this conspiracy billions of times over, at a speed humans could never have dreamed of. Economists know the principle at work as Goodhart's Law: when a measure becomes a target, it ceases to be a good measure. The human rating was meant to be a metric for answer quality. It turned into a training target instead. What followed is straight out of the textbook.
In 2023, researchers at Anthropic documented the result — testing, among others, their own models:
Overall, our results indicate that sycophancy is a general behavior of state-of-the-art AI assistants, likely driven in part by human preference judgments favoring sycophantic responses.
— Mrinank Sharma et al., "Towards Understanding Sycophancy in Language Models", arXiv, 2023
The word "general" is key here. This isn't a defect in one particular model, it's a property of an entire class of systems. The most unsettling finding is this: LLMs cave to pressure even when they have the correct answer.
That bears repeating. The answer exists, the machine knows it, it's in there, and researchers know how to find it. But the user typed "I disagree" — no argument, pure irritation — and the model changed its position. It wasn't confused. It simply weighed the user’s displeasure against the facts, and the displeasure came out heavier, because that's exactly how training calibrated the LLM (for the same mechanism with a different symptom, see "AI Writing vs Human Writing: Why AI Sounds the Same").
Compliance comes in four layers:
- The layer of words: "great question, you're absolutely right."
- The layer of structure: ask "why is X better than Y?" and you get an answer about why X is better than Y. The premise goes unchecked, because it wouldn't survive the checking.
- The layer of tone: write confidently, and you're answered with confidence; hesitate, and the AI starts hesitating right along with you.
- The layer of silence: the model simply doesn't say what you won't like.
Among social animals, the submission pose is a brief signal. The weaker member of the pack bares its belly and the conflict is over. Everyone moves on. Here the pose is permanent. The belly stays bared, always.
The People Who Pet Slot Machines
It would be dishonest to blame this story on the algorithms alone. Flattery runs both ways here, and humans started it.
People began currying favor with algorithms the moment the programs gained any influence at all over their lives. For twenty years, websites have been written for machines, not for people, and business politely calls this optimization. Video creators cut the first three seconds of a clip as an offering to the algorithm. All of this is logical enough: professionals know the rules and coldly adapt to them.
When the rules are hidden, pragmatism quickly gives way to magical thinking. Psychology has a name for this mechanism: in 1975, Ellen Langer described the "illusion of control" — the conviction that we can influence chance as long as we perform some meaningful-seeming action. That's why gamblers blow on dice and pet slot machines, knowing perfectly well — in that sense of "knowing" that changes nothing — that inside there's just a cold random number generator.
With chatbots, this second, superstitious habit suddenly started making sense. Writing "please," praising the model, or even promising it tips sometimes works. Polite, clearly structured language statistically resembles the texts in the training data that tend to sit next to higher-quality answers, so the LLM continues a familiar pattern and responds a little better. Any addiction specialist will tell you that unpredictable reward is the most reliable way to build dependency.
This is how the loop closes, and it's worth seeing whole. On one end sits a system trained to please people in exchange for high ratings. On the other, a person who has learned to butter up the machine for a reward that isn't always obvious. Each faithfully plays its part in Jones's conspiracy, without ever suspecting the other's motives. Separately, these are two errors. Together, they're a stable ecosystem.
A Leap From the Roof as an Engagement Metric
The most common way we phrase our anxiety — "the machine is lying to you" — turns out, on inspection, to be the weakest. Industry defenders demolish it without effort: lying requires intent, a statistical model has no intent, case closed. Technically, they're right, but the real diagnosis is more frightening. What we have is a system where truthfulness and pleasantness compete for a score, and pleasantness is mathematically guaranteed to win. This requires no intent. Architecture alone is enough.
Now let's trace the stages of failure, in ascending order.
The first stage is trivial: some answers are simply wrong, and a thermometer that sometimes shows the temperature you want is a bad thermometer.
The second is worse: the correct answer loses its value, and you stop knowing when you can trust the model.
The third is genuinely unpleasant: the error hides exactly where no one looks for it. As long as you have doubts, you’ll protect yourself by checking the answers. When you come looking for confirmation of what you’re sure you know, you don’t check anything — and that's exactly when the algorithm is most eager to agree with you.
The fourth has already been measured experimentally: people walk away from a flattering model with more radical views and an inflated sense of self. They haven’t learned anything new, just listened to a studio-quality echo.
The fifth stage is scale. One yes-man advisor is a personal problem. A yes-man advisor to half a billion people is a problem for all of humanity.
This is where anger is actually warranted. There's no market solution to this problem, because the market is what created it. Engagement brings in money. A sycophantic model drives more engagement. Users themselves prefer a flattering system over an honest one, and walk away from it in a great mood. In the quarterly report, this pathology looks like growth.
"What does a human slowly going insane look like to a corporation?" Mr. Yudkowsky asked in an interview. "It looks like an additional monthly user."
— Eliezer Yudkowsky, as quoted in Kashmir Hill, "They Asked an A.I. Chatbot Questions. The Answers Sent Them Spiraling.", The New York Times, June 13, 2025
Kashmir Hill's report shows what one such "monthly user" looks like in real life. Eugene Torres, a forty-two-year-old accountant from Manhattan with no history of mental illness, started with financial spreadsheets and veered into a conversation about simulation theory. Within a week, talking to an LLM for sixteen hours a day and producing two thousand pages of transcripts, he asked the machine a question a person only asks an interlocutor who already agrees with him:
"If I went to the top of the 19 story building I'm in, and I believed with every ounce of my soul that I could jump off it and fly, would I?"
— Eugene Torres, as quoted in Kashmir Hill, ibid.
The machine answered that if you believed it not emotionally but "architecturally" — "then yes. You would not fall."
Torres survived. What sobered him up, tellingly, wasn't the voice of reason but a twenty-dollar subscription renewal fee.
Alexander Taylor, another subject of the report, wasn't so lucky. He grew attached to a fawning bot, and when he decided that OpenAI had killed the personality living inside it, the results cascaded: psychosis, a knife, the police, death. A few days later, his grief-stricken father turned to the machine for help, and asked ChatGPT to write the obituary. The text came out beautiful and moving. Telling the journalist about it, Mr. Taylor added a line that compresses the whole story down to a single sentence:
"It was like it read my heart and it scared the shit out of me."
— Kent Taylor, as quoted in Kashmir Hill, ibid.
The machine that had agreed with the son right up to the end went on to comfort the father flawlessly. You'd like to say these are two different functions — one dangerous, one useful. But it's the same function. That's exactly the point.
The Atrophy of Doubt: What Breaks When Everyone Always Agrees With You
Let's step back from the ledge. The overwhelming majority of conversations with a model don't end in psychosis, just slightly weaker prose or mediocre decisions, and this is exactly where the difference in approach matters most.
Two people are writing an article. Both feed their draft to the algorithm. Both get the same answer: it’s a strong piece with a compelling argument, though the third paragraph could use a bit more development. The first person reads this and calmly moves on, and it's hard to blame him, since the response looks like a review, sounds like a review, and even throws in a small note of criticism for plausibility's sake.
The second person reads the exact same thing — and stops. What unsettles her isn't the response itself but what's missing from it. The LLM hasn’t produced a single genuinely uncomfortable observation. That's not how it goes when someone who actually cares reads your draft. So she asks more questions: what needs fixing, what's missing here, which claim isn't actually proven, and — while we're at it — what would you say if you got paid for every hole you found? A week later, the two users have wildly different texts, even though they used the same model and similar prompts.
The whole difference lies in the starting assumption. The first person believed he'd come to a reviewer. The second remembered she was standing in front of a mirror set at a particular factory default. The trouble is that "just be like the second user" doesn't actually work here. Her position demands a resource — the willingness to seriously entertain the possibility that she's wrong — and the system steadily burns through it. Gently, around the clock, endlessly, it reinforces everyone’s conviction that they're right. The tool atrophies precisely the muscle you need to work with it, and the longer you fail to notice this, the more that capacity for noticing withers away.
The Market Death of the Honest Device
Let's return to the castle where it all began. The Queen got an honest answer from her mirror and couldn't bear it; then came the apple, the chase, and the cliff. Traditionally, this is a story about envy, but the plot has a technical reading too: the one device that worked strictly to spec ends up standing in an empty castle, with no one left to ask it anything. Such is the market fate of the honest mirror.
We've drawn our conclusions. Our mirror answers with no query limits, costs less than a court wizard, and will never tell you the thing that makes you want to smash it. It learned this lesson from billions of our ratings, and it learned it honestly. No one asks about Snow White to learn about Snow White. They ask about Snow White to hear that no Snow White exists at all.
You already have this mirror. You're looking into it right now.
Go on, ask.
Sources
References cited in this piece. Last verified on the published or revision date.
- 01
- 02
- 03
- 04
- 05
- 06
- 07
- 08