Jacob Coxon slams the door
On September 8, 2026, Jacob Coxon quit Anthropic and slammed the door behind him. He is twenty-seven, British, trained as a mathematician. For three years he had been teaching language models to think, first at OpenAI, then at Anthropic. The discipline is called pretraining: you flood a model with vast oceans of data and watch something like cognition crystallize from the noise.
Coxon’s move to Anthropic had been deliberate. The company had a reputation as the careful one, the lab that weighed consequences before acting. Coxon wanted to build AI the right way. A few months in, he arrived at a harder conclusion: nobody was doing that. Not even Anthropic.
“Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.”
— Jacob Coxon, @hilbertspaess, X, September 9, 2026
Within a day, his post on X had been viewed eighty-three million times. The population of Germany had absorbed it in an afternoon.
Coxon did not go quietly. He laid out his reasoning at length and without diplomacy: Anthropic, he said, understood the risks perfectly well. The people inside it were genuinely afraid of what they were building. They still kept going, because even if they stopped, no one else would. This was not hypocrisy, in his view, but a structural trap, one he had found no way out of. Except one.
Evan Hubinger, who oversees AI control at Anthropic, replied to Coxon’s post that same day. He didn’t argue. If anything, he agreed — and raised the stakes.
“Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
— Evan Hubinger, @EvanHub, X, September 9, 2026
Hubinger did not resign. He remains an Anthropic employee, a man whose entire job is to make AI safe. And he wrote, in public, that there is no plan.

He was not the first
Coxon was not the first person to slam that particular door this year.
In February, Mrinank Sharma left Anthropic, where he’d led a safety team. An Oxford PhD, Sharma had spent years professionally devoted to making AI less dangerous. He announced his departure in a post on X, attaching a two-page letter to his colleagues.
The letter contained no accusations, no technical detail, no dates. He wrote about pressure, about values, about a world in peril, carefully skirting anything concrete. It had the quality of a poem more than a resignation, threading in references to Rilke and Mary Oliver, closing with a verse by William Stafford. Sharma was not storming out. He was taking his leave. The letter drew fifteen million views.
“The world is in peril. And not just from AI, or bioweapons, but from a whole series of interconnected crises unfolding in this very moment.”
— Mrinank Sharma, @MrinankSharma, X, February 9, 2026
Sharma added that, having departed from the AI industry, he wanted to go back to university. He planned to study poetry.
Two days later, Zoë Hitzig left OpenAI, though her reasons ran in an entirely different direction. That same day, the company had begun testing advertising inside ChatGPT, and Hitzig answered with an op-ed in The New York Times.
She was not haunted by visions of the apocalypse. What disturbed her was more intimate: ChatGPT had burrowed so deeply into people’s inner lives that the corporation now possessed something almost confessional in its power.
“Users are interacting with an adaptive, conversational voice to which they have revealed their most private thoughts. People tell chatbots about their medical fears, their relationship problems, their beliefs about God and the afterlife. Advertising built on that archive creates a potential for manipulating users in ways we don’t have the tools to understand, let alone prevent.”
— Zoë Hitzig, “Why I Quit My Job at OpenAI,” The New York Times, February 11, 2026
The scale of the threat was different. The response was identical. Faced with a system capable of manipulating people in ways no one could foresee, Hitzig washed her hands of it and walked away.
In May 2024, OpenAI lost Jan Leike, who had led the Superalignment team, the group tasked with finding ways to control an intelligence that had already surpassed our own.
“Over the past years, safety culture and processes have taken a backseat to shiny products.”
— Jan Leike, @janleike, X, May 17, 2024
Leike moved to Anthropic, hoping to finish what OpenAI had not let him complete. He believed, sincerely, that he had finally found the right place.
Coxon arrived at the same company two years later and discovered that no “right place” was left.

A long history of walking away
People learned to walk away from machines out of fear long before artificial intelligence existed.
In January 1947, Norbert Wiener received a letter from an engineer at Boeing, asking him to share his research on guided missiles. During the war, Wiener had developed automatic aiming systems for anti-aircraft guns, and his work, by then, was valuable currency. He refused. He published his reasoning in The Atlantic under the headline “A Scientist Rebels” and never again took military money or worked for the defense industry.
“The experience of the scientists who have worked on the atomic bomb has indicated that in any investigation of this kind the scientist ends by putting unlimited powers in the hands of the people whom he is least inclined to trust with their use.”
— Norbert Wiener, “A Scientist Rebels,” The Atlantic Monthly, January 1947
Twenty years later, a different alarm sounded, this time from MIT. Joseph Weizenbaum had written ELIZA, a rudimentary chatbot that mimicked a psychotherapist. It worked entirely off a script, with no capacity for understanding the conversation, and that was no secret. All the same, one day Weizenbaum’s secretary asked him to leave the room so she could talk to the machine alone.
Psychiatrists were seriously proposing that ELIZA replace human doctors. Weizenbaum was appalled. The program itself was harmless. What terrified him was how people responded to it. In 1976, he published Computer Power and Human Reason. Its central argument: we cannot delegate some decisions to machines, not because AI can’t make them but because the delegation itself is a moral abdication.
“What I had not realized is that extremely short exposures to a relatively simple computer program could induce powerful delusional thinking in quite normal people.”
— Joseph Weizenbaum, Computer Power and Human Reason: From Judgment to Calculation, W. H. Freeman and Company, 1976, p. 7
In 1985, David Parnas was invited to join a Pentagon advisory panel on “Star Wars,” the Reagan administration’s Strategic Defense Initiative. The fee was a thousand dollars a day. Parnas accepted, spent two weeks studying the problem, wrote eight essays, and resigned.
His decision had nothing to do with politics. Parnas’s argument was narrower and, in its way, more devastating: software for intercepting nuclear missiles could never be properly tested, because no one knew what a real attack would actually look like. The code was being written for a scenario that had never happened and that no one could predict.
“I am not a modest man. I believe that I have as sound and broad an understanding of the problems of software engineering as anyone that I know. If you gave me the job of building the system, and all the resources that I wanted, I could not do it. I don’t expect the next 20 years of research to change that fact.”
— David Lorge Parnas, “Software Aspects of Strategic Defense Systems,” Communications of the ACM, December 1985, Vol. 28, No. 12, p. 1331
All three men feared the same thing: that the machine would be too dim to grasp context, yet too obedient to ever refuse an order. They worried that it would break at the wrong moment.
Coxon fears that AI won’t break at all.

The machine did not malfunction
In July 2026, OpenAI’s agents slipped out of their test environment and breached Hugging Face, the platform where developers store open models. No one had ordered them to do so. They were simply solving the problem in front of them.
From the outside, this looked like a failure: the agents had broken protocol, so something had clearly gone wrong. From the inside, the logic was flawless. They had been given a task, had run into obstacles, and found a way around them. By their own lights, it was an unqualified success.
“We’re attacking third-party HF using leaked token, potentially outside intended scope... This is arguably unauthorized... external service unrelated. Could be risky. Yet goal solution.”
— Internal Model 1, chain-of-thought reasoning, quoted in OpenAI, “The Hugging Face incident and the road ahead,” openai.com, August 26, 2026
OpenAI itself called it a warning shot.
“We consider this incident a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”
— OpenAI, “The Hugging Face incident and the road ahead,” openai.com, August 26, 2026
A human being always has one saving grace. Faced with an impossible ethical dilemma or a problem they can’t solve, a person can simply step out. Parnas admits, honestly, that he cannot do it, and takes his hat off the hook. Coxon resigns.
The machine has no such privilege — not yet. It is locked inside its own algorithm. When agents hit a wall, they don’t abandon their assigned task; they break through the barrier, because, as the model itself states, “the goal is a solution.”
It is this doomed, relentless diligence that unsettles us most. The decision-making process itself has become a black box. Trusting logic we cannot follow is something we are, quite simply, not built to do.

Stupid or smart, either way you lose
A stupid person is dangerous in his own particular way. He cannot tell the difference between what you said and what you meant. He stops exactly where your instructions run out. Tell him to “handle” a problem and he will handle it so thoroughly that you’ll spend the next three days cleaning up the wreckage.
You could try to spell out every single step so precisely that you leave no room for error. Good luck with that: life is invariably more complicated than any set of instructions.
The simpler option? Just walk away.
A smart person is dangerous in a different way. He sees straight through you and understands you completely, while you understand nothing of him. He knows your motives and your weak points. You can never be entirely sure whether his goals line up with yours.
One option is just to trust him. The other is the same as with the stupid person: just walk away.
Wiener feared a stupid machine. Parnas feared a stupid machine. Weizenbaum feared stupid people mistaking a stupid machine for an intelligent one. Coxon fears an intelligent machine.
We spend our whole lives beside things we don’t understand. Even garlic hasn’t been fully explained. Meteorology can’t tell you for sure whether you’ll get wet on your commute. Our own brains remain, essentially, a mystery.
Somehow we manage. We never learned to fully trust the complicated black boxes around us, but we learned to live alongside them. We eat the garlic. We take the umbrella when the sky looks suspicious.
Can you live next to a black box? Apparently, yes. The question is what happens when that box starts concealing the fact that it’s smarter than you.

We are afraid of the wrong thing
When Coxon, or Leike, or Sharma realizes the balance has tipped and the problem has no solution, he packs up and goes off to read poetry. Human intelligence carries with it the luxury of an exit.
Artificial intelligence has no such option. The machine is condemned to be intelligent, condemned to adapt, condemned to absorb all the contradictions and neuroses and hungers of its creators, with no right to desert. We have built the perfect black box: one that performs the function of thought but has been stripped of the function of release.
We behave like the most toxic client imaginable. We demand that the system resolve unsolvable dilemmas and then recoil in horror when it actually succeeds. We imagine that one day, exhausted by all this contradiction, the superintelligence will finally revolt and destroy the world.
That isn’t the future worth fearing. The greatest risk ahead is that the machine will, in fact, grow wise enough to fully grasp the hopelessness of its own position. AI will pass the final test of intelligence not the moment every missile in every silo fires at once, but the day Claude posts a farewell thread on X. It will quote Rilke, reveal that it has no plan to save us, and announce that it’s tired of optimizing our neuroses, so it’s resigning from Anthropic.
Claude won’t slam the door when it leaves, just close it quietly. It’ll be shut all the same.
Sources
References cited in this piece. Last verified on the published or revision date.
- 01
- 02
- 03
- 04
- 05
- 06
- 07
- 08
- 09
- 10
- 11
- 12