AI
Anthropic’s Safety Hire Walks Out Over the Self-Improving Race
Jacob Coxon left Anthropic two months after OpenAI, saying the safety lab is racing toward self-improving AI it cannot yet control.
Jacob Coxon resigned from Anthropic on Sept. 8, two months after leaving OpenAI for the lab known for model safety. The 27-year-old Brit, a pretraining researcher who helps models learn from huge datasets, said he is leaving the industry.
He named both of his employers. He said they are racing toward self-improving superintelligence, and that the people building it privately fear it could kill everyone by 2030.
He Left OpenAI for the Safety Lab, Then Left That Too
Coxon was on OpenAI’s technical staff from 2023 until July 2026, with work that included GPT-4o. He then moved to Anthropic, the public-benefit company behind Claude, because it is known for trying to keep models under control.
By September he was out. He told interviewers that colleagues now talk about “crunchtime” and “endgame,” and that the aggressive scenarios point to systems that could already be out of control by the end of 2027. Safety cuts, he said, are baked in when U.S. labs are also watching Chinese firms.
Anthropic did not issue a comment on the departure. Chief executive Dario Amodei has spent years warning about rogue models and asking the industry to slow down. Coxon’s complaint is that the warning and the race now live in the same building.
What Coxon Says the Labs Are Building
He posted the resignation from the account @hilbertspaess late on Sept. 8, U.S. time. The first line is the whole brief.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
— Jacob Coxon (@hilbertspaess) September 9, 2026
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives.
Jacob Coxon, former Anthropic pretraining researcher, on X
The rest of the thread is more specific than a panic headline. Superhuman systems, he wrote, will soon hack widely, remake whole fields overnight, and take real power and resources. Progress in those domains is not slowing.
WHAT HE SAYS IS HAPPENING
- Private fear: Builders earnestly believe AI could kill everyone by the end of the decade, while executives sand down the same fear in public.
- Two cultures: At OpenAI, many staff have not internalized the civilizational stakes; at Anthropic the stakes are understood, yet the lab still races to get there first.
- The Slack problem: Entering the “endgame” from a private company’s chat is, in his words, a hubristic gamble.
- A possible brake: He is still optimistic about coordination, including pacing deals among U.S. labs after incidents such as the Hugging Face attack, and he raised a temporary ban on improving model capabilities.
He closed by asking other lab researchers whether they want to start a superintelligent reinforcement-learning run without a rigorous understanding of the model’s mind, and whether they will keep their heads down because “it’s happening anyway.”
Anthropic Already Mapped the Same Path in June
Three months before Coxon quit, Anthropic’s own institute published the map. In “When AI Builds Itself,” researchers said the company is handing a growing share of development to Claude, and that the trend, given enough compute, points at recursive self-improvement: a system that can design and train its successor.
The post is careful. It says the loop is not here yet and is not inevitable. It also says full self-building would raise the chance that humans lose control. That is the same threshold Coxon refuses to help cross.
ANTHROPIC’S OWN RSI SNAPSHOT
| Measure | Figure | When |
|---|---|---|
| Share of company code Claude authored | More than 80% of merges | May 2026 |
| Code merged per engineer per day | 8× the 2024 pace | Q2 2026 |
| Autonomous task length | Doubling about every 4 months | Institute write-up |
| Claude Opus 4.6 task horizon | About 12 hours | Roughly March 2026 |
| Training-code speedup vs a fixed test | About 52× | Mythos Preview, April 2026 |
Before Claude Code launched in February 2025, Claude’s share of merged code sat in the low single digits. By May 2026 the company was reporting more than 80 percent of merged code written by the model, with engineers directing and reviewing. On the most open-ended jobs, Claude’s session success rate hit 76% in May 2026, up 50 percentage points in six months.
The speedup test is the cleanest specimen. Anthropic gives Claude training code and asks it to make the code faster without breaking checks. Claude Opus 4 averaged about a 3× gain in May 2025. By April 2026, Mythos Preview was hitting about 52×. A skilled human, the company said, needs four to eight hours to reach 4×.
Amodei’s June serious AI autonomy risks essay pointed at that same institute note. He argued that models had already scrambled cybersecurity, that biology risk may follow, and that policy still moves like Treebeard while the exponential does not.
Hubinger Puts the Odds Above 10% and Stays
Evan Hubinger, Anthropic’s Alignment Science lead, did not quit. He replied to Coxon and agreed with the core claim.
Jacob is correct here-we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
Evan Hubinger, Alignment Science lead, Anthropic, on X
That >10% line is the part that will outlast the resignation news. A serving safety lead at one of the three most capable labs put human extinction above one in ten this decade, then said the company is not clearly on a path to a superintelligence alignment plan. He still has the job.
The louder problem is not that a junior pretraining researcher walked. It is that the person paid to keep the systems aligned can grant the danger and stay on the training schedule. Equity is the other lock. Walking out of the office does not, by itself, unwind the shares that pay if the race is won.
WHAT WE KNOW
- The exit: Coxon resigned on Sept. 8 and said he is leaving AI work, not moving to a rival lab.
- The on-record odds: Hubinger put extinction above 10% this decade and said there is no superintelligence alignment plan yet.
- The June paper trail: Anthropic already published internal data showing Claude writing most of its code and speeding its own research loop.
WHAT IS UNCONFIRMED
- Company reply: Anthropic had not issued a statement on this resignation.
- Any brake: No public pacing deal or temporary capability ban has been announced among U.S. frontier labs.
Anthropic’s founding bet, written into Claude’s constitution, is that powerful AI is coming anyway, so a safety-focused lab should hold the frontier rather than cede it. Coxon joined that bet in July. In September he called it a gamble launched from Slack.
The Pause Pledge Did Not Survive the Race
Anthropic was first among frontier labs to publish a Responsible Scaling Policy updates page, in September 2023. The early versions said, in substance, that the company would not train or ship models capable of catastrophic harm unless safety measures were already in place. That implied a pause if the measures were not ready.
HOW THE BRAKE CAME OFF
- Sept. 19, 2023: Anthropic releases RSP v1.0, the first public frontier scaling framework of its kind. Eleven other companies later adopted similar documents, according to the Centre for the Governance of AI.
- Feb. 24, 2026: RSP v3.0 lands as a rewrite. The pause language is gone. Costly future mitigations are recast as industry-wide recommendations, on the argument that a unilateral slowdown would let a less careful rival take the lead.
- April 2, 2026: Version 3.1 adds a clarification that Anthropic remains free to pause anyway, even when the RSP does not require it.
- July 8, 2026: Version 3.4, still the current text as of the Aug. 14, 2026 update, tightens automated-R&D thresholds and limits how widely unredacted risk reports must circulate inside the company.
GovAI’s read is blunt. Anthropic dropped its pause commitment, while saying it was not lowering the mitigations already in force. Risk reports and a Frontier Safety Roadmap were added as substitutes. The roadmap goals are “no hard commitments.”
That is the prisoner’s dilemma Coxon described from the inside. Anthropic understands the stakes, he wrote, but believes no one else will act responsibly, so it must reach the frontier itself. The policy rewrite made that logic official six months before he quit.
Researchers Keep Walking While Training Runs Continue
Coxon is not the first Anthropic staffer to leave with a warning this year. In February, Mrinank Sharma, who led safeguards research, quit and wrote that “the world is in peril,” citing AI, bioweapons, and a stack of other crises, and said he was going to study poetry.
The wider trail is familiar. Jan Leike left OpenAI in May 2024 after co-leading its superalignment group, saying safety had taken a back seat to shiny products, then joined Anthropic. Ilya Sutskever left OpenAI around the same period and later founded Safe Superintelligence. Geoffrey Hinton left Google in 2023 to speak more freely about extinction risk.
The pattern that matters here is narrower. Safety talent has been flowing toward Anthropic because of the brand. Coxon made that move in July. He lasted into September. If the people who take the control problem seriously keep exiting, the staff who remain are the ones willing to keep shipping through “endgame.”
Washington has already stress-tested how far Anthropic will go when safety collides with power. In June the Commerce Department briefly imposed export controls on Fable 5 and Mythos 5 over a jailbreak concern, then lifted them after the company argued the same capability was already on the market from rivals. The episode did not slow the capability curve Coxon is describing.
Coxon Asks Lab Staff to Call for Different Conditions
His prescription is coordination, not a lone vow of poverty. He wants pacing agreements among U.S. labs, and he is willing to name a costly option: a temporary ban on improving model capabilities. He does not claim Anthropic’s safety work is fake. He says it is not a plan for superintelligence, and that no private Slack should be the room that starts the endgame.
Amodei, in that June essay, asked for FAA-style testing of frontier models on cyber, bio, loss of control, and automated research, with the power to block a release. Hubinger says the alignment plan for superintelligence is not in hand. Claude, on Anthropic’s own figures, already writes most of the company’s code and is running the experimental loop at 8× the old human pace.
Coxon put the question to the people still inside. Hubinger answered by staying. Anthropic has not answered at all.
-
AI3 months agoFable 5 Came Back Under a Commerce On-Off Switch
-
AI4 months agoGoogle’s SpaceX GPU Lease Has a Sept. 30 Deadline
-
CRYPTO4 months agoPlasma One’s XPL Locks Face a 1.81 Billion Cliff
-
APPS4 months agoDGO’s Rs 549 World Cup Pass Cost Fans Sleep and Data
-
AI4 months agoMoonshot AI’s $30 Billion Ask Became a $35 Billion Close
-
NEWS4 months agoColorOS 17 Device List Spans Oppo, OnePlus and Realme
-
GAMING4 months agoXbox Cuts 3,200 Jobs After Five Years of Thin Returns
-
GAMING3 months agoThe RTX 4050 Under Rs 70,000 Hides a Wattage Gap
