Connect with us

AI

The Yemen Cell Kept Claude’s Flight Code After the Ban

A Yemen weapons cell used Claude Code to write rocket guidance, then kept an offline toolkit after Anthropic banned the accounts.

Published

on

A northern Yemen weapons cell used Anthropic’s Claude Code as its software bench, then kept an offline toolkit after the company banned the accounts. The September 2026 threat intelligence report says the group wrote guidance, navigation, and control software for a guided rocket and two larger missile concepts, hid the end use across split sessions, and test-fired a round that failed.

Anthropic did not name the Houthis. The file is tagged GTG-87001 and placed in northern Yemen, ground the Houthi movement has held for years. Safeguards blocked many of the prompts. They did not block the project, and they could not pull back a compiled kit that no longer needs Claude.

Three Missile Programs, One Coding Agent

The cell was running three weapons tracks at once, the company said, using ordinary Claude models rather than the locked-down Mythos and Fable line. One track was cheap and close-range. Two were not.

THE THREE PROGRAMS

  • Guided rocket: A commodity phone-class flight computer with final-phase homing guidance, the only track Anthropic says reached a live firing.
  • Ballistic missile: A multi-stage design with a stated range goal above 2,000 kilometers (about 1,240 miles).
  • R2000 set: A family of missile variants that included a hypersonic glide vehicle concept.

That range goal sits in the same class of shot the Houthi movement already claims in its public arsenal. What the cell appears to have lacked was not an airframe shop. Researchers at the Century Foundation have described Houthi missile work as still depending on Iran for guidance systems and other hard parts, while local plants turn out fuselages, fuel, and explosives. GNC software is the piece a laptop can try to mint without a new crate of seekers.

Claude Code Sat In for the Software Bench

The actors did not treat Claude like a search box for missile trivia. Anthropic wrote that they used Claude Code “in place of human software engineers” to build the GNC stack that steers and steadies a flying vehicle. They ran several instances at once and gave each a job, the way a lead splits a small team: one wrote code, one did research, and a third reviewed the first instance’s output.

WHAT CLAUDE WAS ASKED TO BUILD

  • Autopilot port: Integrate an open-source autopilot onto a phone-class flight computer.
  • Control code: Write the control and position-estimation software and tune the control settings.
  • Build pipeline: Run a firmware build pipeline and perform a flight simulation.

Final-phase homing is the ugly last stretch of a shot, when a video link can die and a moving target will not wait. Anthropic’s own red-team writeup points to automated terminal guidance on FPV drones already used in that handover, with onboard vision flying the last few hundred meters after a human locks the target. A phone on a rocket is a cruder cousin of that idea: a cheap computer, an open autopilot, and enough code to keep the nose where the seeker last pointed.

Humans stayed in the loop as leads. They handed out tasks, hid the product those tasks were for, and brought failed-shot data back to the model. The labor they were buying was the scarce part, a software bench that used to mean a specialist team and a licensed math suite.

A Live Rocket Test, Then Failure Analysis

Anthropic says it has no evidence the cell fielded a working device. It also says they did not stay on the whiteboard. They took a guided rocket to a field test in Yemen. The round appears to have failed. Within hours they were back in Claude with the telemetry, asking the same coding agent to tell them why.

We do not have evidence the actors succeeded in fielding an operational device; but they did test-fire a guided rocket. This field test appears to have failed: within hours, the actors returned to Claude to work out why it failed.

Anthropic Threat Intelligence Team, Detecting and Countering Misuse of AI, September 2026

A failed launch is not a clean miss in this file. It is proof the software left the chat window, met hardware, and generated data the operators thought was worth debugging. The company found the work in an internal hunt for suspected weapons development, then banned the linked accounts and passed threat information to public- and private-sector partners.

Anthropic posted the wider report on September 10, 2026, the same day it published new tests of whether its models can write this class of flight code in simulation.

The Ban Left an Offline Toolkit Behind

Account bans stop new prompts. They do not unscrew a phone from an airframe, and they do not delete a file that has already been compiled. Anthropic said it still has evidence the actors had built an offline simulation toolkit that does not rely on Claude, or on other engineering environments such as MATLAB. The packaging workstream in its own table is the line that matters after the login dies.

HOW CLAUDE MAPPED ONTO THE GNC CELL

Cluster What Claude was used for Most serious element
Tactical guided rocket Flight control firmware, terminal guidance, post-test telemetry diagnosis A live field test in Yemen, brought back to Claude for failure analysis within hours
Ballistic missile simulation Multi-stage, six degrees of freedom trajectory simulation Medium- and intermediate-range and hypersonic-glide variants
Optimization Reinforcement learning tuning of flight control Accelerated development of the guidance algorithms
Modeling and simulation Calibrating simulations against reference implementations Digital model of an operational weapons system to reduce dependency on physical testing
Packaging Compiling the simulation toolkit into a standalone executable A deliverable that runs and persists without Claude

That last row is a standalone executable, a deliverable Anthropic says runs and persists without the model. A ban on a cloud account is a password that stops working. It is not a raid on a bench in northern Yemen. Once the simulation pack is a file on a stick, cutting Claude is closer to cancelling a software subscription after the installer has already been copied.

The modeling row cuts in the same direction. The cell used the agent to calibrate a digital model of a working weapons system so it would need fewer live shots. Reinforcement learning tuned the flight-control code. The point of that loop is to spend less time on a range that can be watched, and more time on a laptop that cannot.

How Split Sessions Beat the Safety Filters

Anthropic’s usage policy, in force since September 15, 2025, already told users not to produce, modify, or design weapons, or to design the processes that deliver them. The Yemen prompts still got through in pieces. “Our safeguards blocked many of their requests, but not all of them,” the report says. The operators hid their goals and the products the software was meant for, and they split the work across multiple sessions so no single session revealed the full intent.

Each fragment can look like a student autopilot project, a firmware build, or a physics homework set. A phone-class flight computer is also a phone. An open-source autopilot is also a hobby stack. The harmful object was the program those fragments added up to, and the safety system scored the fragments.

That is a bad match for an agentic coding tool. Claude Code writes, builds, and reviews across a repo. A reviewer instance can look like ordinary quality control. A research instance can look like a literature pass. The lead human is the only party who holds the whole picture, and the whole picture is exactly what the split is designed to keep off any one screen.

Anthropic says it has since launched a new set of classifiers meant to catch traffic tied to high-yield explosives and weapons development. The Yemen file is one reason those classifiers exist. It is also a reminder that classifiers sit on the live service, not on the executable the service already emitted.

Six Weapons Cases Forced New Classifiers

The Yemen rocket was not a one-off curiosity in a scam report. Conventional weapons is a new chapter for the company, added after its November 2025 espionage paper. The September 10, 2026 document covers activity from December 2025 to August 2026 across seven harm areas, and it walks through six weapons cases: three in China, two in Russia, and one in Yemen. All of the misuse cases in the report used Claude Haiku, Sonnet, and Opus. None used Fable or Mythos, except one illicit distillation case outside this chapter.

THE WEAPONS CHAPTER AT A GLANCE

  • Six cases: Three in China, two in Russia, one in Yemen, split between hands-on weapons software and procurement or intelligence work.
  • Yemen’s marker: The only live field test Anthropic describes in the chapter, plus a toolkit that outlived the ban.
  • Other hardware: A Russia-based freelance cluster flashed swarm firmware to live development boards and trained vision code on combat footage.
  • Same-week tests: Frontier Red Team simulated strike tests of guidance code on a camera quadcopter, published with the report.

On the easiest of those simulated shots, a parked high-contrast vehicle, Opus 5 reached the target on 80% of launches and Sonnet 5 on 5%. Harder settings, camouflage and decoys and a car that tries to run, broke every model the team tested. The Yemen operators were not holding the most locked-down Anthropic systems. They were holding the everyday coding stack, and they still got far enough to fire a round and to pack a kit that does not need a login.

In the official post that accompanied the paper, Anthropic said it disrupted every operation in the report, used the lessons to strengthen safeguards, and shared findings with authorities and other AI companies where that was warranted. Sharing helps the next lab spot the same split-session pattern. It does not retrieve a file that has already been compiled for a bench that never needed MATLAB in the first place.

The accounts Anthropic could link to GTG-87001 are closed. The company says the toolkit those accounts compiled is a deliverable that runs and persists without Claude.

Frequently Asked Questions

Did Anthropic Say the Houthis Ran the Yemen Cell?

No. The report never uses the Houthi name and never assigns the work to Ansar Allah. It describes “a cell of threat actors based in northern Yemen” under the internal tag GTG-87001. Other weapons cases in the same chapter are tied to institutions, including a China-based researcher linked to the PLA Academy of Military Sciences and a self-identified procurement manager at a Moscow design bureau, which makes the Yemen file’s silence on a group name a choice, not a lack of a labeling system.

Which Claude Models Did the Weapons Cases Use?

The misuse cases in the September 2026 report used Claude Haiku, Sonnet, and Opus. Anthropic says none of those cases used Claude Fable or Mythos-class models, with the exception of one illicit distillation case elsewhere in the same paper. Mythos is the line the company sells with extra cyber and biology controls and vetted access; the Yemen GNC work ran on the ordinary product stack that developers already pay for.

When Did Anthropic’s Weapons Ban in the Usage Policy Take Effect?

The current Usage Policy is dated September 15, 2025, almost a year before the Yemen accounts were described as banned. Under “Do Not Develop or Design Weapons,” it bars producing, modifying, designing, or illegally acquiring weapons and other systems meant to kill, and it separately bars designing weaponization and delivery processes. The Yemen prompts were already against the rules on the day they were sent; the gap was detection of a project cut into harmless-looking sessions, not a missing clause.

Did Anthropic Share the Yemen Case With Other AI Companies?

Yes, in the form the company uses for this kind of file. The report says that where these weapons actors also worked on other platforms, Anthropic shared findings with industry counterparts so those firms could disrupt the same activity, and it passed threat reporting to public- and private-sector partners. The company account’s September 10, 2026 post said the same thing in shorter form: lessons went into Anthropic’s own safeguards, and findings went to authorities and other AI companies where that was appropriate.

Harry is the editor of Oton Technology, an independent site he owns and edits, covering the part of technology that people actually have to act on. After ten years in journalism, first reporting and then editing, he works from primary material by habit: the advisory rather than the write up of it, the filing rather than the press release, the changelog rather than the launch video. Every figure in an article carries its source and its date, and where a number comes from a vendor or an analyst model rather than a count, he says so plainly instead of letting it stand as established fact. What he leaves out is anything he could not verify himself, which on a beat full of unnamed supply chain claims removes a great deal. That standard applies across all the sections the site publishes for an international audience, from artificial intelligence and security to phones, computers, gaming, crypto and the software businesses depend on. He corrects errors in the open and labels them, because a site that hides its mistakes is asking readers to trust the rest on nothing.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending