Yemeni militants turned to artificial intelligence to accelerate the development of advanced weapons, using Anthropic’s Claude AI model as a substitute for human software engineers working on missile guidance and control systems, according to a new threat report released by the US-based AI company.
Anthropic’s September 2026 report says its threat intelligence team identified a weapons engineering cell in northern Yemen that was simultaneously pursuing three weapons programmes. The company did not publicly identify the group behind the operation, although the description of the cell and the territory in which it operated is consistent with areas controlled by Yemen’s Iran-backed Houthi movement.
The discovery offers a striking example of how increasingly capable AI systems can be exploited by actors seeking to shorten the time and expertise required to develop sophisticated military technology.
According to Anthropic, the Yemen-based cell was working on a guided rocket, a multi-stage ballistic missile with a stated range exceeding 2,000 km, and a multi-variant missile programme known as the “R2000” set. The latter included a proposed hypersonic glide vehicle variant.
Claude used as a virtual engineering team
The most significant aspect of the operation was the way the militants used Claude Code.
Rather than merely asking the AI model general questions about missiles, the actors used it to produce and refine the guidance, navigation and control (GNC) software responsible for steering and stabilising a flying vehicle.
Anthropic said the operators integrated an open-source autopilot with a commodity, phone-class flight computer. Claude was then used for software development tasks, including control and position-estimation software, control tuning, firmware builds and flight simulations.
The operators also divided the work among several Claude instances. One instance was tasked with writing code, another with conducting research, while a third reviewed the code produced by the first.
In effect, Anthropic said, the AI was being used much like a small engineering team rather than simply functioning as a conventional chatbot.
The company’s broader assessment is particularly concerning: across the weapons cases it investigated, threat actors generally used Claude to build or refine software for hardware and firmware that they already possessed expertise in and access to.
Guided rocket was actually test-fired
The operation went beyond theoretical research.
Anthropic said the Yemen-based cell conducted a live test of a guided rocket. The test appears to have failed, however. Within hours of the test, the operators returned to Claude and used it to investigate what had gone wrong.
The company stressed that it had no evidence that the actors successfully fielded an operational weapon. Nevertheless, the test demonstrated that the AI-assisted development effort had progressed beyond conceptual discussions and software experimentation.
Anthropic’s assessment of the weapons-development workflow is illustrated on page 114 of its report. The company’s diagram maps the cell’s work through requirements, system architecture, detailed design, implementation, integration, testing and validation. It identifies the tactical guided rocket’s flight-control firmware, terminal guidance and post-test telemetry diagnosis as areas in which Claude was used, with the live Yemeni field test and subsequent failure analysis representing the most serious element.
Militants tried to bypass AI safeguards
The case also demonstrates the difficulty of preventing sophisticated users from circumventing AI safety controls.
Anthropic said its safeguards blocked many, but not all, of the Yemen cell’s requests. The operators attempted to disguise their intentions by concealing what their software would ultimately be used for and dividing their work across multiple conversations.
By splitting the project into separate sessions, no individual conversation necessarily exposed the full purpose of the programme to the safety systems.
That tactic is significant because it highlights a fundamental challenge for AI companies: individual requests may appear benign when viewed in isolation, while the aggregate activity can reveal an effort to construct a sophisticated weapons system.
Anthropic ultimately banned the accounts it associated with the operation and shared intelligence with public- and private-sector partners. But the company acknowledged that shutting down Claude access did not completely terminate the programme.
Investigators found evidence that the actors had already constructed an offline simulation toolkit that could operate without Claude or other conventional engineering environments such as MATLAB.
AI is lowering the barrier to weapons development
The Yemen case forms part of a much broader investigation by Anthropic into the misuse of Claude for conventional weapons.
The company said it had identified and disrupted six weapons-related cases over the past year: three in China, two in Russia and one in Yemen. These involved weapons development, targeting systems, procurement and intelligence gathering.
Anthropic said the cases included the use of Claude for software associated with missiles, armed drones, bombs and other munitions, as well as targeting and control systems.
The company has consequently introduced additional classifiers intended to detect and block activity associated with high-yield explosives and weapons development.
The Yemen investigation therefore points to a new dimension of the AI security problem. The immediate danger is not necessarily that an AI model independently designs a complete weapon. Rather, increasingly capable models can provide specialised technical assistance to people who already possess the hardware, expertise and intent, potentially compressing development timelines and allowing relatively small teams to undertake work that would traditionally require larger engineering staffs.
And in this case, Anthropic’s own investigation shows that the technology was not being used merely to write documents about weapons. It was being used to develop software intended to make a flying weapon steer, stabilise and navigate itself.
That distinction makes the Yemen case one of the clearest warnings yet about the risks posed when frontier AI capabilities reach actors pursuing advanced military programmes.

