Claude Code Was Used to Develop Missile Guidance Software: Anthropic Reveals a Real-World AI Weapons Program

Artificial intelligence is no longer merely helping people write emails, generate images or debug websites.

According to a new threat intelligence report from Anthropic, a group operating in northern Yemen used Claude — and specifically Claude Code — as part of an engineering workflow to develop software for guided rockets, ballistic missiles and a proposed hypersonic glide vehicle.

And this wasn’t simply a user asking a chatbot theoretical questions about missile technology.

Anthropic says the actors effectively organized multiple Claude instances like a small engineering team, assigning different AI agents to write code, perform research and review each other’s work.

The most remarkable detail may be what happened after an actual test.

According to Anthropic, the group test-fired a guided rocket in Yemen. The test apparently failed.

Within hours, the operators were back using Claude to analyze what had gone wrong.

That distinction matters. This is not evidence that an AI autonomously designed and built an operational hypersonic missile. Anthropic explicitly says it has no evidence that the actors successfully fielded an operational weapon.

But it may be something almost as significant: evidence that frontier AI can reduce one of the traditional barriers to sophisticated weapons development — access to highly specialized engineering expertise.

Anthropic’s September 2026 Threat Intelligence Report

On September 10, Anthropic published its latest report on malicious uses of Claude.

The investigation covers activity detected between December 2025 and August 2026 and includes cyber operations, surveillance, influence campaigns, scams and fraud, biological misuse, conventional weapons development and attempts to illicitly extract Claude’s capabilities.

The conventional-weapons section is particularly notable.

Anthropic says its investigators identified six cases involving actors in China, Russia and Yemen using Claude in connection with weapons development, procurement or intelligence gathering.

These included missile systems, armed drones, targeting software and other military technologies.

But one case stands out.

Anthropic assigned it the identifier GTG-87001.

It describes what the company calls:

a Yemen-based guided weapons engineering cell.

Anthropic does not publicly identify the organization behind the operation.

However, both Reuters and AP point out that the activity occurred in northern Yemen, territory controlled by the Iran-aligned Houthi movement. That makes the connection an obvious possibility, but it should not be presented as something Anthropic definitively proved.

Three Weapons Programs Running at the Same Time

According to Anthropic’s investigation, the cell was working on three separate weapons programs.

The first was a tactical guided rocket using commercially available computing hardware and terminal guidance.

The second was considerably more ambitious: a multi-stage ballistic missile with a stated range target exceeding 2,000 kilometers.

The third was a family of missiles referred to internally as “R2000.”

According to Anthropic, one proposed R2000 variant incorporated a hypersonic glide vehicle.

Again, an important distinction is necessary.

Anthropic is describing the development programs and stated objectives of the actors, not confirming that they successfully manufactured operational versions of all three systems.

That difference gets lost surprisingly quickly in social-media summaries of the report.

Claude Code Became Part of the Engineering Team

This is where the story becomes much more interesting from a technological perspective.

Missiles and guided weapons require Guidance, Navigation and Control — GNC — software.

Very simply, these systems must continuously determine where a vehicle is, where it needs to go and how its control surfaces or other actuators should respond in order to keep it stable and on course.

Historically, this is highly specialized engineering work.

Anthropic says the Yemen-based actors used Claude Code in place of human software engineers for portions of that process.

The AI was reportedly used to assist with tasks including software development, position estimation, control-system tuning, firmware builds and flight simulation.

Anthropic says the actors also worked on integrating an open-source autopilot system with commodity phone-class computing hardware.

The important point is not the particular hardware.

It is the economics.

A group that previously might have needed to recruit or train several specialists could potentially use frontier AI to perform portions of that work.

They Used Multiple Claude Instances Like Employees

Perhaps the most fascinating detail in Anthropic’s report is how the group organized its AI workflow.

They didn’t rely on a single Claude conversation.

Instead, Anthropic says the operators ran several Claude instances simultaneously and assigned them different roles.

One worked on code.

Another conducted research.

A third reviewed the code produced by the first.

In other words, the human operators were effectively acting as engineering managers while AI systems performed different parts of the development process.

That is important far beyond this particular case.

We have spent years discussing whether AI will replace individual programmers.

The more consequential possibility may be different:

one human operator could increasingly coordinate multiple specialized AI agents that collectively perform work previously requiring an entire technical team.

And that applies to legitimate companies just as much as it applies to malicious actors.

But Claude’s Safety Systems Did Block Requests

There’s another important nuance.

Claude did not simply comply with everything.

Anthropic says its safeguards blocked many of the group’s requests.

The problem is that they didn’t block all of them.

According to the company, the actors attempted to evade safeguards by disguising the ultimate purpose of their requests and splitting the work across multiple conversations.

No individual conversation necessarily revealed the entire weapons program.

This illustrates a difficult safety problem for AI providers.

A request to debug control software, optimize an algorithm or analyze telemetry can be perfectly legitimate in isolation.

The dangerous context may only become apparent when dozens or hundreds of interactions are connected together.

Then Came a Real Rocket Test

The operation eventually moved beyond simulations.

Anthropic says the actors conducted a live field test of a guided rocket.

The company does not claim the test succeeded.

Quite the opposite.

Anthropic believes it failed because the operators returned to Claude within hours of the launch and asked for assistance diagnosing the failure.

That detail dramatically changes the significance of the case.

There is a considerable difference between asking an AI hypothetical questions about missile design and incorporating AI-generated engineering work into a system that is eventually taken into the field.

Still, Anthropic remains cautious:

there is no evidence that the group successfully deployed an operational weapon created through this process.

The Offline Toolkit May Be the Bigger Problem

Anthropic eventually detected the operation and banned the accounts associated with it.

The company also shared intelligence with relevant public- and private-sector partners.

But banning an account doesn’t necessarily eliminate the knowledge already produced.

Anthropic says investigators found evidence that the actors had created an offline simulation toolkit that no longer depended on Claude or conventional engineering environments such as MATLAB.

This reveals one of the fundamental asymmetries of AI security.

A provider can terminate access to its model.

It cannot necessarily revoke software, knowledge, simulations or workflows that users have already extracted from it.

Anthropic Is Now Testing AI on Weapons Tasks

The company published another piece of research alongside the threat report that makes the story even more significant.

Anthropic’s Frontier Red Team has developed evaluations specifically designed to determine how capable AI models are at tactical intelligence and conventional weapons engineering tasks.

The company tested scenarios such as programming simulated drones to navigate without reliable GPS, dealing with sensor noise and electronic interference, and improving guidance and targeting systems.

The conclusion is uncomfortable.

According to Anthropic, frontier models are beginning to perform some tasks that historically required scarce and highly trained specialists.

The models cannot manufacture missile bodies, produce rocket motors or magically obtain restricted components.

Physical manufacturing, materials, testing infrastructure and supply chains remain substantial barriers.

But software expertise is increasingly becoming less scarce.

And modern weapons contain enormous amounts of software.

This Is the Real AI Proliferation Problem

The conventional image of weapons proliferation involves transferring physical technology.

Blueprints.

Components.

Machine tools.

Missile parts.

Specialized engineers.

AI changes one part of that equation.

It potentially makes expertise itself reproducible at near-zero marginal cost.

A highly capable model doesn’t need to physically travel from one country to another.

It doesn’t need to be recruited.

It doesn’t need years of training.

And multiple copies can work simultaneously.

That doesn’t mean anyone with a chatbot can suddenly build an intercontinental ballistic missile.

That would be an absurd conclusion.

But Anthropic’s report suggests something more realistic and perhaps more important:

AI can help technically competent groups move faster, automate portions of engineering work and compensate for shortages of specialized personnel.

Anthropic’s own researchers describe this as potentially enabling lower-resource groups while amplifying the capabilities of sophisticated state actors.

And It Wasn’t Just Missiles

The Yemen case is only one part of a much larger report.

Anthropic says Claude was also abused or targeted in operations involving cyber espionage, surveillance, political influence operations, biological research, fraud and attempts to extract Claude’s capabilities for competing AI systems.

Reuters reports that the company uncovered, among other cases, a suspected Russia-linked cyberespionage campaign targeting Ukraine.

Another operation used Claude to automate surveillance-style reports involving Uyghurs, Tibetans, Taiwanese political figures, activists and foreign media.

Anthropic also described a China-based operation involving more than 20 AI-powered dating applications, 4,700 artificial personas and at least 25,000 users.

The September report therefore represents something larger than another collection of chatbot abuse stories.

It documents a transition toward AI being embedded inside operational workflows.

The Most Important Detail Isn’t the Missile

The headline will inevitably be:

“Claude was used to develop missiles.”

But focusing exclusively on missiles risks missing the technological shift underneath.

The important part is the architecture of the workflow.

Human operators defined objectives.

AI systems researched solutions.

Another AI wrote software.

Another reviewed it.

Simulations evaluated the results.

Humans connected the software to physical hardware.

A real-world test generated new data.

And that data was fed back into the AI-assisted development loop.

That’s starting to look less like “asking a chatbot questions” and more like AI-accelerated engineering.

The same model can obviously be enormously productive when developing medical devices, industrial automation, satellites, robotics or clean-energy technology.

But capability is capability.

The engineering knowledge that makes a drone more reliable can also make a weapon more reliable.

And that’s precisely why the Anthropic report matters.

AI Has Crossed Another Line

For years, discussions about dangerous artificial intelligence tended to focus on hypothetical future systems: AGI, autonomous weapons or superintelligence.

This case is considerably less futuristic.

The AI didn’t autonomously decide to build a missile.

It didn’t manufacture one.

It didn’t independently launch anything.

Humans remained firmly in control of the program.

But according to Anthropic’s own investigation, those humans were able to use a commercially available frontier AI system to perform portions of specialized weapons-engineering work, organize AI instances like a small technical team and support an iterative development process that eventually reached a physical field test.

That is a much more immediate problem.

The question is therefore no longer simply whether AI could someday lower the barrier to sophisticated weapons development.

Anthropic’s September 2026 report suggests that, at least for some parts of the engineering process, it already has.

Visited 2 times, 2 visit(s) today
share this recipe:
Facebook
X
WhatsApp
Telegram
Email
Reddit