A chatbot's false claim that a Chinese ship carried nuclear parts put armed US troops minutes from a Middle East boarding during the war with Iran.
Armed US service members were minutes from boarding a Chinese-flagged vessel in the Middle East this spring, military aircraft already in the air, when a senior official asked a routine follow-up question: where did the intelligence come from? The answer was an AI chatbot that had fused open-source intelligence with secret signals intelligence. That answer was the first sign the cargo manifest the boarding party was about to intercept did not exist.
According to a CNN exclusive published Thursday, the report originated with an analyst at US Special Operations Command Pacific in Hawaii, who queried a chatbot about a ship's manifest. The bot "fused together open-source intelligence with secret signals intelligence in government holdings," then inaccurately identified what the ship was carrying. The product, a claim that the vessel was transporting components of a Chinese nuclear weapons program, circulated across the US military "in the midst of the war with Iran." The response was immediate: armed personnel prepared to board, military planes went airborne, and the United States swung toward what one source called "almost a war."
Four sources familiar with the episode told CNN the underlying claim was "entirely false." The network could not determine what the ship was actually carrying. Two of the sources, plus a former senior US official familiar with the AI systems used by military and intelligence analysts, described the chatbot's output as the kind of confident, fluent fabrication these systems are known for: a "hallucination," in the AI trade, where a large language model produces text that sounds authoritative but is not grounded in its source material.
It is still unclear whether the chatbot was a commercial product or a US government system. Per the former official, the internal tools are "mostly just copies of the commercial stuff wearing lipstick." Ars Technica independently framed the episode as "one of the more potent and consequential instances of a hallucinating AI ruining the reliability of a professional report."
In January 2026, the War Department launched its AI Acceleration Strategy, declaring the military an "AI-first warfighting force across all domains." The strategy names seven "Pace-Setting Projects," including "Agent Network" for AI-enabled battle management and kill-chain execution, "Open Arsenal" for turning intelligence into weapons "in hours," and "GenAI.mil" for putting frontier models at the defense information system's Impact Level 5, a security tier reserved for controlled unclassified information. Google's Gemini and xAI's Grok are named as deployed models. Secretary of War Pete Hegseth and Under Secretary Emil Michael framed the rollout as a generational shift.
The distance between a chatbot's output and armed US service members in motion against a great-power adversary turned out to be shorter than the public posture suggests. The incident raises a question the strategy does not answer: what verification architecture sits between a large language model and a boarding party? Defenders of the AI-first approach argue that human analysts remain in the loop and that any single product can be checked. In this case, the human in the loop was the senior official who asked the second question. The next incident will arrive faster than a procurement cycle can fix it.
A useful version would require an explicit provenance check on every chatbot output before it can drive a kinetic action: a flag showing which model produced the answer, which inputs it drew on, and which signals intelligence, if any, was human-verified at the source rather than auto-fused. A second layer would separate a chatbot's open-source analysis from any product carrying classified inputs, so a hallucination in one cannot contaminate the other. None of this requires slowing the AI-first strategy. It requires treating the chatbot as an untrusted source, the way any intelligence officer would treat a single tip.
The DoD has not publicly acknowledged the near-boarding. A Hacker News discussion of the CNN story treated it as a near-miss rather than a failure of AI itself, the same framing several defense-AI operators have reached. Both readings are correct. The model did what models do. The institution that trusted the model without a verification layer is the part that nearly started a war.