Code Metal's WarMatrix platform — a layer that orchestrates the Pentagon's legacy war simulations — was awarded an $80M contract to prove the language models querying those simulations won't silently corrupt them.
The Pentagon handed Code Metal an $80 million contract on Thursday to build AI that can query, but never silently rewrite, the military's most trusted war simulations. The award is structured around mathematical proof, not faster code: it pays the startup to verify that new AI planners will not corrupt the decades of modeling the Department of War has already validated.
The contract, an Other Transaction Authority award handled outside the standard federal acquisition regulation, is split into two stages. Code Metal first received $17 million for an initial operating capability on WarMatrix, the platform that orchestrates existing Pentagon simulations, then the full $80 million followed, according to the company's press release distributed via Business Wire and a Fortune exclusive published the same morning.
WarMatrix is not a replacement for the Pentagon's existing simulation stack. It is a layer that lets AI planners ask questions of weapons inventories, aircraft ranges, sensor capabilities, logistics constraints, and basing options, then returns a modeled answer that draws on the legacy models underneath. The open question is the trust boundary: when a language model calls a simulation, how does the Department of War know the model is not silently rewriting the simulation, dropping a constraint, or hallucinating a sensor range that does not exist?
That is where formal verification enters. Formal verification, the mathematical proof that a piece of software behaves as specified, has a decades-long track record in chip design, avionics, and cryptography. Applying it to a stack of legacy simulations and a probabilistic AI planner is the part of the contract no one outside the company has independently validated.
The procurement itself is unusually compressed. From the start of the cycle to the full award, the timeline ran under twelve months, which is fast even for an OTA agreement, where Congress has historically let the Pentagon skip much of the standard procurement overhead. Code Metal chief executive Peter Morales called the pace "really exciting" and told Fortune he had never seen the Department of War move this fast on a contract of this size.
The compression is the falsifier. Formal verification of legacy simulation stacks is the kind of work that historically runs on multi-year timelines, not the procurement's twelve months. The two-stage award structure, $17 million first, $80 million total, looks less like standard contracting and more like the Department of War's own hedge against finding out the verification cannot keep that clock.
WarMatrix's backbone is AFSIM, the Advanced Framework for Simulation, Integration, and Modeling, a C++ environment originally developed by Boeing and now managed by the Air Force Research Laboratory. AFSIM simulates missions from undersea to space at multiple fidelity levels, and it is what the Pentagon uses to model the kinds of scenarios WarMatrix is meant to make accessible to non-specialist planners. A Taiwan scenario: what happens if the Air Force buys 200 Collaborative Combat Aircraft instead of 20 additional fighters, accounting for munitions depletion, carrier vulnerability, base denial, and coalition coordination. That is the kind of question WarMatrix is supposed to let a planner ask in natural language and get a defensible answer back.
Kelly Diaz captured the usability gap in one line to TechTimes: "I took AFSIM training for a week and I still can't do anything in it. It feels like you need a PhD." That is the human-stakes counterweight to the contract's optimism. The Pentagon is buying AI to widen the pool of people who can run these scenarios, and the underlying tool is so hard to use that the top of the planning workforce is already bottlenecked on it.
The wider Pentagon AI conversation has treated the bet skeptically. A Daily Caller News Foundation piece published the same day warned that AI "hallucinations" inside defense workflows remain an unresolved risk. The piece is commentary, not independent validation of Code Metal's approach, and it names the failure mode the formal verification work is meant to eliminate.
The $17 million tranche is the test of whether formal verification can keep the twelve-month clock. If it can, the full $80 million converts a research prototype into a planning tool. If it cannot, the Department of War has paid for a precise lesson in what trust in legacy simulation is actually worth.