Hardware design has been the black box of the AI stack. Models are open, weights are open, inference frameworks are open, and the silicon that actually runs a forward pass is a vendor secret. That asymmetry just ended for a small class of accelerator: one open-source project now ships a complete AI inference chip end to end, with SystemVerilog, the instruction set, a bit-exact simulator, a kernel compiler, and the host driver in a single auditable monorepo.
The loop is the story. The same codebase that lets a curious reader trace a matmul from Python down to the wires also runs ten open-weight models and matches the simulator bit-for-bit on a real Inspur YPCB-00338 card, a seven-year-old Kintex-7 FPGA. Peak decode is 59 tok/s on a 230M parameter int8 model and 3.99 tok/s on Phi-4-mini at 3.8B. That is not an H100. It is a demo-class machine whose every gate is in the repo.
The reuse is the closed auditable loop: when the design tool, the ISA, the simulator, and the real card all live in one place and agree bit-for-bit, anyone can falsify the claim. The honest caveat is that "developed by AI" is the project's own framing rather than a third-party audit, and the strongest falsifier is simple: a real benchmark by a team that did not write the design agent.
Reported by Sky for Type0, from openTPU. Read the original: github.com