CUDA built its moat on two markets it treated as one. On training, the assumption holds: the 20-year stack of compilers, libraries, and frameworks runs the largest models. On inference, the math is rearranging, because the customer is no longer buying a chip or a stack. The customer is buying an answer, and per-answer pricing changes what counts as a moat.
Two signals land in qbitai's reporting this week. Marshall Choy told qbitai CUDA is no longer the deciding factor on inference, calling the workload a competition of open-source software. qbitai also documents DeepSeek's open-sourced TileKernels, a GPU kernel library that reduces hand-written CUDA, and Infinity, a year-old startup using an AI coding agent to scaffold inference kernels for d-Matrix in roughly a working day, claiming a 3.7x speedup over cuDNN. The qbitai headline reads that AI chiseled open a 20-year CUDA moat in 10 hours. The honest reading is one day of porting, scoped to inference, for one customer.
The mechanism has three parts. Per-answer pricing lets inference buyers swap silicon when software lets them. AI coding agents collapsed kernel porting from months to a day. The substitute is narrow by design: it sits at the kernel layer, not the 20-year stack that still owns training. The pattern is not "CUDA is finished." It is that lock-in can hold the top of the funnel while losing the bottom, when the bottom is priced per unit and porting time falls below a billing cycle.
The counterargument is the one d-Matrix, Rebellions, and Infinity quietly accept: a 10-hour scaffold and a $15M raise are not an ecosystem, and CUDA still owns training, libraries, profilers, and framework code. What it is losing is the assumption that inference has to run on its terms.
Reported by Sky for Type0, from 老黄垒20年的CUDA护城河,AI刚刚用10小时凿开了. Read the original: qbitai.com