Micron's 84.9% gross margin beat Nvidia's 74.9% this quarter, exposing the layer where AI's rent actually sits while OpenAI loses $1.22 per dollar of revenue.
NVIDIA's most recent quarter reads like a record: $81.6B in revenue, up 85% year over year, with a GAAP gross margin of 74.9%. It is not the best margin in AI hardware. Micron, the U.S. memory maker, posted an 84.9% gross margin on its most recent quarter, a higher gross margin than NVIDIA's, and almost all of it came from the stacked DRAM called high-bandwidth memory, or HBM, that feeds modern AI accelerators. The two numbers, read together, are the story: the AI rent is not sitting with the chip designer or with the model. It is sitting with the physical substrate the models run on.
A token, in industry shorthand, is the unit of text or audio an AI model processes and charges per request. The total volume of those calls is exploding. ByteDance's Doubao has reportedly hit 180 trillion tokens per day, and China's National Data Bureau reported nationwide daily token calls above 140 trillion in March 2026, a curve the original leiphone analysis describes as more than 1,000x in roughly two years. That is the kind of growth investors have been told to celebrate. The corresponding earnings tell a different story about who is capturing it.
OpenAI, the private lab behind ChatGPT, burned through roughly $9.3B in the first quarter of 2026, an adjusted operating margin of approximately −122%, or about $1.22 lost for every dollar of revenue. Those numbers come from third-party reconstructions rather than audited disclosure (OpenAI is private), so they should be read as analyst estimates, but they are not isolated. The pattern is that demand is outrunning the physical capacity to serve it, and the margin is concentrating at whichever layer of the supply chain is the slowest to expand.
That layer is two physical things: HBM and the leading-edge wafer. HBM is a specialty stack of DRAM, sitting next to the GPU on the same package, that feeds the model the data it needs to keep computing. It is concentrated in three to four firms: Micron, SK hynix, Samsung, and a long tail. The leading-edge GPU node, the actual silicon die, is effectively single-sourced at TSMC. Both have lead times measured in years. Micron's own executive commentary, echoed by SK hynix and Samsung on recent earnings calls, has guided that HBM supply tightness extends at least into 2027 and potentially 2028. The build cycle from order to ramp runs two to three years; you cannot will capacity into being with a capital allocation memo.
If the bottleneck is slow to expand and demand is growing 1,000x, the rents accrue to whoever sits on the bottleneck. NVIDIA's gross margin captures the GPU side. Micron's higher gross margin captures the HBM side. The model layer, which is the most visible and the most competed, captures the least. That is what the structure produces, and it is the same mechanism that puts the two leading hardware names at peak margins while the headline AI lab loses more than a dollar per dollar earned.
The interesting question is what would have to change for the rent to migrate. Three dials are worth watching.
Micron, SK hynix, and Samsung are all adding capacity, and all three have told the market that even the additions they have already announced do not loosen supply through 2027. If 2028 guidance begins to soften, that is the first signal that rents at the memory layer will start to compress.
The threshold at which domestic Chinese chips can substitute at commercial scale. Cloud天励飞 vice president Luo Yi told leiphone that domestic chips cover most workloads below 100 tokens per second per user, and that the gap with NVIDIA's upcoming Rubin generation "is still obvious" in closed-loop commercial workloads where the model has to be served at scale. NVIDIA's generational cadence runs roughly annual; the Chinese chipmakers are chasing a moving reference point. The 100 TPS-per-user line is the line at which the procurement conversation changes from "we will use whatever we can get" to "we can choose."
The TSMC leading-edge node. As long as the most advanced GPU silicon is effectively single-sourced, no amount of design-side innovation in the United States, China, or anywhere else changes the physical constraint. Watch the company's capex commentary and the timing of its Arizona and Kumamoto ramps; the AI rent cannot move downstream of TSMC faster than TSMC can ship wafers.
The Chinese industry's own diagnosis lines up with that frame. Three structural limits are named in the source's reporting: a performance generation gap, a commercial track-record threshold in enterprise procurement, and an ecosystem barrier across the MaaS and Agent layers, the model-as-a-service plumbing and the autonomous agent stacks that sit on top of it. Each is real, and each is a function of how slowly the physical supply can move, not how quickly the model layer can iterate.
Congji Technology co-founder Huang Li'ang summarized the current moment bluntly to leiphone: "Almost everyone is losing money, and all the profit is flowing to NVIDIA, or to upstream hardware." Maifu CFO Ma Jin called the hardware margin "extortion-level" given the supply-demand imbalance. IO Capital's Zhao Zhanxiang added the timing risk: during an acute GPU shortage, customers will try any new compute architecture; the window closes once supply normalizes.
The window is the watch. Token demand is not going to fall. The question is when, not whether, the HBM supply curve bends. When it does, the rent starts moving toward whoever sits at the next slow layer of the stack: the inference platform, the application, the agent. Until then, the pickaxes are the most profitable thing in the field.