General purpose AI models grab the headlines, but the real enterprise value is moving to small focused models, and the data companies beneath them are growing faster than the labs.
Picture a long-running procurement agent, the kind frontier models still struggle to operate for months on end, sitting on a narrow model trained for one workflow. The data company underneath it is now worth more than the model it serves, and the rest of the AI stack is starting to look the same way.
The general-purpose "frontier" models built by OpenAI, Anthropic, and a handful of well-funded peers dominate the news cycle. Underneath them, a second tier of specialized models, trained for a single procurement workflow, a single underwriting task, a single code review, is starting to carry the actual enterprise spend. The infrastructure is bifurcating, and the data layer sitting between the two tiers is the part of the stack scaling fastest.
Mercor, an AI training and expert-data marketplace that supplies frontier labs with the human-labeled data and RL environments used to train their models, crossed $2 billion in gross annualized revenue in June, The Information reported this month. The figure was first put on the record by Mercor chief product officer Osvald Nitski on 20VC's 20 Product podcast (Harry Stebbings' long-running founder-interview show), and the briefing independently confirms it. Mercor doubled from $1 billion to $2 billion in four months. The company last raised at a $10 billion valuation; it is reportedly in discussions for a round at $20 billion.
Mercor's revenue is heavily concentrated on a handful of frontier-model customers. The official episode notes put the question bluntly on the agenda: "Can Mercor escape its dependence on a handful of frontier-model customers?" He did not dispute the framing on the transcript.
That shape is unusual for a data company and familiar for a frontier lab. It tells you where the demand is: a small number of well-funded model builders burning cash to win the general-purpose race, with most Fortune 500 buyers routing around them. The same episode includes the agenda question "Are enterprises still terrified of working with frontier model companies?" and the answer he offers is yes. Large enterprises, he argues, are scared to partner directly with frontier labs and prefer to channel work through a data intermediary. That routing is the wedge.
Two further claims from the same conversation, summarized in the teahose.com recap and the BigGo Finance notes, are worth holding separately.
The first: open-source models raise the floor rather than the ceiling. The argument: open weights compress the commodity end of the market and force frontier labs to keep spending on data to defend the top end. That keeps demand for expert data load-bearing, because the addressable market for "frontier" capabilities keeps expanding. The version of the claim on the transcript is blunt: "We can't spend money fast enough to service all of the demand that we have. We end every week with so much more money in the bank. I don't think there's an ROI problem right now." Treat that as a single-source, on-the-record remark, not a market fact.
The second: every company ends up with its own focused AI model. The argument is that general-purpose models plateau on long-horizon workflows, where the BigGo notes say frontier systems reach only around 50% success on tasks like a fully autonomous procurement agent running for months. Once a workflow is too long and too specific for a general model, the only path to production is a narrow one trained on expert data. That is the layer Mercor sells into.
The counterargument, which the 20VC interview does not address, is that frontier labs are not standing still on the specialized tier. OpenAI, Anthropic, and Google all ship vertical products and increasingly ship tools that let enterprises fine-tune on their own data. If a frontier lab can absorb the narrow-model value itself, the bifurcation collapses back into a single tier, the data intermediary's margin compresses, and "built on a frontier model" turns out to be a passing label rather than a moat. Mercor's $20B-round talks are, in part, a bet that the bifurcation holds.
The watch item for the rest of 2026 is concrete. RL environments, synthetic training grounds where models learn by trial and error, are now the fastest-growing data type Mercor sells, and the BigGo recap expects physical and robotics data to become a meaningful revenue line within three years. The Information's briefing puts the $2 billion in gross annualized revenue on the record as of June. The next data point to read is whether that revenue mix is diversifying away from frontier-lab buyers, or doubling down on them. The next time an enterprise AI product lands in your inbox, the question worth asking is which tier it competes in and who owns the data underneath.