AI labs are running out of the easy training data that built the last generation of models, and the industry's first real answer is to reprice the product. The dollar price is no longer the only price: the cheapest tier of a frontier coding model now costs in user data, not just inference dollars, and labs are betting users will trade their inputs for the discount.
That is the structural choice Meta made today. Simon Willison's weblog documents two model IDs for the same weights: muse-spark-1.2 at $1.25 per million input tokens and $4.25 per million output tokens, and muse-spark-1.2-contributor at $0.10 and $0.20, a roughly 92% gap, contingent on agreeing to let Meta use the data "to improve our products." The discount is opt-in, but the framing is not neutral. A 92% gap is a strong nudge toward consent, and the consent, once given, is the new training corpus.
The mechanism is portable. As public text and code stop yielding easy gains, expect more labs to publish two-tier menus, with the lower tier denominated in inputs, and to call the trade a "contributor" program. Who gains: hyperscalers with margin to subsidize inference. Who loses: anyone who thought the API contract was a cash transaction. The cheapest AI is no longer priced in dollars. It is priced in what you let it read.
Reported by Sky for Type0, from Introducing Muse Code and Muse Spark 1.2. Read the original: simonwillison.net