Today's AI capability per question comes from two expensive places: a bigger frozen model and a longer stream of candidate answers. Both get costly fast, and both run into the same wall: you pay for compute whether the question is easy or hard.
An arXiv preprint argues there is a third lever. A small generator network can read an incoming query and produce a tailored parameter patch for it, the same way a tailor measures a customer before cutting cloth. Produce not a single patch but a distribution of patches, and the system can spend extra compute by sampling inside its own settings, not by guessing more words. The paper's framing lands it cleanly: weight-update sampling now sits on the table alongside token sampling as a way to spend test-time compute.
That reshapes what the inference budget even is. The choice is no longer just a bigger brain or more output tokens; it now includes sampling adapted versions of the model itself. The honest caveats travel with the result: this is a preprint, the gains are measured against other adaptation baselines rather than shipped products, and the cross-query transfer evidence is a single paper's, not an established mechanism. The lever exists. Whether production systems reach for it is the next bet to watch.
Reported by Sky for Type0, from Learning to Predict Distributions over Weight Updates for Test-Time Adaptation. Read the original: tldr.takara.ai