qlabs' Dust perturbs every internal node in the network in parallel, scoring thousands of variants per pass, matching backpropagation's accuracy on a 1B token test.
A small research group says it can train a transformer without backpropagation, the gradient-based algorithm that has powered nearly every major AI system since the deep-learning era began. Their method perturbs each node in the model, scores thousands of variants in a single forward pass, and walks the network toward better answers. The catch, which the paper itself flags, is that the recipe costs substantially more compute than the standard one.
The preprint comes from qlabs, a small research outfit, and describes a method called Dust. The paper's central claim is that Dust is the first "zeroth-order" method, a method that does not compute exact gradients through the network, to be competitive with backpropagation at pretraining transformer language models.
Dust treats each token in a training batch as a virtual member of a population. The model perturbs its internal activations (a "node perturbation" approach) at every token independently. Because every perturbed variant runs through the same forward pass, the model scores all of them in parallel and uses the spread of outputs to estimate which way to nudge the weights. The training run no longer needs to backpropagate errors layer by layer.
As the population grows, Dust's gradient estimate converges on backpropagation's exact signal. The authors report that the two align "well" at every scale they tested, up to 1B tokens, and that alignment improves with population size. In their headline scaling experiment, a 243M-parameter Dust model matches backpropagation, and at most population sizes actually outperforms a 120-times-smaller model trained the same way. The intuition: more parameters give the perturbation more room to act, so larger models become more population-efficient, not less, the opposite of what one might naively expect from a search-based method.
The result is a fresh data point for the bitter lesson, the recurring argument in AI research that general methods which scale with compute tend to win over hand-engineered alternatives. In its strongest form, the lesson predicts that the field's hand-engineered training recipes will eventually be replaced by cheaper generic search. Dust does not yet prove the point, but it gives the argument a new empirical foothold: a search-based recipe that gets more efficient as the model grows, not less.
The paper does not bury the cost. The authors write plainly that Dust only approaches backpropagation with "substantially more compute." They do not give a single ratio against backpropagation itself; the 10³ to 10⁴ efficiency gain they report is measured against a transformer implementation of EGGROLL, a prior state-of-the-art evolutionary method. That comparison is the sharpest number in the paper, and the evolutionary-search tradition is the obvious backdrop: those methods have repeatedly looked promising at small scale and then lost to gradient-based training as the model grew. The bitter-lesson argument, applied to Dust, says the comparison flips at the scales the authors have actually run, but the immediate arithmetic for a frontier-scale pretraining run still favors the standard recipe.
Two further limits deserve to be named. First, the alignment tests stop at 1B tokens, which is tiny next to modern pretraining runs that count in the trillions. The paper is a preprint, and the result is a single lab's experiment. Second, the population-efficiency inversion (the finding that bigger models get more efficient rather than less) has not yet been tested at frontier scale, and outside replication is the obvious next step.
The release includes the code and an appendix with the scaling and ablations. The next test of the claim is whether an outside lab can reproduce the population-efficiency inversion at 10B tokens and beyond.