A new paper proposes 'grafting': learn the update on an earlier saved version of the model, then scale it onto the latest version of the model, with an optional filter (a mask) on the most sensitive update directions.
A new arXiv paper from Chen Henry Wu argues that one of the more expensive habits in modern AI post-training, retraining a model on its own outputs every time you want to teach it something new, is not actually required. The recipe the paper proposes, called grafting, learns the update on an earlier checkpoint, applies it to the post-trained model, and scales the result the way model merges do. Across three continual-learning settings, the author claims grafting beats both vanilla fine-tuning and the costly retrain-on-your-own-outputs route on what is called a Pareto frontier, better at the new task without giving up more of the old one.
The term "continual learning" describes the problem of keeping a model current with new knowledge without breaking what it already knows. The term "on-policy" refers to the now-standard fix: have the model generate its own answers, treat those answers as training data, and retrain on them. That pipeline is the backbone of reinforcement-learning post-training for reasoning models, and it is expensive. Every round burns fresh inference, fresh labels, and fresh GPU time. The paper's headline claim is that this cost is optional for the continual-learning case, not foundational.
Grafting, in the paper's plain form, is three steps. First, take an earlier version of the model, ideally from before the end of pretraining, and run ordinary supervised fine-tuning on the new data against that older checkpoint. Second, take the resulting weight update and graft it onto the post-trained model. Third, scale the graft as a model merge, written in the paper as θ_merge = (1 − λ) θ_post + λ θ_post', with an optional mask on the most sensitive update directions when the new data is far from what the model already knows. The framing matters: rewriting the merge as θ_merge = θ_post + λ(θ_post' − θ_post) decouples where the update is learned from how it is applied, and the paper's task-arithmetic experiments show the donor model for the update need not be the model it is applied to.
The empirical claim, per the paper's main results, is that grafting Pareto-dominates both vanilla SFT and on-policy self-distillation (OPSD) across three settings: distilling from expert traces, self-improvement with the STaR and Pedagogical RL recipes, and injecting knowledge after a model's pretraining cutoff. On the donor-checkpoint question, the most decision-relevant knob, Figure 1 reports that for OLMo-3-7B (thinking), the best donor is even before the end of pretraining; for Qwen3-1.7B (thinking), model merging alone already improves generalization, and grafting dominates from there. The supporting finding underneath those numbers is the mechanism: SFT on new data is not broken because the signal is missing, it is broken because the update interferes with the existing model. Scaling the update down via a merge, by choosing a smaller λ, is what turns ordinary SFT into a strong baseline.
The framing against prior work is part of the story. On-policy self-distillation, as practiced in recent 2026 work from Zhao et al., Hubotter et al., and Shenfeld et al., converts off-policy data into on-policy updates by distilling from a solution-conditioned copy of the model. The paper's argument is that this can suppress exploration, leak privileged information into the teacher's reasoning, and degrade out-of-domain performance or collapse reasoning length. Grafting is offered as a way to keep the off-policy data off the policy entirely: learn the update on a cheaper donor, merge it back in, and skip the teacher loop. The practical implication, the paper argues, is avoiding the expensive on-policy sampling that is the recurring bottleneck in RL post-training pipelines.
What is and is not verified at the paper's current version matters. The preprint is a v1 from 5 October 2026 with no independent peer review or replication; the Pareto-dominance numbers are self-reported, and the strongest settings to check against an independent run are STaR/Pedagogical RL and the post-cutoff knowledge-injection case, where the paper itself flags that the experimental details need table-level verification before any quantitative claim is locked in. The savings framing is also relative, not absolute. The paper argues grafting avoids expensive on-policy sampling, but it does not yet publish a dollar or GPU-hour figure, and it would be premature to assert one. The masking step in the recipe is described as optional, and the paper's own framing treats it as a marginal add-on rather than a load-bearing part of the result.
The falsifier worth watching is concrete: whether the Pareto-dominance holds on a third, frontier-scale RL-trained model under independent evaluation, with the donor-checkpoint curve and the scaling parameter λ both reported. If it does, the cheaper recipe stops being a single-paper result and becomes a reusable post-training pattern. If it does not, the technique still has a clean story: SFT plus a merge is a strong baseline, and the on-policy detour is not the only way to keep a model current.