Allen AI released the 8 billion parameter model under Apache 2.0, alongside its training data and evaluation harness, so labs can run a fast cited report drafter on their own machines.
Allen AI has open-sourced AstaBrief, an 8-billion-parameter model that drafts cited scientific reports from a research question and a stack of retrieved papers. The most interesting part is what the institute had to learn about its own training signal to make a small model work at all.
The model, released under Apache 2.0, is built on Qwen3-8B and trained to do one job: read a research question plus a set of retrieved papers, then write a report in which every meaningful claim carries a citation. The training data and evaluation harness are published alongside it under their own separate licenses. In Ai2's own measurements, the full Asta pipeline now averages 51.1 seconds per report in the new fast mode, compared with 178.5 seconds in the Claude-powered mode it joins. That is roughly a 3.5x speedup, with the smaller model doing the drafting.
That number is also a research claim, not just a product claim. To make an 8-billion-parameter model write citations that hold up, Ai2 ran ablations on what to filter for during training. The strongest single signal was citation density: the share of sentences in a training report that carried at least one reference. More elaborate combinations of quality filters did not meaningfully beat that one metric. The team also chose to write the whole report in a single pass, rather than generating it section by section, and Ai2's measurements suggest that choice did not hurt quality. It is part of the reason a small model can finish a report in the time the previous, larger pipeline took to start one.
The training recipe is the part other labs can copy. Ai2 filtered 90,000 real Asta queries down to 47,000 supervised fine-tuning examples, then ran 6,000 direct-preference-optimization pairs on top, with no reinforcement learning. The SFT mix and DPO mix are published alongside the model, so other teams can audit, reuse, or rerun the recipe. The data and prompts carry their own licenses, separate from the model's Apache 2.0 terms, and are worth checking before redistribution.
The evaluation harness is also public. AstaBrief is benchmarked on SQA-CS2, Ai2's own scientific question-answering set, with the model card reporting 55% win rate over the Claude-powered Asta ScholarQA pipeline on the dev split and 72% on the test split under LLM-as-judge scoring. Ai2 reports 95% agreement between that judge and human preferences in a small meta-evaluation, but both the meta-eval and the headline win rates are LLM-judged rather than human-ranked, and should be read as a directional read on quality, not a settled ranking.
Ai2 is candid about what the small model does not settle. In its own 14-question human study, DR Tulu, a larger Ai2 research model, still wins on overall preference, with AstaBrief preferred on citation accuracy by two of three researchers. The evaluation window is also 2025: Ai2 explicitly notes it has not rerun the full comparison against today's frontier models, so the win rates are not a current state-of-the-art claim. A reader who wants a "best open-weights report writer right now" ranking will have to wait for an independent rerun.
The case for downloading the artifact is not that it is the best at any one benchmark. It is that a lab can run the model, the data, and the eval pipeline on its own infrastructure, which matters when the research question touches unpublished work, draft papers, or sensitive clinical or industrial data that should not leave the building. Ai2's own 374-user early sample is small and directional, but it shows the same pattern: 29.1% of users tried the new fast mode on two or more days, averaging 3.67 threads each, with 84.2% positive feedback against 85.2% for the slower mode. A faster, local option gets used even when a larger hosted model is also available.
AstaBrief is licensed Apache 2.0. The training mixes and prompts carry their own terms, and the eval code is open under its own repository license, so anyone planning to redistribute the pipeline should read the licenses before publishing. The next thing worth watching is whether an independent group reruns the comparison against the current frontier, and whether the citation-density finding holds up as a transferable signal for other small, citation-bearing writing tasks.