A paired study of two small classifiers that decide which model, tool, or action an LLM agent uses next, and the three errors the authors found in their own analysis.
A research team published a paired evaluation of two small, cheap classifiers that score the per-decision work inside LLM agent stacks, then turned the same rigor on their own pipeline and corrected the most-cited deployment number by roughly a factor of five.
The paper, a single-author preprint by Jiawei Li, evaluates what the author calls System-1 classifiers: small models that score one decision in a single forward pass instead of paying for a full LLM call. The name borrows from Kahneman's fast/automatic mode. Eleven such decision points show up in any LLM agent stack: which model to use on a given turn, which tool to call, whether a retrieved passage is relevant, whether a prompt looks like an injection, whether personally identifiable information is about to leave the system, and several more. The cost of running an agent is dominated by thousands of these small per-decision calls, so a fast, accurate classifier is the difference between an agent that is cheap to operate and one that burns money on every turn.
The two classifiers in the study are Laya, an open-weight model that can be run locally, and Jev, a hosted API. On the headline accuracy number, Jev is meaningfully more accurate than Laya on 9 of the 11 decision points, with paired deltas ranging from +10.8 to +46.0 percentage points. On a 400-case test of tool selection with random distractors, Laya scored 84.0 and Jev scored 99.5, a 15.5-point gap. On a 200-case test with 60 to 77 candidate intents, the gap widened to 46 points (35.0 vs. 81.0). On a 400-case prompt-injection guard, the gap was 14.8 points (78.2 vs. 93.0). On a 200-case groundedness test, the gap was 11.0 points (79.0 vs. 90.0).
The same numbers were then audited by the author and corrected. The self-audit surfaced three analysis errors and one design confound. A deployment cost saving originally reported as 23.9% is actually 4.3%, the result of an omitted pre-screen cost on the cheap classifier itself. A "gate accuracy" metric was being reported as if it were end-to-end quality; the actual end-to-end number is 58%, not 98%. An in-sample 5% miss-rate threshold translated to up to 17% held-out misses once the same threshold was applied to data the classifier had not seen. A suspected "channel effect" on prompt-injection false positives disappears once the comparison uses channel-native content rather than foreign-channel content. The author flagged two other suspected confounds in the abstract that did not change the headline conclusions.
The two classifiers also tie on RAG relevance gating, and neither beats chance on zero-shot model routing, the very decision most often cited as the killer app for cheap per-decision classifiers. The cheap-routing pitch is not yet safe.
Laya, the open-weight classifier, is brittle in two specific ways the leaderboard obscures. It changes 30% of its answers when the option order is reversed, a kind of positional fragility that is hard to detect without a paired test. Its accuracy also drops sharply as the number of candidate tools grows. On items with many or similar candidates competing, Laya hits 31% accuracy. On items with a unique correct tool, Jev hits 98%, because the open-weight classifier cannot distinguish near-duplicate tools at scale. The order in which candidate tools are presented materially changes the answer, and a tool-selection metric in a vendor benchmark is a poor proxy for production behavior.
The full benchmark is public. The GitHub repository ships 7,283 base cases and 6,640 robustness variants drawn from 18 public sources, byte-identical inputs, paired tests, cross-hardware and cross-day reproducibility checks, raw outputs, and analysis code. Anyone who disagrees with the corrections can rerun them, and the repository also includes a Chinese-language technical report and an ACL preprint PDF in addition to the arXiv version.
For a buyer or builder of agent tooling, the takeaway is that the next bottleneck in cheap per-decision infrastructure is not the model. It is the audit. The 23.9% to 4.3% correction, the 98% to 58% gate-versus-quality confusion, and the 5% to 17% in-sample-versus-held-out threshold are the three numbers worth asking about before a classifier goes in front of a real agent. The paper is a single-author preprint, not peer-reviewed, and the corrections are the author's own retrospective review of their own pipeline, not an external review. The fact that the corrections are public, and the code to re-check them is public, is the part worth copying. The cheap-classifier layer still ships with an asterisk: 4.3%, 58%, and 17% are the numbers to ask about.