Same data, same model — 65% to 77% on the same money laundering detection task came from rewriting what each label meant.
A small, open-weights model on the same anti-money-laundering dataset scored 65% one day and 77% the next. Nothing about the model changed. The labels did.
The author, an engineer publishing a public benchmark of two small confidence-scoring models called Jev and Laya, kept the same model, the same training data, and the same evaluation split, and rewrote the label descriptions so each one described what the transaction looked like in the input rather than what a criminal was trying to do. Accuracy jumped 12 points. The lesson generalizes: on classification work, the label text you give a model is the task you are actually asking it to learn, and most teams write that label text for themselves, not for the model.
Confidence-scoring models, the kind that return a probability over a fixed set of labels instead of generating text, are a real and expanding capability tier. They are cheap to run, easy to evaluate, and let a small team ship a working "decide what to do with this input" service in days rather than months. The author's experiment covers 31,369 Jev test samples across 12 tasks and 23 Laya runs totaling 65,078 requests, and it is a working audit of what actually moves the numbers on this kind of system. Three reusable findings emerge.
First, the label wording is the task. On the AML data, a separate account-level experiment reported in the same blog saw the same pattern: once a per-account time-window bias in the labels was removed, account-level laundering classification rose from 63% to 75%, with the model unchanged. The fix was not a new architecture. It was a description of the labels that matched the inputs the model could actually see.
Second, calibration is what makes a confidence filter shippable. The same author reports an aggregate calibration error (the average gap between a model's predicted confidence and the actual hit rate) of 0.013 for Jev against 0.486 for untuned Laya, a roughly 35x gap. That gap is the difference between a 90%-and-up confidence filter you can actually trust and one that mostly filters out the easy cases. On the blog's support-intent run, that filter is the lever: support-intent accuracy rose from 92.3% to 97.6% at 90%+ confidence, on the 82% of questions the model felt sure about. A released Laya Experts checkpoint later refit its temperature on a held-out split and reported a Banking77 calibration error of 0.009 against Jev's 0.089, with Banking77 accuracy at 91.7% versus 79.8%. The pattern is the same: when the model is calibrated, the confidence number means what it says.
Third, some tasks have no signal. The same author tried single-transaction laundering classification, the kind of "is this one wire suspicious" question teams often want a model to answer, and got 54% accuracy, essentially chance. That is not a model failure. It is the author naming a task with no learnable structure in the data. The model is doing exactly what the data permits, and the right response is to redesign the task or refuse to ship it, not blame the model.
Two practical warnings. The author is the author, the runner, and the tuner: thresholds were set per task, and nobody outside the project has reproduced these numbers. That does not make the findings wrong, but it caps what they prove. The lesson is the method, not the leaderboard. Treat the 65% to 77% lift, the 0.013 to 0.486 calibration gap, and the 92.3% to 97.6% confidence-filter gain as a worked example of how small decision systems actually behave under a careful build, and rerun the same checks on your own data before you trust the same numbers.
Second warning: cost. A hosted Jev call from Singapore to a US endpoint is not the same cost as a local Laya run on a lab GPU, even when the per-call API price looks like zero. Local inference has hardware, power, and operations costs that have to be priced in. Calling local-as-free is a usable shortcut, not a full P&L.
The practical recipe, then, is short. Write the labels in the same vocabulary the input uses. Refit the temperature on a held-out split and check expected calibration error before you trust any confidence threshold. Filter on confidence only after calibration, and only on the share of inputs the model actually scores high. If a task sits at chance after a fair attempt, treat that as a finding, not a failure. Ship the parts that work, name the parts that don't, and keep the evaluator in the loop.