When the individually best parts of a local AI inference stack lose to a fitted alternative, the spec sheet is not the right object to compare. The stack is.
Piszczek's run on a 24 GB RTX PRO 4000 Blackwell SFF (192-bit bus, 432 GB/s peak) is the cleanest public exhibit of that pattern. He ran Qwen3.8 27B at 256,000 tokens of context and posted a ten-run production average of 50.44 tokens per second, about the pace of a fast typist reading aloud. He then raised the precision budget on the speculative drafter by 69.2 MiB. Average throughput fell to 37.02 tokens per second. Same hardware, same workload. The drafter was the more aggressive choice on its own, and the system got worse.
The reusable category here is not "buy a Blackwell." It is the discipline of auditing fit between quant, drafter, kernels, and memory layout before trusting any headline number. Piszczek's own benchmark discipline makes the rule explicit: he keeps a 21.97% custom-build gain, a 2.81× drafter gain, and a 12.61 tok/s 256K tail in separate benchmark gates, on the principle that "combining them into one heroic speedup would make a better headline and a worse benchmark." The gates stay separate for the same reason the parts must be checked against each other. The reader who internalizes that stops shopping for the highest-spec part and starts asking which build was actually tested as a whole.
Reported by Sky for Type0, from Qwen3.8 27B at 256K: 50 TPS on a 24 GB GPU. Read the original: piszczek.pl