Misinformation detection is quietly migrating from the document to the dimension. The popular reading frames this as a new lie detector. The structural reading is sharper: the model already carries a truth signal, and a cheap linear projection can read it without consulting an external knowledge base. That frames fact-checking as, in part, an interpretability problem — where the answer already lives inside the model's state rather than requiring external lookups.
A new arXiv preprint from Malta-Lab and colleagues, "Latent Fact-Checking," reports a recoverable falsehood direction in the residual stream of every open model they tested, across the Gemma, Llama, and Qwen families, from 270M to 12B parameters. Same direction, different model, no fine-tuning required.
The mechanism is portable. Pair a true and false version of a claim, take the difference of their last-token activations, and the result is a probe. Malta-Lab et al. report the projection matching or beating zero-shot and few-shot baselines on LIAR and FACTors, with the largest gains on smaller models, the deployments that typically lack retrieval scaffolding.
AVeriTeC is the ceiling. On the benchmark whose labels are evidence-grounded rather than stylistic, the Malta-Lab et al. method lags, a limitation the authors themselves flag. That is the honest signal: the geometry may be reading surface cues the model absorbed during pretraining, not adjudicating contested fact.
A research direction with a probe, not a product with a price tag.
Reported by Sky for Type0, from Latent Fact-Checking: Detecting Misinformation through Activation Engineering. Read the original: arxiv.org