Researchers want the term pinned to a rule based process plus a communication checklist, so audits, regulators, and buyers can verify the same claim.
Two labs claim their systems can reason. Each runs a different test, publishes a different score, and uses the word "reasoning" to mean a slightly different thing. When the press, regulators, or downstream users ask which one is right, nobody can answer them on the same evidence.
That gap is the subject of a new position paper, Position: Reasoning is a Learnable Rule-Based Process, posted to arXiv this week. Its authors argue the AI field's most stubborn obstacle to trustworthy "reasoning" systems is not capability but measurement. Until the term "reasoning" has a shared operational definition, claims about which model is better at it cannot be externally audited, and the question of whether any of these systems is actually reasoning becomes unverifiable.
The paper's constructive payload is concrete. First, it offers a synthesized operational definition: valid and sound reasoning, treated as a learnable rule-based process. In plain terms, the authors want "reasoning" pinned to something a third party can check, the way a proof checker checks a proof. Second, the paper ships a communication checklist for AI reasoning research: a set of best practices for how papers, labs, and reviewers should describe reasoning claims so others can verify them.
Why this matters outside the lab: "reasoning" has become the AI industry's favorite trust word. Frontier labs publish reasoning benchmarks, enterprise buyers ask whether a model can reason before deployment, and regulators increasingly require evidence of reasoning ability for high-stakes uses. If the word keeps shifting, every audit, regulation, or buyer's checklist built on it is also shifting. The authors call this the construct-validity problem: the underlying idea cannot be measured because the field has not agreed on what would count as evidence for it.
Historically, reasoning was the territory of symbolic AI, where each step could be checked against a hand-built rule. The recent surge in reasoning benchmarks has come from deep probabilistic generative models, the family of large language models now used in production, where the same word now covers everything from chain-of-thought arithmetic to multi-step planning to olympiad-style problem solving. The paper argues the field has not reckoned with the gap between the two traditions.
The strongest counterargument is pragmatic. Benchmarks work fine without a shared definition; labs ship models, users adopt them, and progress shows up in deployment. The authors' reply, in effect, is that benchmark progress and real reasoning progress can drift apart when the construct is not pinned, and that the cost of the drift shows up exactly where the public notices: in contested trust claims, failed audits, and post-hoc explanations that do not survive third-party review. The checklist is meant to be the smallest intervention that closes the gap without slowing the field down.
When a lab says a new model reasons, the next useful question is not how well but by what rule, checked how. The checklist is designed to make that question answerable in the same words across labs, audits, and reporting. The first major reasoning benchmark to grade itself against it will be the real test.
Author identities and affiliations are not in the arXiv abstract; both the Microsoft Research publication page and the full HTML preprint carry them.