Anthropic and outside collaborators built the first public scorecard for the philosophical, low feedback reasoning that AI safety work depends on.
The questions that matter most for AI risk are exactly the ones current training pipelines are worst at rewarding. Math problems come with answer keys. Code comes with unit tests. The reasoning AI safety work actually depends on (long-horizon argument, philosophical framing, judgment under irreducible uncertainty) does not. The Conceptual Reasoning Index, published this week on Anthropic's alignment blog, is the first public attempt to grade that gap.
CRI is an aggregate of three benchmarks targeting conceptual reasoning: tasks without empirical feedback loops, where the model has to reason about values, intent, or systems rather than optimize a verifiable score. The work is described as "in collaboration with Anthropic," meaning external researchers shipping on the lab's alignment blog, and the index is viewable at conceptualreasoning.ai, with a stated plan to update as new models and benchmarks appear.
AI safety, the field trying to reduce catastrophic risk from advanced AI, depends on reasoning that math and coding benchmarks do not capture. Math benchmarks, coding evals, and some agentic harnesses all give a model something to optimize against. The reasoning that goes into deciding whether a deployment is safe, whether a governance structure concentrates too much power, or whether a misalignment scenario is plausible does not come with a gold label. Anthropic's post states the gap directly: "current training pipelines are typically worse at tasks that cannot be empirically or mathematically verified, i.e., the very tasks that matter most for AI risk." CRI is the bet that this reasoning is at least measurable.
Anthropic's post names three structural problems. First, AI will increasingly be asked to reason about AIs more capable than any human. Second, some of those decisions are one-shot, with no rollback or second try. Third, many of the relevant questions are value questions with no ground truth, where two reasonable people can land on different answers. The post's own projection: "Once models can perform work that reduces AI risk at the level of human experts, AI(-assisted) output in the area might dwarf unassisted human output." If that is right, the reasoning capacity behind that work is load-bearing, and unverifiable reasoning capacity is the riskiest part of the stack to leave unmeasured.
The dataset underneath, LMCA, is not publicly downloadable. Researchers who want to inspect the prompts and scoring rubric have to request access through a form. That is not unusual for evaluation work that could be contaminated by training-set leakage, but it does mean the broader research community cannot independently audit what CRI is actually measuring. The published index reports a score; the underlying artifacts are gated. Until outside researchers can audit the data, every public reading of the scorecard is a reading of Anthropic's version of the scorecard.
On Hacker News, several commenters pushed back on the idea that a frontier lab, one whose commercial model depends on shipping increasingly capable systems, is the right institution to build a public scorecard for the very risks its deployment is accused of accelerating. The objection targeted the incentives around CRI, not the methodology. A lab that built the benchmark is also the lab whose strategy the benchmark is implicitly being used to evaluate. That tension is real, and the post does not address it.
CRI is genuinely the first aggregated public scorecard aimed at the slice of reasoning (argumentative, philosophical, low-feedback) that alignment work relies on and that gradient-based training is least equipped to produce. It is also a frontier-lab artifact, with the conflict of interest that implies. Treating those as competing narratives misses the structural point: the gap CRI is trying to measure is exactly the gap where lab incentives and public-interest evaluation are least aligned. Whether the scorecard outlives that tension is the watch item.
The next test is whether outside evaluators can get under the hood. The gated dataset, the form-based access, and the "in collaboration with Anthropic" credit on the post are the friction points. If independent researchers replicate the index's findings on held-out reasoning tasks, CRI becomes a useful instrument. If they cannot get to the data, the scorecard stays a company publication.