A public benchmark is forcing the vision-language field to climb Pearl's causal ladder on image-and-text inputs, and the climb is now scored on a shared ruler. The CausalVLBench paper, accepted at EMNLP 2025 Main, splits "visual causal reasoning" into the same three task families the rest of causal inference has used for decades: read the causal graph from a picture, name the lever to pull, and answer the "what if" question about a scene that did not occur. The hierarchy is the point. These three rungs separate association from intervention from counterfactual reasoning in any formalism, and the benchmark institutionalizes that hierarchy for LVLMs.
The mechanism is older than the benchmark: a shared ruler arrives, the field gets measured against it, and every subsequent model release either posts its three-family score or quietly steps around the yardstick. Once the tasks are public, "our model reasons about cause and effect in images" stops being a marketing line and becomes a per-rung number other labs can falsify. That pressure is the contribution.
The field gains a referee for visual-causal claims. It loses the freedom to define "visual causal" on a model's own terms. The model cards worth reading next quarter will be the ones that publish all three family scores at once.
Reported by Sky for Type0, from CausalVLBench: Benchmarking Visual Causal Reasoning in Large Vision-Language Models. Read the original: arxiv.org