A new benchmark forces every large vision-language model through a three-step test for causal reasoning in images — type0 | type0