Position: VLM Causal Reasoning Benchmarks Should Probe Temporal Understanding, Not Presume It
This position paper argues that vision-language model (VLM) benchmarks for causal reasoning rely on two under-examined assumptions. First, benchmarks presuppose temporal constitution, the understanding of time as the medium through which causes produce effects, without testing it as a prerequisite. …