Causal Benchmarks for Multimodal AI Should Measure Categorical Difference, Not Capability Gap
Abstract
This position paper argues that causal benchmarks for multimodal AI should measure categorical difference, not capability gap. The field reports machine causal performance against human performance and interprets progress as gap-closing along a shared scale. The shared-scale assumption rests on a flattened reading of Hume's regularity theory in which causation reduces to constant conjunction across instances and the temporal experience of the observer has no role. When understood in terms of its presuppositions, Hume's regularity is inseparable from a temporally extended subject who undergoes the conjunctions over time. Human and machine causal cognition therefore have different temporal architectures. The two are grounded in time as substrate and time as dimension, a categorical difference that scale or additional intervention machinery cannot close. Evaluation paradigms that treat the two as commensurable produce inflated assessments of machine causal competence that independently scored probes cannot capture. We propose \textit{Anthropomorphic Competence Inflation} as a construct the field can develop into a broader measurement instrument and formalize it through the \textit{Inflation Score}. Preliminary observations on two recent vision-language models match the pattern predicted by the categorical claim. The implications affect evaluation practice, deployment in causal-decision domains, and the next phase of AI causality research.