AI Behavioral Evaluation Should Be Grounded in Psychophysics Across Marr’s Levels
Jonas Mueller ⋅ Bernhard Egger ⋅ Bjoern Eskofier ⋅ Pasquale Minervini ⋅ ANTONIO RIZZO ⋅ Marco Valentino ⋅ Dario Zanca
Abstract
This position paper argues that AI evaluation research is fragmented across incompatible frameworks (e.g., adversarial robustness, attention probing, bias analysis, corruption tolerance, interpretability), each operating at a different explanatory level without a shared grammar to make this explicit. We propose unifying these efforts under \textbf{AI psychophysics}, a behavioral evaluation methodology grounded in classical psychophysics and Marr's three levels of analysis. The core contribution is a formalization showing that all these threads instantiate the same psychometric function $\psi(\mathbf{S})=p(\mathbf{J}|\mathbf{S})$, but at different levels of explanatory granularity (computational, algorithmic, implementational). We present a taxonomy of 21 evaluation paradigms organized by this framework and show that the absence of a shared grammar causes category errors, where metrics from one level are routinely used to support claims at another. We conclude with a concrete research agenda for standardizing AI evaluation around this framework.
Chat is not available.
Successful Page Load