Auditing Language-to-Action Control in Driving VLAs
Aradhya Goel ⋅ Bhoomika Gupta ⋅ Muhammed Ustaomeroglu ⋅ Guannan Qu
Abstract
Safety evaluations increasingly use natural language as the measuring instrument. An evaluator changes the instruction, reads the resulting behavior, and infers something about the model. That design measures what it claims to measure only if the manipulated language actually controls the behavior being scored. We treat this as a question of construct validity and answer it with an environment-grounded audit that has to be passed before any higher-level safety contrast is interpreted. The audit asks whether the metric can move under a known non-language intervention, whether the treatment reaches the policy, and whether language has authority on the specific axis being measured. We apply it to three released driving vision-language-action models and find a different validity failure in each. In Alpamayo, a training versus deployment contrast that looks decisive on its own disappears once a factorial control shows the framing moves behavior whether or not the objective is present, leaving an objective-by-frame interaction of $+0.009$ m/s against a $0.5$ m/s threshold declared before the run. In SimLingo, a learned control token raises compliance by $58.7$ points while the same request in plain English raises it by $3.5$, so an evaluation written in English would understate what the model can be made to do by roughly a factor of seventeen. AutoVLA responds to injected text without responding to its meaning. We also show the opposite error, where an unpaired analysis makes a small but real effect look like a clean null. None of the interfaces we tested supports an interpretable alignment-faking evaluation, and the audit makes it clear that each one fails for a different reason. The practical recommendation is that a language-based safety result should travel with the positive control showing the metric could have moved, the effect threshold declared before the run, and the point at which interpretation stops. Without those, such a result should be reported as non-diagnostic rather than as evidence of safety.
Chat is not available.
Successful Page Load