Evaluating Multimodal Benefits in Closed-Loop Robot Control
Abstract
A better aggregate score does not necessarily mean that an added sensing modality improved physical behavior. We study this problem in a simulated Unitree H1 carrying task by comparing separately trained policies with and without egocentric video. The aggregate timing and force statistics initially favor video, but closer analysis changes their interpretation. The timing detector usually begins searching after the robot has already started moving, while the video policy loses contact more often; once contact breaks, the simulator records zero wrench. Supporting controls reveal two related gaps: future partner motion changes the controller’s action but almost never in the predicted direction, and rendered images retain about 120 ms of visual history even when current geometry is fixed. These results motivate claim-conditioned evaluation: the comparison must answer the intended question, the measurement must remain valid and observable, and the added information must be isolated and used in the expected way. In a constructed known-effect setting, a synthetic positive control demonstrates that the timing analysis can recover a known benefit. The study shows how plausible multimodal gains can appear without establishing faster, gentler, or more appropriately informed closed-loop behavior.