Evaluating Egocentric Cues in Humanoid Collaboration
Abstract
Spatial information is useful to an embodied agent only if it is represented faithfully, changes action in the expected direction, and remains useful through physical interaction. We study these requirements in a simulated Unitree H1 carrying task with separately trained policies with and without egocentric video. Two controlled tests expose gaps before outcome evaluation. Giving a teacher 0.5 s of future partner motion changes its action in all 24 policy–motion combinations, but the change follows the predicted turning direction in only one. Holding current geometry fixed still leaves a coherent partner silhouette in rendered images for about 120 ms, showing that visual history remains part of the observation. Downstream measures introduce two further gaps: the timing detector usually starts after robot motion is already underway, while more frequent contact loss sets the recorded wrench to zero and makes aggregate force appear smaller. A constructed known-effect control recovers the expected directional and timing benefit. Together, the results show why spatial sensitivity, spatial correctness, and closed-loop physical value must be evaluated separately in embodied collaboration.