LessMimic: Versatile Humanoid-Object Interaction with Unified Distance Field Representations
Abstract
Humanoid robots capable of versatile whole-body interaction with everyday objects represent a central goal of embodied intelligence. Existing approaches often rely on motion-reference inputs or task-specific rewards, coupling policies to particular motion scripts, object geometries, and contact timings. This limits geometric generalization and flexible interaction control, where simple commands specify motion intent while object geometry determines how contact should be executed. We introduce LessMimic, a framework for versatile motion-free-at-inference humanoid–object interaction. At inference, a single whole-body policy is steered by root commands and an interaction-type flag instead of motion-reference inputs, and is conditioned on a compact interaction representation built from short histories of distance-field-derived surface distances and local surface directions. This geometry-conditioned interface enables command-steered control while adapting contact behavior to local object geometry. Through visual distillation, LessMimic further enables egocentric-depth-based deployment for MoCap-free sensing. Across object scales from 0.4× to 1.6×, LessMimic maintains robust performance over four interaction tasks, including PickUp, SitStand, Push, and Carry, where motion-conditioned baselines degrade sharply away from the training scale. Beyond single-task evaluation, LessMimic further supports sequential composition of heterogeneous interaction skills, attaining 62.1% success on five-task sequences and retaining 23.5% success at 15 task instances. By grounding interaction control in local geometry rather than motion references, LessMimic provides a path toward geometry-generalizable, command-steerable humanoid–object control with heterogeneous skill composition.