Learning When Visual Context Matters for Mouse Behavior Analysis
Abstract
Understanding mouse behavior at scale is a foundational problem in neuroscience, with two core tasks: classifying frames into known behaviors and segmenting recordings into recurring patterns without labels. Pose-based methods are the standard approach but face fundamental limitations: pose coordinates capture only skeletal information, discard environmental context, and degrade under frequent occlusions in social scenarios. Mouse videos can provide the visual context that pose lacks, yet not all frames benefit equally, and adding visual features uniformly to every frame is suboptimal. We find that a pose-based motion encoder, trained with auxiliary supervision from videos, can identify where it struggles: with a discrete representation, the encoder learns a codebook of recurring patterns, and frames whose motion does not fit this codebook produce large quantization residuals. We show that these high-residual frames are where visual context tends to help most: residual-guided frame selection consistently outperforms uniform selection across budgets, and often exceeds dense visual processing. We therefore propose ResiFuse, a framework that pretrains a motion encoder on pose with an auxiliary cross-modal loss and integrates vision context only on high-residual frames at fine-tuning time. We evaluate ResiFuse on three mouse behavior datasets and show improvements over pose- and video-based baselines on most tasks, with the largest gains on the most socially complex dataset.