Mighty Mouse: A Pose Foundation Model That Remembers What You Never Labeled
Abstract
Pose estimation from animal video is typically deployed one lab at a time: each group annotates its own frames, trains a dedicated model de novo, and estimates only the keypoints it labeled. We introduce Mighty Mouse, a pose estimation model trained across five independently annotated head-fixed mouse datasets harmonized into a 36-keypoint vocabulary, comprising 57,986 labeled keypoint observations across twelve camera views. A lab can adapt it with a small number of annotated frames while inheriting keypoints it never labeled. However, standard adaptation can silently damage these inherited outputs, which prior benchmarks cannot detect because they evaluate only target-annotated keypoints. We expose this failure with a masked-label protocol: withhold a normally labeled keypoint during adaptation, then evaluate it against the held-back truth. Across diverse keypoints and datasets, full fine-tuning and low-rank adaptation (LoRA) perform comparably on annotated keypoints but often severely degrade withheld ones; LoRA can erase isolated channels entirely. Anchored LoRA instead supervises annotated channels with ground truth and unannotated channels with confidence-weighted heatmaps from the frozen base model. It retains withheld keypoints near base-model accuracy while matching or improving performance on annotated keypoints, using a small adapter. Together, Mighty Mouse and Anchored LoRA make heterogeneous pose annotations reusable across labs while preserving the inherited outputs that make a shared model valuable.