ChildPose: Foundation for Children Pose Modeling
Abstract
Human pose estimation has advanced rapidly driven by large-scale adult-centric models. However, children remain significantly underrepresented due to data scarcity, distinct body morphology, and unique motion dynamics. We introduce ChildPose, a modular video foundation model designed for child-centric pose estimation. ChildPose transforms pretrained pose encoders into a streaming pediatric model via memory-augmented temporal attention and age-aware co-training. By maintaining a memory bank of adjacent frame representations, the model achieves robust localization under motion blur, occlusion, and atypical poses, while an auxiliary age-prediction objective enforces representations that capture developmental morphology. We curate a diverse, video-based pediatric dataset spanning various age groups and capture conditions, evaluating out-of-distribution generalizability on entirely unseen subjects. Across multiple 2D baselines, ChildPose consistently outperforms both off-the-shelf models and child-data fine-tuned variants, while ensuring stable sequence-level predictions. Our results underscore that effective pediatric pose estimation necessitates child-specific modeling beyond standard adult-centric estimators.