The Oracle Knows: Learning to Aggregate Temporal Flow Matching Hypotheses for Monocular 3D Pose Estimation
Abstract
Markerless 3D pose from a single camera is among the most scalable behavioral biosignals, and is increasingly applied to clinical gait assessment. However, depth ambiguity makes it intrinsically ill-posed, motivating probabilistic methods that generate multiple plausible hypotheses. Yet converting these hypotheses into a single accurate prediction remains an open problem. We identify and quantify this limitation as the oracle gap: the difference between an ideal selector that picks the best hypothesis per joint and practical aggregation methods. Critically, this gap widens with more hypotheses, showing that the bottleneck lies in generation as much as in aggregation. We propose KNOWPOSE,whichbridges single-frame and long-sequence approaches, addressing this gap from both sides. On the generation side, we introduce a lightweight per-joint temporal encoder that conditions the flow matching velocity field on short 2D pose sequences rather than single frames. On the aggregation side, we propose a Learned Residual Aggregation Module (LRAM), a lightweight readout that predicts per-joint corrections to the mean hypothesis, sidestepping explicit selection when hypotheses are too similar to distinguish. Our method outperforms single-frame methods with far fewer hypotheses at real-time speed. Because the corrections are body-proportion-specific, calibrating only LRAM per subject recovers much of the remaining gap while the generator stays frozen. The same module transfers across datasets without architectural changes.