Feed-Forward 3D Gaussian Splatting for High-Fidelity Animatable Hand Avatar Reconstruction from a Single Image
Abstract
Recent hand avatar reconstruction methods achieve high-quality results through multi-view observations or per-subject optimization, limiting scalability and practical applicability. We present \emph{FF3DGS-Hand}, a feed-forward framework that reconstructs animatable hand avatars from a single monocular image using 3D Gaussian Splatting. Our method establishes a stable canonical geometry for thin and articulated hand structures, then reconstructs appearance by leveraging source image evidence instead of relying solely on global latent features. We further introduce a rendering-aware, source-conditioned appearance framework that selectively updates Gaussians according to their rendering contribution, recovering input-specific details while suppressing artifacts in unobserved regions. The resulting representation is animatable via standard hand pose parameters and supports efficient rendering. Experiments demonstrate strong visual quality and quantitative performance in the single-image setting.