EgoSurgHands: An Egocentric 3D Hand Pose Dataset & Benchmark for Open-Surgery Training
Ankit Pal ⋅ Guillaume Kugener ⋅ Sofia Volpi ⋅ Joshua Schwartz ⋅ Alex Idarraga ⋅ Pranav Rajpurkar ⋅ Gabriel A Brat
Abstract
We introduce EgoSurgHands, an egocentric 3D hand-pose dataset and benchmark for open-surgery training tasks captured with Meta Project Aria glasses on standard surgical training pads. To our knowledge, EgoSurgHands is the first public benchmark for surgical training that is simultaneously egocentric, sensor-grounded, and human-validated. EgoSurgHands contains 41,909 training hand-instance rows across 55 procedure recordings and 9,476 inter-annotator-agreement (IAA)-validated rows across 13 held-out procedure recordings from 4 wearers, 13 distinct tasks, 3 glove colors, and 3 stages of training. Each frame includes Aria intrinsics, SLAM extrinsics, MPS (Machine Perception Services) landmarks, per-joint confidence, timestamps, and eye gaze. Validation 3D keypoints are initialized from Aria MPS and independently corrected by four expert annotators under a multi-rater IAA workflow, yielding evaluation ground truth independent of any tested model's training signal. Using EgoSurgHands, we benchmark six state-of-the-art hand reconstructors under zero-shot transfer: HaMeR, WiLoR, WildHands, JointTransformer, SimpleHand, and MeshGraphormer. Ego-pretrained models consistently outperform non-egocentric baselines: WildHands improves over the non-ego HaMeR reference by 4.51 mm PA-MPJPE, while JointTransformer achieves the best zero-shot result at 14.95 mm PA-MPJPE. These results show that generalist egocentric pretraining transfers to surgical egocentric tasks, but also reveal a substantial remaining surgical-domain gap. As a lightweight diagnostic adaptation experiment, a 58K-parameter two-headed (pose + translation) residual corrector further reduces error on frozen backbones, reaching 10.40 mm PA-MPJPE and 75.8 mm absolute MPJPE on WildHands ($-$49.2\% and $-$85.8\% vs zero-shot). We release the dataset, benchmark code, checkpoints, and multi-rater annotation framework at [URL withheld for double-blind review].
Chat is not available.
Successful Page Load