Embodied AI Cannot Scale Without Open-Source Distributed Geometric Optimization Backends for Life-Scale Egocentric Data
Abstract
This position paper argues that the embodied AI community must invest in open-source, distributed geometric backends for egocentric pose estimation, and that the current bias toward end-to-end neural solutions is creating a data infrastructure deficit that will bottleneck the next generation of Vision-Language-Action (VLA) models and radiance field reconstruction. While neural frontends (DUST3R, VGGT, DepthAnythingV2) achieve remarkable local spatial accuracy, we show that no open source existing system for ego centric data, neural or classical, delivers metrically consistent global pose trajectories over the 100,000 frame "life scale" sequences that modern embodied AI applications require. This gap is structural, not incidental: fixed window neural architectures cannot enforce global consistency by construction, rolling shutter distortion compounds systematically over long horizons in ways feedforward networks cannot model, and every production system that does work at this scale is proprietary and hardware locked. We formally define the Open Distributed Geometric Optimization Backend, a hybrid architecture combining uncertainty weighted neural priors with distributed bundle adjustment and spline continuous trajectory parameterization, and argue it is the necessary open source infrastructure to generate the metrically grounded training data on which future end to end models depend.