Video Stitching from Multiple Moving Cameras
Abstract
We address geometric video alignment for the same scene captured by multiple freely moving cameras. Existing methods, which are based on image stitching or video stabilization, are typically limited to two cameras or must assume the restrictive setting of a fixed linear ordering of the cameras. We propose MViS, a lightweight test-time optimization framework for joint multi-view alignment with temporal consistency. MViS uses a spatial-transformer-inspired net to jointly process frames from all views within a short temporal window and predict per-view, per-time homographies. Optimization proceeds in two stages: geometry-guided alignment followed by appearance-based refinement, with both stages enforcing multi-view and temporal consistency across the sequence. The method requires no scene-specific hyperparameter tuning, handles arbitrary and time-varying view-intersection graphs, and remains robust under rapid camera motion. On a standard two-view video benchmark, MViS matches or surpasses the state of the art. To facilitate evaluation in more complex scenarios, we introduce an annotated dataset with more than two freely moving cameras and establish the first benchmark for this setting. Our code and dataset will be released upon acceptance.