CoANeRV: Coordinate-Aware Token-Space Neural Video Representation
Abstract
Neural representations for videos (NeRV) have shown strong reconstruction fidelity by storing video-specific information in network weights. However, existing formulations typically require either costly per-video optimization or video-specific weight generation, making it difficult to scale to efficient amortized video representation. We propose \textbf{CoANeRV}, a coordinate-aware token-space neural video representation framework that shifts video-specificity from decoder weights to compact latent video tokens. Instead of optimizing or generating an instance-specific network, CoANeRV uses a shared coordinate-conditioned decoder to reconstruct videos by querying tokenized video content at continuous spatio-temporal coordinates. This formulation decouples video representation from parameter instantiation, enabling feed-forward video encoding while retaining coordinate-level reconstruction flexibility. To make token-space reconstruction effective, CoANeRV introduces a coordinate-aware decoding architecture that aligns spatio-temporal queries with video tokens through axis-adaptive positional encoding and temperature-modulated cross-attention. Block-wise coordinate querying further reduces peak attention memory, making high-resolution reconstruction practical. Experiments on diverse video datasets show that CoANeRV consistently improves reconstruction quality over prior feed-forward NeRV and INR baselines, reduces peak memory compared with attention-based coordinate decoders, and provides efficient amortized encoding without per-video optimization. These results suggest that token-space neural representation is a scalable alternative to conventional weight-space NeRV formulations. The code is available at https://anonymous.4open.science/r/CoANeRV-50F7/.