NIKA: Efficient Neural Video Representation via Structured Latent Diversity
Slater R Victoroff ⋅ Madison May
Abstract
Implicit neural video representations offer compact, continuous alternatives to conventional codecs, but higher reconstruction fidelity typically requires more computation per decoded frame. We introduce NIKA, which shifts video-specific capacity from the decoder into a structured latent state, allowing reconstruction quality to scale with latent expressivity while keeping active decoding lightweight. NIKA constructs this state from complementary spatial, spectral, and temporal bases, then decodes it with lightweight ConvNeXt-style convolutional upsampling. On UVG, a 2.91M-parameter NIKA model achieves 33.33 dB PSNR at 462 decoding FPS with 4.5G MACs on an RTX A5000, outperforming comparable single-resolution NeRV-family baselines while using 39--51$\times$ fewer MACs. Ablations show that diversifying latent components improves reconstruction more reliably than reallocating capacity within a single component, and qualitative analysis reveals specialization among components. Together, these results identify structured latent diversity as a practical alternative to scaling decoder complexity for high-fidelity, efficient neural video representation.
Chat is not available.
Successful Page Load