Rate-Constrained Edge Metadata for Sender–Receiver Generative Video Super-Resolution
Abstract
Real-world television and streaming pipelines often deliver compressed HD video to increasingly capable displays, making video super-resolution (VSR) both practically important and fundamentally ambiguous. Existing generative VSR methods are largely receiver-only, requiring the model to infer missing high-frequency structures from degraded pixels alone. We propose MetaVSR, a sender--receiver collaborative framework that transmits compact structural edge metadata as bitrate-constrained side information for generative VSR. MetaVSR builds on a one-step video diffusion transformer and encodes both the low-quality video and transmitted metadata through the native 3D VAE, enabling token-level fusion within the original DiT attention blocks without additional control branches. We further introduce a rate-aware Canny edge metadata selection strategy that allocates metadata bits to structurally challenging regions, and evaluate reconstruction quality under matched total transmission budgets that include both compressed video and metadata. Across five public VSR benchmarks, MetaVSR shows that budgeted sender-side structural metadata can improve reconstruction over the receiver-only DOVE-5B baseline, increasing PSNR by 1.03 dB and SSIM by 0.055 on average while using a smaller 2B backbone. On MetaSPS, our rate--distortion evaluation shows approximately 20\% rate saving under light-noise degradation and more than 50\% under heavy-noise degradation compared with a no-metadata generative baseline. These results demonstrate that compact sender-side edge metadata can reduce reconstruction ambiguity and improve bitrate-efficient VSR when its transmission cost is explicitly accounted for.