MedCache: Training-Free Spatially Aware Caching for Accelerated Medical Video Generation
Abstract
Video world models are now being adapted to medical applications such as surgical simulation, procedural rehearsal, and synthetic data generation, but their inference cost remains a major obstacle to practical use. Training-free caching is attractive because it can be applied to existing checkpoints, yet current methods face a clear trade-off. Lightweight rules track only the average change in the video and treat all regions of the frame as equally important, which is a poor fit for medical video where the meaningful motion is often confined to a small region such as a tool tissue interaction. Richer rules recover spatial awareness by partially running the network on every step to inform the skip decision, but this decision cost is paid whether the step is skipped or not, and it eats into the speedup. We introduce \textbf{MedCache}, a training-free and probe-free cache for rectified-flow video generation. MedCache reads a spatial risk signal directly from the latent using cheap arithmetic, with no help from the network itself. It combines this risk signal with a prompt-derived domain prior and a lazy consistency check that runs only when reuse is no longer clearly safe. Across three open-world model families in Text2World and Image2World settings, MedCache improves both speed and quality over the strongest training free baselines in the medical regime it targets. On Cosmos-H-Surgical 2B Text2World, it raises Total quality from 0.767 to 0.837 over EasyCache while reducing latency from 185 seconds to 95 seconds. Ablations show that risk-weighted drift drives the runtime gain, while the prior and guard recover fidelity.