Spark: Path-Aware Experiential Self-Evolution for VLMs Spatiotemporal Reasoning
Abstract
While vision-language models (VLMs) perform well on general tasks, their spatiotemporal reasoning remains limited. Existing distillation and Reinforcement Learning (RL) methods are often computationally expensive and poorly generalizable. Recent studies suggest that model self-improvement by experience distilled from success and failure trajectories is a promising direction. However, in VLM tasks where reasoning must be grounded in visual information, valuable signals of reasoning quality arise not only from final outcomes, but also from how faithfully the reasoning path interacts with visual information. We examine the dependency between reasoning paths and visual information, and observe that high-quality paths exhibit stronger visual anchoring and greater visual dependency. Building on this observation, we introduce the Spatiotemporal Path-Aware Reasoning paradigm based on experiential Knowledge (Spark) that encourages the model to explore and exploit valuable experiences derived from paths with subtle quality differences. To enable diverse and efficient path exploration, we integrate Monte Carlo Tree Search (MCTS) with fine-grained visual rewards and submodular optimization, allowing the search process to prune redundant branches while preserving critical reasoning paths. We further present a training-free experience construction mechanism that converts contrastive pairs of distinct-value paths into structured experiences, which are dynamically injected into the VLM through similarity-based retrieval during inference. Extensive experiments demonstrate that Spark enables Qwen3.5-9B to achieve an average accuracy of 61.21% across three benchmarks, outperforming Gemini-3-Pro (57.20%). Further analyses verify the flexibility and cross-model generalization of our method, highlighting the crucial role of path-aware experience in advancing spatiotemporal intelligence.