Evidence Guided Dual Expert Memory with Joint Routing for VLLMs Online Correction
Abstract
Vision-large language models (VLLMs) editing aims to efficiently update model responses while retaining generalization and unrelated knowledge. Recent online editors store previous edits as reusable experts, forming a mixture-of-experts editing memory for scalable VLLM correction. By retrieving relevant experts at inference time, this paradigm avoids repeatedly overwriting model parameters and enables continual reuse of accumulated editing experience. However, existing sample-level residual experts usually encode each edit instance as a whole, entangling target-relevant visual evidence with irrelevant context. This coarse expert construction weakens expert reusability, leads to inaccurate routing, and increases unintended side effects during continual editing. To address this limitation, we reformulate online VLLM editing as evidence guided expert construction and routing. A fine mask identifies edit-relevant visual evidence, from which support-suppress experts are constructed to strengthen target consistent cues and weaken edit conflicting cues. Each expert is stored with routing signals and reliability scores, and later retrieved through joint sparse routing over visual evidence, textual context, and expert quality. Experiments demonstrate consistent improvements in editing reliability, generalization, and interpretability over competitive online editing baselines.