Decomposed Graded Verifier for Generative World Modeling
Abstract
Generative visual verifiers play a crucial role in advancing multimodal intelligence systems. In this paper, to enhance the robustness of visual verifiers in complex and realistic open-world scenarios, we move beyond conventional binary verifiers that produce only True/False discrete judgments. Instead, we introduce a verifier that automatically decomposes complex prompts or environmental rules into fine-grained components and assigns pointwise scores, thereby providing a continuous and more informative critique signal. Under a general reinforcement learning framework, we first empirically demonstrate that the proposed pointwise decomposed verifier significantly outperforms binary verifiers in traditional text-to-image scenarios. We then extend our investigation to a novel paradigm: generative world modeling, where generative models are leveraged for world simulation and world reasoning. From a theoretical standpoint, we further show that when ground-truth images are easy to obtain, a pairwise verification paradigm yields more accurate critiques than the pointwise formulation. Empirical results on diverse world-modeling tasks, such as Maze and Sudoku, further validate the effectiveness of our approach. Overall, our findings suggest a key insight: the transition from generative models to world-modeling agents critically hinges on the availability of accurate pairwise decomposed visual verifiers.