Benchmarking and Optimizing Multimodal Structured Generation: The OracleGraph Dataset and PRISM Framework
Zhihong Sun ⋅ Qi Fan ⋅ Pan Liu ⋅ Yang Yi ⋅ Chen Ye
Abstract
Structured generation remains a critical bottleneck for Multimodal Large Language Models (MLLMs). Current post-training paradigms struggle to resolve this: Supervised Fine-Tuning (SFT) hits a rigid ceiling, while scalar-reward Reinforcement Learning (e.g., DPO, GRPO, SimPO) suffers from a policy-level signal gap. Compressing long routing trajectories into sparse signals induces a short-completion bias and traps RL near the SFT baseline, whereas vanilla self-distillation actively regresses. To systematically quantify this joint visual-structural-relational reasoning, we introduce OracleGraph, a 10,000-page 2.5D historical document benchmark featuring a 12-dimensional diagnostic VQA suite and a structured-hallucination taxonomy. To overcome these optimization pathologies, we propose PRISM, a Policy-Reward Integrated Self-distillation framework. PRISM leverages multimodal information asymmetry via a Hindsight Teacher guided by a Diagnostic Graph Report (DGR). By synergizing Softmax Reward-Weighted Aggregation (preserving Graph-F1 rankings) with a Length-Normalized Auxiliary Preference Loss, PRISM explicitly neutralizes catastrophic truncation, yielding a $+3.81$pp Graph Reward gain and a $+2.3$pp schema-parsing improvement over SDPO ($p < 0.001$). While PRISM's nominal $+1.02$pp mean gain over SFT is not statistically significant ($p = 0.12$), it drastically stabilizes optimization, achieving $2.4\times$ lower cross-seed variance. Crucially, on a cross-book out-of-distribution (OOD) probe, PRISM's stability advantage amplifies to a $14\times$ variance reduction over SFT, confirming robust structural stabilization. Code and datasets are available at https://anonymous.4open.science/r/PRISM2026.
Chat is not available.
Successful Page Load