VisConf: Quality-Aware Consensus for Best-of-N in Vision-Language Model Reasoning
Abstract
Best-of-N methods improve text-only LLM reasoning by sampling multiple trajectories before selecting an answer. For Vision-Language Models (VLMs), answer frequency and model confidence alone can be insufficient, as sampled generations also vary in their engagement with the visual input. We find that distributional confidence and visual attention are complementary indicators of rollout quality. Each is limited on its own, while trajectories that show strong signals in both are more likely to be correct. Motivated by this, we introduce VisConf, a rollout-quality score that multiplicatively combines visual engagement with candidate-calibrated Self-Certainty, penalizing rollouts that are weak in either signal. We incorporate Rank-Weighted Consensus (RWC), a tuning-free aggregation rule that converts within-question VisConf ranks into answer-level vote weights. Across three VLMs on MathVista, MMMU-Pro, and MMStar, VisConf with RWC consistently outperforms strong BoN baselines. Our analysis shows particularly large gains when correct trajectories are present but underrepresented in the candidate pool. Code: https://anonymous.4open.science/r/VisConf-VLM-BoN/.