Cross-Channel Agreement Beats Consensus: Compositional Verification for Geometry Reasoning
Abstract
Multimodal geometry reasoning requires models to jointly perceive diagrams and perform symbolic derivations. Test-time verification methods such as majority voting assume voter errors are roughly independent, but in geometry this assumption breaks down: all candidates perceive the same diagram, so perceptual biases cascade into correlated errors that produce confident but wrong consensus. We observe that neuro-symbolic candidates contain compositional structure that can serve as a verification resource. We introduce SKETCH, a format decomposing each candidate into a neural pathway (natural-language claim) and a symbolic pathway (visual grounding → typed DSL → deterministic execution). These pathways are governed by different error processes—symbolic errors are correlated across candidates through shared perception, while neural errors are more dispersed. Our method, Compositional Consensus Verification (CCV), gates consensus on cross-channel agreement: because the two pathways fail for different reasons, accidental agreement is rare, making each concordant vote far more precise than a raw majority vote. Across four geometry benchmarks, CCV achieves 81.0% accuracy (+10.4 pp over majority voting) without external scorers or additional inference. Seven non-compositional baselines—including trained PRMs—cluster at 70–71%, while compositional cross-layer features jump to 79–81%, confirming that structural diversity is the core verification resource, and that a principled gating algorithm effectively unlocks its potential.