ParaLLax: Disentangling Difficulty from Validity in Reasoning Verifiers
Akshitveena Singh
Abstract
Large language models frequently reach correct answers through unsound reasoning, and the process reward models built to catch this are increasingly deployed as gates and reward signals for mathematical agents. Their evaluation, however, is never controlled for a variable that predicts the label almost as well as the reasoning does: problem difficulty. We show that a null model given no solution text--only length, LaTeX density, step count and dataset--reaches $f1_B = 0.515$ against a base rate of $0.306$, two-thirds of the lift the best text-reading detector achieves. We call this the ParaLLax effect and give a protocol that measures it. Under control, a pooled autoencoder falls to the raw-SBERT floor and both PRMs deflate by $\approx 0.10$ AUROC. A genuine signal nonetheless survives: per-step structure reaches $0.591 \pm 0.050$ (5 seeds), holds flat from 22M to 335M encoders, and beats an all-positive baseline within every dataset where difficulty-only detection does not. Probing a 7B verifier locates the confound causally--length is decodable at $R^2 \approx 0.92$, and ablating a $\approx$16-dimensional subspace reproduces our statistical control--but no linear operator removes it without also removing step-error competence. Validity is detectable, but measuring it requires per-step structure and explicit confound control.
Chat is not available.
Successful Page Load