Measuring Weak-to-Strong Legibility of Reasoning Models
Abstract
Language models are increasingly trained to "reason" before answering users' queries, outputting hundreds of intermediate tokens before a final answer. While reasoning is designed with direct model capabilities in mind, we argue that reasoning language models (RLMs) should also be assessed by interaction effects between their reasoning traces and weaker models. In contexts like safety monitoring and distillation, the performance of weak models increasingly depends on legibility---how thoroughly and precisely a reasoning trace externalizes an RLM's actions. Existing efficiency-based metrics for legibility fail to capture "thoroughness", instead focusing on conciseness. Thus, we introduce transfer utility, a method for measuring trace legibility derived from interactions where weak models finish tasks given incomplete traces. Evaluating 85k traces from 12 RLMs across three tasks, we find that reasoning traces generated by the highest-performing and the most efficient models rank the lowest for transfer utility. We also find evidence that transfer utility predicts the ability of weak monitors to verify procedurally dense traces for math or logical problem-solving, though the effect moderates when verifying fact-driven traces. Finally, we show that open-source reward models decouple from transfer utility signals on correctness. Together, these findings surface status quo trade-offs in weak-to-strong legibility, motivating an interaction-first framework for future applications.