Consolidating Reasoning with Test-Time Learning
Abstract
Parallel thinking scales inference-time compute by generating multiple reasoning traces for a query, and aggregating them to reach a final answer. However, most aggregation algorithms treat each reasoning trajectory independently, ignoring the complementary reasoning sub-processes across the full ensemble of trajectories. To learn from these parallel trajectories, we propose CORAL (Consolidating Reasoning with Test-Time Learning), a novel framework for aggregating trajectories by jointly consolidating them into parametric memory using test-time training. Our algorithm meta-learns how to consolidate using nested optimization at training time, where an inner loop encodes reasoning trajectories into a lightweight LoRA adapter, and an outer loop learns meta-parameters for the inner loop such that the consolidation is helpful for the model to synthesize a final solution. As no standard datasets exist for training aggregation methods, we construct a novel OpenParallelThinking corpus composed of 12,810 math problems paired with a mixture of diverse trajectories from six LLMs. Evaluated on challenging mathematics and STEM benchmarks, CORAL consistently outperforms best aggregation baselines by 20% overall. Importantly, we curate reasoning trajectories at multiple quality levels to simulate the reality of noisy parallel reasoning and demonstrate that CORAL is especially robust to the quality of thinking trajectories and able to learn useful knowledge from entirely wrong solutions. Additionally, this weak-to-strong generalization feature scales with the number of incorrect trajectories available for aggregation, while RL can further incentivize such reasoning consolidation capability. We will release CORAL and OpenParallelThinking for use in future work on reasoning aggregation.