Merging RLVR-Trained Experts via Policy-Shift-Guided Spectral Alignment
Abstract
We propose Policy-Shift-Guided Spectral Alignment (PSA), a retraining-free method for merging RLVR post-trained language-model experts by guiding spectral subspace selection with token-level expert-base policy shifts. We first show that existing SFT-oriented merging techniques under-preserve the sparse, directional changes in next-token probabilities that distinguish RLVR experts from the pretrained base. To address this challenge, PSA scores calibration tokens by expert--base probability difference under the same prefix, and converts probability-shift-weighted input activations into column weights. These weights define a weighted SVD objective that prioritizes update directions active on high-shift tokens. After low-rank truncation, PSA applies cross-expert polar alignment to the truncated task-vector bases and restores each expert's per-layer Frobenius scale before aggregation. Experiments on Qwen2.5-7B and Qwen3-1.7B across three RLVR tasks show that PSA consistently outperforms strong merging baselines while preserving expert-level capabilities in a unified model.