Rethinking Personalized Generation: Test-time Alignment via Factorized Ranking Models
Qiyao Ma ⋅ Junshan Zhang ⋅ Zhe Zhao
Abstract
Aligning large language models (LLMs) to diverse user preferences is fundamentally hindered by standard alignment paradigms that optimize for monolithic users. In this work, through empirical studies, we first discover a massive, untapped performance headroom for personalized generation through test-time scaling. We demonstrate that personalized generation is uniquely suited for test-time sampling methods like Best-of-$N$ (BoN) because it can be viewed primarily as a candidate matching problem rather than a generator capability bottleneck. While standard reward models can theoretically exploit this headroom, their massive parameter counts introduce a prohibitive computational bottleneck. To overcome this limitation, we propose a parameter-efficient framework utilizing million-parameter scale Multi-Layer Perceptron (MLP) ranking models. Our personalized ranking model directly reuses the internal embeddings of the base generator with minimal overhead. By scaling train-time data to provide fine-grained personalized preferences, this million-parameter ranking model accurately scores large candidate pools and can seamlessly guide generation to reduce the cost of materializing $N$ candidates. Extensive experiments on nine datasets across three different personalized generation settings demonstrate that our framework effectively exploits the discovered headroom. Remarkably, our million-parameter MLP performs competitively with billion-parameter reward models explicitly finetuned on the same tasks, while requiring only a negligible fraction of the inference cost.
Chat is not available.
Successful Page Load