Distributionally Robust Token Optimization in RLHF
Yeping Jin ⋅ Jiaming Hu ⋅ Yannis Paschalidis
Abstract
Large Language Models (LLMs) tend to respond correctly to prompts that align well with the data they were trained and fine-tuned on. Yet, small shifts in wording, format, or language can trigger surprisingly large failures, especially on multi-step reasoning problems. To address this problem, we propose a $\textbf{Distributionally Robust Token Optimization (DRTO)}$ approach, which combines token-level Reinforcement Learning from Human Feedback (RLHF) with $\textit{Distributionally Robust Optimization (DRO)}$. DRTO constructs f-divergence ambiguity sets over span-level actor losses, providing a principled way to emphasize difficult response segments during policy optimization. Empirically, DRTO enhances consistency under distribution shifts in multiple reasoning benchmarks among different tasks.
Chat is not available.
Successful Page Load