DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training
Tianhao Hu ⋅ Xiangcheng Liu ⋅ Yuchun Miao ⋅ Youshao Xiao ⋅ Hong Yu Zang ⋅ Yang Zheng ⋅ huang xuan ⋅ Jinrui Ding ⋅ Yufei Zhang ⋅ Yu Yang ⋅ Yi-Kai Zhang ⋅ Yueqing Sun ⋅ Chengcheng Han ⋅ Xiandi Ma ⋅ Wei Wang ⋅ Qi GU ⋅ Yerui Sun ⋅ Yuchen Xie ⋅ Xunliang Cai
Abstract
Asynchronous reinforcement learning (RL) has become a critical paradigm for accelerating large-scale LLM post-training in industrial settings, yet it faces a structural \emph{long-tail dilemma}: rollout efficiency is bottlenecked by a small number of the longest trajectories, which are precisely the most valuable ones for RL training. Existing approaches alleviate this dilemma at the cost of either system overhead (e.g., re-prefill in partial-rollout methods) or algorithmic compromises (e.g., discarded long trajectories in replication-based methods). We trace these tradeoffs to a common assumption: the rollout cluster serves a single policy version at any moment. We propose \textbf{DORA} (\textbf{D}ynamic \textbf{OR}chestration for \textbf{A}synchronous Rollout), which breaks this assumption by maintaining multiple policy versions concurrently within the rollout cluster. DORA combines three mechanisms: \emph{multi-version streaming training} that decouples trajectory completion from batch boundaries, a centralized \emph{load-balancing orchestrator} that re-partitions resources across versions, and \emph{zero-re-prefill migration} that transfers KV-Cache directly across same-version instances. Experiments on open-source benchmarks show that DORA achieves up to $2.12\times$ end-to-end throughput improvement and $8.2\times$ rollout-stage acceleration over synchronous training while preserving convergence parity. \textbf{In real-world production deployment} with thousands of accelerators, DORA achieves up to $6.2\times$ rollout speedup and produces competitive open-source LLMs.
Chat is not available.
Successful Page Load