Latency-Conditioned Selection Bias in RLHF
PRAKUL S HIREMATH
Abstract
Asynchronous Partial Rollouts (APR) accelerate RLHF by training on the first $k$ of $N$ parallel completions, but this speed hides a critical flaw: because autoregressive generation time scales linearly with token count, a wall-clock cutoff acts as an invisible length filter. On reasoning tasks where reward correlates positively with trajectory length, APR systematically deletes the most valuable, multi-step reasoning chains from the training batch. We formalize this phenomenon as **Latency-Conditioned Selection Bias** (LCSB). We prove that whenever reward and selection probability negatively correlate ($\mathrm{Cov}_{p_\theta}(R, \alpha) < 0$), the APR objective structurally attenuates and can even invert the on-policy gradient. Crucially, standard importance-sampling corrections (e.g., V-trace) cannot resolve this because all trajectories strictly originate from the behavior policy. After empirically validating this latency-length filter in a deployed vLLM system and demonstrating catastrophic, monotonic performance degradation on reasoning benchmarks, we provide two actionable solutions. We introduce **Effective Gradient Information** (EGI), a diagnostic metric that detects LCSB thousands of steps before reward curves diverge, and **Selection-Probability Weighting** (SPW), a near-zero-overhead correction that mathematically neutralizes the bias to recover on-policy reasoning performance without sacrificing asynchronous throughput.
Chat is not available.
Successful Page Load