A$^2$IQL: Adaptive Asymmetric Implicit Q Learning for Automated Warehouse Consolidation
Guangyi Liu ⋅ Andrea Angiuli ⋅ Mirko Ristivojevic ⋅ Joseph W Durham ⋅ Michael Caldara ⋅ Michael Zavlanos
Abstract
We apply offline reinforcement learning to large-scale warehouse consolidation, an application that requires ranking tens of thousands of candidates at each decision epoch with an architecture that scales gracefully with variable-cardinality action spaces. Since the available reward signal is only an approximation of the true target-domain objective, the policy must also transfer robustly. We propose a two-stage framework, Asymmetric Implicit Q-Learning (AIQL) and its adaptive cross-domain extension A$^2$IQL. First, AIQL learns a scoring policy from fixed offline consolidation logs by decoupling a per-candidate scoring actor from an aggregate-statistic critic, keeping value estimation tractable and within the support of logged data. Second, A$^2$IQL adapts the source-domain-trained policy to a new target objective using scarce target-domain data: a Bayesian reward predictor supplies an uncertainty-penalized conservative bonus that modifies only the advantage-weighted actor extraction, leaving the source critic unchanged. On a single-domain benchmark, AIQL improves throughput by 20% over a heuristic baseline using offline data alone. When adapted to a new target objective via A$^2$IQL, the method outperforms source-only training, data pooling, and off-dynamics baselines by up to 23% under a $30{:}1$ source-to-target data ratio while achieving the lowest variance across all methods.
Chat is not available.
Successful Page Load