Jump Start Your Policy Learning with Lessons from 145,000 Training Runs
Abstract
Reliable progress in offline policy learning depends on a small set of methodological practices, including careful reporting, well-tuned baselines, and evaluation across diverse conditions. Prior work has noted that results can be sensitive to reporting choices, hyperparameter tuning, and dataset properties independently, but these sources of variability have not been systematically investigated at the scale needed to understand how they shape conclusions. To address this gap, we present a large-scale empirical study of offline reinforcement and imitation learning, training over 145,000 policies across 114 datasets. At this scale, no algorithm dominates: aggregate performance across top methods is often close, but the leaders differ substantially across environments. We find that proper hyperparameter tuning frequently reshuffles perceived algorithm rankings; and simple baselines, including behavior cloning, are often stronger than is commonly assumed after extensive tuning. We also study hyperparameter transfer and sensitivity across environments, identifying a simple strategy that generalizes well. From these analyses we distill practical recommendations, and release JumpStart: a resource suite of trained models, per-model hyperparameter and reward data, strong baselines across all environments, and a website to make retrieval and analysis trivial. Together, these resources aim to make offline policy learning research more reliable and to open new directions for work beyond the scope of this study.