MINT: Meeting-time INdicators for Truncation in Multi-Step Off-Policy RL
Seungyub Han ⋅ Taehyun Cho ⋅ Dohyeong Kim ⋅ Kyungjae Lee ⋅ Jungwoo Lee
Abstract
Multi-step off-policy TD methods correct the behavior--target mismatch through a \emph{product} of per-step trace coefficients, yet this product inherently amplifies variance or discards long-horizon signal, making the cap length a sensitive hyperparameter. We propose \textbf{MINT} (Meeting-time INdicators for Truncation), which replaces coefficient products with a binary coupling indicator: one until the behavior and target trajectories meet, zero thereafter. Because the indicator is idempotent, variance depends only on the meeting-time survival probability---not on ratio products---and a single action sample estimates the per-step coupling probability without mixing-time knowledge. \texttt{MINT} contracts in supremum norm to $Q^\pi$ under arbitrary behavior policies; empirically, \texttt{MINT} outperforms handcrafted multi-step baselines across MuJoCo and DeepMind Control tasks without cap-length tuning.
Chat is not available.
Successful Page Load