Split-RL: Local Conflict Resolution in Reinforcement Learning
Abstract
We introduce Split-RL, a reinforcement learning (RL) algorithm based on decision trees. It uses guidance labels to explicitly partition the state-space and isolate localized conflicting signals. RL policies are often trained by mixing contradictory feedback that is localized in space and time. For example, a robot may prioritize speed in open areas versus precision in tight gaps. These regions are often easily detectable from sensor data or a human operator. Standard neural networks (NNs) lack the structural mechanism to isolate these signals. While Gradient Boosting Trees (GBT) provides the inductive bias to explicitly partition the state-space, standard tree-fitting procedures fit only aggregated gradients, merging contradictory updates before they reach the leaves. Split-RL leverages this bias to integrate guidance labels into tree construction, routing these updates into disjoint regions of the state-space to prevent signal interference during training. Such labels can arise from simple heuristics, sensor-derived events, or existing rule-based systems, allowing Split-RL to incorporate rule-derived structure while retaining a learned policy. We theoretically show how early gradient aggregation loses objective-specific information and characterize when the Split-RL score isolates conflicting updates. Finally, we evaluate Split-RL on constrained tasks and offline imitation learning from mixed-quality datasets. Across constrained domains with spatial, rare-event, and temporal conflicts, Split-RL consistently outperforms existing methods by achieving low-cost, feasible solutions while maintaining competitive rewards. In offline settings, Split-RL localizes the tradeoff between imitation and policy improvement, outperforming other methods and achieving near-expert performance across varying dataset sizes.