When Actions Matter: Causal Affordances for Long-Horizon Credit Assignment in World Models
Pradeep Kumar Banerjee ⋅ Frank Röder ⋅ Nihat Ay
Abstract
Long-horizon tasks require identifying which past decisions actually mattered for distant outcomes. In deep achievement chains, a single early decision can determine success thousands of steps later, far beyond the reach of short-horizon imagination. We introduce WHAM, a world-model agent for causal credit assignment in long-horizon tasks. WHAM ascends Pearl's causal hierarchy within a learned world model to identify decision-critical bottleneck states, prioritizes replay and imagination around them, and propagates value through sparse bottleneck chains using replay-derived returns. A hierarchical bottleneck critic bridges credit across thousands of timesteps with provably controlled error. On Crafter, WHAM achieves $67.1\%$ ($+8.2$ over a state-of-the-art baseline), with gains that grow systematically with achievement-chain depth: diamond collection rises $17\times$ to $8.5\%$, and the best seed reaches $12.5\%$, exceeding the human expert rate. Long-horizon credit flows where actions shape distant outcomes.
Chat is not available.
Successful Page Load