Behavioral Indecision, Not Task Performance, Gates When a Learning Agent's Reward Can Be Inferred
Abstract
Inverse reinforcement learning infers an agent's reward from its behavior, typically assuming a fixed policy. Its extensions to learning agents assume a learning rule. For animals, children, and machines trained by others, neither the reward nor the rule is known to the observer. It has not been measured when the reward can be read from states and actions alone, or what sets the limits. Here we measure this directly. Seven canonical learning rules, spanning temporal-difference, policy-gradient, and model-based families, learn a small gridworld with a goal and hazard cells from scratch, while observers that see only states and actions try to name the hazards. We find that a fixed-policy observer reads the reward only inside a window that opens once the agent has learned to avoid the hazards and closes once its behavior commits to one path. The closing edge is set by the agent's decisiveness, how strongly it prefers its best action, and its learning rule matters only through it. Task performance predicts readability in the wrong direction: learners that reach the optimal path are almost never read, and learners that still wander almost always are. An observer that instead models the learning process keeps reading after the window has closed, even without assuming the right rule. Behavioral indecision, not task performance, thus gates when a still-learning agent's reward can be inferred by any observer that treats behavior as fixed, and modeling learning lifts the gate.