From Internal Evidence to Behavioral Commitment in Embodied Skill Selection
Abstract
Agent Skills expose a name and a short description for every skill before any procedure is loaded, and an agent selects a skill from this metadata alone. In this study, we examine where skill selection fails. Across three vision language models and two skill libraries grounded in embodied tasks, selectors choose the correct skill for only 43% of the tasks on average. A wrong selection may arise because the selector misunderstands the task or because it misreads the skill library, and the two failures call for different remedies. To separate them, we first probe the hidden state immediately before the skill name is written and find the correct skill linearly decodable for 64% to 86% of tasks. The task is thus understood when a wrong name is produced. Since the information is present, we then amplify it with a steering vector along the correct selection direction, which changes accuracy by only 1.6 points on average. The failure therefore lies in how the skill library is read. To find out how, we edit the library while holding everything else fixed. Removing the skills that can never be correct raises accuracy by 13 to 24 points. Swapping the name of the correct skill with that of a closely related skill moves most selections to whichever skill carries the correct name. In other words, the selector reads names first. Finally, we ask whether training repairs this reading. After training with token-level or representation-level supervision, selection accuracy improves, the correct skill remains just as decodable from the hidden states, and the selector relies more on descriptions, following a moved description twice as often as before. Therefore, it improves selection by changing how the metadata is read, which agrees with the diagnosis.