When Do Agent Skills Help? Failure-to-Recovery Pathways in Skill–Harness Systems
Abstract
Agent skills are commonly evaluated as if their value were intrinsic to the skill, yet in deployment that value is mediated by the surrounding harness. We ask when a frozen skill translates into realized agent capability by crossing four skill conditions with a five-level harness on 1,872 tasks from nine benchmarks, with three repeats using a frozen Qwen3.5-9B model (112,320 executions). We find that stronger harnesses do not uniformly help and recovery is most useful when failures are observable, actionable, and locally correctable. A matched decomposition of the change produced by augmenting standardized interface support with execution feedback and bounded recovery separates feedback-exposed gains, gains without recorded exposure, and regressions. Moreover, pre-recovery behavioral history strongly diagnoses whether the stronger recovery harness will succeed, suggesting that recovery can be invoked selectively rather than applied by default. Skill utility is therefore a property of the skill--harness--task system, not of the skill alone.