Silent Semantic Failures Below the Policy: A Historical Benchmark for Robot-Learning Software
Abstract
Claims about generalist or zero-shot robot policies depend on a software layer that loads demonstrations, types actions and observations, computes losses, manages episode boundaries, and evaluates policies. That layer can continue running while changing the behavior being learned or measured. We construct a historical benchmark of 20 pinned defects from five robot-learning codebases. Each task includes a pre-fix revision, maintainer fix, provenance-labelled prompt, and hidden fail-before/pass-after semantic oracle. Eighteen defects are classified as likely silent semantic failures, and six exercise training, evaluation, rollout, or reset consequences. An isolated cloud matrix reproduces every historical failure and gold fix. For the six behavioral tasks, each oracle also accepts the gold repair while rejecting a pre-declared plausible incomplete repair. We freeze a 12-task symptom-only primary set, a three-task temporal holdout, and an outcome-neutral coding-agent study protocol. A task-level audit finds that 16 of 20 maintainer fixes added no focused automated regression. This paper contributes construction evidence and an evaluation protocol—not agent performance or physical-safety claims. It targets a measurement gap below the policy: whether repository-level repairs preserve the contracts on which zero-shot Physical AI evaluations rely.