Constraint Satisfaction Is Not Correctness: Instruction-Blind Rewards for Meta-Agent Supervision
Dipankar Sarkar
Abstract
Meta-agents that supervise tool-using agents can rely on deterministic checks over emitted payloads, because such checks are cheap, uniform across subordinates, and auditable. We show a structural limitation of these payload-only rewards. For two instructions to the same tool with disjoint correct sets, any aggregate whose components depend only on the payload and the tool's fixed constraints induces the same ordering over payloads under both instructions, and therefore cannot strictly separate correctness for both. We demonstrate the consequence with an instruction-blind constructor that reads only a tool's schema, assertions and ontology. Across $18$ tasks in three domains it emits just three distinct payloads, scores $1.00$ on every reward axis and on their aggregate, and is scored $0.15$ by an independent semantic check; scrambling every instruction leaves its output unchanged byte for byte. A prompted language model reaches the same saturation on all $72$ constraint-aware candidates. The result gives a supervisor a simple audit: if no reward component receives instruction semantics or an execution outcome, adding further payload-only axes cannot repair the missing signal.
Chat is not available.
Successful Page Load