Adversarially Attacking Symbolic Vocabulary Vulnerabilities In LLM Planners
Stephen Obadinma ⋅ Xiaodan Zhu
Abstract
As large language models (LLMs) become increasingly deployed as autonomous agents, the ability of these models on automated planning tasks has become essential. However, despite their strengths, LLMs' planning ability is not robust, particularly when they are not allowed to exploit commonsense cues in the symbolic vocabulary (e.g., names of actions, predicates, objects) used to define planning tasks. This makes them unreliable compared to symbolic planners. With recent reasoning models showing improved planning performance on tasks with obfuscated symbolic vocabularies, we utilize an adversarial attack framework to reveal to what extent these crucial weaknesses of LLM agents remain and whether they can reason properly when presented with severe cases. As such, our main contribution is devising and optimizing an LLM-based attack model which we call $\texttt{Symbolic-Swapper}$ to automatically generate obfuscated domains with non-standard symbolic vocabularies that systematically break strong reasoning models' ability to plan. In doing so, we reveal by how much LLMs planning abilities can be further degraded under more advanced obfuscation schemes, allowing us to ascertain their success when having to rely on their pure planning ability. We find that attacks can decrease the planning performance of base models and even robust planning frameworks by over 70\%. We further analyze the factors behind how their planning abilities break down, and find that successful obfuscation models to significantly degrade in reasoning quality despite them showing an ability to understand the domain, revealing a gap in how model's perceive a task domain and how they are actually able to plan under it, with attacks often bypassing model ability to successfully self-correct.
Chat is not available.
Successful Page Load