Do LLM Agents Actually Differentiate? Auditing the Gap Between Textual and Behavioral Heterogeneity in Traffic Simulation
Abstract
Large language model (LLM) agents increasingly represent heterogeneous individuals through natural-language personas, perceptions, memories, and external knowledge. Yet whether such natural-language differences are actually used by the LLM to produce differentiated agent behavior remains insufficiently verified. Consequently, textual heterogeneity alone does not reveal the behavioral influence of each contextual component. We develop a validated component-level counterfactual auditing framework based on paired deterministic replay, with NLI-validated semantic interventions and embedding-space perturbation characterization, and apply it to an LLM-driven traffic simulation. The behavioral audit reveals a pronounced imbalance across contextual components. Persona and external knowledge exert substantially stronger influence than perception and memory. Removing shared external knowledge increases the behavioral sensitivity of all three individual components, showing that shared grounding can substantially reshape how other prompt components affect decisions. Multi-layer residual-stream analysis shows that representation shifts follow the same overall ordering as behavioral sensitivity, yet perception and memory retain a larger fraction of their span-level shifts toward the final prompt token despite their weak behavioral influence. Self-reported influence remains stable across prompt-ordering changes but diverges from intervention-based behavioral evidence, consistently assigning high influence to perception. Together, these findings show that contextual components contribute unevenly and interactively to LLM agent behavior. Prompt-level heterogeneity is therefore an unreliable proxy for behavioral differentiation. Our framework provides a direct way to audit whether contextual differences introduced to represent heterogeneous LLM agents are actually expressed in their decisions.