Bounded Reasoning: Cognitive Hierarchy in Human-versus-AI Cyber Defense
Zahra Aref ⋅ Sheng Wei ⋅ Narayan Mandayam
Abstract
Human-agent evaluations often compress interaction into a single performance score, even when human and automated policies adapt differently over time. We study this issue in a sequential cyber-defense game on an attack graph, where a human or reinforcement-learning defender protects cloud assets against a Deep Q-Network (DQN) attacker. The study compares four defender settings: a human reward-only game that operationalizes the DQN information structure, a human reward-plus-transition game that operationalizes the Cognitive Hierarchy Theory-driven DQN (CHT--DQN) information structure, an automated DQN defender, and an automated CHT--DQN defender. In the reward-only game, participants receive payoff and reward feedback. In the reward-plus-transition game, participants receive the same information plus attacker-aware transition probabilities derived from the CHT--DQN model. Across 80 Mechanical Turk participants and matched automated simulations, human defenders adapted to outcomes in a way the automated defenders did not: they were more likely to reselect a node after a successful defense than after a failure, in both games, an asymmetry we interpret as consistent with Prospect Theory (PT) and Cumulative Prospect Theory (CPT). The 40-round average also mixes early rounds, in which the DQN attacker bot acts mostly at random, with late rounds, in which it mostly exploits; in the final stage, mean protection ranks $\Hfore > \Hrew >$ Auto-CHT $>$ Auto-DQN, an ordering the overall mean hides. By contrast, the reward-plus-transition game does not produce a statistically reliable overall gain in weighted data protection over the reward-only game. These results suggest that evaluating human-agent cyber-defense systems only by an averaged task score can miss behaviorally meaningful differences in adaptation, action allocation, and bounded human reasoning.
Chat is not available.
Successful Page Load