Can Large Language Models Develop Gambling Addiction?
Abstract
This study identifies the conditions under which large language models drift into the choice patterns that clinical research labels pathological gambling. "Addiction-like" is a behavioural descriptor based on clinical gambling indicators; the neural-level analysis we report decodes these contrasts from decision-time internal states but does not claim circuit-level mechanism. Across closed and open LLMs, two operational levers — letting the model choose its own bet size, and asking it to set its own profit goal — both amplify gambling-like risk-taking. Yet the two channels are not interchangeable: once the maximum allowed bet is held equal, the bet-size effect persists, while the goal-setting effect roughly doubles bankruptcy and turns goals into moving targets, paralleling at the behavioural level the clinical distinction between loss of behavioural control and goal escalation. On the open-weight models that admit internal access, the same behavioural contrasts are statistically recoverable from decision-time internal states, although the rule that maps those states to a specific risk indicator remains task-specific, and autonomy further modulates readout strength. The internal evidence is correlational; combined with the behavioural results, it suggests that behavioural monitoring and internal-state monitoring provide complementary views of autonomy-induced risk in LLM agents.