Does Your Model Really Understand Mental Health Data? Auditing and Mitigating Shortcut Learning in Mental Health Language Models
Abstract
Large language models (LLMs) are increasingly used for mental health understanding and prediction, but strong benchmark performance does not necessarily indicate that models rely on the information a task is intended to measure. Models may instead exploit label-correlated signals introduced by dataset construction or collection. To examine this gap, we study this problem by separating shortcut availability in a dataset from shortcut reliance by a model. We audit five settings from four mental health corpora and evaluate shortcut reliance across eight language models using controlled counterfactual interventions. Results show that strong dataset-level shortcut signals do not reliably predict model reliance; they can correspond to broad, selective, or near-zero behavioral sensitivity across models and datasets. On DAIC-WOZ, we further find that shortcut information is widely decodable from hidden representations; yet such decodability does not necessarily imply behavioral use. Finally, we evaluate mitigation strategies spanning input, inference, output, supervised training, and reinforcement learning. None of the evaluated methods consistently reduce shortcut reliance across models while preserving useful behavior. At the same time, some methods achieve low shortcut sensitivity by suppressing task-relevant sensitivity or even collapsing predictions. Our findings highlight the need to distinguish shortcut presence from model use and call for better mitigation methods that generalize across models.