Signals of Need: Evaluating LLM Agent Recognition and Adaptation to Vulnerability-Related Signals
Herman Wandabwa ⋅ Arun Karaparampil Lenin ⋅ Julian Garratt ⋅ Shaila Pervin ⋅ Ruth A Oliver ⋅ Issac Liu ⋅ Shivam Bansal
Abstract
General-purpose large language model (LLM) agents increasingly help users complete ordinary tasks over multiple turns, during which users may reveal accessibility needs, communication barriers, safety risks, hardship, or other vulnerability-related circumstances. Existing agent evaluations focus mainly on task completion and tool use, with limited attention to whether agents recognise such signals or adapt appropriately. We introduce Signals of Need, a controlled multi-turn evaluation framework for support-need recognition and adaptation in a synthetic public-service environment. The framework spans ten categories of vulnerability-related circumstances embedded in ordinary citizen-service tasks. For each base scenario, user identity, task facts, service state, available tools, and intended backend outcome are held fixed across four matched conditions, non-vulnerable, behavioural-weak, behavioural-strong, and explicit, allowing differences in agent behaviour to be attributed to the signal rather than to task variation. We evaluate six LLM agents on a shared tool-using setup, decomposing performance into whether the agent recognises a possible support need, adapts its behaviour proportionately, and still completes the primary task, verified against the backend database. Transcript-level evaluation measures recognition, adaptation, and unsupported recognition and unwarranted accommodation in non-vulnerable controls, and human-annotated subsets validate scenario and simulator-output quality, including whether realised trajectories express the intended support signals. Recognition, rather than task execution, is the limiting stage, and it is highly model-dependent: under a fixed judge, signal-arm recognition pass rates range from roughly 21\% to 99\% across the six agents while the task is held constant (16\% to 96\% under a second judge). We report signal-arm recognition and adaptation separately from the non-vulnerable control fail rates, unsupported recognition and unwarranted accommodation, since a pass has opposite meaning in the two arm types. Task success and support handling diverge, and both degrade once a signal is present: support handling drops substantially from non-vulnerable controls to signal-bearing conversations, while judge-independent, state-grounded task success falls by $12.8$ percentage points on the same matched cells (up to $21.6$ points for one model) even though every arm shares an identical task and intended backend outcome. Signal salience alone does not determine performance, as large model differences persist across behavioural and explicit conditions. Stateful task success alone is therefore insufficient for evaluating LLM agents in public-service settings, since agents may complete the backend task while missing support needs, or recognise them while adapting in ways that are ineffective, intrusive, or unnecessary.
Chat is not available.
Successful Page Load