Small and Smaller: How the Behaviour and Errors of Sub-10B Language-Model Agents Change with Size
Abstract
Small language models are proposed as the default agents for routine work, and the evidence is almost entirely success rates. We ask what changes in an agent's behaviour and in its errors when one model family grows from 0.8B to 9B parameters and nothing else changes. We evaluate four Qwen3.5 checkpoints, each with thinking on and off on the same weights, in the three τ²-bench domains, 6,672 conversations in all. Two independent behavioural analysis frameworks are applied to the first of the three trials, 2,224 conversations, and their outputs joined turn by turn: the ACT-onomy action taxonomy, applied by an LLM coder, and the τ²-bench error reviewer, run with two judge models from different families. We found that the mix of behaviours holds at the level of ten action groups (bias-corrected Cramér's V at most 0.055) but shifts at the level of 46 sub-actions: reflecting, reading memory and deciding under uncertainty fall with size in all three domains, planning and acting on the environment rise. The kind of error shifts more than the amount: comprehension errors fall by an order of magnitude, guideline violations do not fall and are present in six of ten of the 9B's successful conversations. Joined at the turn, comprehension errors concentrate on reflecting and deciding turns under both judges, the sub-actions that fall with size are the ones that carry them, and reflection becomes both rarer and less error-prone as size grows. With thinking on, the agent does much the same things in a different place: over nine tenths of its planning, evaluating, reflecting and deciding is written in the reasoning channel, and its comprehension errors still concentrate on the deciding and reflecting turns, now hidden from the judge. The two judges disagree on individual turns and agree on each of these findings: the mix of tags, its direction along the ladder, and which behaviours carry which errors.