Too Smart to Teach: The Articulability Ceiling in Language Models
Abstract
Larger language models solve harder problems, but do they also produce proportionally better explanations? We study in-context explanation transfer across 40 open-weight models (360M--72B) and 5 receivers (360M--7B), using problems drawn from 13 benchmark sources and evaluated with domain-specific metrics. Our central finding is an \emph{articulability ceiling}: under equal token budgets, mid-range models are better teachers than frontier models. Transfer peaks at 170--220 words and then drops sharply beyond 300 words, falling below even the shortest explanations. Large models partly compensate through verbosity: longer explanations yield higher total transfer on average, but each additional word carries less pedagogical value, so a 360M model delivers more transfer value per word than a 72B model. We also find that RLHF improves pedagogical clarity: instruct models score better on LLM-judge evaluation than base models despite lower lexical overlap, and teacher correctness dominates all other factors. Code and logs are released with the paper.