Jagged Competence, Jagged Adaptation: What In-Context Experience Reuse Adds to Frontier Multimodal Agents
Abstract
Agents now adapt without training by reusing a prior episode in context, as retrieval-based procedural-memory systems do. What such a system reports as improvement mixes acquired skill with retrieved answers, and current evalua- tions do not separate the two. Holding the task prompt fixed, we vary only the demonstration block across five multimodal agentic domains, drawn from Visual- WebArena, VLABench, Assembly101, IndEgo and Bench2Drive-VL, and record per item whether the demonstration contained that item’s answer. Competence is uneven inside every benchmark’s own categories, not only across benchmarks. Reusing an episode raises the score in every domain and for all five models, by +2.8 to +14.1pp on the web tasks, and by amounts as uneven as the competence itself; the two patterns do not coincide, so where a model is weak does not predict where reuse helps. For the model with the largest lift, six of the 27 web tasks it newly solved were baseline refusals rather than new capability, and removing them drops its gap closed from 40.3% to 30.6%. Once the items whose demonstra- tion held the answer are removed, the lift survives in two of the four constructed domains. Better retrieval therefore moves such a system toward an answer-present ceiling rather than toward more skill.