MetaLoop: Benchmarking the Full Metacognitive Loop in LLMs
Abstract
We introduce MetaLoop, a seven-task benchmark that evaluates the full metacognitive loop in large language models: monitoring one's own uncertainty, translating that signal into action (abstaining, switching strategies, correcting errors), and updating from feedback. Existing evaluations test these components in isolation; MetaLoop produces integrated per-model profiles across all three. Evaluating 12 frontier models alongside 460 human participants, we find that (1) accuracy does not predict metacognitive ability (high-accuracy models routinely approve their own errors); (2) models universally produce calibration signals but fail to act on them; (3) forced self-explanation distinguishes genuine self-monitoring from brittle pattern-matching; and (4) humans and LLMs both show monitoring–control gaps but with different failure modes (humans are loss-averse, models overconfident), with clarification detection as the one clear human advantage. We release the benchmark, scoring code, model outputs, and human data.