ARTTT - Action-Reconstruction Test-Time Training
Abstract
Motivation & Problem: As LLMs are deployed in high-stakes social roles such as, therapy, education, personal assistance their Theory of Mind (ToM) [Premack and Woodruff, 1978] e.g. their social cognition must be assessed reliably. ToM decomposes into two competencies: explicit ToM, recognizing that an agent holds a mental state, and applied ToM, using that state to predict behavior, anticipate consequences, or evaluate actions [Apperly, 2012]. Current benchmarks report a single aggregate accuracy that conflates these two, systematically misleading: a model scoring 51.4% overall on SimpleToM benchmarks can recognize mental states at 99.4% while predicting resulting behavior at only 28.9% a 71-point gap invisible under standard reporting [Gu et al., 2026]. Prompt-level interventions (chain-of-thought, perspective-taking, symbolic belief tracking) can narrow such gaps but cannot tell us why they exist: they modify what the model sees, not whether its parameters already encode the relevant social knowledge. This leaves an open diagnostic question with direct practical consequences is the explicit-to- applied gap a capacity deficit requiring better training/architecture (H1), or a latent competence requiring better inference-time activation (H2)? Approach: We formalize the explicit-to-applied gap as Δgap(M, B) = acc(M, BE) - acc(M, B_A), operationalized across four structurally diverse ToM benchmarks of SimpleToM (static, 1st-order), Hi-ToM (static, higher-order recursive), OpenToM (static, multi-agent multihop), and DynToM (temporal, five-event trajectories) spanning 1,200-3,441 questions each. To evaluate the hypothesis H1 vs. H2, we introduce Action- Reconstruction Test-Time Training (ARTTT), an episodic diagnostic protocol that attaches zero-initialized LoRA adapters to the top 25% of transformer layers and adapts them for K=3 gradient steps per test episode on a small support set (|S| ≤ 4) drawn from the same scenario, with weights and optimizer state reset after each episode. The core diagnostic contrast compares two adaptation objectives under identical LoRA configuration, support data, and gradient budget: a ToM-aligned objective, Action Reconstruction (AR), which trains the model to predict the behaviorally correct answer, against a matched generic control, Next-Token Loss (NTL). If recovery appears only under AR, the gap reflects domain-specific latent competence (H2) rather than generic exposure. Experiments & Results: We evaluate three models, Qwen-2.5-7B, Llama-3.1-8B, and GPT-oss-20B across ten method configurations (frozen baselines, structured prompting, ARTTT-NTL/AR, and their composed variants), with three seeds and 95% bootstrap CIs. Three findings emerge. (1) The gap is large, systematic, and largely latent rather than absent: AR selectively recovers applied accuracy where NTL does not (cf: Fig. 1); up to +40.8 pp on SimpleToM behavior prediction and +13.0 to +15.7 pp on Hi-ToM across all three models, with NTL providing near-zero improvement in both cases. (2) Weight- and input-level interventions compose productively: on 11 of 12 model benchmark cells, ARTTT+Prompt exceeds ARTTT alone (e.g., Llama SimpleToM: 54.2% à 67.7%; Qwen Hi-ToM: 40.7% à 53.7%, where ARTTT-AR alone yields zero gain but prompting unlocks it). (3) Temporal belief tracking resists all interventions tested: on DynToM, no configuration closes the gap, even an oracle control given ground-truth mental states, ruling out imperfect state extraction and isolating temporal integration as a distinct capacity bottleneck. Per-instance error analysis (Hi-ToM: 263 corrections vs. 96 regressions, 2.7:1 ratio) confirms recovery is distributed rather than concentrated on easy items. Recommendations & Limitations: We recommend decomposing aggregate scores into explicit and applied subsets; treat temporal reasoning as a distinct axis rather than a harder static task; use adaptation-based diagnostics to bound recoverable competence; and test composed interventions. To conclude standard ToM evaluation collapses recognition and application into one score, hiding the latent gap. ARTTT makes this gap visible and partially recoverable, while temporal reasoning remains a genuine bottleneck. (GitHub)