Deployment-Memory LLM Test-Time Training Should Require Behavioral Evidence Beyond Perplexity
Abstract
Large language model test-time training (TTT) is often evaluated through local proxy metrics: models are updated on recent tokens, retrieved context, target-domain data, or verifiable task attempts, and then judged by perplexity, future-token loss, long-context performance, or reward. These metrics are well matched to claims about stream adaptation, domain adaptation, context compression, and reward-backed test-time improvement. This position paper argues that TTT has a distinct claim-calibration problem: proxy evidence for local adaptation can migrate into stronger claims about deployed assistant memory, personalization, or sparse post-deployment learning. Such claims require behavioral evidence: later recall, paraphrase robustness, retention, locality, conflict handling, and use in downstream actions after the original support context is removed. We propose a claim-calibrated evidence ladder for LLM TTT, audit recent work through this lens, and provide a diagnostic counterexample in which proxy losses improve without behavioral recall. Our goal is to give authors and reviewers a standard for aligning TTT memory claims with the evidence actually reported.