Do LLMs Bind Episodes? Probing Cross-Episode Parametric Retrieval Through Shared Cues
Abstract
Large Language Models (LLMs) fail at a core aspect of human cognition: retrieving related episodes through shared elements **. When asked which clubs Ronaldo played for from 2008 to 2010, Claude Opus 4.6 answers Real Madrid, missing Manchester United, despite answering correctly when each period is queried separately. We study this failure through two complementary lenses drawn from cognitive neuroscience: retrieval-based inference (test-time recombination) and integrative encoding (encoding-time reactivation). "In vivo'', we benchmark frontier models (Claude, GPT, Gemini, DeepSeek, Grok, Mistral) and document 4--62\% binding failures on verified knowledge. Extended thinking, a proxy for retrieval-based inference, reduces but does not eliminate failures; multiple-choice probing recovers most of them, suggesting knowledge is encoded but not easily retrievable. "In vitro'', controlled experiments across model families and narrative lengths show that standard remedies offer partial progress: scaling improves binding but still falls short at 70B, extended training enhances binding but requires impractical redundancy, and transfer learning slightly helps with extensive supervision. Mechanistic analysis confirms correct answers appear in the model's top logits yet fail to surface during generation. Paralleling integrative encoding, we design Recall-based Synthetic Supervision (RSS), which prompts the model to recall past episodes and synthesize binding supervision online. RSS covers 92\% of evaluation queries, confirming the information is recoverable from parametric memory. Under realistic sequential training, however, hallucination, forgetting, and evolving supervision signals prevent this coverage from translating into accuracy. We establish episodic binding as an open problem in parametric memory ** Distinct from the perceptual binding problem; our usage draws on the associative inference literature on memory integration (Bunsey & Eichenbaum, 1996; Zeithamova & Preston, 2010).