Explainability-Guided History Distillation for Efficient Sequential Recommendation
Abstract
Sequential recommendation models such as SASRec and BERT4Rec rely on self-attention over a user's full interaction history, incurring an inference cost that scales quadratically with sequence length. While efficient attention variants and heuristic truncation mitigate this architectural footprint, they overlook internal model attribution signals and neglect the efficiency gains achievable by distilling the raw interaction sequence itself. To bridge this gap, we propose explainability-driven sequence distillation, a retraining-free framework that leverages post-hoc explainability methods to score each interaction's predictive contribution. By retaining only the highest-scoring time steps, our framework compresses user histories offline, enabling the resulting shorter sequences to accelerate online inference without modifying the underlying model. Across diverse backbones, attention mechanisms, and dataset configurations, our approach preserves or improves ranking performance while substantially reducing inference GFLOPs relative to recency-based truncation, achieving ranking quality improvements exceeding 25\% on datasets lacking strong recency bias. Crucially, even in settings with extreme temporal recency bias, our approach matches or closely tracks recency-based truncation in efficiency and performance at equal computational budgets. These findings demonstrate that post-hoc explainability, conventionally restricted to model interpretation, serves as a general, retraining-free principle for compressing raw input contexts and reducing online serving costs in sequential recommendation.