From Clicks to Intent: Distilling LLM Intent Understanding into Compact, Deployable Models in Financial Services
Abstract
Large Language Models (LLMs) can extract meaning from noisy, unstructured data that traditional methods struggled to interpret, including the raw user clickstream data that industrial platforms collect at scale. Such models can infer user intent from behavioral logs, but their size, latency, and inconsistency make them impractical to deploy directly in high-volume, regulated settings. In addition, existing production systems for capturing user intent are often ad hoc and narrow, lacking the flexibility to support both quantitative downstream recommendations and qualitative understanding at scale. In this work, we present a system that distills the intent-understanding capability of an LLM into lightweight, deployable models suited to these constraints and demonstrate its applicability for personalization in financial services. A self-supervised Transformer encodes raw web clickstreams into a 64-dimensional session level embedding. An LLM then builds an interpretable intent taxonomy and labels each session. These labels are distilled into a lightweight classifier that runs directly on the embedding, with no LLM calls or text encoding at inference. The model produces two outputs: the dense embedding drives quantitative prediction, while the distilled labels add human-readable intent for explainability and governance. In production, the distilled pipeline runs orders of magnitude faster than direct LLM inference, cutting the compute cost of large-scale annotation with only a 7\% drop in labeling quality, and it scales to millions of sessions per day. On downstream tasks, the session embedding improves macro Recall@1 by 1.88\% and cuts Log Loss by 13.38\% over baselines on the mobile homepage tile ranking task, and it outperforms the LLM's labels by 4.3\% micro F1 on the user conversion prediction task. This results in two compact models that run at production scale while keeping the accuracy and interpretability that a regulated financial services deployment requires.