Listening Before Leaving: Learning Customer Relationship Deterioration from Voice and Behavioral Event Sequences
Abstract
Listening Before Leaving: Learning Customer Relationship Deterioration from Voice and Behavioral Event Sequences Customer churn is typically treated as a prediction problem over transactional histories, yet churn is often the endpoint of a relationship that has been deteriorating for weeks or months. Customer-support conversations offer a qualitatively different observation of this process by exposing dissatisfaction, unresolved problems, and changing intent. Recent banking foundation models (Ostroukhov et al., 2026; Fadeev et al., 2025) achieve strong performance on transactional sequences but ignore customer-agent interactions, while prior work incorporating call data relies on pre-transformer NLP or static fusion (Vo et al., 2021; Rudd et al., 2023). We ask a more fundamental question than whether voice improves churn prediction: can the temporal interaction between what customers say and what they do reveal relationship deterioration that neither modality captures independently? We formulate churn as latent relationship-state estimation from asynchronous multimodal observations and introduce a temporal cross-modal attention architecture for this setting. Client transaction histories are encoded with a GRU while per-transaction states serve as attention keys; 768-dimensional pre-computed dialogue embeddings are projected into a shared 64-dimensional space and used as queries. Cross-modal attention is restricted to a ±14-day neighborhood, motivated by an empirical property of our data: 91% of customer-support dialogues occur within seven days of a transaction. Learned attention pools these aligned representations into a relationship-aware signal, combined with the client-level behavioral state for classification and naturally reverting to behavioral evidence when dialogue is unavailable. We evaluate on MBD-mini, an industrial-scale open banking dataset containing approximately 950M anonymized transactions, 5M dialogue embeddings, and over 1.5M clients across two years, of which 46,006 have both modalities (median three dialogues per client). Under strict temporal separation between observation and prediction windows, we compare logistic regression, XGBoost, transaction-only GRU, dialogue-only encoders, and naive multimodal fusion. Preliminary validation results show temporal alignment achieves the strongest observed AUC (0.847); dialogue-only models remain below 0.60, and naive multimodal fusion improves over transaction-only baselines by less than 0.002 AUC. Simply adding customer voice provides little predictive benefit as the useful signal appears to depend on when conversational evidence occurs relative to behavioral change. The results expose a second, less obvious property of the problem: customers with observed dialogue churn at 1.7%, versus 4.1% for customers without, suggesting dialogue presence itself is subject to strong selection effects. Rather than interpreting conversation as intrinsically protective, we investigate which conversational trajectories break this apparent protection and whether voice-behavior disagreement can serve as an early-warning signal of latent deterioration. Our contributions are threefold: (1) a formulation of churn as temporal latent-state estimation from asynchronous voice and behavioral observations; (2) a cross-modal temporal alignment mechanism that extracts signal missed by static multimodal fusion; and (3) empirical characterization of when customer voice becomes predictive, including dialogue sparsity, temporal proximity, and voice-behavior disagreement. Ongoing work scales the model to the full MBD population, evaluates cross-dataset transfer, and studies segment-level deterioration trajectories and intervention lead time moving churn modeling beyond who will leave toward detecting how a customer relationship is deteriorating while there is still time to intervene.