Cross-Distribution Generalization in Longitudinal Behavioral Data Through Frozen Coherence Constraints
Abstract
Cross-distribution generalization in longitudinal behavioral data remains a persistent obstacle. Models trained on one cohort degrade on new populations, and prior approaches operating through feature-distribution alignment or model-capacity scaling have not substantially improved cross-distribution performance. We diagnose the limitation as insufficient representational constraints during training. We propose holding the language model frozen, so that the coherence criterion is defined by its pre-trained language prior and cannot shift during training. We introduce PRISM, a frozen-backbone framework that decomposes each behavioral trajectory into temporal, spectral, and semantic streams and integrates them through directed cross-attention with instance-specific gating. Training uses a dual-path objective combining discriminative and coherence constraints on a shared representation. On GLOBEM, a widely used cross-cohort benchmark for behavioral health prediction, PRISM achieves 79.93\% out-of-distribution accuracy, 27.13 points above the previous best, with an in-distribution to out-of-distribution accuracy gap of 1.24 points. PRISM also achieves the highest OOD accuracy on three additional datasets (LifeSnaps anxiety, MFAFY engagement, CrossCheck schizophrenia symptoms), with consistently smaller ID-to-OOD gaps than fine-tuned VLM baselines. Ablations identify the frozen coherence constraint as the component responsible for the distribution-invariance pattern.