CURE: Coupled User-Grouped Reinforcement Learning for Cross-Domain Recommendation with Non-Overlapping Users
Abstract
Cross-Domain Recommendation is essential for enabling cross-sell in multi-service platforms; however, limited user overlap and strict cross-service data-sharing constraints often result in little to no ground-truth supervision for cross-domain learning. To bridge this gap, we propose an LLM-based framework for cross-domain recommendation in the non-overlapping user scenario. First, we propose an agentic pipeline that constructs cross-domain pseudo supervision, which is unavailable in the non-overlapping user scenario, by leveraging a user’s source-domain history to generate a target-domain item recommendation. Using this dataset, we further propose \underline{C}oupled \underline{U}ser-G\underline{R}ouped R\underline{E}inforcement Learning (CURE), which induces coupling between in- and cross-domain gradient updates via user-level joint normalization, thereby calibrating synthetic cross-domain updates against reliable in-domain signals. We provide a theoretical analysis showing that user-level coupling stabilizes policy updates under imperfect pseudo supervision and improves cross-domain generalization. Experiments on Amazon Reviews and MovieLens show consistent gains over baselines in cross-domain recommendation, while also improving in-domain performance. An online A/B test demonstrates a statistically significant 56\% lift in click-through rate over existing ML methods.