Improving Neural Processes in the Low-Data Regime via Context-Subset Training and Self-Distillation
Abstract
Neural processes (NPs) are flexible uncertainty-aware meta-learners, but often fail when meta-training tasks are scarce, losing generalization and collapsing into task memorization. We propose Context-Subset Self-Distillation Neural Processes (CSSDNP), a training framework that trains a student on random context subsets while distilling the full-context predictive distribution of an EMA teacher from the same task. This context-asymmetric objective combines reduced-context predictive learning with forward-KL distillation, enforcing conditional consistency across different context views without architectural or test-time changes. Theoretically, we show that context-subset training learns the reduced-context prediction problem while augmenting each task, and that calibrated full-context distillation preserves the asymptotic Bayes target while improving finite-sample estimation via teacher-guided shrinkage. Empirically, CSSDNP improves robustness and generalization across NP variants and tasks, especially in low-meta-data regimes.