AgentSSL: Can MLE Agents Leverage Unlabeled Data?
Abstract
Acquiring high-quality labeled data and engineering effective learning pipelines are two critical bottlenecks to applying machine learning systems in new domains. Although recent MLE agents have shown promise in automating pipeline design, their behavior has been studied almost exclusively in fully supervised settings. In many domains, however, labeled data is expensive while unlabeled data is abundant. In this work, we study agents in the semi-supervised learning (SSL) regime by framing SSL as a program search problem and exploring how agents navigate this space under limited supervision. Our analysis focuses on three questions: (1) whether agents can effectively leverage unlabeled data, (2) how they leverage it and what strategies they discover, and (3) whether they can adapt these strategies to specific domains. We find that agents can leverage and synthesize SSL strategies spanning decades of SSL research, achieving performance comparable to, and in some cases outperforming, established human-engineered SSL methods. We further validate these observations on real-world ecological datasets, highlighting both the challenges and opportunities of transferring agent-discovered strategies beyond controlled benchmarks.