Do Scientific Inductive Biases Help? Validation for Low-Data Mixture Property Prediction
Abstract
In scientific machine learning, encoding domain knowledge as inductive biases is a widely adopted strategy to enhance generalization in low-data regimes. However, a fundamental methodological ambiguity persists: do predictive gains stem genuinely from the hypothesized scientific mechanisms, or merely from generic representation enrichment that introduces additional nonlinear capacity? In this work, we investigate this distinction in low-data chemical mixture property prediction. We formulate a hierarchy of hypothesis-induced representations and develop a controlled validation method to verify the impact of scientific inductive biases beyond generic enrichment. Across synthetic and real datasets evaluated on Ridge regression and TabPFN, we show that: (1) on synthetic data with known mechanisms, aligned scientific inductive biases yield clear predictive advantages over generic enrichment; (2) on real datasets, performance gains become marginal, noisy, and highly task-dependent; and (3) the modern tabular foundation models like TabPFN exhibit less reliance on handcrafted scientific representations, suggesting that their learned priors partially substitute explicit domain inductive biases.