Stop Assuming Every Dataset Is Governed by a Useful Equation: Rethinking Symbolic Regression for Scientific Discovery
Abstract
Symbolic regression (SR) is increasingly evaluated on benchmarks that combine synthetic equation discovery tasks with ''black-box'' real-world datasets whose governing laws are unknown. Despite this welcome shift, many SR algorithms still implicitly assume that a concise, human-readable equation that can support scientific reasoning and discovery underlies the data. This position paper argues that more research is needed on symbolic regression applied to settings where the optimal model is not in the form of an analytically useful closed-form expression—we call those non-canonical settings. We argue that in many practical settings, this premise may fail: real-world problems may involve non-elementary or experimentally determined functions, high-dimensional inputs, and categorical factors. Moreover, these non-canonical settings have distinctive challenges that are still mostly overlooked by the community. Ignoring them may lead to misguided inductive biases, misleading complexity metrics, brittle extrapolation, and poor support for categorical or high-dimensional inputs. We also outline research directions that could address those challenges: better complexity measures linked to cognitive load or specific tasks, native handling of categorical features, semantic (behavior-level) inductive biases, and shape- or physics-informed constraints. Our aim is to redirect the field toward ensuring that symbolic models provide analytically useful and reliable representations for scientific discovery in realistic settings.