Machine Collective Intelligence with Canonical Syntax Representations for Explainable Scientific Discovery
Abstract
Deriving governing equations from empirical observations is a longstanding challenge in science, and the discovery of explainable and extrapolatable equations remains a central bottleneck for AI-driven scientific discovery. Despite notable progressed recent LLM-based symbolic regression methods, they typically suffer from sensitivity to initial seeds and difficulty on explainability quantification. In this paper, we propose collective reasoning intelligence for symbolic optimization (CRISO), a collective decision process for fully autonomous discovery of symbolic equations through evolutionary scientific reasoning across multiple reasoning agents. CRISO represents scientific equations as abstract syntax trees (ASTs), enabling a principled quantification of explainability, while group-level knowledge propagation mitigates the local-optima sensitivity inherent to single-agent symbolic regression. Across ten benchmark problems spanning deterministic, stochastic, and previously uncharacterized dynamics, CRISO autonomously recovered the underlying governing equations and achieved state-of-the-art accuracy without any human feedback or additional finetuning. The resulting equations reduced extrapolation error by up to six orders of magnitude relative to deep neural networks, while condensing 0.2--1 million model parameters into just 5--40 constants. Finally, the forward symbolic relation autonomously discovered by CRISO for our in-house chemical reactor was empirically validated on held-out experimental observations, achieving state-of-the-art predictive accuracy.