One Operator, Larger Graphs: Learning to Compare Symbolic Expressions
Abstract
Can rewriting mathematics with a single operation make it easier for a neural model to learn? We study expressions built from E(a, b) = exp(a) − log(b), an operation that can express a broad range of elementary functions. A uniform vocabulary comes at a structural cost: under the compiler studied here, expression trees become about 41 times larger at the median. Sharing repeated expressions reduces this cost, but compiled graphs still contain about ten times as many nodes on average. We then ask whether graph models can learn to recognize equivalent expressions. Small controlled experiments separate missing input information from unsuccessful training. Giving nodes information about their position in the whole graph allows models that previously failed to fit to learn their training examples. In the main comparison, all nine final models fit their training sets, yet their mean accuracy remains near chance on a test that swaps variable roles while matching measured structural features. A separate compiler for bounded rational polynomials ensures that the construction stays within the real-valued domains of its primitive operations; numerical checks provide additional, finite evidence about its implementation. Together, these results show why representation size, successful training, and generalization must be assessed separately when designing learned components for symbolic reasoning. Neural equivalence scores remain suggestions for an independent checker, not proofs.