Do Molecular Agents Reason, or Merely Follow Representations? Evaluating Scientific Workflow Equivalence and Invariant Error Recovery
Mehmet K ALBAYRAK
Abstract
Autonomous Large Language Model (LLM) agents are rapidly transitioning from conversational assistants to co-scientists in the molecular sciences---orchestrating synthesis planners, docking simulators, and property predictors. Yet, a fundamental question remains unaddressed: do chemical agents execute stable, representation-invariant scientific reasoning, or are their procedural workflows fragile artifacts of surface string representations? In physical chemistry, graph-isomorphic transformations (e.g., SMILES atom re-indexing, re-rooting, synonymous nomenclature, or candidate reordering) define identical physical problem instances ($x \sim x'$). A scientifically grounded agent must exhibit Scientific Workflow Equivalence ($\tau(x) \equiv_{\text{sci}} \tau(x')$): maintaining invariant hypothesis structures, causal tool dependencies, and nominative conclusions across equivalent representations. We introduce the Molecular Workflow Isomorphism Benchmark (MWIB), the first benchmark, to our knowledge, that lifts molecular graph equivalence to agentic workflow equivalence and evaluates procedural invariance alongside biophysical error recovery. Evaluating open-source foundation agents (Llama-3.1-8B, Mistral-7B, Qwen2.5-7B) alongside 2026 Frontier Reasoning Models (Gemini-3.7-Flash, Claude-Sonnet-5, GPT-5.4-Nano) across 240+ multi-step trajectories reveals widespread Procedural Divergence: shifting from canonical to randomized SMILES degrades Scientific Workflow Consistency (SWC) down to 42.1% ($p < 0.001$, Cliff's $\delta = 0.68$) and triggers up to 51.7% decision flips on chemically identical matter. Furthermore, when exposed to controlled tool corruptions (e.g., unphysical molecular weights or inverted lipophilicity), default agents exhibit severe Error Propagation (Error Recovery Rate $< 25\%$), blindly accepting corrupted observations without invoking invariant verification. To address this fragility, we propose a Verification-First Equivariant Agent Architecture that enforces canonical graph projection and invariant check gates, restoring workflow consistency to $>96.5\%$ and error recovery to $>95.0\%$. MWIB provides an open, deterministic testbed for building rigorous, representation-invariant molecular co-scientists.
Chat is not available.
Successful Page Load