Functionally Equivalent or Not? Graph-Grounded Differential Surrogate Execution for Code Equivalence
Abstract
Determining whether two programs are functionally equivalent is central not only to code modernization, patch validation, and refactoring, but also to code-generation evaluation in general. Yet the usual signals are incomplete: tests cover only finite inputs, textual similarity confuses implementation with behavior, and unconstrained language model judgments are difficult to audit. Direct execution is often impossible when a program depends on an obsolete, licensed, unavailable, or unsafe environment. We introduce FEAgent, a selective equivalence assessor agent that combines typed program-graph evidence with differential surrogate execution. FEAgent first aligns public interfaces and behaviorally relevant graph anchors, then issues bounded queries over call-flow, control-flow, data-flow, type, import, and effect relations. With this information in context, a branch-aware generator agent proposes discriminating inputs, and two blinded language-model surrogates independently reason about and predict source and target observables. Every claim and predicted divergence is recorded in an evidence ledger. A deterministic reconciler then returns EQUIVALENT, INEQUIVALENT, or UNCLEAR rather than forcing a verdict when paths are uncovered or evidence conflicts. We specify a pre-registerable evaluation on function-level equivalence, cross-language migration, and repository-level bug patches, including execution-backed adjudication, risk-coverage analysis, component ablations, robustness tests, and a study of false positives and false negatives hidden by unit-test-only scoring. On public benchmarks, FEAgent identifies a candidate false-negative rate of 27% on EquiBench; on SWE-bench Verified, 28.4% of test-passing agent patches exhibited a concrete behavioral divergence from the reference patch. These results position FEAgent as an audit layer between testing and formal verification: it exposes behavioral differences missed by existing oracles, retains evidence for independent review, and makes uncertainty explicit without claiming a universal proof of equivalence.