Same Algorithms, Different Scientist: Design Space Bias in AI Discovery
Abstract
AI-scientist evaluations treat human-written design spaces—menus of methods, code templates, or configuration grammars—as the opportunities a model could discover. We ask whether measured scientific competence is invariant to rewritings that leave every executable option unchanged. We introduce the executable behavioural quotient, which groups written policies whose execution on frozen probe states is identical, and use it both as an invariance test and as a unit of discovery credit. Across 12 combinatorial protein-fitness landscapes, 40–41 executable policies collapse to 21–22 distinct campaigns; 19 written options implement one algorithm, and 35% of the menu is behaviourally redundant. Nevertheless, equivalent rewritings move model choices by a median total-variation distance of 0.417, against a 0.144 same-prompt replicate floor, and reverse 56 of 255 credit verdicts. A separately sealed knapsack grammar reproduces order sensitivity across 24 permutations and 6 problem families. Mechanism analyses identify separable positional and textual biases: probability falls with menu position even after controlling for utility, while the shared textual ordering is negatively associated with utility. Quotient normalisation removes the representation instability but does not improve discovery. Scientific-agent evaluations should therefore define search, comparison, and credit over executable behavioural classes rather than the written menu that names them.