Auditing Pairwise Equivalence Judgments: Self-Critique Effects and Diversity Measurement in Multi-Agent Hypothesis Generation
Abstract
Multi-agent systems built on large language models (LLMs) are increasingly applied to scientific discovery and hypothesis generation. Both the effect of refinement and the diversity of the delivered set are hard to interpret before experimental ground truth exists, and both are typically reported by deciding, for pairs of generated hypotheses, whether they describe the same underlying mechanism. We study two evaluation questions that rest on this pairwise equivalence judgment: (1) how much self-critique changes delivered hypotheses beyond run-to-run variability, and (2) how the equivalence rule used to group hypotheses affects measured diversity. Our study workflow generates mechanistic explanations represented as causal chains from a molecular initiating event through intermediate biological steps to an observed outcome, revised through self-critique rounds before a final consolidation step. Across four independent proprietary instances, we hold opening hypotheses fixed, rerun the downstream workflow with 0, 1, and 5 critique rounds, and score matched hypothesis pairs with an LLM-as-a-judge. Relative to matched same-depth reruns, moving from 0 to 1 round produces 34.5 percentage points (pp) of additional mechanism-level divergence, whereas 1 to 5 rounds adds 1.3 pp. We then compare three equivalence rules: term frequency--inverse document frequency (TF--IDF) similarity, dense embeddings, and the same LLM-as-a-judge. To characterize how these rules respond to specific edits, we construct controlled hypothesis pairs that either alter only wording or equivalent biological terminology while preserving the causal explanation, or replace one component of the causal chain while holding the remaining structure fixed. All three rules are invariant to meaning-preserving edits, but when the initiating event is replaced, the LLM-as-a-judge identifies 83\% of valid pairs as different mechanisms, versus 0\% for TF--IDF and 8\% for embeddings; varying only the rubric that defines same mechanism moves this figure from 38\% to 96\%. Together, these results show that pairwise equivalence judgments are a measurement choice: how mechanism equivalence is defined affects both the estimated effect of self-critique and the measured diversity of generated hypotheses.