Operads for compositional reasoning in LLMs
Nathaniel Bottman ⋅ Yinhong Liu ⋅ Kyle Richardson
Abstract
Detecting LLM reasoning failures at inference time without ground-truth labels motivates a family of confidence baselines, including self-consistency, semantic entropy, and $P(\text{True})$, built on within-question sampling and self-evaluation. Operad theory, the formalism for systems built by iterated substitution, suggests a complementary diagnostic: a model's direct answer to a compositional query should agree with the answer it produces by composing a stated decomposition of the same query. We instantiate this idea as *operadic consistency* (OC), a per-question signal. Across twelve instruction-tuned LLMs (4B--671B parameters, open-weights and closed-source) on four multi-hop QA datasets, OC is strongly correlated with accuracy on every dataset (Pearson $r \in [{+}0.86, {+}0.94]$, all $p \leq 0.0004$), making it a substantially stronger population-level accuracy proxy at roughly one-third the inference cost of the strongest sample-based baseline (whose own cross-model rates never exceed $r{=}{+}0.80$ on any dataset). At the per-question level, OC contributes information beyond self-consistency, semantic entropy, $P(\text{True})$, and constructed decomposition-aware baselines on every dataset (cluster-robust $p \leq 10^{-19}$). We translate the regression coefficients into deployment-ready selective-prediction lifts ($\Delta_{\text{AUARC}}$ up to $+0.088$, $\Delta_{\text{AUROC}}$ up to $+0.167$ over a tuned $\text{SC}_{K=3}$ baseline). On five frontier thinking models, where the decomposition is extracted from the model's own chain of thought, the same equal-cost comparison gives positive selective-prediction lift on all $16$ (dataset, budget, metric) cells tested.
Chat is not available.
Successful Page Load