FindStatBench: Evaluating Large Language Models on Combinatorial Code Synthesis
Abstract
We introduce FindStatBench, an execution benchmark for evaluating large language models on combinatorial code synthesis. Built from FindStat, the benchmark contains 2,329 tasks across 24 collections and 5.52M hidden instances in two executable families: statistic synthesis (object → integer) and map synthesis (object → object). Each task provides a mathematical description and at most five public input–output examples; a model must emit a single Python solve function, with no retrieval, tool use, execution feedback, voting, or reranking. Submissions are scored by exact sandboxed execution on held-out combinatorial objects. We evaluate eleven systems under this protocol: four closed-source production models and seven open-weight models served through a common inference provider. Three findings motivate FindStatBench as a stress test for symbolic program induction rather than general software engineering. First, the strongest open- and closed-source systems converge within ~1 pp instance accuracy, while an oracle over all eleven systems improves task accuracy by only ~10 pp over the best single model; five-way sampling from one mid-tier model reaches the same ceiling. Second, examples are not uniformly helpful: on several classical named bijections, zero-example prompts produce perfect implementations while five-example prompts collapse to near-zero hidden accuracy and fail even their public examples, suggesting prompt-induced regression away from canonical algorithms. Third, apparent coverage gaps can reflect output-budget mechanics rather than inability: hidden reasoning can consume the visible response budget before code is emitted, a failure mode observed in both closed-source and open-weight reasoning models. Across all evaluated systems, statistic synthesis remains substantially easier than map synthesis, set partitions and binary trees remain near-zero-accuracy collections, and long prompts induce a sharp accuracy cliff. Cost-normalised, the strongest open-weight systems dominate the closed-source production models on this benchmark. FindStatBench exposes a distinctive capability gap: current models can often write plausible mathematical code, but exact symbolic rule induction over structured objects remains brittle across scale, specialization, and weight provenance.