RxnHaystack: Separating Scale from Chemical Reasoning in Reaction Corpora
Abstract
Modern chemical reaction databases far exceed the context windows of current LLMs, yet no existing benchmark tests whether models can reason reliably over structured chemical data at this scale. We introduce a tiered long-context benchmark grounded in 122,456 cleaned USPTO reactions, with 100 questions spanning structural lookup, property aggregation, reaction classification, and relational reasoning over reaction graphs. We evaluate a plain LLM, a CodeAct agent with Python execution, and a Recursive Language Model (RLM) based on openai/gpt-5-mini. RLM remains near-perfect on lookup and property computation as context grows, whereas LLM and CodeAct degrade sharply. More importantly, decomposition does not remove the need for chemical reasoning: at full-corpus scale RLM reaches only 0.39 on mechanism classification and 0.05 on prospective and multi-constraint tasks despite scoring 1.00 on mechanical graph traversal. The benchmark therefore separates failures caused by scale from failures of chemical abstraction and forward synthesis reasoning.