CIPHER: A Decoupled Exploration-Selection Framework for Test-Time Scaling of Data Science Agents
Maxime Heuillet ⋅ Sharadind Peddiraju
Abstract
Enterprise data science and analytics spans closed-ended information extraction with verifiable answers and open-ended insight discovery. Agents are evaluated on one or the other, so there is no evidence on whether one system can serve both. We present \texttt{CIPHER}, a data science agent built on a Decoupled Exploration--Selection (DES) framework that applies test-time scaling at the plan level: generating $N$ candidate analysis plans, selecting $M{<}N$ for parallel execution, and aggregating the results. We evaluate $48$ distinct design choices for the DES framework (spanning generation, selection, budget and aggregator capacity) jointly on $357$ tasks from Infi-DA-Bench (closed-ended) and InsightBench (open-ended), on a proprietary and an open-weights base model. A single zero-tuning configuration is competitive against purpose-built specialist baselines under matched models, within $1.2$pp of \texttt{DataWise} and $1.3$pp of \texttt{Agent-Poirot}. The joint sweep exposes that generation and aggregation are most influential performance levers, while random selection is a strong strategy. These results offer new insights on the design of test-time scaling strategies to improve the performance of agents on enterprise analytics tasks.
Chat is not available.
Successful Page Load