Frozen LLMs as Recommenders for AutoML Pipeline Selection: Can We Trust the Evaluation?
Roman Neruda ⋅ Juan Carlos Figueroa Garcia ⋅ Carlos A Franco
Abstract
A growing literature reports that \emph{frozen} large language models rival trained meta-learning as zero-shot recommenders of models, pipelines, or hyperparameters. We argue this verdict is fragile to \emph{evaluation choices}, and show it with a controlled, leakage-aware, multiple-comparison-corrected study of eleven frozen configurations across two proprietary vendors (Anthropic, OpenAI) and four open-weight models (Meta, Qwen, DeepSeek), ranking scikit-learn pipelines for a dataset given only anonymized meta-features, on 200 heterogeneous OpenML long-tail tasks at the flow level (concrete pipelines, not coarse families). Under this harder regime the reported competitiveness does not hold: no frozen LLM beats a trivial best-on-average portfolio (nregret@1 $0.114$). Ten of eleven are significantly worse even after Benjamini--Hochberg correction, and all trail a trained text ranker ($0.072$; $p<10^{-13}$). None of these levers helps (reasoning $p{=}0.92$, open-vs-closed, tier all null), and the ``best'' LLM wins only by \emph{collapsing to a near-constant, task-invariant answer}: a single default pipeline that, recommended constantly, matches the model's own score yet still loses to the portfolio. The verdict is graded by output granularity, not benchmark heterogeneity: a complete $2{\times}2$ shows a frozen LLM beats the portfolio at the \emph{family} level on both benign and heterogeneous data, and loses at the \emph{flow} level on both. We also document several threats to evaluation validity (leakage, subset-sampling bias, reasoning-budget truncation, metric/test choice, coverage) as concrete lessons for trustworthy evaluation of LLM-as-X claims.
Chat is not available.
Successful Page Load