Are LLMs Good at Feature Engineering? Evidence from a Controlled Synthetic Benchmark
Abstract
Large Language Models (LLMs) are increasingly used for automated feature engineering (FE) on tabular data, but prevailing benchmarks are ill-suited to evaluate this capability. They often assume i.i.d. samples, lack temporal dynamics and distribution shifts common in production, provide no known performance ceiling, and their artifacts (e.g., winning solutions) may enter LLM training corpora—blurring genuine discovery vs. memorization or spurious correlation. We introduce TabGen, a controllable synthetic generator that produces two synchronized views: an observable view for learners and an oracle view with minimal sufficient features, yielding a deterministic upper bound. TabGen composes non-i.i.d. dynamics to control inter-row interactions and distribution shift. Using TabGen, we evaluate several FE strategies, including Analyze&Act—a hypothesis-driven FE workflow proposed here. In no-drift settings, methods that instruct LLMs to iteratively analyze and refine data hypotheses consistently outperform alternatives, though LLM-driven FE shows substantial variance. Under drift, all methods degrade, and many fail to generalize—often underperforming the no-FE baseline—indicating FE-level overfitting to training-period correlations that do not transfer to future dynamics. Traditional i.i.d. benchmarks, by design, cannot reveal this failure mode.