ICICLE: In-Context Inference for Curve-Learning Extrapolation
Abstract
Freeze-thaw Bayesian optimization allocates training budget by predicting how partially observed learning curves will continue. LC-PFN first applied prior-data fitted networks (PFNs) to learning-curve extrapolation, and FT-PFN extended this approach to configuration-aware freeze-thaw optimization with a ten-slot hyperparameter encoder. We introduce ICICLE, an in-context surrogate with a width-free configuration encoder and a synthetic training prior covering typed hyperparameters, heavy-tailed noise, quantization, and divergence. ICICLE has 1.4 million parameters (9.5% of FT-PFN's count) and produces similar mean-regret trajectories to FT-PFN on LCBench, PD1, and TaskSet. On the 86-task Meta-Album benchmark with 33 hyperparameters, ICICLE reaches lower mean normalized regret than Quick-Tune overall and on mini and extended, with comparable results on micro. It does so without pre-training or meta-training on Meta-Album tasks, whereas Quick-Tune is meta-trained on related Meta-Album tasks.