Treating Hyperparameters as Interventions: Task-Invariant Representation Learning for Transferable HPO
Mengyang Li ⋅ Ou Wu
Abstract
We argue that hyperparameter optimization (HPO) is best approached not as a black-box search problem but as the inverse of an intervention: when training dynamics are observed, $\lambda$ is the controlled cause and the task is what is invariant to it, and the cost of HPO is largely the cost of failing to separate the two. We propose CA-HPO (Configuration-invariAnt HPO), a transfer HPO framework that learns a task representation explicitly constrained to be invariant under hyperparameter interventions and predictively sufficient for performance. We define this target representation as the solution to a constrained variational problem rather than as an identified latent variable, so the framework avoids the strong identifiability assumptions that previous causal representation learning approaches demand. We derive three guarantees in this setting: a population-level result showing that minimizing a predictive loss together with an invariance penalty recovers the target representation, a finite-sample bound on the estimation error of the empirical minimizer, and a counterfactual prediction risk bound that decomposes into representation error, meta-generalization error, and aleatoric noise. Experiments on Llama-3-8B fine-tuning and twelve classical benchmarks show 5 to 15 times speedup over Bayesian optimization and recent meta-HPO baselines, while the learned representations remain stable across a 100-fold range of learning rates, transfer across domains with different hyperparameter spaces, and pass a direct conditional independence test against $\lambda$.
Chat is not available.
Successful Page Load