Test-time Risk Adaptation with Mixture of Agents
Abstract
Deployed reinforcement learning agents often face safety requirements that are specified only after training: new hazard maps, revised risk thresholds, or behavioral alignment constraints. We study zero-update deployment-time adaptation in which a fixed library of risk-neutral source policies must be reused under a newly specified reward–risk tradeoff. We propose TRAM (Test-time Risk Adaptation via Mixture of Agents), a source-scored composition rule that evaluates each source under the target reward and an occupancy-based deployment risk, then selects actions using risk-adjusted source scores. Unlike training-time risk-sensitive methods tied to a fixed surrogate such as return variance, TRAM supports spatial barrier exposure, divergence to a reference behavior, and local volatility risks specified at test time. We make the surrogate nature of the method explicit: TRAM is not claimed to solve the full occupancy-control problem of the stitched policy, but it admits a measurable source-hull mismatch term connecting source-scored risk to realized risk. Experiments in gridworlds, MuJoCo Reacher, Safety-Gymnasium, and an LLM alignment setting show that TRAM improves deployment risk while preserving reward and requires no parameter updates at test time.