EAU: A Meta-Agent for Environment-Agent-User Co-evolution
Abstract
Training LLM agents, e.g., in retail scenarios, to negotiate with and persuade users toward successful task execution is far from trivial. It faces three main parts: environments are often malformed and cannot represent every state faithfully; simulated users are biased and stochastic; and optimization relies on hand-crafted reward functions that impose a single objective, whereas real tasks are inherently multi-faceted and dynamic. Together, these factors make user-centric multi-turn reinforcement learning (RL) highly unstable. Prior work addresses this from the algorithm side, through reward shaping or denser supervision, while leaving the environment and the user simulator fixed. We instead make them objects of optimization, and propose a meta-agent that supervises training and adapts all three components, achieving an environment-agent-user co-evolution (eau). Crucially, this co-evolution operates across two temporal horizons: intra-run, where the meta-agent dynamically adjusts the training setup on the fly; and inter-run, where experience inherited from prior training episodes, such as which user behaviors caused failures or which reward signals proved most effective, are carried over to bootstrap subsequent runs, enabling cumulative learning rather than restarting from scratch each time. We instantiate this co-evolution through a neuro-symbolic architecture. The environment is likewise represented symbolically, making its dynamics transparent and modifiable. The agent itself remains a neural LLM optimized via RL. The user is governed by a symbolic rule-based module whose internal state, including intent, disclosed information, and preferences, is fully observable and editable. This unified symbolic backbone allows the meta-agent to read training trajectories and directly adapt all three components, covering environment rewards, RL hyper-parameters, and user rules, via hot-swapped modifications without pausing the ongoing run. Across 2 benchmarks, eau improves over the previous state of the art by up to 36.7\% success rate, and shows that adaptation of the training setup can happen intra-run rather than only between runs.