From LLM Control to Environment-Adaptive Agents: Sample-Efficient RL for Real-World HVAC Control
Toki Miyake ⋅ Hiroaki Murakami ⋅ Keiichiro Taniguchi ⋅ Katsuya Koike ⋅ Changyo Han ⋅ Shizuku Iida ⋅ Yoshihiro Kawahara
Abstract
Large language models (LLMs) are increasingly used as autonomous agents for real-world decision making, yet their deployment in physical enterprise environments faces a key challenge: LLM controllers provide useful prior knowledge but do not naturally learn persistent policies through interaction with a specific environment. Reinforcement learning (RL) agents can adapt to environment-specific dynamics, but from-scratch exploration is costly and sample-inefficient. We study this tension in HVAC control, where exploratory actions directly affect occupant comfort and energy use. We propose an LLM-to-agent handoff framework in which an LLM temporarily guides an environment-adaptive RL agent that learns an environment-specific control policy. LLM actions bootstrap early exploration, while critic-based filtering determines when teacher guidance remains beneficial and when the learned policy should take over. In EnergyPlus simulations across two climates, our agent surpasses the LLM teacher with up to 27$\times$ fewer environment interactions than standard SAC. In a real-room deployment, after 15 days of online learning, it improves comfort control over both rule-based and direct LLM control. These results position LLM-to-agent handoff as a practical approach to sample-efficient environment adaptation for autonomous agents in physical enterprise control.
Chat is not available.
Successful Page Load