TrustFlow: Adaptive Trust Calibration for Language Model Guided Reinforcement Learning
Abstract
Reinforcement learning (RL) often suffers from low sample efficiency and inefficient exploration in complex environments. Recent work leverages large language models (LLMs) as sources of prior knowledge for sequential decision-making. However, existing LLM-guided RL approaches typically rely on static or loosely coupled integration, failing to account for the evolving competence of the agent during training. As a result, LLM guidance may become redundant or even detrimental, leading to suboptimal learning dynamics, unnecessary dependence at inference time, and increased latency due to frequent LLM interaction. In this work, we propose \textbf{TrustFlow}, a unified framework that formulates LLM–RL integration as adaptive trust calibration. The key idea is to dynamically regulate the influence of LLM guidance based on learning progress, enabling a gradual transition from prior-driven exploration to LLM-free decision-making. Concretely, a preference representation is first constructed to connect sparse LLM rankings with policy representations. Trust calibration then balances LLM guidance with RL exploration. Building on this trust signal, value alignment is enforced, and a trust-modulated optimization scheme is introduced, incorporating preference structure into the learning process. Empirical results show that TrustFlow consistently improves sample efficiency and enables a smooth transition to fully LLM-free policies, eliminating the need for LLM access at inference time across diverse environments.