Learning to Ask: Metacognitive Action Policy for Large and Small Language Model Collaboration
Abstract
Edge-cloud collaboration between large language models (LLMs) and small language models (SLMs) offers a promising way to leverage the problem-solving capability of LLMs while keeping private data on device. Existing methods typically have the LLM generate guidance from an initial user query, which the on-device SLM combines with local private data to produce the final response. However, the LLM's general guidance often does not align with the user's personalized needs, limiting the utilization of the LLM's capabilities. It has been observed that the on-device SLM exhibits metacognitive capabilities, which allow it to recognize its limitations and adjust its reasoning process. Motivated by this, we transform the on-device SLM into a metacognitive inquirer agent that can assess its internal state and actively decide when, what, and how to query the LLM for personalized guidance. Specifically, we propose Active Inquiry via Metacognitive Actions (AIMA). AIMA learns an on-device policy that jointly optimizes the selection of high-level metacognitive actions and the action-conditioned query refinement. We further introduce a privacy-constrained rewriting mechanism that detects and eliminates sensitive information leakage in the refined query. Extensive experiments and theoretical analysis demonstrate that our method outperforms existing methods.