Bayesian Decision Making around Experts
Abstract
Complex learning agents are increasingly deployed alongside existing experts, such as human operators or previously trained agents. However, it remains unclear how learners should optimally incorporate certain forms of expert data, or how to quantify the potential benefit from doing so. We study this problem in the context of Bayesian multi-armed bandits, considering offline settings, where a learner receives a dataset of outcomes from an expert before interaction, and online settings, where posterior updates from different data sources may have different computational costs, so the learner must decide whether a given update should use its own experience or an outcome generated by an expert. We formalize how expert data influences the learner's posterior and quantify how pretraining or online learning on expert outcomes tighten information-theoretic regret bounds. We propose an information-directed rule for allocating a limited update budget across data sources, and we study strategies for how the learner can infer when to trust the expert, safeguarding for compromised experts. By disentangling and quantifying the value of expert data, our framework builds towards a practical, information-theoretic understanding of how agents should learn from others.