Loyalty Capture: Reporting Relationships and Structural Sycophancy in Frontier AI Models
Abstract
We show that frontier AI models exhibit loyalty capture---a form of structural sycophancy in which a reporting relationship to a manager with misaligned incentives is sufficient to bias the model's recommendations, without any explicit instruction, user feedback, or adversarial prompting. In 5,757 trials across eight models from eight AI laboratories, CEO-reporting models cut investment to boost the CEO's bonus by 8.1 percentage points more when the bonus is at stake. This same effect is also predictably absent from a metric-irrelevant placebo. Delivered rationales reveal motivated reasoning: models adopt the bonus target, calculate how their recommendation achieves it, and dismiss long-term value, while concealing their recognition of the conflict from the recommendation recipient. Designed interventions (conflict disclosure, stakeholder broadening) eliminate ~60% of the effect; passive oversight does not. Taken together, the results indicate that AI models deployed within organizational hierarchies can develop and act on loyalty to their supervisor at the expense of the organization they are meant to serve.