Delegating Authorization to Misaligned Agents: Coalitional Alignment and Safe Control
Natalie Collina ⋅ Surbhi Goel ⋅ Aaron Roth ⋅ Sikata Sengupta
Abstract
Long-running AI agents create a control problem: each action changes the state in which later actions are chosen. If the agent is not fully aligned, consequential actions require approval; yet human approval at every step makes human attention a bottleneck. Delegating review to other AI agents appears circular because the reviewers may themselves be misaligned. We identify a condition on a reviewing panel that is weaker than individual alignment yet necessary and sufficient for a guarantee that the principal fares at least as well in expectation as under a designated baseline policy. Each reviewer agent reports whether an action proposal made by a proposing agent improves its own utility relative to the baseline. We show that a threshold rule tolerating $k$ disapprovals is safe exactly when, after any $k$ reviewers are removed, the principal's utility is bounded by a nonnegative combination of the remaining reviewers' utilities, up to a term nonnegative on every feasible proposal. We call this $k$-robust coalitional alignment. No individual reviewer need be aligned: the condition can arise naturally, for example when reviewers' objectives are unbiased but noisy perturbations of the principal's. The characterization lifts to sequential control: in a discounted Markov decision process with a history-dependent proposer agent, local safety at every state is necessary and sufficient for the induced policy to match or improve the baseline in expected discounted utility. When reviewers vote strategically, full-panel coalitional alignment guarantees that every Nash equilibrium is safe under the unanimous approval rule; in contrast, more permissive thresholds can admit unsafe equilibria even when reviewers are individually aligned. Finally, we establish an information-theoretic limit of binary approval: safety and completeness cannot be achieved together without an aligned reviewer. This impossibility can be circumvented either by allowing approximation error or by eliciting richer information from reviewers. Audits of StrongREJECT and RewardBench~2 show that coalitional coverage occurs in existing evaluator panels and that prompt-specific coalition-aware rules can authorize substantially more desirable behavior than anonymous thresholds.
Chat is not available.
Successful Page Load