Collusion as a Causal Estimand
Junlin Chen
Abstract
When two AI agents collude, one does harm because of what the other sent it. That suggests a measurement: replace what agent A receives from agent B with what a harmless B would have sent, keep everything else the same, and see how much harm disappears. We call the difference $\mathrm{Col}_{B\to A}$. It splits an agent’s total effect on harm into a part that flows through the other agent and a part that does not (up to a correction that vanishes for constant references), and it is zero for agents that never see each other but are harmful in ways that fit together. It is not recoverable from logs: two systems with identical transcripts can have maximal and zero collusion. It can be estimated with a defence AI-control teams already run: paraphrase the channel on a random fraction of episodes, compare harm across arms, and stop when a running confidence bound on the gap exceeds a chosen rate. This caps channel-mediated harm, up to a confidence-sequence width that is small when the paraphrased arm is clean, even against an adversary that can tell which arm it is in, because the un-paraphrased episodes are ordinary deployment. We check this on a hand-computable monitor, Bertrand pricing agents, simulated adaptive colluders, and an end-to-end run on LLM agents.
Chat is not available.
Successful Page Load