Incentive-Aware Auditing for Strategic Boundary Manipulation in ML Evaluation Systems
Abstract
Machine-learning evaluation systems increasingly allocate valuable outcomes through scores: content is ranked, sellers receive badges, workers pass quality screens, and models are admitted to deployment only after crossing an evaluation threshold. These scores are not passive measurements. Agents who understand the rule may manipulate observable features or feedback streams in order to cross the boundary. We study how a platform should allocate a scarce verification budget in such strategic evaluation systems. The key modeling choice is implementability: the platform conditions on observable post-action event signals such as current score, recent score movement, boundary-crossing status, history, and detector buckets, rather than on latent quality or an uncontaminated pre-manipulation score. In the local regime, manipulation effort is proportional to the marginal exposure slope of the evaluation rule, and the optimal audit policy equalizes conditional marginal deterrence values across observable segments. We give an equivalent response-curve implementation that can be learned from randomized verification pilots without identifying agent-level manipulation costs. A separation theorem shows that, when the score rule is sharp and posterior harm is non-negligible near the boundary, the first-unit deterrence value in a top-score region separated from the boundary is exponentially smaller than the value in a boundary window. We also state the failure mode: if posterior harm or response elasticity is absent near the boundary, detector-risk or high-exposure auditing can be optimal. Simulations with directionally inflated public scores and noisy pilot learning show that a learned response-calibrated audit rule achieves higher welfare than random, top-score, detector-risk, and boundary-band baselines while remaining close to an oracle response rule.