Absorbed, Not Handled: Final Accuracy Hides Unhandled Batch Effects in a Tool-Using Histopathology Pipeline
Abstract
Agentic imaging pipelines are judged on their final answer, so an intermediate failure that does not move the answer is absorbed: it passes validation and resurfaces under a new distribution. We measure this for batch effects. On Camelyon17-WILDS we score every tool call of a tumor-detection pipeline against hidden labels: a cross-slide hospital-signature probe tells whether a step removed the site effect, an oracle accuracy check whether it destroyed the biology. Over 300 episodes and eight policies, a pipeline that never handles the shift leaves the hospital signature intact in 100% of held-out batches yet answers 72% of them correctly: every correct held-out answer, and 45% [38–51%] of all its correct answers, carries an unhandled shift. Macenko normalization never removes the signature and harms 10% of episodes; a probe-driven agent removes it everywhere but over-corrects 58%, because no probe tells a new slide from a new hospital; only a correction that is harmless when unnecessary is safe. Class-conditional (PRPS-style) alignment is safe: it removes the signature everywhere with zero harm, at the same 0.80 final-answer score as doing nothing. An LLM orchestrator (Claude Opus 5) also removes it everywhere, yet harms 22% of episodes, finishes below doing nothing (0.70 versus 0.80), and fabricates tool observations in 21.5% of steps. Three seeds agree. Final accuracy cannot tell a handled shift from an absorbed one; a site score and harm flag recomputed from released traces can.