Ablate-and-Measure Breaks When the Measurement Depends on the Ablation
ABHIMANYU PRASAD
Abstract
Attribution methods are validated by ablation: remove a component, re-measure model behavior, and check whether the predicted change occurred. This loop assumes the measurement is available whether or not the ablation was applied. On agentic benchmarks that assumption fails. Per-episode metrics are defined only for episodes that reach a scoreable state, and ablations change how often episodes get there. We document this with a 144-episode context-ablation study on $\tau^2$-bench that we could not interpret. The fraction of episodes yielding a defined grounding rate ranges from 50.0% to 91.7% across conditions; for the benchmark's own action-level score it ranges from 4.2% to 33.3%. Because paired testing retains only cells defined in both arms, our control condition's measured rate ranges from 0.096 to 0.304 depending on which treatment it is compared against, a 21-point swing in a condition that never changed, exceeding the effect the study was powered to detect. One treatment reverses sign between two defensible analyses of the same episodes. We report this as a failed experiment with a transferable lesson: as data attribution extends to agentic behaviors, the metrics being predicted are exactly the outcome-gated kind that breaks this loop, and four cheap diagnostics reveal whether it is happening.
Chat is not available.
Successful Page Load