Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
Abstract
Unlike code completion, debugging requires localizing faults and applying targeted edits. We observe that frontier LLMs often hack unit tests by regenerating correct but over-edited solutions during debugging. To evaluate how far LLMs are from precise debugging, we introduce Precise Debugging Benchmarking (PDB), a general, dataset-agnostic framework that converts any coding dataset into a debugging benchmark with precision-aware evaluation. PDB generates buggy programs by synthesizing verified atomic bugs and composing them into multi-bug programs. We define two novel metrics, edit-level precision and bug-level recall, which measure how many necessary edits are made and how many bugs are resolved. We release two evaluation benchmarks: PDB-Single on single-line bugs and PDB-Wild on multi-line and repository-level bugs. Experiments show that frontier models, such as GPT-5.1-Codex and DeepSeek-V3.2-Thinking, achieve unit-test pass rates above 76% but exhibit precision below 45%, even when explicitly instructed to perform minimal debugging on single-line bugs. Iterative and agentic debugging strategies do not substantially improve precision or recall, highlighting the need to rethink post-training pipelines for coding models.