Learning to Commit: Next-Commit Prediction via Online Supervised Contrastive Reflection
Abstract
Large language model (LLM)-based coding agents achieve strong results on controlled benchmarks yet routinely produce pull requests that real maintainers reject. The cause is not functional incorrectness but a lack of organicity: generated code ignores project conventions, duplicates internal APIs, and violates implicit architectural constraints. The latest repository snapshot is insufficient-it reveals the final codebase state, not the change patterns that shaped it. We make three contributions. (1) Paradigm: we propose next-commit prediction as a paradigm for learning agentic coding skill from a repository's own history-each historical commit is a self-supervised target whose oracle diff supplies dense supervision. (2) Evaluation: we establish Organicity Evaluation as a measurable objective for coding agents and contribute the Learning to Commit benchmark-5 curated open-source GitHub repositories under strict per-repository temporal splits-scoring patches along file localisation, internal API reuse, patch bloat, and code-style consistency. (3) Method: we instantiate the paradigm with the Learning to Commit framework, in which the agent performs supervised contrastive reflection: it blindly attempts each historical commit, contrasts its prediction against the oracle, and incrementally distils a reusable skill document that conditions subsequent generations. On the Learning to Commit benchmark and on SWE-bench Pro reframed under our temporal-split protocol, our framework consistently improves organicity on held-out future tasks and further lifts test-pass rates on SWE-bench Pro, narrowing the gap between benchmark success and real-world mergeability.