A Self-Improving Rubric-Generation Agent for Machine Learning Paper Reproduction
zheng zhang ⋅ Dewei Feng ⋅ Kefei Wang ⋅ Ruijie Xi ⋅ Balakrishnan Varadarajan
Abstract
Evaluating agents on long-horizon scientific workflows requires detailed, task-specific criteria that credit meaningful intermediate progress. PaperBench provides such criteria for machine learning paper reproduction through expert-written hierarchical rubrics, but creating them is expensive and difficult to scale. A rubric-generation agent could automate this process, yet its quality would ordinarily depend on a manually designed prompt and workflow. We ask whether such an agent can instead improve itself from experience. Starting from a minimal single-pass generator, an editor proposes revisions to its shared generation policy while keeping model weights fixed. A validation gate accepts only candidates that improve an objective combining Code Development projected score calibration, full rubric paper-derived requirement coverage, and Code Development matched-leaf agreement. The selected agent is frozen before evaluation on papers held out from optimization. On a fixed split of 20 PaperBench papers, optimization accepts four updates within 30 rounds. We additionally compare against a substantially more detailed author-designed policy, created with AI assistance but without access to development metrics, optimization outcomes, or held-out results. Across 330 held-out submissions, the final agent reduces MAE from $0.1276$ to $0.0943$, raises Spearman correlation from $0.8269$ to $0.8950$, and raises hierarchical coverage from $0.6868$ to $0.7627$. These results show that validation-gated self-improvement can improve a reusable rubric-generation policy that transfers to papers held out from optimization.
Chat is not available.
Successful Page Load