PyMETA: Evaluating Student Code Diagnosis on and Beyond the First Execution Error
Abstract
Code-diagnosis evaluation depends on how error supervision is constructed and what that supervision treats as a correct diagnosis. We introduce PyMETA, a dataset of 48,646 Python submissions with execution-derived single-error labels, and a targeted subset of 97 submissions annotated by experts through repair and re-execution. Under the first-outcome protocol, four recent prompted models obtain 87.5-93.8\% macro F1, all above the strongest finetuned baseline at 80.6\%. This reverses the comparison obtained with the earlier set of prompted models. The expert subset gives a less favorable account of complete diagnosis: exact-set match ranges from 43.3\% to 48.5\%, although sample F1 is approximately 79-81\%. The evaluation setup also changes which behavior is visible. On the same 45 audited submissions without an expert-annotated Logic Error, none of the four recent models predicts that label when restricted to one output, whereas 46.7-57.8\% include it under multi-error prompting, usually after an explicit-error label. These findings show that model comparisons and error-bias claims depend on what the gold label represents, which outputs the model is allowed to return, and how those outputs are scored.