Can We Trust Code Paraphrases?
Abstract
Paraphrasing public code benchmarks can help test whether model performance depends on memorized wording, but the audit is meaningful only if the rewrite preserves the original task. We study whether scalable static signals can certify that preservation. Across CodeLingua and LiveCodeBench, three open paraphrase models, and five rewrite degrees, we audit 18,825 original--paraphrase pairs using execution, Qwen3-Embedding-4B cosine similarity, and 56,475 cross-model judge scores. CodeLingua provides an execution-backed reference and reveals 63 broken paraphrases. Yet Qwen, Gemma, and GLM still assign the maximum judge score to 22, 23, and 43 failures, and all three judges unanimously certify the same 13 broken rewrites. Judge scores are strongly saturated, embedding similarity is not a proof of correctness, and both signals are weak predictors of downstream failures. Reliable paraphrase auditing therefore requires execution-grounded calibration and explicit quality gates. More broadly, the results complement recent work on benchmark contamination, robustness to rephrasing, and memorization advantage in code LLMs by showing that paraphrase validity itself must be audited before such claims are interpreted.