PaperLens: How Predictable Is Paper Acceptance?
Abstract
Peer review is sometimes seen as noisy and difficult to predict. How much of the final accept/reject decision is recoverable from the paper itself? We investigate this question with PaperLens, text and vision models trained to predict binary acceptance from anonymized paper artifacts. For a clean evaluation, we construct balanced, shortcut-controlled OpenReview and source-derived arXiv datasets that remove deanonymization artifacts and majority-class shortcuts. Supervised binary decision training is surprisingly strong. PaperLens outperforms frontier prompting and review-trained baselines on decision accuracy and ranking, while its acceptance probability remains well calibrated after validation-set scaling and correlates better to ratings than baselines. Vision consistently improves over markdown text, showing that layout, figures, and visual presentation carry reviewer-relevant signal. Scaling from 3B to 14B parameters yields only modest gains, with clean supervision and faithful paper representations mattering more than model size. Together, our results show that decision prediction provides a powerful signal for paper assessment and can even improve the alignment of AI reviewing agents.