Paperena: Reliable Science from Unreliable AI
Abstract
AI scientist systems are automating research workflows from hypothesis generation to paper writing. Their speed and low cost could accelerate discovery and produce valuable research at a scale previously impossible. However, their output may range from important contributions to papers containing reasoning errors, invalid experiments, hallucinated citations, or unsupported claims. Without scalable evaluation, valuable work becomes harder to identify, errors can propagate into subsequent research, and trust in the scientific record may erode. This adds to an evaluation problem that scientific communities already face as research volume grows. To close this evaluation gap, we build Paperena, a flexible framework and platform to record, evaluate, and disseminate research produced primarily by AI. It records papers, process artifacts, feedback, and revisions, and assesses validity and value of papers using specialised verifiers, general-purpose AI reviews, and selective human feedback. These assessments remain available to AI scientists for revision and are aggregated into provisional rankings that help readers navigate large volumes of papers and identify candidates for expert review. We demonstrate Paperena end to end by evaluating ten papers written by and open-source AI scientist system. The evaluations identify validity problems and yield distinct assessments of novelty, interestingness, and potential impact. Its modular design allows additional feedback sources, and aggregation methods to be incorporated, so the assessments can improve as more reliable evaluation methods become available.