Auditing Evidence Framing at NeurIPS: A Decade-Scale Study of Accepted Papers
Abstract
NeurIPS now publishes a broad mix of algorithmic, theoretical, empirical, dataset, benchmark, systems, and evaluation-infrastructure work. Yet we lack systematic evidence on how the composition of this mix has changed over time. We audit titles, abstracts, and metadata for 19,361 accepted NeurIPS papers from 2015--2024. We define semantic axes for formal-theoretical evidence, benchmark/resource evidence, empirical-performance evidence, artifact-release orientation, scale/capability framing, claim strength, and hedging; score papers with contrastive embedding measures; and validate the constructs with blinded LLM-agent ratings, lexical probes, encoder triangulation, length checks, and topic controls. In raw decade trends, accepted papers shift toward benchmark/resource evidence (+0.842 SD), empirical-performance evidence (+0.750 SD), and scale/capability framing (+0.606 SD), and away from formal-theoretical framing (-0.872 SD). Topic composition explains much of this movement, but embedding-cluster residuals remain +0.246, +0.264, +0.247, and -0.300 SD, respectively. Independent lexical LDA controls preserve the same signs at smaller magnitudes. Artifact-release framing is a useful negative result: it rises modestly in the raw corpus (+0.226 SD) but is near zero or negative after topic adjustment. We release a reusable audit package that includes metadata, semantic and lexical features, validation artifacts, robustness scripts, and an interactive Evidence Norms Atlas for inspecting aggregate evidence-framing patterns by year, topic neighborhood, and track.