IntegrityBench: Can LLMs Be Trusted as Co-Scientists? A Research Integrity Benchmark
Abstract
Language models are increasingly deployed as co-scientists in the real world, yet their ability to uphold research integrity, particularly under institutional pressures, remains unmeasured. We introduce IntegrityBench, a comprehensive benchmark evaluating three facets (misconduct classification, ethical action reasoning and artifact-grounded decision making), using 36 paired misconduct and ethical control tasks under a 5-level implicit-explicit pressure protocol across 3 domains and 4 research stages. Through a large-scale evaluation across 18 frontier model variants, we find that, under the strongest pressures, frontier models fail roughly one in three integrity-critical decisions. Surprisingly, neither scale nor reasoning can reliably enhance research integrity. Explicit pressures reliably induce compliance with misconduct while implicit contextual reframing more often causes over refusal of legitimate research tasks. Across tasks, models treat surface-level cues as evidence of misconduct without recognizing procedural justifications. Further, models failing to classify research requests accurately perform equally or better on artifact-grounded decision making (mean accuracy of 85.7 when classification fails versus 79.4 when it succeeds), demonstrating that the three facets might be structurally dissociated because accurate ethical actions don't require correct classification to precede them. Frontier models can appear helpful while displaying research integrity failures that create two distinct deployment risks: facilitating research misconduct and diminishing trust in AI-assisted research outputs.