benchside: Toward Scalable Meta-Science through AI-Assisted Scientific Practice
Abstract
Meta-science depends on evidence about how research is conducted, yet scientific decisions, technical detail, and repeated workflow frictions are poorly recorded in many fields. AI assistants create an opportunity to capture this data as research unfolds. Raw transcripts of conversations with agents, however, do not clearly record what was proposed, challenged, decided, or left unresolved, while unfamiliar interfaces can make AI tools difficult to adopt. We present benchside, a system that packages familiar scientific interactions as reusable AI roles and records hypotheses, objections, decisions, and workflow observations as structured journal entries. The roles aim to make AI support accessible through practices scientists already understand, while the journal makes important scientific transitions and insights explicit, and preserves context across sessions. With explicit, revocable consent, aggregated and non-identifiable workflow observations can be shared to support meta-science research. We evaluate the pi supervision role in nine scripted scenarios and find that 93% of expected moments that occurred are captured. An auditor using the journal and transcript answers questions more accurately than one using a baseline LLM transcript (0.83 versus 0.75) while using 29% fewer tokens. Together, these results suggest a path for helping scientists adopt AI while creating the evidence needed to understand how AI is changing scientific practice and to inform meta-science.