MavaLM: In Search of a Verifiable Automated Database Harness for AI Scientists
Abstract
Research agents are becoming closed-loop scientific systems, but their reliability depends on more than the underlying model. Their actions need a controlled harness, and verifiers must be calibrated rather than treated as ground truth. We introduce MavaLM, a literature-grounded AI Lab for benchmarking a foundational scientific-discovery capability: turning papers into faithful, auditable experimental states. MavaLM combines a generator acting as a research agent with versioned annotation and verification Skills, typed validation, field-level decisions, and provenance. On expert-curated perovskite solar-cell collections, we evaluate record reconstruction, verifier fidelity to experts, and verification policy across model families and scales. Research agents remain least reliable on procedural and condition-dependent evidence. Calibrated verifiers can screen routine fields but should defer ambiguous or policy-sensitive cases to experts. MavaLM evaluates the complete generator--verifier--harness system and provides an auditable release boundary for agent-produced scientific records.