Extraction Is Representation: Auditing LLM-Generated Causal Maps of Science
Caleb Tan
Abstract
Large language models (LLMs) can turn scientific papers into structured claims at scale, but may represent a literature differently from expert curators. We introduce AIClaim, a source-linked evaluation framework that separates four properties of “extraction quality”: evidence traceability, repeated-run stability, convergence with expert curation, and downstream graph structure. Three frontier model families each processed 375 full papers three times under a common schema, producing 52,348 causal-claim records. Evidence quotations could be relocated for 75.3–83.3% of claims. Relation-compatible semantic $F_1$ across repeat runs was 0.516–0.634, but only 24.4–33.0% of claim clusters appeared in all three runs. Convergence with the expert-curated CHIELD database was 0.054–0.068; once endpoints aligned, however, direction agreed in 85.7–90.5% of cases. Post-hoc vocabulary alignment improved literal compatibility without resolving the representational gap. Text-proximate, paper-specific variables produced more fragmented maps than CHIELD’s reusable constructs. Some adjudicated differences reflected purpose-relative levels of abstraction, while others remained errors or annotation-policy disagreements. We argue that evaluating scientific AI requires provenance checks, repeated extraction, and graph-shape metrics to expose how scientific knowledge is transformed into machine-generated representations.
Chat is not available.
Successful Page Load