Agentic Retrieval for Bio-Based Recovery of Critical and Near-Critical Minerals
Abstract
Biological extraction routes are promising for critical-minerals recovery. We present a Large Language Model (LLM)-assisted autonomous literature mining agent that converts a large local PDF corpus into a schema-validated evidence dataset. The agent maintains persistent state in SQLite, applies staged gating, and invokes Claude LLM only on a small shortlist for figure/table-aware extraction. In one run, the pipeline produced 56 validated records. The dataset is dominated by bioleaching and biosorption, frequently targeting Cu, Ni, Zn, Co, and Fe. In the bioleaching of Cu subset (N=16), an analysis shows a dominant ore/tailing–bacterial Fe/S oxidizers–inorganic acid/ferric pathway plus smaller waste-context pathways with organic-acid signatures. This work demonstrates a cost-efficient pathway from unstructured PDFs to a machine-readable knowledge base that can accelerate evidence synthesis and experiment design for bio-based critical materials recovery.