LLM-Generated Investment Research Memos Need Validation
Abstract
Large language models (LLMs) are increasingly used to generate long-form analytical memos, yet their reliability is often summarized through claim-level factuality. This paper studies whether that level of evaluation adequately represents the risk faced by a reader who consumes a memo as a whole. We propose a source-bounded evaluation protocol that extracts material claims from LLM-generated investment research memos, aligns them with company-period evidence, and verifies them against SEC filings, earnings releases, and other official financial disclosures. The evaluation covers 30 public companies across 12 quarterly reporting periods and three memo-generation agents, yielding 1,080 memos and 58,093 extracted claims. Only 1.77\% of claims in Codex-generated memos, 1.28\% in Claude-generated memos and 1.17\% in DeepSeek-generated memos are explicitly refuted. However, these claims occur in 36.94\%, 40.56\%, and 40.28\% of the corresponding memos, respectively. Thus, a low local refutation rate can coexist with substantial memo-level exposure to factual error. The result highlights the importance of evaluating long-form LLM outputs at both the claim level and in terms of the likelihood of encountering at least one material contradiction in a complete artifact.