From Accuracy to Sendability: Evaluating an Enterprise AI Agent in Financial Index Client Service
Abstract
Evaluations of document-grounded LLM agents usually measure retrieval and factual accuracy. In a regulated expert workflow, those metrics miss the question: can the answer be shared with customers directly? We report a practitioner's study of an agent that drafts client replies for a client-service team at a financial-index provider. It reads published index methodology but has no access to live data or client configurations, and an expert reviews every draft before a client sees it. Over two months we collected 376 real client inquiries and decomposed a screened subset into 363 subquestions, of which 70 could be answered from published methodology alone. Three experts graded overlapping subsets of the drafts on three questions: did the agent find the right document, the right provision, and could the draft be sent with only minor edits? Retrieval proved to be the lesser challenge. The agent identified the correct document in 85--95\% of cases and the correct provision in 82--92\%, yet only 26--64\% of drafts were sendable with minor edits. The agent hedged conclusions the firm had already settled, substituted rules from a different methodology when the governing document was silent, and answered hypothetical or third-party questions an expert would decline. Our reviewers agreed closely on whether retrieval was correct, but diverged considerably on whether a draft required major or minor edits before it could be shared with a customer. In our current workflow the agent therefore serves as an internal tool, where the expert judges its output and contains its errors.