Beyond Leaderboards: Continual, Task-Specific LLM Evaluation for Enterprise Finance
Abstract
Selecting an LLM can determine whether an enterprise AI application becomes a reliable production system or remains a proof of concept. Yet rapid releases, updates, and deprecations make model selection an ongoing challenge. We compare 19 recent proprietary LLMs from four leading providers on four financial sentiment datasets, evaluating task accuracy, run-to-run stability, and sensitivity to entity identity. Our black-box audit replaces company names with neutral placeholders or same-sector alternatives to test whether predictions reflect textual evidence or entity-specific priors. General-purpose leaderboard rankings provide only limited guidance: the newest, largest, or nominally most capable models are not consistently best, and newer versions do not reliably improve accuracy or robustness. Repeated identical evaluations disagree on 8.5\% of predictions on average and shift model rankings by up to seven positions. Same-sector substitutions reduce accuracy by 1.64 percentage points on average, with significant drops for all 19 models, whereas masking produces a smaller, less consistent effect. Across conditions, rankings shift by 3.7 positions on average. These findings show that one-shot evaluations can overstate ranking precision and robustness. Enterprise model selection should therefore combine representative task data, multiple runs, entity interventions, cost, and latency, with requalification following provider changes.