Style Is Not Selection Skill: Style-Proof Evaluation of LLM Trading Agents
Abstract
Financial LLM agents are often evaluated by testing whether their realized returns exceed those of a baseline. We show that a zero-centered paired test can confuse style—how many baseline positions the agent changes and in which directions—with selection skill—whether it changes the right events. In our sample of 723 earnings announcements, 547 baseline-flat events averaged +67.4 bps of factor-adjusted return over three days. A no-skill agent can therefore earn a positive payoff by turning randomly chosen flat positions into long positions, which a zero-centered test may mistake for skill. Because a developer can choose the policy's style, we call a rule style-proof if its false-positive rate does not exceed the stated significance level for any style. Our matched rule fixes the numbers and types of position changes and randomly reassigns them among eligible events. It then asks whether the agent's actual choices outperform random choices with the same style. If a no-skill agent chooses events at random within prespecified pools, this guarantee is exact in finite samples. At a stated rate of 5%, no-skill simulations on the realized returns produce false-positive rates as high as 17.1% for the zero-centered test, while the matched rule remains near 5%. With the same five-point increase in directional accuracy, the zero-centered test labels a long policy skilled 70.9% of the time but a short policy only 6.8%. After calibration to the same false-positive rate, this asymmetry largely disappears and the two rules have similar power. Across eight real-system comparisons, payoffs attributable to style range from −137 to +133 bps per intervention, while none shows detectable selection at the available power. Because selected and unselected events may not be fully comparable and the audit has limited power, we treat these results as descriptive, not as evidence that the systems lack selection skill.