Do LLM Alpha-Mining Leaderboards Survive Search-Adjusted Sharpe Stress Tests?
Abstract
Because leaderboard winners are selected after search, their raw Sharpe-like statistics cannot be interpreted as prespecified single tests. Conventional Deflated Sharpe Ratio calculations generally cannot be identified from published summaries because the required cross-candidate Sharpe distribution is not reported. We therefore use a transparent DSR-style stress test under explicitly labeled assumptions. In a fixed, purposive frame of 13 surfaced alpha-mining source families, trial-count evidence appears in 10/13, while a multiple-testing-aware correction appears in 1/13. Under each of four assumed moment profiles, 6/7 provenance-eligible top-1 rows first do not pass by tested assumed search effort N≤6. The claim that all seven eventually fail somewhere through N=10,000 is not robust to an alternative annualization convention for the 24/7 cryptocurrency row. In a known-truth simulation with N=1000, T=1008, and true annual Sharpe 0.5, a true signal is selected in 74/1000 repetitions; neither the DSR-style proxy nor conventional DSR retains any of those 74 selected true winners. With candidate equicorrelation ρ=0.9, the proxy is substantially more conservative than conventional DSR. These results constitute a sensitivity audit, not a reconstruction of hidden search histories or a judgment of investment value. They support disclosure of search effort, selection rules, sample information, return moments or returns, correction procedures, and complete machine-readable candidate results.