Evaluating Searched Trading Strategies by Betting
Abstract
When a researcher tests n trading strategies on one history and keeps the best back2 test, the reported statistics of the winner are inflated by selection, while the quality of the selected strategy may improve, stay flat, or degrade. Standard multiplicity corrections repair the report with constants that depend on an effective trial count, and we show on real data that defensible choices of that count change the correction by nearly a factor of two. We propose evaluating the searched strategy itself by betting: a sign-bet mixture e-process on its out-of-sample epoch returns, merged by averaging across search draws, which is valid under optional continuation of the research program and under arbitrary dependence across candidates. On 1,000 CRSP single names and 56 years of daily diversified panels the betting evidence tracks decision quality and ignores report inflation: reported Sharpe rises with search width everywhere, while the merged e-value rises with width only where wider search actually selects better strategies. We preregister this e-process as the scoring rule for a sealed 2026–2028 forward test.