The Wrong Clock for Transaction Streams: Why Equal-Time Windows Fail as Trading Thins
Abstract
Irregularly sampled event streams are usually discretised onto equal wall-clock windows before a model is trained or scored; we ask whether that choice decides what the model can learn. On fourteen complete NASDAQ sessions of order-by-order data, over two thousand securities logging 15,000 to 240,000 book events a session, we hold one predictor fixed, order-flow imbalance (net buying minus selling at the best quotes), and measure how often it calls the direction of the next window's mid-price move, against a baseline that sees only the previous move on the same grid. We vary only the clock, the rule that cuts a session into windows: equal time (calendar), equal numbers of book updates (event), or equal displayed volume. Order flow on the equal-time grid is no better than the previous move, at the 600 windows the thinnest names can support. On the equal-volume grid it adds 5.1 percentage points over that baseline, pooled, positive on every session and larger as trading thins; the volume clock's lead over the calendar clock still stands at 2.2 points with the three most damaging alternatives applied at once. Most of the effect is boundary placement: the predictable move sits in the first tens of milliseconds after a volume boundary, where an equal-time grid lands only by coincidence. The advantage needs sight of enough trades to locate the boundaries: it is absent on IEX and BX, two feeds seeing a fraction of NASDAQ's trades, and thinning NASDAQ's own feed reproduces part of that decay. We split every gain into the predictor improving and the baseline weakening; on PSX, another small venue, the gain is positive only because the baseline weakened, a trap whenever clocks are compared. Trained tree ensembles amplify the clock's effect, while frozen sequence representations, pretrained or not, carry under a percentage point on either clock. We make no trading claim. For sparse event streams the discretisation is part of the benchmark: report it, and cut by activity before comparing models.