Speedrunning Relational Foundation Model Pretraining
Gabriel Veloso ⋅ Francisco Galuppo Azevedo ⋅ Denis O Correa ⋅ Luan B Sena
Abstract
Progress on foundation model pretraining is gated by who can afford a trial, and relational pretraining is no exception. Recipes are developed inside individual labs, reported under private protocols, and do not accumulate. Speedrun leaderboards remove that barrier by decentralizing the search rather than the training run. The problem is fixed, the stack is open, entries are ranked by wall-clock time to a quality target, and one GPU is enough to hold the record. We introduce modded-nanoRT, the first such leaderboard for Relational Transformers. Which database to hold out decides whether the benchmark can resolve improvements at all, so we fix it by calibration, ranking six candidates by how sharply a 45-minute learning curve separates improvement from seed noise. We open the leaderboard with typed trilinear attention, a relational operator that conditions attention scores and values on foreign-key semantics, together with Muon and SIGReg. The recipe reaches the target in $305.63$ seconds on one NVIDIA H200, $8.83\times$ inside the calibration horizon, and wins 13 of 15 pairs outside the development database. We release the benchmark, its frozen artifacts, and the record runs.
Chat is not available.
Successful Page Load