Scaling Context Down for On-Device Tabular Foundation Models
Abstract
Tabular foundation models (TFMs) make predictions through in-context learning by conditioning on labelled examples supplied at inference time. Processing this context can strain compute and memory on constrained devices. We ask whether a TFM needs its full context, or whether a small fraction can retain most predictive performance. We evaluate three pretrained TFMs (TabPFN v2.6, TabPFN v3, and TabICL v2) on 32 TabArena classification tasks. Reducing the context to 128 examples keeps a majority of model-task evaluations within five AUC percentage points of full context, using less than 2\% of the full context for the median task. The three models agree strongly on which tasks are most sensitive. We profile these 128-example runs on a MacBook Air M2 (24GB). All complete, while 12 full-context runs hit memory or time limits. For TabICL v2, among completed comparisons, the reduced context runs faster, in only 14.4\% of the full-context time and 32.1\% of its peak memory (medians). These results reframe deployment from how much context a model can accept to how much a task needs, while motivating the next question of which examples are worth retaining.