Where You Put the Space Matters: Boundary-Conditioned Whitespace Robustness in Thai Dense Retrieval
Abstract
Dense retrieval systems are expected to remain robust to small variations in how a query is written. However, such variation is not linguistically uniform across languages because writing systems encode boundaries differently. Recent work shows that robustness to alternative tokenization varies substantially across languages. Whitespace provides a particularly revealing example. In Thai, whitespace does not consistently delimit lexical words, so inserting a space at different positions can interact differently with the structure of the written query. This makes Thai dense retrieval a useful setting for asking whether the same-sized whitespace perturbation is associated with different retrieval outcomes depending on where it is inserted. We study explicit whitespace insertion in Thai queries using matched perturbations with fixed magnitude and varying boundary status. D1 inserts one whitespace at a boundary admitted by a frozen boundary-selection rule, whereas NB1 inserts one whitespace at a Thai-internal Thai Character Cluster (TCC) position not proposed as a boundary by any segmentation engine. The boundary-selection rule is developed using VISTEC-TP-TH-2021 human word segmentation, evaluated on a held-out subset constructed from the human-segmented Wisesight-1000 data, and frozen before perturbation generation on MIRACL-Thai. All non-whitespace characters remain unchanged, and both conditions have an edit distance of one from the original query. We evaluate BGE-M3, multilingual-E5-large, and Qwen3-Embedding-0.6B on 200 MIRACL-Thai queries, of which 195 support the matched single-insertion D1/NB1 comparison. We report conditional failure at 5 (CF@5), defined as the proportion of baseline Top-5 successes for which the relevant passage falls outside the Top-5 after perturbation. CF@5 is descriptively higher for NB1 than D1 across all three models: 2.17% versus 0.54% for BGE-M3, 5.82% versus 0.00% for multilingual-E5-large, and 3.91% versus 0.56% for Qwen3-Embedding-0.6B. With 195 matched queries and few failure events, we report these results descriptively rather than as tested differences. Because both conditions use the same one-whitespace edit, the observed contrast is not captured by perturbation magnitude alone. The study provides a boundary-conditioned robustness evaluation protocol that distinguishes where an input is perturbed from how much it is perturbed.