Diagnosing Information Dependence in Time-Series–Text Alignment
Ranyi Luo ⋅ Yuanyuan Deng ⋅ Songgaojun Deng
Abstract
Time-series–language models are evaluated by aggregate retrieval and question-answering scores, which report how often a system succeeds but not which information supports that success. In a reproduced CLaSP evaluation, the auxiliary judge accepts $99.7\%$ of query–candidate pairs, yielding near-perfect retrieval scores for both trained and randomly initialized retrievers ($0.997$ and $0.999$), which the metric cannot separate. We introduce a diagnostic framework for component sensitivity, temporal-order dependence, and statistical sufficiency, with conclusions bounded by the available controls. We apply it to CLaSP and TRACE, two time-series–text retrievers, and ChatTS, a multimodal language model. The diagnostics reveal distinct information-dependence profiles. On the same items, CLaSP and ChatTS exhibit different semantic sensitivities: CLaSP is most affected by fluctuation type, whereas ChatTS is more sensitive to trend family. Shuffling substantially degrades all three systems, but only one corpus provides the order-invariant descriptions required to isolate temporal-order dependence. Information reduction further reveals different sufficiency frontiers: residual performance is supported by empirical distribution shape in one case, whereas in another it is explained by sequence length in the candidate pool. Measured alignment is therefore better characterized by an information-dependence profile than by a scalar score. Such profiles also identify when evaluation design restricts what can be inferred. We released our code and diagnostic design details to support future search:https://anonymous.4open.science/r/ts-text-diagnostics-778A/.
Chat is not available.
Successful Page Load