Counterfactual Explanations for Time-Series Classification via Constrained Flow Matching
Abstract
Counterfactual explanations for time-series classification generate synthetic instances that flip a prediction to a desired class while remaining plausible and making small, local changes. Existing methods often rely on classifier gradients, and many do not naturally extend to one-class settings. To bridge this gap, we propose CTCF, a model-agnostic method for hard-decision classifiers in supervised and one-class settings. CTCF decouples generation from classifier-specific control by reusing an unconditional Flow Matching generator and learning a flow-time-dependent desired-class region from interpolation states labeled by endpoint hard decisions, without classifier gradients or class probabilities. At inference time, CTCF uses a dual-control mechanism: fused-lasso endpoint-objective steering encourages small, segment-local changes, while linearized projection correction reduces violations of a conservative desired-class region. Theoretically, under ideal marginal matching, we show that the surrogate risk on interpolation states equals the corresponding ODE-trajectory risk. Experiments on UCR datasets demonstrate CTCF's effectiveness across quantitative metrics and qualitative case studies.