DiA: Directional Adapter
Ashish Singh ⋅ Prakash Chandra Chhipa
Abstract
Fine-grained video action recognition often depends on how an action unfolds over time rather than on appearance alone. Visually similar classes may differ in preparation, execution, or completion phases, making uniformly aggregated temporal context suboptimal. We study this problem under a practical objective: i) improving recognition performance with few tunable parameters, ii) being compute efficient, and iii) using only the readily available video modality. Motivated by this missing temporal observation and the accuracy--compute efficiency--unimodality objective, we propose **Directional Adapter (DiA)**, a parameter-efficient adapter on top of CLIP. DiA formulates temporal specialization into causal and anti-causal directions, allowing past-to-present and future-to-present cues to be modeled differently. To keep this directional modeling parameter- and compute-efficient, both directions use depth-wise temporal convolutions in a compact bottleneck space and are combined through a learnable fusion. DiA achieves state-of-the-art performance across 7 public benchmarks with only $\sim$3.5M tunable parameters, while also delivering higher throughput and lower latency than prior methods. Code is available at https://anonymous.4open.science/r/diafullsupervised-8CC3/README.md.
Chat is not available.
Successful Page Load