AQFormer: severity-aware transformer with aphasia-specific CAM for spoken keyword classification in aphasic speech
Abstract
Spoken language expression in individuals with aphasia is highly variable due to phonological errors, unexpected pauses, and word-boundary deviations. This variability degrades standard keyword classification performance and limits clinical adoption because deep neural models remain uninterpretable. We propose a severity-aware, explainable transformer framework optimized for spoken keyword classification in aphasic speech. Our architecture fine-tunes a self-supervised speech encoder (WavLM-Large) coupled with a task-specific Transformer head, conditioning intermediate speech features directly on clinical severity scores via Feature-wise Linear Modulation (FiLM). Alongside the classifier, an Aphasia-specific Class Activation Mapping (A-CAM) mechanism separates decision-driving evidence from speech impairment patterns via a custom multimodal filter tracking pauses and deviations. Tested under speaker-disjoint splits, our framework achieves 96.61\% classification accuracy. Furthermore, our attribution maps outperform standard Grad-CAM variants on faithfulness metrics (0.38 deletion AUPC) and strongly align with objective phoneme error boundaries (IoU = 0.807) without requiring expert target labels during training. .