Cross-Attentive Bayesian Low-Rank Adaptation for Multimodal Uncertainty Estimation
Habibeh Naderi Khorshidi ⋅ Behrouz Soleimani ⋅ Stan Matwin
Abstract
Parameter-efficient fine-tuning (PEFT) enables efficient adaptation of frozen language models, but most PEFT methods remain deterministic or unimodal, limiting their reliability in low-resource audio-text settings where uncertainty depends on both linguistic evidence and acoustic conditions. We introduce SPECTRA (Stochastic Posterior Estimation with Cross-modal Token-level Rank Adaptation), a multimodal Bayesian low-rank adaptation framework for uncertainty-aware audio-text learning. SPECTRA keeps the text and audio backbones frozen and confines stochasticity to a compact rank-$r$ latent matrix inside each LoRA adapter. At each transformer layer, text-derived low-rank token features query frame-level audio embeddings through lightweight cross-attention; the resulting token-specific acoustic context parameterizes the mean and variance of an amortized variational posterior over the adapter latent. This design treats audio not merely as an additional feature stream, but as a localized reliability signal that modulates both adaptation and confidence while preserving the scalability of PEFT. Posterior prediction is performed with Monte Carlo adapter samples, enabling a Bayesian uncertainty analysis that decomposes normalized predictive uncertainty into total, aleatoric, and adapter-space epistemic components and evaluates whether uncertainty distinguishes correct from misclassified predictions. Across IEMOCAP and clinical interview prediction tasks, SPECTRA is consistently competitive with or improves upon text-only Bayesian PEFT and conventional multimodal transfer-learning baselines, with token-level cross-attention yielding the most reliable gains. Additional modality-disagreement stress tests show that mismatched audio causes substantially larger degradation in AUC, likelihood, calibration, and Brier score than noisy but matched audio, highlighting the importance of modeling cross-modal reliability. These results suggest that Bayesian cross-modal conditioning in low-rank adapter space provides an efficient and principled mechanism for calibrated multimodal adaptation.
Chat is not available.
Successful Page Load