Clustering-Free End-to-End Spoof Diarization with Encoder-Decoder Attractors
Bongsu Jung ⋅ Donghee Kim ⋅ Wooil Kim
Abstract
Spoof diarization aims to jointly localize spoofed regions and determine which spoofing method generated each segment within a partially spoofed utterance. Existing approaches rely on modular pipelines that separate localization and clustering, introducing training--inference mismatch and requiring oracle knowledge of the number of spoofing classes at inference time. We propose \textbf{E2E-SD}, the first \emph{clustering-free} end-to-end framework for spoof diarization, which reformulates the task as variable-cardinality temporal set prediction. E2E-SD jointly learns localization and class assignment using encoder-decoder attractors with Hungarian-matched training, eliminating the need for post-hoc clustering and oracle class counts. Our framework dynamically generates class-specific attractors through an LSTM decoder and estimates the number of active spoofing classes via attractor existence probabilities, requiring no external class-count information at any stage. On the PartialSpoof benchmark, E2E-SD achieves a $\JER$ of 15.76\%, a 41.6\% relative improvement over the previous best result (26.99\%), while operating fully oracle-free with 88.7\% class-count estimation accuracy. The largest gain is observed on unknown attacks, where $\JER$ drops from 49.02\% to 21.00\% (57.1\% relative improvement), demonstrating that end-to-end optimization with dynamic attractors enables more attack-agnostic generalization than modular clustering pipelines.
Chat is not available.
Successful Page Load