2026
FINE-GRAINED FRAME MODELING IN MULTI-HEAD SELF-ATTENTION FOR SPEECH DEEPFAKE DETECTION
ICASSP 2026poster
Transformer-based models have shown strong performance in speech deepfake detection, largely due to the effectiveness of the multi-head self-attention (MHSA) mechanism. MHSA provides frame-level attention scores, which are particularly valuable because deepfake artifacts often occur in small, locali…