ICASSP 2026poster0 citations

ENVSSLAM-FFN: LIGHTWEIGHT LAYER-FUSED SYSTEM FOR ESDD 2026 CHALLENGE

Xiaoxuan Guo, Jiayi Zhou, Haonan Cheng

Abstract

Recent advances in generative audio models have enabled high-fidelity environmental sound synthesis, raising serious concerns for audio security. The ESDD 2026 Challenge therefore addresses environmental sound deepfake detection under unseen generators (Track 1) and black-box low-resource detection (Track 2) conditions. We propose EnvSSLAM-FFN, which integrates a frozen SSLAM self-supervised encoder with a lightweight FFN back-end. To effectively capture spoofing artifacts under severe data imbalance, we fuse intermediate SSLAM representations from layers 4-9 and adopt a class-weighted training objective. Experimental results show that the proposed system consistently outperforms the official baselines on both tracks, achieving Test Equal Error Rates (EERs) of 1.20% and 1.05%, respectively.

BibTeX
@inproceedings{icassp2026_envsslamffnlight,
  title = {ENVSSLAM-FFN: LIGHTWEIGHT LAYER-FUSED SYSTEM FOR ESDD 2026 CHALLENGE},
  author = {Xiaoxuan Guo and Jiayi Zhou and Haonan Cheng},
  booktitle = {ICASSP 2026},
  year = {2026}
}
ENVSSLAM-FFN: LIGHTWEIGHT LAYER-FUSED SYSTEM FOR ESDD 2026 CHALLENGE · ICASSP 2026