Deep Enhancement Spotting Network for Low-complexity Keyword Spotting in Noisy Environments
Yongqiang Chen, Qianhua He, Yanxiong Li, Zunxian Liu, Mingru Yang, Jinxin Huang
Abstract
Keyword Spotting (KWS) is crucial for hands-free voice-activated systems, requiring a balance between accuracy and complexity, especially in noisy environments. While Speech Enhancement (SE) can improve KWS accuracy, existing methods often lack the ability to effectively utilize the rich features produced during enhancement. In this paper, we design a low-complexity network to address the challenges of KWS in noisy environments. We integrate the tasks of both SE and KWS into a unified network that learns a shared representation from both tasks. The proposed network features two blocks: a Residual Full-band and Sub-band Fusion (RFSF) block, and a Deformable Transition (DT) block. Our dual-task network surpasses existing KWS models in accuracy with low complexity, making it suitable for deployment on edge devices.
BibTeX
@inproceedings{icassp2025_deepenhancements,
title = {Deep Enhancement Spotting Network for Low-complexity Keyword Spotting in Noisy Environments},
author = {Yongqiang Chen and Qianhua He and Yanxiong Li and Zunxian Liu and Mingru Yang and Jinxin Huang},
booktitle = {ICASSP 2025},
year = {2025}
}