← Search

Jingyong Hou

5 accepted papers

2023

Wekws: A Production First Small-Footprint End-to-End Keyword Spotting Toolkit

ICASSP 2023accepted

Keyword spotting (KWS) enables speech-based user interaction and gradually becomes an indispensable component of smart devices. Recently, end-to-end (E2E) methods have be-come the most popular approach for on-device KWS tasks. However, there is still a gap between the research and deployment of E2E…

Cited by 0SourceScholar
2020

Mining Effective Negative Training Samples for Keyword Spotting

ICASSP 2020accepted

Max-pooling neural network architectures have been proven to be useful for keyword spotting (KWS), but standard training methods suffer from a class-imbalance problem when using all frames from negative utterances. To address the problem, we propose an innovative algorithm, Regional Hard-Example (RH…

Cited by 0SourceScholar
2019

Adversarial Examples for Improving End-to-end Attention-based Small-footprint Keyword Spotting

ICASSP 2019accepted

In this paper, we explore the use of adversarial examples for improving a neural network based keyword spotting (KWS) system. Specially, in our system, an effective and small-footprint attention-based neural network model is used. Adversarial example is defined as a misclassified example by a model,…

Cited by 0SourceScholar
2019

Domain Adversarial Training for Improving Keyword Spotting Performance of ESL Speech

ICASSP 2019accepted

A second language (L2) learner usually cannot speak L2 well in both pronunciations and forming-of-words. Hence his/her L2 speech cannot be well recognized by a recognizer trained with native data. Domain adversarial training (DAT), capable of reducing the acoustic mismatch between training and testi…

Cited by 0SourceScholar
2016

Approximate search of audio queries by using DTW with phone time boundary and data augmentation

ICASSP 2016accepted

Dynamic Time Warping (DTW) is widely used in language independent query-by-example (QbE) spoken term detection (STD) tasks due to its high performance. However, there are two limitations of DTW based template matching, 1) it is not straightforward to perform approximate match of audio queries; 2) DT…

Cited by 0SourceScholar