← Search

Yao Tian

7 accepted papers

2024

A Dual-Path Framework with Frequency-and-Time Excited Network for Anomalous Sound Detection

ICASSP 2024accepted

In contrast to human speech, machine-generated sounds of the same type often exhibit consistent frequency characteristics and discernible temporal periodicity. However, leveraging these dual attributes in anomaly detection remains relatively under-explored. In this paper, we propose an automated dua…

Cited by 0SourceScholar
2024

Efficient Personal Voice Activity Detection with Wake Word Reference Speech

ICASSP 2024accepted

Personal voice activity detection (PVAD) is gradually used in speech assistants. Traditional PVAD schemes extract the target speaker’s embedding from existing query reference speech through a pre-trained speaker verification model. Consequently, the performance of the PVAD model may suffer if the qu…

Cited by 0SourceScholar
2022

BiFSMN: Binary Neural Network for Keyword Spotting

IJCAI 2022poster

The deep neural networks, such as the Deep-FSMN, have been widely studied for keyword spotting (KWS) applications. However, computational resources for these networks are significantly constrained since they usually run on-call on edge devices. In this paper, we present BiFSMN, an accurate and extre…

2022

The Volcspeech System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge

ICASSP 2022accepted

This paper describes our submission to ICASSP 2022 Multi-channel Multi-party Meeting Transcription (M2MeT) Challenge. For Track 1, we propose several approaches to make the clustering-based speaker diarization system enable to handle overlapped speech. Front-end dereverberation and the direction-of-…

Cited by 0SourceScholar
2021

Improving RNN Transducer Modeling for Small-Footprint Keyword Spotting

ICASSP 2021accepted

The recurrent neural network transducer (RNN-T) model has been proved effective for keyword spotting (KWS) recently. However, compared with cross-entropy (CE) or connectionist temporal classification (CTC) based models, the additional prediction network in the RNN-T model increases the model size an…

Cited by 0SourceScholar
2020

Adaptation of RNN Transducer with Text-To-Speech Technology for Keyword Spotting

ICASSP 2020accepted

With the advent of recurrent neural network transducer (RNN-T) model, the performance of keyword spotting (KWS) systems has greatly improved. However, the KWS systems, employed for wake-word detection, still rely on the availability of keyword specific training data for achieving reasonable performa…

Cited by 0SourceScholar
2017

Deep neural networks based speaker modeling at different levels of phonetic granularity

ICASSP 2017accepted

Recently, a hybrid deep neural network/i-vector framework has been proved effective for speaker verification, where the DNN trained to predict tied-triphone states (senones) is used to produce frame alignments for sufficient statistics extraction. In this work, in order to better understand the impa…

Cited by 0SourceScholar