← Search

Yonghong Yan

22 accepted papers

2025

Automatic Text Pronunciation Correlation Generation and Application for Contextual Biasing

ICASSP 2025accepted

Effectively distinguishing the pronunciation correlations between different written texts is a significant issue in linguistic acoustics. Traditionally, such pronunciation correlations are obtained through manually designed pronunciation lexicons. In this paper, we propose a data-driven method to au…

Cited by 0SourceScholar
2025

Debiased Training For Semi-supervised Sound Event Detection

ICASSP 2025accepted

Recently, semi-supervised sound event detection has attracted increasing research interest due to the scarcity of labeled data. However, traditional semi-supervised learning methods can lead to training instability and confirmation bias because of potentially incorrect pseudo labels. To address this…

Cited by 0SourceScholar
2025

Rainbow Delay Compensation: A Multi-Agent Reinforcement Learning Framework for Mitigating Observation Delays

NeurIPS 2025poster

In real-world multi-agent systems (MASs), observation delays are ubiquitous, preventing agents from making decisions based on the environment's true state. An individual agent's local observation typically comprises multiple components from other agents or dynamic entities within the environment. Th…

Cited by 0SourcecodeScholar
2024

One-Epoch Training with Single Test Sample in Test Time for Better Generalization of Cough-Based Covid-19 Detection Model

ICASSP 2024accepted

The outbreak of COVID-19 has raised researchers’ attention to audio-based rapid disease detection. Most of the previous studies have obtained competitive detection performance. However, these results are usually obtained by testing data from the same source offline. When making cross-dataset testing…

Cited by 0SourceScholar
2024

Snore Sound Features Based on Percussive Enhancing and Positional Encoding Combined with Multi-Task Learning for Osahs Detection

ICASSP 2024accepted

Obstructive sleep apnea hypopnea syndrome (OSAHS) is a serious sleep disorder. As the typical symptom of OSAHS, snoring has been proved effective in OSAHS diagnosis and potential to replace the current laborious and expensive polysomnography. However, the lack of analysis on the characteristics of p…

Cited by 0SourceScholar
2023

Piecewise Position Encoding in Convolutional Neural Network for Cough-Based Covid-19 Detection

ICASSP 2023accepted

A fast and efficient COVID-19 detection method is of vital importance to control the spread of the epidemic. Many studies have achieved good performance on cough-based COVID19 detection in the past two years. However, the effect of position information in time-frequency features of cough audio has b…

Cited by 0SourceScholar
2021

Decomposing Complex Questions Makes Multi-Hop QA Easier and More Interpretable

EMNLP 2021finding

Multi-hop QA requires the machine to answer complex questions through finding multiple clues and reasoning, and provide explanatory evidence to demonstrate the machine’s reasoning process. We propose Relation Extractor-Reader and Comparator (RERC), a three-stage framework based on complex question d…

2021

History Utterance Embedding Transformer LM for Speech Recognition

ICASSP 2021accepted

History utterances contain rich contextual information; however, better extracting information from the history utterances and using it to improve the language model (LM) is still challenging. In this paper, we propose the history utterance embedding Transformer LM (HTLM), which includes an embeddin…

Cited by 0SourceScholar
2021

Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Text Data

ICASSP 2021accepted

This paper presents a method to pre-train transformer-based encoder-decoder automatic speech recognition (ASR) models using sufficient target-domain text. During pre-training, we train the transformer decoder as a conditional language model with empty or artifical states, rather than the real encode…

Cited by 0SourceScholar
2020

Transformer-Based Online CTC/Attention End-To-End Speech Recognition Architecture

ICASSP 2020accepted

Recently, Transformer has gained success in automatic speech recognition (ASR) field. However, it is challenging to deploy a Transformer-based end-to-end (E2E) model for online speech recognition. In this paper, we propose the Transformer-based online CTC/attention E2E ASR architecture, which contai…

Cited by 0SourceScholar
2019

A Deep Learning Based Binaural Speech Enhancement Approach with Spatial Cues Preservation

ICASSP 2019accepted

The studies of binaural hearing indicated considerable benefits of the spatial information of sound sources in speech understanding in noise. In this paper, we propose a binaural speech enhancement approach based on deep neural network. In this approach, the signals at the left and right channels ar…

Cited by 0SourceScholar
2019

A Subband Energy Modification Method for Elevation Control in Median Plane

ICASSP 2019accepted

Elevation perception is crucial for binaural reproduction. A recent study proposed an elevation control method by modifying the energy of HRTFs in each auditory scale subband, such as the ERB and Mel subband. However, this subband division is designed based on auditory excitation patterns and may no…

Cited by 0SourceScholar
2019

An Audio Scene Classification Framework with Embedded Filters and a DCT-based Temporal Module

ICASSP 2019accepted

Deep convolutional neural network (DCNN) has recently improved the performance of acoustic scene classification. However, the input features of the network are usually based on predefined hand-tailored filters, which may not apply to the specific tasks. To overcome this, we propose a hybrid framewor…

Cited by 0SourceScholar
2019

Multiple Temporal Scales Based Speaker Embeddings Learning for Text-dependent Speaker Recognition

ICASSP 2019accepted

To extract high speaker-sensitive embeddings from deep neural networks is still a challenge in the field of speaker recognition. This paper proposes a novel network that learns speaker embeddings from multiple temporal scales. This idea comes from the recent biological research that the human audito…

Cited by 0SourceScholar
2019

Self-attention Based Prosodic Boundary Prediction for Chinese Speech Synthesis

ICASSP 2019accepted

Predicting prosodic boundaries from input text plays an important role in Chinese text-to-speech (TTS) system, which directly influences the naturalness and intelligibility of synthesized speech. In this paper, we propose to combine self-attention with multitask learning for prosodic boundary predic…

Cited by 0SourceScholar
2018

A Deep Neural Network Based Method of Source Localization in a Shallow Water Environment

ICASSP 2018accepted

This paper applies deep neural network (DNN) to source localization in a shallow water environment because of its powerful modeling capability and the little dependence on the prior knowledge of environmental parameters. The classical two-stage scheme is adopted, in which feature extraction and DNN…

Cited by 0SourceScholar
2018

Improving Multichannel Speech Recognition with Generalized Cross Correlation Inputs and Multitask Learning

ICASSP 2018accepted

Acoustic signals from microphone arrays are used to improve performance in distant speech recognition due to the availability of spatial information. And multichannel automatic speech recognition (ASR) systems often separate speech enhancement module from acoustic modeling, which may be not optimal…

Cited by 0SourceScholar
2018

On SDW-MWF and Variable Span Linear Filter with Application to Speech Recognition in Noisy Environments

ICASSP 2018accepted

Neural network based spectral mask estimation for acoustic beamforming, which consists of linear filtering and mask estimation, has shown to be a promising approach for robust speech recognition in noisy environments. Nevertheless, few improvements are made on the linear filtering. In this paper, we…

Cited by 0SourceScholar
2018

Semi-Supervised Learning with Deep Neural Networks for Relative Transfer Function Inverse Regression

ICASSP 2018accepted

Prior knowledge of the relative transfer function (RTF) is useful in many applications but remains little studied. In this paper, we propose a semi-supervised learning algorithm based on deep neural networks (DNNs) for RTF inverse regression, that is to generate the full-band RTF vector directly fro…

Cited by 0SourceScholar
2016

Robust multiple speech source localization using time delay histogram

ICASSP 2016accepted

Spatial aliasing and spatial resolution are the two issues faced by most multiple speech source localization methods. The histogram of time delays is a simple but effective method to deal with these two issues on linear arrays. But few methods were capable of applying the time delay histogram to dir…

Cited by 0SourceScholar