← Search

Yu Tsao

45 accepted papers

2026

CONDITION-INVARIANT FMRI DECODING OF SPEECH INTELLIGIBILITY WITH DEEP STATE SPACE MODEL

ICASSP 2026poster

Clarifying the neural basis of speech intelligibility is critical for computational neuroscience and digital speech processing. Recent neuroimaging studies have shown that intelligibility modulates cortical activity beyond simple acoustics, primarily in the superior temporal and inferior frontal gyr…

Cited by 0SourcePDFScholar
2026

GAME-TIME: EVALUATING TEMPORAL DYNAMICS IN SPOKEN LANGUAGE MODELS

ICASSP 2026oral

Conversational Spoken Language Models (SLMs) are emerging as a promising paradigm for real-time speech interaction. However, their capacity of temporal dynamics, including the ability to manage timing, tempo and simultaneous speaking, remains a critical and unevaluated challenge for conversational f…

Cited by 0SourcePDFScholar
2026

Tracking Listener Attention: Gaze-Guided Audio-Visual Speech Enhancement Framework

ICASSP 2026poster

This paper presents a Gaze-Guided Audio-Visual Speech Enhancement (GG-AVSE) framework to address the cocktail party problem. A major challenge in conventional AVSE is identifying the listener's intended speaker in multi-talker environments. GG-AVSE addresses this issue by exploiting gaze direction a…

Cited by 0SourcePDFScholar
2025

A Study on Zero-shot Non-intrusive Speech Assessment using Large Language Models

ICASSP 2025accepted

This work investigates two strategies for zero-shot non-intrusive speech assessment leveraging large language models. First, we explore the audio analysis capabilities of GPT-4o. Second, we propose GPT-Whisper, which uses Whisper as an audio-to-text module and evaluates the text’s naturalness via ta…

Cited by 0SourceScholar
2025

Leveraging Joint Spectral and Spatial Learning with MAMBA for Multichannel Speech Enhancement

ICASSP 2025accepted

In multichannel speech enhancement, effectively capturing spatial and spectral information across different microphones is crucial for noise reduction. Traditional methods, such as CNN or LSTM, attempt to model the temporal dynamics of full-band and sub-band spectral and spatial features. However, t…

Cited by 0SourceScholar
2025

MSECG: Incorporating Mamba for Robust and Efficient ECG Super-Resolution

ICASSP 2025accepted

Electrocardiogram (ECG) signals play a crucial role in diagnosing cardiovascular diseases. To reduce power consumption in wearable or portable devices used for long-term ECG monitoring, super-resolution (SR) techniques have been developed, enabling these devices to collect and transmit signals at a…

Cited by 0SourceScholar
2025

MSEMG: Surface Electromyography Denoising with a Mamba-based Efficient Network

ICASSP 2025accepted

Surface electromyography (sEMG) recordings can be contaminated by electrocardiogram (ECG) signals when the monitored muscle is closed to the heart. Traditional signal processing-based approaches, such as high-pass filtering and template subtraction, have been used to remove ECG interference but are…

Cited by 0SourceScholar
2025

QualiSpeech: A Speech Quality Assessment Dataset with Natural Language Reasoning and Descriptions

ACL 2025long

This paper explores a novel perspective to speech quality assessment by leveraging natural language descriptions, offering richer, more nuanced insights than traditional numerical scoring methods. Natural language feedback provides instructive recommendations and detailed evaluations, yet existing d…

2024

AV-SUPERB: A Multi-Task Evaluation Benchmark for Audio-Visual Representation Models

ICASSP 2024accepted

Audio-visual representation learning aims to develop systems with human-like perception by utilizing correlation between auditory and visual information. However, current models often focus on a limited set of tasks, and generalization abilities of learned representations are unclear. To this end, w…

Cited by 0SourceScholar
2024

Hierarchical Cross-Modality Knowledge Transfer with Sinkhorn Attention for CTC-Based ASR

ICASSP 2024accepted

Due to the modality discrepancy between textual and acoustic modeling, efficiently transferring linguistic knowledge from a pretrained language model (PLM) to acoustic encoding for automatic speech recognition (ASR) still remains a challenging task. In this study, we propose a cross-modality knowled…

Cited by 0SourceScholar
2024

Multi-Task Pseudo-Label Learning for Non-Intrusive Speech Quality Assessment Model

ICASSP 2024accepted

This study proposes a multi-task pseudo-label learning (MPL)-based non-intrusive speech quality assessment model called MTQ-Net. MPL consists of two stages: obtaining pseudo-label scores from a pretrained model and performing multitask learning. The 3QUEST metrics, namely Speech-MOS (S-MOS), Noise-M…

Cited by 0SourceScholar
2024

RankUp: Boosting Semi-Supervised Regression with an Auxiliary Ranking Classifier

NeurIPS 2024poster

State-of-the-art (SOTA) semi-supervised learning techniques, such as FixMatch and it's variants, have demonstrated impressive performance in classification tasks. However, these methods are not directly applicable to regression tasks. In this paper, we present RankUp, a simple yet effective approach…

2024

SDEMG: Score-Based Diffusion Model for Surface Electromyographic Signal Denoising

ICASSP 2024accepted

Surface electromyography (sEMG) recordings can be influenced by electrocardiogram (ECG) signals when the muscle being monitored is close to the heart. Several existing methods use signal-processing-based approaches, such as high-pass filter and template subtraction, while some derive mapping functio…

Cited by 0SourceScholar
2024

Scalable Ensemble-Based Detection Method Against Adversarial Attacks For Speaker Verification

ICASSP 2024accepted

Automatic speaker verification (ASV) is highly susceptible to adversarial attacks. Purification modules are usually adopted as a pre-processing to mitigate adversarial noise. However, they are commonly implemented across diverse experimental settings, rendering direct comparisons challenging. This p…

Cited by 0SourceScholar
2024

Self-Supervised Speech Quality Estimation and Enhancement Using Only Clean Speech

ICLR 2024poster

Speech quality estimation has recently undergone a paradigm shift from human-hearing expert designs to machine-learning models. However, current models rely mainly on supervised learning, which is time-consuming and expensive for label collection. To solve this problem, we propose VQScore, a self-su…

2023

D4AM: A General Denoising Framework for Downstream Acoustic Models

ICLR 2023poster

The performance of acoustic models degrades notably in noisy environments. Speech enhancement (SE) can be used as a front-end strategy to aid automatic speech recognition (ASR) systems. However, existing training objectives of SE methods are not fully effective at integrating speech-text and noise-c…

2023

ECG Artifact Removal from Single-Channel Surface EMG Using Fully Convolutional Networks

ICASSP 2023accepted

Electrocardiogram (ECG) artifact contamination often occurs in surface electromyography (sEMG) applications when the measured muscles are in proximity to the heart. Previous studies have developed and proposed various methods, such as high-pass filtering, template subtraction and so forth. However,…

Cited by 0SourceScholar
2023

Interpretations of Domain Adaptations via Layer Variational Analysis

ICLR 2023poster

Transfer learning is known to perform efficiently in many applications empirically, yet limited literature reports the mechanism behind the scene. This study establishes both formal derivations and heuristic analysis to formulate the theory of transfer learning in deep learning. Our framework utiliz…

2023

On the Robustness of Non-Intrusive Speech Quality Model by Adversarial Examples

ICASSP 2023accepted

It has been shown recently that deep learning based models are effective on speech quality prediction and could outperform traditional metrics in various perspectives. Although network models have the potential to be a surrogate for complex human hearing perception, they may contain instabilities in…

Cited by 0SourceScholar
2023

Prefallkd: Pre-Impact Fall Detection Via CNN-ViT Knowledge Distillation

ICASSP 2023accepted

Fall accidents are critical issues in an aging and aged society. Recently, many researchers developed "pre-impact fall detection systems" using deep learning to support wearable-based fall protection systems for preventing severe injuries. However, most works only employed simple neural network mode…

Cited by 0SourceScholar
2023

T5lephone: Bridging Speech and Text Self-Supervised Models for Spoken Language Understanding Via Phoneme Level T5

ICASSP 2023accepted

In Spoken language understanding (SLU), a natural solution is concatenating pre-trained speech models (e.g. HuBERT) and pretrained language models (PLM, e.g. T5). Most previous works use pre-trained language models with subword-based tokenization. However, the granularity of input units affects the…

Cited by 0SourceScholar
2022

Analyzing The Robustness of Unsupervised Speech Recognition

ICASSP 2022accepted

Unsupervised speech recognition (unsupervised ASR) aims to learn the ASR system with non-parallel speech and text corpus only. Wav2vec-U [1] has shown promising results in unsupervised ASR by self-supervised speech representations coupled with Generative Adversarial Network (GAN) training, but the r…

Cited by 0SourceScholar
2022

Conditional Diffusion Probabilistic Model for Speech Enhancement

ICASSP 2022accepted

Speech enhancement is a critical component of many user-oriented audio applications, yet current systems still suffer from distorted and unnatural outputs. While generative models have shown strong potential in speech synthesis, they are still lagging behind in speech enhancement. This work leverage…

Cited by 0SourceScholar
2022

MetricGAN-U: Unsupervised Speech Enhancement/ Dereverberation Based Only on Noisy/ Reverberated Speech

ICASSP 2022accepted

Most of the deep learning-based speech enhancement models are learned in a supervised manner, which implies that pairs of noisy and clean speech are required during training. Consequently, several noisy speeches recorded in daily life cannot be used to train the model. Although certain unsupervised…

Cited by 0SourceScholar
2022

Partially Fake Audio Detection by Self-Attention-Based Fake Span Discovery

ICASSP 2022accepted

The past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly implemented biometric identification models and can be harnessed by in-the-wild attackers for illegal uses. The ASVspoo…

Cited by 0SourceScholar
2022

Speech Recovery For Real-World Self-Powered Intermittent Devices

ICASSP 2022accepted

The incompleteness of speech inputs severely degrades the performance of all the related speech signal processing applications. Although many researches have been proposed to address this issue, they controlled the data missing conditions by simulation with self-defined masking lengths or sizes. Bes…

Cited by 0SourceScholar
2022

When BERT Meets Quantum Temporal Convolution Learning for Text Classification in Heterogeneous Computing

ICASSP 2022accepted

The rapid development of quantum computing has demonstrated many unique characteristics of quantum advantages, such as richer feature representation and more secured protection on model parameters. This work proposes a vertical federated learning architecture based on variational quantum circuits to…

Cited by 0SourceScholar
2022

XDBERT: Distilling Visual Information to BERT from Cross-Modal Systems to Improve Language Understanding

ACL 2022short

Transformer-based models are widely used in natural language understanding (NLU) tasks, and multimodal transformers have been effective in visual-language tasks. This study explores distilling visual information from pretrained multimodal transformers to pretrained language encoders. Our framework i…

Cited by 3SourcePDFScholar
2021

Unsupervised Neural Adaptation Model Based on Optimal Transport for Spoken Language Identification

ICASSP 2021accepted

Due to the mismatch of statistical distributions of acoustic speech between training and testing sets, the performance of spoken language identification (SLID) could be drastically degraded. In this paper, we propose an unsupervised neural adaptation model to deal with the distribution mismatch prob…

Cited by 0SourceScholar
2021

Unsupervised Noise Adaptive Speech Enhancement by Discriminator-Constrained Optimal Transport

NeurIPS 2021poster

This paper presents a novel discriminator-constrained optimal transport network (DOTN) that performs unsupervised domain adaptation for speech enhancement (SE), which is an essential regression task in speech processing. The DOTN aims to estimate clean references of noisy speech in a target domain,…

2020

Self-Supervised Denoising Autoencoder with Linear Regression Decoder for Speech Enhancement

ICASSP 2020accepted

Nonlinear spectral mapping-based models based on supervised learning have successfully applied for speech enhancement. However, as supervised learning approaches, a large amount of labelled data (noisy-clean speech pairs) should be provided to train those models. In addition, their performances for…

Cited by 0SourceScholar
2019

MetricGAN: Generative Adversarial Networks based Black-box Metric Scores Optimization for Speech Enhancement

ICML 2019oral

Adversarial loss in a conditional generative adversarial network (GAN) is not designed to directly optimize evaluation metrics of a target task, and thus, may not always guide the generator in a GAN to generate data with improved metric scores. To overcome this issue, we propose a novel MetricGAN ap…

2019

Reinforcement Learning Based Speech Enhancement for Robust Speech Recognition

ICASSP 2019accepted

Conventional deep neural network (DNN)-based speech enhancement (SE) approaches aim to minimize the mean square error (MSE) between enhanced speech and clean reference. The MSE-optimized model may not directly improve the performance of an automatic speech recognition (ASR) system. If the target is…

Cited by 0SourceScholar
2018

A Novel LSTM-Based Speech Preprocessor for Speaker Diarization in Realistic Mismatch Conditions

ICASSP 2018accepted

In this study, we investigate on the effects of deep learning based speech enhancement as a preprocessor to speaker diarization in quite challenging realistic environments involving the background noises, reverberations and overlapping speech. To improve the generalization capability, the advanced l…

Cited by 0SourceScholar
2018

Enhancement and Analysis of Conversational Speech: JSALT 2017

ICASSP 2018accepted

Automatic speech recognition is more and more widely and effectively used. Nevertheless, in some automatic speech analysis tasks the state of the art is surprisingly poor. One of these is “diarization”, the task of determining who spoke when. Diarization is key to processing meeting audio and clinic…

Cited by 0SourceScholar
2018

Speech Dereverberation Based on Integrated Deep and Ensemble Learning Algorithm

ICASSP 2018accepted

Reverberation, which is generally caused by sound reflections from walls, ceilings, and floors, can result in severe performance degradation of acoustic applications. Due to a complicated combination of attenuation and time-delay effects, the reverberation property is difficult to characterize, and…

Cited by 0SourceScholar
2017

A locally linear embbeding based postfiltering approach for speech enhancement

ICASSP 2017accepted

This paper presents a novel postfiltering approach based on the locally linear embedding (LLE) algorithm for speech enchantment (SE). The aim of the proposed LLE-based postfiltering approach is to further remove the residual noise components from the SE-processed speech signals through a spectral co…

Cited by 0SourceScholar
2017

Discriminative autoencoders for speaker verification

ICASSP 2017accepted

This paper presents a learning and scoring framework based on neural networks for speaker verification. The framework employs an autoencoder as its primary structure while three factors are jointly considered in the objective function for speaker discrimination. The first one, relating to the sample…

Cited by 0SourceScholar
2016

Nonnegative matrix factorization-based frequency lowering technology for Mandarin-speaking hearing aid users

ICASSP 2016accepted

Frequency lowering technologies have demonstrated effectiveness in English speech recognition for English-speaking people with high-frequency hearing loss. Their effect on Mandarin speech has not been well investigated. This paper serves two important purposes: it 1) examines the effect of frequency…

Cited by 0SourceScholar
2015

A discriminative post-filter for speech enhancement in hearing aids

ICASSP 2015accepted

For hearing aid (HA) devices, speech enhancement (SE) is an essential unit aiming to improve signal-to-noise ratio (SNR) and quality of speech signals. Previous studies, however, indicated that user experience with current HAs was not fully satisfactory in noisy environments, suggesting that there i…

Cited by 0SourceScholar