← Search

Chouchang Yang

11 accepted papers

2025

Better Exploiting Spatial Separability in Multichannel Speech Enhancement with an Align-and-Filter Network

ICASSP 2025accepted

Multichannel speech enhancement (SE) techniques combine multiple microphone signals to extract clean speech from noisy mixtures based on spatial filtering. As the target speech may come from arbitrary, unknown directions, current deep learning-based SE systems could suffer from performance bottlenec…

Cited by 0SourceScholar
2025

MIB: Mixed Information Bottleneck for Out-of-Distribution Keyword Spotting

ICASSP 2025accepted

Deep Keyword Spotting (KWS) systems continuously process audio streams to detect keywords. However, performance of deep neural networks degrade when the input data diverges from the training data; referred to as Out-of-Distribution (OOD) data problem. In this paper, we show performance degradation o…

Cited by 0SourceScholar
2025

RestoreGrad: Signal Restoration Using Conditional Denoising Diffusion Models with Jointly Learned Prior

ICML 2025poster

Denoising diffusion probabilistic models (DDPMs) can be utilized to recover a clean signal from its degraded observation(s) by conditioning the model on the degraded signal. The degraded signals are themselves contaminated versions of the clean signals; due to this correlation, they may encompass ce…

Cited by 0SourcePDFScholar
2024

An MVDR-Embedded U-Net Beamformer for Effective and Robust Multichannel Speech Enhancement

ICASSP 2024accepted

In multichannel speech enhancement (SE) systems, deep neural networks (DNNs) are often utilized to directly estimate the clean speech for effective beamforming. This approach, however, may not generalize adequately to new acoustic or noise conditions. Alternatively, DNNs can indirectly perform SE by…

Cited by 0SourceScholar
2024

CIFD: Controlled Information Flow to Enhance Knowledge Distillation

NeurIPS 2024poster

Knowledge Distillation is the mechanism by which the insights gained from a larger teacher model are transferred to a smaller student model. However, the transfer suffers when the teacher model is significantly larger than the student. To overcome this, prior works have proposed training intermediat…

Cited by 2SourcePDFScholar
2024

End-To-End Personalized Cuff-Less Blood Pressure Monitoring Using ECG and PPG Signals

ICASSP 2024accepted

Cuffless blood pressure (BP) monitoring offers the potential for continuous, non-invasive healthcare but has been limited in adoption by existing models relying on handcrafted features from ECG and PPG signals. To overcome this, researchers have looked to deep learning. Along these lines, in this pa…

Cited by 0SourceScholar
2024

Leveraging Self-Supervised Speech Representations for Domain Adaptation in Speech Enhancement

ICASSP 2024accepted

Deep learning based speech enhancement (SE) approaches could suffer from performance degradation due to mismatch between training and testing environments. A realistic situation is that an SE model trained on parallel noisy-clean utterances from one environment, the source domain, may fail to perfor…

Cited by 0SourceScholar
2024

Zero-Shot Intent Classification Using a Semantic Similarity Aware Contrastive Loss and Large Language Model

ICASSP 2024accepted

Zero-shot systems can reduce the cost of collecting data and training in a new domain since they can work directly with the test data without further training. In this paper, we build zero-shot systems for intent classification, based on Semantic Similarity-aware Contrastive Loss (SSCL) that address…

Cited by 1SourceScholar
2023

CWCL: Cross-Modal Transfer with Continuously Weighted Contrastive Loss

NeurIPS 2023poster

This paper considers contrastive training for cross-modal 0-shot transfer wherein a pre-trained model in one modality is used for representation learning in another domain using pairwise data. The learnt models in the latter domain can then be used for a diverse set of tasks in a 0-shot way, similar…

Cited by 8SourcePDFScholar
2023

Improved Mask-Based Neural Beamforming for Multichannel Speech Enhancement by Snapshot Matching Masking

ICASSP 2023accepted

In multichannel speech enhancement (SE), time-frequency (T-F) mask-based neural beamforming algorithms take advantage of deep neural networks to predict T-F masks that represent speech and noise dominance. The predicted masks are subsequently leveraged to estimate the speech and noise power spectral…

Cited by 0SourceScholar
2023

To Wake-Up or Not to Wake-Up: Reducing Keyword False Alarm by Successive Refinement

ICASSP 2023accepted

Keyword spotting systems continuously process audio streams to detect keywords. One of the most challenging tasks in designing such systems is to reduce False Alarm (FA) which happens when the system falsely registers a keyword despite the keyword not being uttered. In this paper, we propose a simpl…

Cited by 0SourceScholar