← Search

Mingjiang Wang

10 accepted papers

2025

CSMT: Combining Snoring and Metadata-based Text for Sleep Apnea Severity Classification

ICASSP 2025accepted

Sleep apnea is a common sleep disorder that, if untreated, can lead to serious health issues. Snoring is a typical symptom of sleep apnea and can be utilized to develop a noncontact automatic detection method for sleep apnea severity classification (SASC). However, due to patient heterogeneity, the…

Cited by 0SourceScholar
2025

Dual Position Attention Time-Frequency Network for Binaural Audio Synthesis

ICASSP 2025accepted

In applications such as virtual reality and augmented reality, binaural audio provides listeners with a more immersive experience. To synthesize binaural audio with enhanced spatial localization, especially in scenarios involving moving sound sources, accurate phase estimation is crucial. However, e…

Cited by 0SourceScholar
2025

Joint Training Framework for Accent and Speech Recognition Based on Conformer Low-Rank Adaptation

ICASSP 2025accepted

In real-world scenarios, accent variations often reduce Automatic Speech Recognition (ASR) accuracy. Addressing this typically involves a multi-task ASR and Accent Recognition (ASR-AR) framework, but there is limited research on optimizing task-specific feature extraction and enhancing ASR with AR i…

Cited by 0SourceScholar
2025

PriorSinger: Singing Voice Synthesis Model with Prior Condition Cross Attention

ICASSP 2025accepted

The singing voice synthesis system is designed to generate realistic and expressive singing based on a given musical score. Generative Adversarial Networks (GANs) or diffusion models generate acoustic features, such as Mel-spectrograms, which are subsequently reconstructed into waveforms by a vocode…

Cited by 0SourceScholar
2024

Hybrid Attention Time-Frequency Analysis Network for Single-Channel Speech Enhancement

ICASSP 2024accepted

The time-frequency domain remains central to the speech signal analysis. Enhancing the efficacy of neural network-based speech models demands a detailed multi-scale analysis of time-frequency features. This study presents the Hybrid Attention Time-Frequency Analysis Network (HATFANet), an innovative…

Cited by 0SourceScholar
2024

Lightweight Multi-Axial Transformer with Frequency Prompt for Single Channel Speech Enhancement

ICASSP 2024accepted

Time-frequency analysis in single-channel speech enhancement has received considerable attention. While Transformer-based architectures are gaining traction, their computational burden can be substantial, especially when dealing with longer speech samples. To address this, our research introduces th…

Cited by 0SourceScholar
2023

Half-Temporal and Half-Frequency Attention U2Net for Speech Signal Improvement

ICASSP 2023accepted

During communication, volume changes, noise, and reverberation can disturb speech signals, significantly affecting the quality and intelligibility of speech. In the context of the ICASSP 2023 Signal Processing Grand Challenge, the first Speech Signal Improvement Grand Challenge (SIG) is organized to…

Cited by 0SourceScholar
2023

Two-Stage UNet with Multi-Axis Gated Multilayer Perceptron for Monaural Noisy-Reverberant Speech Enhancement

ICASSP 2023accepted

In denoising and de-reverberation tasks, the dominant methods are complex spectral masking and complex spectral mapping. To combine advantages and improve speech enhancement performance, we propose a two-stage UNet (TSUNet) to estimate complex spectral masking and complex spectral mapping. We use a…

Cited by 0SourceScholar
2022

Design of Real-Time System Based on Machine Learning for Snoring and OSA Detection

ICASSP 2022accepted

Obstructive sleep apnea (OSA) is a common sleep disorder. The diagnosis of OSA based on snoring is low-cost, convenient and non-invasive. In this study, we place a microphone under the patient’s bed and combined with full-night polysomnography to record audio signals. Five machine learning models an…

Cited by 0SourceScholar
2022

FB-MSTCN: A Full-Band Single-Channel Speech Enhancement Method Based on Multi-Scale Temporal Convolutional Network

ICASSP 2022accepted

In recent years, deep learning-based approaches have significantly improved the performance of single-channel speech enhancement. However, due to the limitation of training data and computational complexity, real-time enhancement of full-band (48 kHz) speech signals is still very challenging. Becaus…

Cited by 0SourceScholar