← Search

Meng Ge

21 accepted papers

2025

A Chinese Expressive Long-dialogue Speech Dataset with Scripts

ICASSP 2025accepted

With the advancement of large-scale models, the demand for emotionally rich, long-context, and highly natural communication in human-computer interaction increases. However, the exploration of long-context or script-level speech conversation tasks remains limited due to the lack of specific supervis…

Cited by 0SourceScholar
2025

Augmenting Short Enrollment Speech via Synthesis for Target Speaker Extraction

ICASSP 2025accepted

A high-quality enrollment speech is crucial to target speaker extraction (TSE), since it provides essential cues for identifying the target speaker in the mixture. However, real applications usually only permit a short enrollment speech, e.g. a wakeup word for a mobile device, that provides limited…

Cited by 0SourceScholar
2025

Mamba-SEUNet: Mamba UNet for Monaural Speech Enhancement

ICASSP 2025accepted

In recent speech enhancement (SE) research, transformer and its variants have emerged as the predominant methodologies. However, the quadratic complexity of the self-attention mechanism imposes certain limitations on practical deployment. Mamba, as a novel state-space model (SSM), has gained widespr…

Cited by 0SourceScholar
2025

Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging Module

ICASSP 2025accepted

The information loss or distortion caused by single-channel speech enhancement (SE) harms the performance of automatic speech recognition (ASR). Observation addition (OA) is an effective post-processing method to improve ASR performance by balancing noisy and enhanced speech. Determining the OA coef…

Cited by 0SourceScholar
2025

Time-Graph Frequency Representation with Singular Value Decomposition for Neural Speech Enhancement

ICASSP 2025accepted

Time-frequency (T-F) domain methods for monaural speech enhancement have benefited from the success of deep learning. Recently, focus has been put on designing two-stream network models to predict amplitude mask and phase separately, or, coupling the amplitude and phase into Cartesian coordinates an…

Cited by 0SourceScholar
2025

Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis

NeurIPS 2025spotlight

While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modelin…

Cited by 0SourceScholar
2024

An Empirical Study on the Impact of Positional Encoding in Transformer-Based Monaural Speech Enhancement

ICASSP 2024accepted

Transformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to enable Transformers to distinguish the order of elements in a sequence. However, it remains unclear how positional encod…

Cited by 0SourceScholar
2024

Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-Talker Speech

ICASSP 2024accepted

Target speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario where the target speech is highly overlapped with the interfering speech. However, this scenario only accounts for a small…

Cited by 0SourceScholar
2024

FUG: Feature-Universal Graph Contrastive Pre-training for Graphs with Diverse Node Features

NeurIPS 2024poster

Graph Neural Networks (GNNs), known for their effective graph encoding, are extensively used across various fields. Graph self-supervised pre-training, which trains GNN encoders without manual labels to generate high-quality graph representations, has garnered widespread attention. However, due to t…

2024

Gradient Weighting for Speaker Verification in Extremely Low Signal-to-Noise Ratio

ICASSP 2024accepted

Speaker verification is hampered by background noise, particularly at extremely low Signal-to-Noise Ratio (SNR) under 0 dB. It is difficult to suppress noise without introducing unwanted artifacts, which adversely affects speaker verification. We proposed the mechanism called Gradient Weighting (Gra…

Cited by 0SourceScholar
2024

Improving Distinguishability of Class for Graph Neural Networks

AAAI 2024technical

Graph Neural Networks (GNNs) have received widespread attention and applications due to their excellent performance in graph representation learning. Most existing GNNs can only aggregate 1-hop neighbors in a GNN layer, so they usually stack multiple GNN layers to obtain more information from larger…

Cited by 3SourcePDFScholar
2024

Multi-Modal Sarcasm Detection Based on Dual Generative Processes

IJCAI 2024poster

With the advancement of the internet, sarcastic sentiment expression on social media has grown increasingly diverse. Consequently, multimodal sarcasm detection has emerged as a valuable tool for users to comprehend and interpret sarcastic expressions. Previous research suggests that effectively inte…

Cited by 3SourcePDFScholar
2024

SVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural Networks

ICASSP 2024accepted

Speech applications are expected to be low-power and robust under noisy conditions. An effective Voice Activity Detection (VAD) front-end lowers the computational need. Spiking Neural Networks (SNNs) are known to be biologically plausible and power-efficient. However, SNN-based VADs have yet to achi…

Cited by 0SourceScholar
2024

Unveiling Implicit Deceptive Patterns in Multi-Modal Fake News via Neuro-Symbolic Reasoning

AAAI 2024technical

In the current Internet landscape, the rampant spread of fake news, particularly in the form of multi-modal content, poses a great social threat. While automatic multi-modal fake news detection methods have shown promising results, the lack of explainability remains a significant challenge. Existing…

Cited by 12SourcePDFScholar
2023

Stream Attention Based U-Net for L3DAS23 Challenge

ICASSP 2023accepted

Machine learning applications of 3D audio are gaining increasing interest in recent years. In this paper, we propose a stream attention based U-Net to remove background noise and reverberation based on ICASSP Signal Processing Grand Challenge 2023: L3DAS23 Challenge<sup xmlns:mml="http://www.w3.org/…

Cited by 0SourceScholar
2022

Compressing Transformer-Based ASR Model by Task-Driven Loss and Attention-Based Multi-Level Feature Distillation

ICASSP 2022accepted

The current popular knowledge distillation (KD) methods effectively compress the transformer-based end-to-end speech recognition model. However, existing methods fail to utilize complete information of the teacher model, and they distill only a limited number of blocks of the teacher model. In this…

Cited by 0SourceScholar
2022

L-SpEx: Localized Target Speaker Extraction

ICASSP 2022accepted

Speaker extraction aims to extract the target speaker’s voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction of the target speaker. However, these studies assume that the target speaker’s…

Cited by 0SourceScholar
2022

RAW-GNN: RAndom Walk Aggregation based Graph Neural Network

IJCAI 2022poster

Graph-Convolution-based methods have been successfully applied to representation learning on homophily graphs where nodes with the same label or similar attributes tend to connect with one another. Due to the homophily assumption of Graph Convolutional Networks (GCNs) that these methods use, they ar…

Cited by 49SourcePDFScholar
2021

Multi-Stage Speaker Extraction with Utterance and Frame-Level Reference Signals

ICASSP 2021accepted

Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full advantage of short reference speech sample. The extracted s…

Cited by 0SourceScholar
2021

Robust Voice Activity Detection Using a Masked Auditory Encoder Based Convolutional Neural Network

ICASSP 2021accepted

Voice activity detection (VAD) based on deep learning has achieved remarkable success. However, when the traditional features (e.g., raw waveforms and MFCCs) are directly fed to the deep neural network model, the performance decreases because of noise interference. Here, we propose a robust VAD appr…

Cited by 0SourceScholar
2020

Spectrograms Fusion with Minimum Difference Masks Estimation for Monaural Speech Dereverberation

ICASSP 2020accepted

Spectrograms fusion is an effective method for incorporating complementary speech dereverberation systems. Previous linear spectrograms fusion by averaging multiple spectrograms shows outstanding performance. However, various systems with different features cannot apply this simple method. In this s…

Cited by 0SourceScholar