← Search

Andong Li

16 accepted papers

2026

GOMPSNR: Reflourish the Signal-to-Noise Ratio Metric for Audio Generation Tasks

AAAI 2026technical

In the field of audio generation, signal-to-noise ratio (SNR) has long served as an objective metric for evaluating audio quality. Nevertheless, recent studies have shown that SNR and its variants are not always highly correlated with human perception, prompting us to raise the questions: Why does

Cited by 0SourcePDFScholar
2026

SLD-L2S: Hierarchical Subspace Latent Diffusion for High-Fidelity Lip to Speech Synthesis

AAAI 2026technical

Although lip-to-speech synthesis (L2S) has achieved significant progress in recent years, current state-of-the-art methods typically rely on intermediate representations such as mel-spectrograms or discrete self-supervised learning (SSL) tokens. The potential of latent diffusion models (LDMs) in thi

Cited by 0SourcePDFScholar
2026

Tool-Grasp: A 6-DoF Functional Grasping Framework for General-Purpose Hand Tools

ICRA 2026poster

Detecting functional grasp poses for tool operation is critical for robots in complex real-world tasks, yet existing methods lack this capability. Key challenges are: 1) Scarce realworld datasets with fine-grained functional labels and task-valid grasp annotations, as their construction requires dom…

Cited by 0Scholar
2025

BSDB-Net: Band-Split Dual-Branch Network with Selective State Spaces Mechanism for Monaural Speech Enhancement

AAAI 2025technical

Although the complex spectrum-based speech enhancement (SE) methods have achieved significant performance, coupling amplitude and phase can lead to a compensation effect, where amplitude information is sacrificed to compensate for the phase that is harmful to SE. In addition, to further improve the…

Cited by 0SourcePDFScholar
2025

BridgeVoC: Neural Vocoder with Schrödinger Bridge

IJCAI 2025

While previous diffusion-based neural vocoders typically follow a noise-to-data generation pipe-line, the linear-degradation prior of the mel-spectrogram is often neglected, resulting in limited generation quality. By revisiting the vocoding task and excavating its connection with the signal restora

Cited by 0SourcePDFScholar
2025

DSINet: Towards Real-Time Target Speaker Extraction with Dynamic Speaker Information Fusion

ICASSP 2025accepted

Target speaker extraction (TSE) aims to directly extract the desired speech given enrollment utterances of the target speaker. Despite significant progress in recent years, most existing methods remain non-causal and computationally intensive. This paper introduces DSINet, a real-time time-frequency…

Cited by 0SourceScholar
2025

Learning Neural Vocoder from Range-Null Space Decomposition

IJCAI 2025

Despite the rapid development of neural vocoders in recent years, they usually suffer from some intrinsic challenges like opaque modeling, and parameter-performance trade-off. In this study, we propose an innovative time-frequency (T-F) domain-based neural vocoder to resolve the above-mentioned chal

2024

All Neural Kronecker Product Beamforming for Speech Extraction with Large-Scale Microphone Arrays

ICASSP 2024accepted

Existing frame-wise neural beamformers for speech extraction can obtain promising performance in relatively high signal-to-noise ratio (SNR) scenarios using small microphone arrays, while they still suffer from performance degradation in relatively low SNR environments, e.g., SNR<-5 dB. As an attemp…

Cited by 0SourceScholar
2024

Opine: Leveraging a Optimization-Inspired Deep Unfolding Method for Multi-Channel Speech Enhancement

ICASSP 2024accepted

Proximal gradient theory has demonstrated its superiority in the compressive sensing field for complex signal recovery. As an early trial in the speech front-end field, we propose OPINE, an optimization-inspired deep unfolding framework to simulate traditional iterative optimization process for mult…

Cited by 0SourceScholar
2023

Gesper: A Unified Framework for General Speech Restoration

ICASSP 2023accepted

This paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by spee…

Cited by 0SourceScholar
2022

Dual-Branch Attention-In-Attention Transformer for Single-Channel Speech Enhancement

ICASSP 2022accepted

Curriculum learning begins to thrive in the speech enhancement area, which decouples the original spectrum estimation task into multiple easier sub-tasks to achieve better performance. Motivated by that, we propose a dual-branch attention-in-attention transformer dubbed DB-AIAT to handle both coarse…

Cited by 0SourceScholar
2022

Embedding and Beamforming: All-Neural Causal Beamformer for Multichannel Speech Enhancement

ICASSP 2022accepted

Standing upon the intersection of traditional beamformers and deep neural networks, we propose a causal neural beamformer paradigm called Embedding and Beamforming, and two core modules are devised accordingly, namely EM and BM. For EM, instead of estimating spatial covariance matrix explicitly, the…

Cited by 0SourceScholar
2022

Joint Magnitude Estimation and Phase Recovery Using Cycle-In-Cycle GAN for Non-Parallel Speech Enhancement

ICASSP 2022accepted

For the lack of adequate paired noisy-clean speech corpus in many real scenarios, non-parallel training is a promising task for DNN-based speech enhancement methods. However, because of the severe mismatch between input and target speeches, many previous studies only focus on the magnitude spectrum…

Cited by 0SourceScholar
2022

Taylor, Can You Hear Me Now? A Taylor-Unfolding Framework for Monaural Speech Enhancement

IJCAI 2022poster

While the deep learning techniques promote the rapid development of the speech enhancement (SE) community, most schemes only pursue the performance in a black-box manner and lack adequate model interpretability. Inspired by Taylor's approximation theory, we propose an interpretable decoupling-style…

2021

ICASSP 2021 Deep Noise Suppression Challenge: Decoupling Magnitude and Phase Optimization with a Two-Stage Deep Network

ICASSP 2021accepted

It remains a tough challenge to recover the speech signals contaminated by various noises under real acoustic environments. To this end, we propose a novel system for denoising in the complicated applications, which is mainly comprised of two pipelines, namely a two-stage network and a post-processi…

Cited by 0SourceScholar