← Search

Wei Rao

13 accepted papers

2026

Chain-of-Glimpse: Search-Guided Progressive Object-Grounded Reasoning for Video Understanding

ICML 2026poster

Video understanding requires identifying and reasoning over semantically discriminative visual objects across frames, yet existing object-agnostic solutions struggle to effectively handle substantial object variations over time. To address this, we introduce Chain-of-Glimpse, a search-guided progres…

Cited by 0SourceScholar
2025

Aesthetic Perception Prompting for Interpretable Image Aesthetics Assessment with MLLMs

ICASSP 2025accepted

Image Aesthetic Assessment (IAA) aims to rate the aesthetic quality of images and has many practical applications. However, existing methods typically rely on limited annotated data for training, leading to two key issues: 1) score-only predictions lack interpretability, making it hard for users to…

Cited by 0SourceScholar
2024

Hierarchical Speaker Representation for Target Speaker Extraction

ICASSP 2024accepted

Target speaker extraction aims to isolate a specific speaker’s voice from a composite of multiple sound sources, guided by an enrollment utterance or called anchor. Current methods predominantly derive speaker embeddings from the anchor and integrate them into the separation network to separate the…

Cited by 0SourceScholar
2023

Distance-Based Weight Transfer for Fine-Tuning From Near-Field to Far-Field Speaker Verification

ICASSP 2023accepted

The scarcity of labeled far-field speech is a constraint for training superior far-field speaker verification systems. In general, fine-tuning the model pre-trained on large-scale near- field speech through a small amount of far-field speech substantially outperforms training from scratch. However,…

Cited by 0SourceScholar
2023

Gesper: A Unified Framework for General Speech Restoration

ICASSP 2023accepted

This paper describes the legends-tencent team’s real-time General Speech Restoration (Gesper) system submitted to the ICASSP 2023 Speech Signal Improvement (SSI) Challenge. This newly proposed system is a two-stage architecture, in which the speech restoration is performed, and then followed by spee…

Cited by 0SourceScholar
2023

Inter-Subnet: Speech Enhancement with Subband Interaction

ICASSP 2023accepted

Subband-based approaches process subbands in parallel through the model with shared parameters to learn the commonality of local spectrums for noise reduction. In this way, they have achieved remarkable results with fewer parameters. However, in some complex environments, the lack of global spectral…

Cited by 0SourceScholar
2023

Speech Enhancement with Intelligent Neural Homomorphic Synthesis

ICASSP 2023accepted

Most neural network speech enhancement models ignore speech production mathematical models by directly mapping Fourier transform spectrums or waveforms. In this work, we propose a neural source filter network for speech enhancement. Specifically, we use homomorphic signal processing and cepstral ana…

Cited by 0SourceScholar
2023

TEA-PSE 3.0: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System For ICASSP 2023 Dns-Challenge

ICASSP 2023accepted

This paper introduces the Unbeatable Team’s submission to the ICASSP 2023 Deep Noise Suppression (DNS) Challenge. We expand our previous work, TEA-PSE, to its upgraded version – TEA-PSE 3.0. Specifically, TEA-PSE 3.0 incorporates a residual LSTM after squeezed temporal convolution network (S-TCN) to…

Cited by 0SourceScholar
2022

TEA-PSE: Tencent-Ethereal-Audio-Lab Personalized Speech Enhancement System for ICASSP 2022 DNS Challenge

ICASSP 2022accepted

This paper describes Tencent Ethereal Audio Lab – Northwestern Polytechnical University personalized speech enhancement (TEA-PSE) system submitted to track 2 of the ICASSP 2022 Deep Noise Suppression (DNS) challenge. Our system specifically combines the dual-stage network which is a superior real-ti…

Cited by 56SourceScholar
2019

Optimization of Speaker Extraction Neural Network with Magnitude and Temporal Spectrum Approximation Loss

ICASSP 2019accepted

The SpeakerBeam-FE (SBF) method is proposed for speaker extraction. It attempts to overcome the problem of unknown number of speakers in an audio recording during source separation. The mask approximation loss of SBF is sub-optimal, which doesn't calculate direct signal reconstruction error and cons…

Cited by 0SourceScholar
2018

Single Channel Speech Separation with Constrained Utterance Level Permutation Invariant Training Using Grid LSTM

ICASSP 2018accepted

Utterance level permutation invariant training (uPIT) technique is a state-of-the-art deep learning architecture for speaker independent multi-talker separation. uPIT solves the label ambiguity problem by minimizing the mean square error (MSE) over all permutations between outputs and targets. Howev…

Cited by 0SourceScholar
2018

Unsupervised Domain Adaptation via Domain Adversarial Training for Speaker Recognition

ICASSP 2018accepted

The i-vector approach to speaker recognition has achieved good performance when the domain of the evaluation dataset is similar to that of the training dataset. However, in realworld applications, there is always a mismatch between the training and evaluation datasets, that leads to performance degr…

Cited by 0SourceScholar