← Search

Weiping Tu

21 accepted papers

2026

DeformTrace: A Deformable State Space Model with Relay Tokens for Temporal Forgery Localization

AAAI 2026technical

Temporal Forgery Localization (TFL) aims to precisely identify manipulated segments in video and audio, offering strong interpretability for security and forensics. While recent State Space Models (SSMs) show promise in precise temporal reasoning, their use in TFL is hindered by ambiguous boundaries

Cited by 0SourcePDFScholar
2026

GEM-TFL: Bridging Weak and Full Supervision for Forgery Localization through EM-Guided Decomposition and Temporal Refinement

CVPR 2026

Temporal Forgery Localization (TFL) aims to precisely identify manipulated segments within videos or audio streams, providing interpretable evidence for multimedia forensics and security. While most existing TFL methods rely on dense frame-level labels in a fully supervised manner, Weakly Supervised

Cited by 0SourceScholar
2025

Attention Weighting and Conditional Entropy-driven Quantization Loss for Neural Audio Codecs

ICASSP 2025accepted

Existing end-to-end neural codecs have made great progress in preserving audio quality. Despite their success, they still face challenges in achieving accurate and efficient quantization. Specifically, these codecs often overlook which features have a greater impact on perceptual audio quality durin…

Cited by 0SourceScholar
2025

FIRING-Net: A filtered feature recycling network for speech enhancement

ICLR 2025poster

Current deep neural networks for speech enhancement (SE) aim to minimize the distance between the output signal and the clean target by filtering out noise features from input features. However, when noise and speech components are highly similar, SE models struggle to learn effective discrimination…

Cited by 0SourcePDFScholar
2025

FreqSense: Universal and Low-Latency Adversarial Example Detection for Speaker Recognition with Interpretability in Frequency Domain

ICASSP 2025accepted

Speaker recognition (SR) systems are particularly vulnerable to adversarial example (AE) attacks. To mitigate these attacks, AE detection systems are typically integrated into SR systems. To overcome the limitations of low detection accuracy, poor generalization, and high latency in existing schemes…

Cited by 0SourceScholar
2025

HAPG-SAQAM: Human Auditory Perception Guided Spatial Audio Quality Assessment Metric

ICASSP 2025accepted

Spatial audio quality evaluation is essential for applications like virtual and augmented reality, where accurate sound reproduction enhances user immersion. While subjective listening tests are the gold standard, they are costly and time-consuming. To address this, we propose HAPG-SAQAM, an objecti…

Cited by 0SourceScholar
2025

Improving Speech Enhancement by Cross- and Sub-band Processing with State Space Model

ICASSP 2025accepted

Recently, the state space model (SSM) represented by Mamba has shown remarkable performance in long-term sequence modeling tasks, including speech enhancement. However, due to substantial differences in sub-band features, applying the same SSM to all sub-bands limits its inference capability. Additi…

Cited by 0SourceScholar
2024

An Efficient and Interpre Table Speech Enhancement Network Via Deep Dictionary Learning

ICASSP 2024accepted

Speech enhancement is a vital and highly ill-posed problem for many speech downstream tasks. While currently existing deep learning based speech enhancement methods have held state-of-the-art results, they still possess apparent shortcomings in that most of the deep learning based models lack interp…

Cited by 0SourceScholar
2024

Curricular Contrastive Regularization for Speech Enhancement with Self-Supervised Representations

ICASSP 2024accepted

Existing deep learning-based speech enhancement methods only adopt clean speech as positive samples to guide the training of speech enhancement networks while negative samples, i.e., noisy speech, are unexploited. In this paper, we adopt contrastive regularization (CR) built upon contrastive learnin…

Cited by 0SourceScholar
2024

EMALG: An Enhanced Mandarin Lombard Grid Corpus with Meaningful Sentences

ICASSP 2024accepted

This study investigates the Lombard effect, where individuals adapt their speech in noisy environments. We introduce an enhanced Mandarin Lombard grid (EMALG) corpus with meaningful sentences, enhancing the Mandarin Lombard grid (MALG) corpus. EMALG features 34 speakers and improves recording setups…

Cited by 0SourceScholar
2024

Improving Acoustic Echo Cancellation by Exploring Speech and Echo Affinity with Multi-Head Attention

ICASSP 2024accepted

Deep learning-based approaches formulate acoustic echo cancellation (AEC) as a supervised speech separation task, where the mixture signal and the far-end signal are combined directly before or after the encoding stage. However, the mixture signal and the far-end signal are not integrated sufficient…

Cited by 0SourceScholar
2024

SuperCodec: A Neural Speech Codec with Selective Back-Projection Network

ICASSP 2024accepted

Neural speech coding is a rapidly developing topic, where state-of-the-art approaches now exhibit superior compression performance than conventional methods. Despite significant progress, existing methods still have limitations in preserving and reconstructing fine details for optimal reconstruction…

Cited by 0SourceScholar
2023

Improving Acoustic Echo Cancellation by Mixing Speech Local and Global Features with Transformer

ICASSP 2023accepted

We propose MiT-Net, a novel mix-transformer neural network with a pyramid encoder operating in the time domain, for the task of acoustic echo cancellation. The MiT-Net formulates acoustic echo cancellation as a supervised speech separation problem, in which near-end speech is separated from a single…

Cited by 0SourceScholar
2023

Learning From Single-Expert Annotated Labels for Automatic Sleep Staging

ICASSP 2023accepted

Existing automatic sleep staging algorithms rely on accurately labeled data. However, due to the subjectivity of sleep experts, accurate labels must be obtained through joint labeling by multiple experts, which results in high time and labor costs. In this work, we treat labels mislabeled by a singl…

Cited by 0SourceScholar
2023

PMMSD: Development of the Matrix Sentence Intelligibility Dataset for Mandarin with Lombard Effect

ICASSP 2023accepted

This paper presents a Paired Mandarin Matrix Sentence Dataset (PMMSD), which will be available after publication. PMMSD is the first Mandarin matrix sentence intelligibility dataset containing both plain and Lombard speech for scientific research. The results verify that different Lombard styles wou…

Cited by 0SourceScholar
2023

Selector-Enhancer: Learning Dynamic Selection of Local and Non-local Attention Operation for Speech Enhancement

AAAI 2023technical

Attention mechanisms, such as local and non-local attention, play a fundamental role in recent deep learning based speech enhancement (SE) systems. However, a natural speech contains many fast-changing and relatively briefly acoustic events, therefore, capturing the most informative speech features…

Cited by 11SourcePDFScholar
2019

Kullback-Leibler Divergence Frequency Warping Scale for Acoustic Scene Classification Using Convolutional Neural Network

ICASSP 2019accepted

Most of current best performing Acoustic Scene Classification (ASC) systems utilize Mel scale spectrograms with Convolutional Neural Networks (CNNs). Mel scale is a common way to suit frequency warping of human ears, with strict decreasing frequency resolution on low to high frequency range. However…

Cited by 0SourceScholar
2017

Sound physical property matching between non central listening point and central listening point for NHK 22.2 system reproduction

ICASSP 2017accepted

NHK has proposed a famous 3D audio system: 22.2 multi-channel system, but its loudspeakers are too many and are troublesome to put in home. Ando and Wang has proposed two simplification methods to reduce its channel number, but only 3D sound field at the central listening point can be recovered well…

Cited by 0SourceScholar
2015

A down-mixing method for 22.2 multichannel system reproduction

ICASSP 2015accepted

This paper proposes a general multichannel system reproduction method. Firstly, relative to original multichannel system, a general global model is build up by guaranteeing sound pressure and the direction of particle velocity at the receiving point constant, and making the square error of particle…

Cited by 0SourceScholar