← Search

Shan Liang

9 accepted papers

2025

Adversarial Training and Gradient Optimization for Partially Deepfake Audio Localization

ICASSP 2025accepted

Partially deepfake audio localization is important in audio forensics. However, existing localization models for partially deepfake audio face two major challenges: distribution shifts between training and testing data as well as insufficient utilization of information from both manipulated regions…

Cited by 0SourceScholar
2025

OV-MER: Towards Open-Vocabulary Multimodal Emotion Recognition

ICML 2025poster

Multimodal Emotion Recognition (MER) is a critical research area that seeks to decode human emotions from diverse data modalities. However, existing machine learning methods predominantly rely on predefined emotion taxonomies, which fail to capture the inherent complexity, subtlety, and multi-apprai…

2022

A Robust Deep Audio Splicing Detection Method via Singularity Detection Feature

ICASSP 2022accepted

There are many methods for detecting forged audio produced by conversion and synthesis. However, as a simpler method of forgery, splicing has not attracted widespread attention. Based on the characteristic that the tampering operation will cause singularities at high-frequency components, we propose…

Cited by 0SourceScholar
2022

ADD 2022: the first Audio Deep Synthesis Detection Challenge

ICASSP 2022accepted

Audio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three t…

Cited by 0SourceScholar
2019

Adaptive Dereverberation Using Multi-channel Linear Prediction with Deficient Length Filter

ICASSP 2019accepted

In almost all adaptive dereverberation algorithms based on the multi-channel linear prediction (MCLP) model, it is assumed that the filter length can cover the reverberation time. However, in many practical situations, a deficient length filter, whose length is less than the reverberation time, is e…

Cited by 0SourceScholar
2019

Loss and Double-edge-triggered Detector for Robust Small-footprint Keyword Spotting

ICASSP 2019accepted

Keyword spotting (KWS) system constitutes a critical component of human-computer interfaces, which detects the specific keyword from a continuous stream of audio. The goal of KWS is providing a high detection accuracy at a low false alarm rate while having small memory and computation requirements.…

Cited by 0SourceScholar
2018

Boosting Noise Robustness of Acoustic Model via Deep Adversarial Training

ICASSP 2018accepted

In realistic environments, speech is usually interfered by various noise and reverberation, which dramatically degrades the performance of automatic speech recognition (ASR) systems. To alleviate this issue, the commonest way is to use a well-designed speech enhancement approach as the front-end of…

Cited by 0SourceScholar
2016

Exploiting spectro-temporal structures using NMF for DNN-based supervised speech separation

ICASSP 2016accepted

The targets of speech separation, whether ideal masks or magnitude spectrograms of interest, have prominent spectro-temporal structures. These characteristics are very worthy to be exploited for speech separation, however, they are usually ignored in previous works. In this paper, we use nonnegative…

Cited by 0SourceScholar
2015

Cross-domain cooperative deep stacking network for speech separation

ICASSP 2015accepted

Nowadays supervised speech separation has drawn much attention and shown great promise in the meantime. While there has been a lot of success, existing algorithms perform the task only in one preselected representative domain. In this study, we propose to perform the task in two different time-frequ…

Cited by 0SourceScholar