← Search

Yi-Wen Liu

5 accepted papers

2021

Attacking and Defending Behind A Psychoacoustics-Based Captcha

ICASSP 2021accepted

This paper proposes a novel audio CAPTCHA system that requires a user to respond immediately after hearing a short and easy-to-remember cue in its mixture with background music. Potential attacking paths based on cross correlation (CC) and sound event detection (SED) are implemented to test the secu…

Cited by 0SourceScholar
2021

Sound Event Detection by Consistency Training and Pseudo-Labeling With Feature-Pyramid Convolutional Recurrent Neural Networks

ICASSP 2021accepted

Due to the high cost of large-scale strong labeling, sound event detection (SED) using only weakly-labeled and unlabeled data has drawn increasing attention in recent years. To exploit large amount of unlabeled in-domain data efficiently, we applied three semi-supervised learning strategies: interpo…

Cited by 0SourceScholar
2019

Stereo Source Separation in the Frequency Domain: Solving the Permutation Problem by a Sliding K-means Method

ICASSP 2019accepted

Blind source separation (BSS) has been widely utilized for recovering a set of source signals from their mixtures. When the mixture is convolutive, source separation can be solved in the frequency domain but involves several challenges including the scaling uncertainty and the permutation indetermin…

Cited by 0SourceScholar
2016

Posterior probabilistic modeling for inter-channel phase and time difference estimation in audio signals

ICASSP 2016accepted

A method is proposed for the estimation of the time difference of arrival (TDOA) from a sound source to a pair of microphones. Given noisy observations of the source, the magnitude spectrum of the source is first estimated via smoothing across time frames. Then, a probability density function (PDF)…

Cited by 0SourceScholar
2015

Joint estimation of vocal tract and nasal tract area functions from speech waveforms via auto-regression moving-average modeling and a pole assignment method

ICASSP 2015accepted

Nasal resonance is utilized in certain languages to differentiate word meanings. The joint filtering effect by the vocal tract and the nasal tract can be modeled by the auto-regression moving-average (ARMA) approach. However, unlike all-pole (i.e., AR) modeling, it has been difficult to derive the e…

Cited by 0SourceScholar