← Search

Tai-Shih Chi

11 accepted papers

2021

Extending Music Based On Emotion And Tonality Via Generative Adversarial Network

ICASSP 2021accepted

We propose a generative model for music extension in this paper. The model is composed of two classifiers, one for music emotion and one for music tonality, and a generative adversarial network (GAN). Therefore, it can generate symbolic music not only based on low level spectral and temporal charact…

Cited by 0SourceScholar
2020

A Multi-Dilation and Multi-Resolution Fully Convolutional Network for Singing Melody Extraction

ICASSP 2020accepted

Each human cognitive function involves bottom-up and top-down processes. Several methods have been proposed for singing melody extraction by emphasizing either the bottom-up or top-down processes. For hearing, the bottom-up processes include spectral and spectro-temporal decomposition of the sound b…

Cited by 0SourceScholar
2019

Autoencoding HRTFS for DNN Based HRTF Personalization Using Anthropometric Features

ICASSP 2019accepted

We proposed a deep neural network (DNN) based approach to synthesize the magnitude of personalized head-related transfer functions (HRTFs) using anthropometric features of the user. To mitigate the over-fitting problem when training dataset is not very large, we built an autoencoder for dimensional…

Cited by 0SourceScholar
2019

CNN Based Two-stage Multi-resolution End-to-end Model for Singing Melody Extraction

ICASSP 2019accepted

Inspired by human hearing perception, we propose a two-stage multi-resolution end-to-end model for singing melody extraction in this paper. The convolutional neural network (CNN) is the core of the proposed model to generate multi-resolution representations. The 1-D and 2-D multi-resolution analysis…

Cited by 0SourceScholar
2019

Reinforcement Learning Based Speech Enhancement for Robust Speech Recognition

ICASSP 2019accepted

Conventional deep neural network (DNN)-based speech enhancement (SE) approaches aim to minimize the mean square error (MSE) between enhanced speech and clean reference. The MSE-optimized model may not directly improve the performance of an automatic speech recognition (ASR) system. If the target is…

Cited by 0SourceScholar
2018

A Generative Auditory Model Embedded Neural Network for Speech Processing

ICASSP 2018accepted

Before the era of the neural network (NN), features extracted from auditory models have been applied to various speech applications and been demonstrated more robust against noise than conventional speech-processing features. What's the role of auditory models in the current NN era? Are they obsolet…

Cited by 0SourceScholar
2018

A Hybrid Neural Network Based on the Duplex Model of Pitch Perception for Singing Melody Extraction

ICASSP 2018accepted

In this paper, we build up a hybrid neural network (NN) for singing melody extraction from polyphonic music by imitating human pitch perception. For human hearing, there are two pitch perception models, the spectral model and the temporal model, in accordance with whether harmonics are resolved or n…

Cited by 0SourceScholar
2015

A hearing model to estimate mandarin speech intelligibility for the hearing impaired patients

ICASSP 2015accepted

A hearing model, which is parameterized by hearing thresholds, degrees of loudness recruitment and reductions of frequency resolution of a hearing-impaired (HI) patient, is proposed in this paper. The model is developed in the filter-bank framework and is flexible for fitting hearing-loss conditions…

Cited by 0SourceScholar
2015

Modulation Wiener filter for improving speech intelligibility

ICASSP 2015accepted

This paper presents a single-channel high-dimensional Wiener filter in the spectro-temporal modulation domain. Unlike other conventional noise reduction techniques, the proposed algorithm not only reduces noise but also enhances the “textures” of the speech signal. A non-iterative decision-directed…

Cited by 0SourceScholar