← Search

Chung-Hsien Wu

10 accepted papers

2021

Assessment of Bipolar Disorder Using Heterogeneous Data of Smartphone-Based Digital Phenotyping

ICASSP 2021accepted

In mental health disorder, Bipolar Disorder (BD) is one of the most common mental illness. Using rating scales for assessment is one of the approaches for diagnosing and tracking BD patients. However, the requirement for manpower and time is heavy in the process of evaluation. In order to reduce the…

Cited by 0SourceScholar
2020

Combining Deep Embeddings of Acoustic and Articulatory Features for Speaker Identification

ICASSP 2020accepted

In this study, deep embedding of acoustic and articulatory features are combined for speaker identification. First, a convolutional neural network (CNN)-based universal background model (UBM) is constructed to generate acoustic feature (AC) embedding. In addition, as the articulatory features (AFs)…

Cited by 0SourceScholar
2020

Statistics Pooling Time Delay Neural Network Based on X-Vector for Speaker Verification

ICASSP 2020accepted

This paper aims to improve speaker embedding representation based on x-vector for extracting more detailed information for speaker verification. We propose a statistics pooling time delay neural network (TDNN), in which the TDNN structure integrates statistics pooling for each layer, to consider the…

Cited by 0SourceScholar
2019

Speech Emotion Recognition Using Deep Neural Network Considering Verbal and Nonverbal Speech Sounds

ICASSP 2019accepted

Speech emotion recognition is becoming increasingly important for many applications. In real-life communication, non-verbal sounds within an utterance also play an important role for people to recognize emotion. In current studies, only few emotion recognition systems considered nonverbal sounds, su…

Cited by 0SourceScholar
2018

Attention-Based Dialog State Tracking for Conversational Interview Coaching

ICASSP 2018accepted

This study proposes an approach to dialog state tracking (DST) in a conversational interview coaching system. For the interview coaching task, the semantic slots, used mostly in traditional dialog systems, are difficult to define manually. This study adopts the topic profile of the response from the…

Cited by 0SourceScholar
2018

Locality-Preserving Complex-Valued Gaussian Process Latent Variable Model for Robust Face Recognition

ICASSP 2018accepted

Learning a low-dimensional image representation yields effective and efficient face recognition. The use of such a representation helps to weaken the curse of dimensionality. However, the traditional facial representation method is not robust against partial occlusions or variations of expression. T…

Cited by 0SourceScholar
2017

Fully complex deep neural network for phase-incorporating monaural source separation

ICASSP 2017accepted

Deep neural network (DNN) have become a popular means of separating a target source from a mixed signal. Most of DNN-based methods modify only the magnitude spectrum of the mixture. The phase spectrum is left unchanged, which is inherent in the short-time Fourier transform (STFT) coefficients of the…

Cited by 0SourceScholar
2017

Mood detection from daily conversational speech using denoising autoencoder and LSTM

ICASSP 2017accepted

In current studies, an extended subjective self-report method is generally used for measuring emotions. Even though it is commonly accepted that speech emotion perceived by the listener is close to the intended emotion conveyed by the speaker, research has indicated that there still remains a mismat…

Cited by 0SourceScholar
2015

Affective structure modeling of speech using probabilistic context free grammar for emotion recognition

ICASSP 2015accepted

A complete emotional expression typically contains a complex temporal course in a natural conversation. Related research on utterance-level and segment-level processing lacks understanding of the underlying structure of emotional speech. In this study, a hierarchical affective structure of an emotio…

Cited by 0SourceScholar