← Search

Jen-Tzung Chien

32 accepted papers

2026

Risk-Aware Bilingual Spoken Dialogue for Campus Mental Health Support

AAAI 2026technical

This work presented a web-based system which introduces an active-listening strategy in a spoken dialogue for self-disclosure to support mental health of a campus user. To enhance the system usability and safety, this demo is developed to conduct the bilingual (Mandarin/English) spoken dialogue wher

Cited by 0SourcePDFScholar
2025

Attention Disentanglement for Semantic Diffusion Modeling in Text-to-Image Generation

ICASSP 2025accepted

Text-to-image model has been recently improved to generate the semantically rich high-quality images by strengthening natural language processing via transformer in a stable diffusion process. However, the challenges are still remained in accurately rendering the objects, colors and compositions, an…

Cited by 0SourceScholar
2024

Asymmetric Clean Segments-Guided Self-Supervised Learning for Robust Speaker Verification

ICASSP 2024accepted

Contrastive self-supervised learning (CSL) for speaker verification (SV) has drawn increasing interest recently due to its ability to exploit unlabeled data. Performing data augmentation on raw waveforms, such as adding noise or reverberation, plays a pivotal role in achieving promising results in S…

Cited by 7SourceScholar
2024

Attention-Guided Adaptation for Code-Switching Speech Recognition

ICASSP 2024accepted

The prevalence of the powerful multilingual models, such as Whisper, has significantly advanced the researches on speech recognition. However, these models often struggle with handling the code-switching setting, which is essential in multilingual speech recognition. Recent studies have attempted to…

Cited by 0SourceScholar
2022

Augmentation Strategy Optimization for Language Understanding

ICASSP 2022accepted

This paper presents a new language processing and understanding where an adaptive data augmentation strategy for individual documents is proposed instead of using one universal policy for the whole dataset. Importantly, a reinforcement learning and understanding method is exploited for document clas…

Cited by 0SourceScholar
2020

Information Maximized Variational Domain Adversarial Learning for Speaker Verification

ICASSP 2020accepted

Domain mismatch is a common problem in speaker verification. This paper proposes an information-maximized variational domain adversarial neural network (InfoVDANN) to reduce domain mismatch by incorporating an InfoVAE into domain adversarial training (DAT). DAT aims to produce speaker discriminative…

Cited by 0SourceScholar
2019

Semi-supervised Nuisance-attribute Networks for Domain Adaptation

ICASSP 2019accepted

How to overcome the training and test data mismatch in speaker verification systems has been a focus of research recently. In this paper, we propose a semi-supervised nuisance attribute network (SNAN) to reduce the domain mismatch in i-vectors and x-vectors. SNANs are based on the idea of nuisance a…

Cited by 0SourceScholar
2016

Automatic speech recognition for acoustical analysis and assessment of cantonese pathological voice and speech

ICASSP 2016accepted

This paper describes the application of state-of-the-art automatic speech recognition (ASR) systems to objective assessment of voice and speech disorders. Acoustical analysis of speech has long been considered a promising approach to non-invasive and objective assessment of people. In the past the t…

Cited by 0SourceScholar
2016

Discriminative deep recurrent neural networks for monaural speech separation

ICASSP 2016accepted

Deep neural network is now a new trend towards solving different problems in speech processing. In this paper, we propose a discriminative deep recurrent neural network (DRNN) model for monaural speech separation. Our idea is to construct DRNN as a regression model to discover the deep structure and…

Cited by 0SourceScholar
2015

Modulation Wiener filter for improving speech intelligibility

ICASSP 2015accepted

This paper presents a single-channel high-dimensional Wiener filter in the spectro-temporal modulation domain. Unlike other conventional noise reduction techniques, the proposed algorithm not only reduces noise but also enhances the “textures” of the speech signal. A non-iterative decision-directed…

Cited by 0SourceScholar