← Search

Chunlei Zhang

18 accepted papers

2026

Disentangle-then-Align: Non-Iterative Hybrid Multimodal Image Registration via Cross-Scale Feature Disentanglement

CVPR 2026

Multimodal image registration is a fundamental task and a prerequisite for downstream cross-modal analysis. Despite recent progress in shared feature extraction and multi-scale architectures, two key limitations remain. First, some methods use disentanglement to learn shared features but mainly regu

Cited by 0SourcecodeScholar
2025

Preference Alignment Improves Language Model-Based TTS

ICASSP 2025accepted

Recent advancements in text-to-speech (TTS) have shown that language model (LM)-based systems offer competitive performance to their counterparts. Further optimization can be achieved through preference alignment algorithms, which adjust LMs to align with the preferences of reward models, enhancing…

Cited by 0SourceScholar
2024

Make-A-Voice: Revisiting Voice Large Language Models as Scalable Multilingual and Multitask Learners

ACL 2024long

Large language models (LLMs) have successfully served as a general-purpose interface across multiple tasks and languages, while the adaptation of voice LLMs is mostly designed for specific purposes (either single-task or monolingual), where the advantages of LLMs especially for low-resource language…

2024

uSee: Unified Speech Enhancement And Editing with Conditional Diffusion Models

ICASSP 2024accepted

Speech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we propose a Unified Speech Enhancement and Editing (uSee) model with conditional…

Cited by 15SourceScholar
2023

Prosody-TTS: Improving Prosody with Masked Autoencoder and Conditional Diffusion Model For Expressive Text-to-Speech

ACL 2023findings

Expressive text-to-speech aims to generate high-quality samples with rich and diverse prosody, which is hampered by dual challenges: 1) prosodic attributes in highly dynamic voices are difficult to capture and model without intonation; and 2) highly multimodal prosodic representations cannot be well…

2022

Joint Modeling of Code-Switched and Monolingual ASR via Conditional Factorization

ICASSP 2022accepted

Conversational bilingual speech encompasses three types of utterances: two purely monolingual types and one intra-sententially code-switched type. In this work, we propose a general framework to jointly model the likelihoods of the monolingual and code-switch sub-tasks that comprise bilingual speech…

Cited by 0SourceScholar
2022

Robust Disentangled Variational Speech Representation Learning for Zero-Shot Voice Conversion

ICASSP 2022accepted

Traditional studies on voice conversion (VC) have made progress with parallel training data and known speakers. Good voice conversion quality is obtained by exploring better alignment modules or expressive mapping functions. In this study, we investigate zero-shot VC from a novel perspective of self…

Cited by 0SourceScholar
2022

Towards end-to-end Speaker Diarization with Generalized Neural Speaker Clustering

ICASSP 2022accepted

Speaker diarization consists of many components, e.g., front-end processing, speech activity detection (SAD), overlapped speech detection (OSD) and speaker segmentation/clustering. Conventionally, most of the involved components are separately developed and optimized. The resulting speaker diarizati…

Cited by 0SourceScholar
2021

Improving RNN Transducer with Target Speaker Extraction and Neural Uncertainty Estimation

ICASSP 2021accepted

Target-speaker speech recognition aims to recognize target-speaker speech from noisy environments with background noise and interfering speakers. This work presents a joint framework that combines time-domain target-speaker speech extraction and Recurrent Neural Network Transducer (RNN-T). To stabil…

Cited by 0SourceScholar
2021

Self-Supervised Text-Independent Speaker Verification Using Prototypical Momentum Contrastive Learning

ICASSP 2021accepted

In this study, we investigate self-supervised representation learning for speaker verification (SV). First, we examine a simple contrastive learning approach (SimCLR) with a momentum contrastive (MoCo) learning framework, where the MoCo speaker embedding system utilizes a queue to maintain a large s…

Cited by 0SourceScholar
2020

Speaker-Aware Target Speaker Enhancement by Jointly Learning with Speaker Embedding Extraction

ICASSP 2020accepted

Deep learning based speech separation approaches have received great interest, among which the recent speaker-aware speech enhancement methods are promising for solving difficulties such as arbitrary source permutation and unknown number of sources. In this paper, we propose a novel training framewo…

Cited by 0SourceScholar
2019

Semi-supervised Learning with Generative Adversarial Networks for Arabic Dialect Identification

ICASSP 2019accepted

Dialect Identification (DID) refers to the process of identifying different dialects within the same language class. Compared with more general language identification (LID), DID is a more challenging task because of the substantial similarity between dialects. For an i-vector based LID/DID, prior s…

Cited by 10SourceScholar
2019

UTD-CRSS Systems for 2018 NIST Speaker Recognition Evaluation

ICASSP 2019accepted

In this study, we present systems submitted by the Center for Robust Speech Systems (CRSS) from UTDallas to NIST SRE 2018 (SRE18). Three alternative front-end speaker embedding frameworks are investigated, that includes: (i) i-vector, (ii) x-vector, (iii) and a modified triplet speaker embedding sys…

Cited by 21SourceScholar
2017

A study of speaker verification performance with expressive speech

ICASSP 2017accepted

Expressive speech introduces variations in the acoustic features affecting the performance of speech technology such as speaker verification systems. It is important to identify the range of emotions for which we can reliably estimate speaker verification tasks. This paper studies the performance of…

Cited by 0SourceScholar
2016

Joint information from nonlinear and linear features for spoofing detection: An i-vector/DNN based approach

ICASSP 2016accepted

Sustaining automatic speaker verification(ASV) systems from spoofing attacks remains an essential challenge, even if significant progress in ASV has been achieved in recent years. In this study, an automatic spoofing detection approach using an i-vector framework is proposed. Two approaches are used…

Cited by 20SourceScholar
2016

Language recognition using deep neural networks with very limited training data

ICASSP 2016accepted

This study proposes a novel deep neural network (DNN) based approach to language identification (LID) for the NIST 2015 Language Recognition (LRE) i-Vector Machine Learning Challenge. State-of-the-art DNN based LID systems utilize large amounts of labeled training data. The 2015 LRE i-Vector Machine…

Cited by 0SourceScholar
2016

UTD-CRSS system for the NIST 2015 language recognition i-vector machine learning challenge

ICASSP 2016accepted

In this paper, we present the system developed by the Center for Robust Speech Systems (CRSS), University of Texas at Dallas, for the NIST 2015 language recognition i-vector machine learning challenge. Our system includes several subsystems, based on Linear Discriminant Analysis - Support Vector Mac…

Cited by 0SourceScholar