← Search

You Zhang

12 accepted papers

2026

EMOTIONAL DIMENSION CONTROL IN LANGUAGE MODEL-BASED TEXT-TO-SPEECH: SPANNING A BROAD SPECTRUM OF HUMAN EMOTIONS

ICASSP 2026poster

Emotional text-to-speech (TTS) systems sturggle to capture the full spectrum of human emotions due to the inherent complexity of emotional expressions and the limited coverage of existing emotion labels. To address this, we propose a language model-based TTS framework that synthesizes speech across…

Cited by 0SourcePDFScholar
2026

How Does Instrumental Music Help SingFake Detection?

ICASSP 2026poster

Although many models exist to detect singing voice deepfakes (SingFake), how these models operate, particularly with instrumental accompaniment, is unclear. We investigate how instrumental music affects SingFake detection from two perspectives. To investigate the behavioral effect, we test different…

Cited by 0SourcePDFScholar
2025

Multi-Attribute Multi-Grained Adaptation of Pre-Trained Language Models for Text Understanding from Bayesian Perspective

AAAI 2025technical

Current neural networks often employ multi-domain-learning or attribute-injecting mechanisms to incorporate non-independent and identically distributed (non-IID) information for text understanding tasks by capturing individual characteristics and the relationships among samples. However, the extent…

2025

VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

NAACL 2025system demonstrations

In this work, we introduce VERSA, a unified and standardized evaluation toolkit designed for various speech, audio, and music signals. The toolkit features a Pythonic interface with flexible configuration and dependency control, making it user-friendly and efficient. With full installation, VERSA of…

2024

Improving Personalized Sentiment Representation with Knowledge-enhanced and Parameter-efficient Layer Normalization

COLING 2024main

Existing studies on personalized sentiment classification consider a document review as an overall text unit and incorporate backgrounds (i.e., user and product information) to learn sentiment representation. However, it is difficult when these methods meet the current pretrained language models (PL…

2024

Learning Arousal-Valence Representation from Categorical Emotion Labels of Speech

ICASSP 2024accepted

Dimensional representations of speech emotions such as the arousal-valence (AV) representation provide a continuous and fine-grained description and control than their categorical counterparts. They have wide applications in tasks such as dynamic emotion understanding and expressive text-to-speech s…

Cited by 0SourceScholar
2024

Personalized LoRA for Human-Centered Text Understanding

AAAI 2024technical

Effectively and efficiently adapting a pre-trained language model (PLM) for human-centered text understanding (HCTU) is challenging since user tokens are million-level in most personalized applications and do not have concrete explicit semantics. A standard and parameter-efficient approach (e.g., Lo…

2023

Domain Generalization via Switch Knowledge Distillation for Robust Review Representation

ACL 2023findings

Applying neural models injected with in-domain user and product information to learn review representations of unseen or anonymous users incurs an obvious obstacle in content-based recommender systems. For the generalization of the in-domain classifier, most existing models train an extra plain-text…

2022

Multi-Thread CTAEA-Based Workstation Reconfiguration for Multi-Stage Automobile Engine Flow Shop Considering Performance Deterioration

RA-L 2022

In the automobile engine flow shop (AEFS), the equipment has more chance to work for a long time, so the manufacturing performance may deteriorate rapidly, thus causing operating unbalance and inefficiency. To tackle this problem at minimum cost, this study proposes a multi-thread constrained two-ar

Cited by 6SourceScholar