← Search

Xulong Zhang

21 accepted papers

2026

CARE: Multi-Task Pretraining for Latent Continuous Action Representation in Robot Control

ICASSP 2026poster

Recent advances in Vision-Language-Action (VLA) models have shown promise for robot control, but their dependence on action supervision limits scalability and generalization. To address this challenge, we introduce CARE, a novel framework designed to train VLA models for robotic task execution. Unli…

Cited by 0SourcePDFScholar
2026

DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement

ICML 2026poster

Unified Multimodal models (UMMs) built on a single architecture have shown impressive performance in both understanding and generation. We identify a fundamental challenge lies in inductive biases induced by distinct supervision signals: generation branch prefers high-fidelity, fine-grained represen…

Cited by 0SourceScholar
2026

MIRRORTALK: FORGING PERSONALIZED AVATARS VIA DISENTANGLED STYLE AND HIERARCHICAL MOTION CONTROL

ICASSP 2026poster

Synthesizing personalized talking faces that uphold and highlight a speaker's unique style while maintaining lip-sync accuracy remains a significant challenge. A primary limitation of existing approaches is the intrinsic confounding of speaker-specific talking style and semantic content within facia…

Cited by 0SourcePDFScholar
2025

CycleFlow: Leveraging Cycle Consistency in Flow Matching for Speaker Style Adaptation

ICASSP 2025accepted

Voice Conversion (VC) aims to convert the style of a source speaker, such as timbre and pitch, to the style of any target speaker while preserving the linguistic content. However, the ground truth of the converted speech does not exist in a non-parallel VC scenario, which induces the train-inference…

Cited by 0SourceScholar
2025

Graph Contrastive Learning with Decoupled Augmentation

ICASSP 2025accepted

Graph contrastive learning based on augmentation strategies has recently demonstrated remarkable performance. Existing methods typically jointly leverage attribute and structural augmentations to generate graph views, learning data invariance information through contrasting sample pairs. However, th…

Cited by 0SourceScholar
2025

Homogeneous Graph Extraction: An Approach to Learning Heterogeneous Graph Embedding

ICASSP 2025accepted

Heterogeneous Graph Neural Networks (HGNNs) aim to embed rich structural and semantic information of heterogeneous graphs into low-dimensional node representations. While HGNNs extend the foundational work of homogeneous Graph Neural Networks, the methodology for effectively transforming heterogeneo…

Cited by 0SourceScholar
2024

ED-TTS: Multi-Scale Emotion Modeling Using Cross-Domain Emotion Diarization for Emotional Speech Synthesis

ICASSP 2024accepted

Existing emotional speech synthesis methods often utilize an utterance-level style embedding extracted from reference audio, neglecting the inherent multi-scale property of speech prosody. We introduce ED-TTS, a multi-scale emotional speech synthesis model that leverages Speech Emotion Diarization (…

Cited by 0SourceScholar
2024

EmoTalker: Emotionally Editable Talking Face Generation via Diffusion Model

ICASSP 2024accepted

In recent years, the field of talking faces generation has attracted considerable attention, with certain methods adept at generating virtual faces that convincingly imitate human expressions. However, existing methods face challenges related to limited generalization, particularly when dealing with…

Cited by 0SourceScholar
2024

IDEAW: Robust Neural Audio Watermarking with Invertible Dual-Embedding

EMNLP 2024main

The audio watermarking technique embeds messages into audio and accurately extracts messages from the watermarked audio. Traditional methods develop algorithms based on expert experience to embed watermarks into the time-domain or transform-domain of signals. With the development of deep neural netw…

2024

Learning Disentangled Speech Representations with Contrastive Learning and Time-Invariant Retrieval

ICASSP 2024accepted

Voice conversion refers to transferring speaker identity with well-preserved content. Better disentanglement of speech representations leads to better voice conversion. Recent studies have found that phonetic information from input audio has the potential ability to well represent content. Besides,…

Cited by 0SourceScholar
2023

Dynamic Alignment Mask CTC: Improved Mask CTC With Aligned Cross Entropy

ICASSP 2023accepted

Because of predicting all the target tokens in parallel, the non-autoregressive models greatly improve the decoding efficiency of speech recognition compared with traditional autoregressive models. In this work, we present dynamic alignment Mask CTC, introducing two methods: (1) Aligned Cross Entrop…

Cited by 0SourceScholar
2023

Improving EEG-based Emotion Recognition by Fusing Time-Frequency and Spatial Representations

ICASSP 2023accepted

Using deep learning methods to classify EEG signals can accurately identify people’s emotions. However, existing studies have rarely considered the application of the information in another domain’s representations to feature selection in the time-frequency domain. We propose a classification networ…

Cited by 0SourceScholar
2023

Improving Music Genre Classification from multi-modal Properties of Music and Genre Correlations Perspective

ICASSP 2023accepted

Music genre classification has been widely studied in past few years for its various applications in music information retrieval. Previous works tend to perform unsatisfactorily, since those methods only use audio content or jointly use audio content and lyrics content inefficiently. In addition, as…

Cited by 0SourceScholar
2023

Learning Speech Representations with Flexible Hidden Feature Dimensions

ICASSP 2023accepted

Non-parallel many-to-many voice conversion is a kind of style transfer task in speech. Recently, AutoVC has been applied in this field as a popular solution, as it can achieve distribution-matching style transfer by training only the re- construction loss. However, in order to strike a good balance…

Cited by 0SourceScholar
2023

QI-TTS: Questioning Intonation Control for Emotional Speech Synthesis

ICASSP 2023accepted

Recent expressive text to speech (TTS) models focus on synthesizing emotional speech, but some fine-grained styles such as intonation are neglected. In this paper, we propose QI-TTS which aims to better transfer and control intonation to further deliver the speaker’s questioning intention while tran…

Cited by 0SourceScholar
2023

VQ-CL: Learning Disentangled Speech Representations with Contrastive Learning and Vector Quantization

ICASSP 2023accepted

Voice Conversion(VC) refers to converting the voice characteristics of audio to another one as it is said by other people. Recently, more and more studies have focused on disentangle-based VC, which separates the timbre and linguistic content information from an audio signal to effectively achieve V…

Cited by 0SourceScholar
2022

Avqvc: One-Shot Voice Conversion By Vector Quantization With Applying Contrastive Learning

ICASSP 2022accepted

Voice Conversion(VC) refers to changing the timbre of a speech while retaining the discourse content. Recently, many works have focused on disentangle-based learning techniques to separate the timbre and the linguistic content information from a speech signal. Once successful, voice conversion will…

Cited by 0SourceScholar
2022

DRVC: A Framework of Any-to-Any Voice Conversion with Self-Supervised Learning

ICASSP 2022accepted

Any-to-any voice conversion problem aims to convert voices for source and target speakers, which are out of the training data. Previous works wildly utilize the disentangle-based models. The disentangle-based model assumes the speech consists of content and speaker style information and aims to unta…

Cited by 0SourceScholar
2022

nnSpeech: Speaker-Guided Conditional Variational Autoencoder for Zero-Shot Multi-speaker text-to-speech

ICASSP 2022accepted

Multi-speaker text-to-speech (TTS) using a few adaption data is a challenge in practical applications. To address that, we propose a zero-shot multi-speaker TTS, named nnSpeech, that could synthesis a new speaker voice without fine-tuning and using only one adaption utterance. Compared with using a…

Cited by 0SourceScholar
2021

Singer Identification Using Deep Timbre Feature Learning with KNN-NET

ICASSP 2021accepted

In this paper, we study the issue of automatic singer identification (SID) in popular music recordings, which aims to recognize who sang a given piece of song. The main challenge for this investigation lies in the fact that a singer’s singing voice changes and intertwines with the signal of backgrou…

Cited by 0SourceScholar