← Search

Xie Chen

65 accepted papers

2026

AHAMask: Reliable Task Specification for Large Audio Language Models Without Instructions

AAAI 2026technical

Although current large audio language models (LALMs) extend text large language models (LLMs) with generic acoustic understanding abilities, they usually suffer from prompt sensitivity, where different instructions of the same intention can yield drastically different outcomes. In this work, we pro

Cited by 0SourcePDFScholar
2026

AUV: TEACHING AUDIO UNIVERSAL VECTOR QUANTIZATION WITH SINGLE NESTED CODEBOOK

ICASSP 2026poster

We propose AUV, a unified neural audio codec with a single codebook, which enables a favourable reconstruction of speech and further extends to general audio, including vocal, music, and sound. AUV is capable of tackling any 16 kHz mixed-domain audio segment at bit rates around 700 bps. To accomplis…

Cited by 0SourcePDFScholar
2026

MEASURING PROSODY DIVERSITY IN ZERO-SHOT TTS: A NEW METRIC, BENCHMARK, AND EXPLORATION

ICASSP 2026poster

Prosody diversity is essential for achieving naturalness and expressiveness in zero-shot text-to-speech (TTS). However, frequently used acoustic metrics capture only partial views of prosodic variation and correlate poorly with human perception, leaving the problem of reliably quantifying prosody di…

Cited by 10SourcePDFScholar
2026

Measuring Audio's Impact on Correctness: Audio-Contribution-Aware Post-Training of Large Audio Language Models

ICLR 2026poster

Large Audio Language Models (LALMs) represent an important frontier in multimodal AI, addressing diverse audio tasks. Recently, post-training of LALMs has received increasing attention due to significant performance improvements over foundation models. While single-stage post-training such as reinfo…

Cited by 0SourcecodeScholar
2026

Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception

ICLR 2026poster

Fine-grained perception of multimodal information is critical for advancing human–AI interaction. With recent progress in audio–visual technologies, Omni Language Models (OLMs), capable of processing audio and video signals in parallel, have emerged as a promising paradigm for achieving richer unde…

Cited by 0SourcecodeScholar
2026

Position: Towards Responsible Evaluation for Text-to-Speech

ICML 2026poster

Recent advances in text-to-speech (TTS) technology have enabled systems to generate speech that is often indistinguishable from human speech, bringing benefits to accessibility, content creation, and human-computer interaction. However, current evaluation practices are increasingly inadequate for ca…

Cited by 0SourceScholar
2026

STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence

ICLR 2026poster

Despite rapid progress in Multi-modal Large Language Models and Large Audio-Language Models, existing audio benchmarks largely test semantics that can be recovered from text captions, masking deficits in fine-grained perceptual reasoning. We formalize audio 4D intelligence that is defined as reasoni…

Cited by 0SourcecodeScholar
2026

THE SJTU X-LANCE LAB SYSTEM FOR MSR CHALLENGE 2025

ICASSP 2026poster

This report describes the system submitted to the music source restoration (MSR) Challenge 2025. Our approach is composed of sequential BS-RoFormers, each dealing with a single task including music source separation (MSS), denoise and dereverb. To support 8 instruments given in the task, we utilize…

Cited by 0SourcePDFScholar
2026

WaveEx: Accelerating Flow Matching-based Speech Generation via Wavelet-guided Extrapolation

AAAI 2026technical

Flow matching-based generative models offer a principled approach to modeling continuous-time dynamics in speech generation. However, inference is often computationally expensive due to repeated neural network evaluations required by ODE solvers. We propose WaveEx, a training-free and plug-in accele

Cited by 0SourcePDFScholar
2026

YuE: Scaling Open Foundation Models for Long-Form Music Generation

ICLR 2026poster

We tackle the task of long-form music generation, particularly the challenging \textbf{lyrics-to-song} problem, by introducing \textbf{YuE (乐)}, a family of open-source music generation foundation models. Specifically, YuE scales to trillions of tokens and generates up to five minutes of music while…

Cited by 0SourcecodeScholar
2025

Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video

ICCV 2025poster

We propose a novel and general framework to disentangle video data into its dynamic motion and static content components. Our proposed method is a self-supervised pipeline with less assumptions and inductive biases than previous works: it utilizes a transformer-based architecture to jointly generate…

Cited by 0SourcePDFScholar
2025

DRCap: Decoding CLAP Latents with Retrieval-Augmented Generation for Zero-shot Audio Captioning

ICASSP 2025accepted

While automated audio captioning (AAC) has made notable progress, traditional fully supervised AAC models still face two critical challenges: the need for expensive audio-text pair data for training and performance degradation when transferring across domains. To overcome these limitations, we prese…

Cited by 0SourceScholar
2025

ELLA-V: Stable Neural Codec Language Modeling with Alignment-Guided Sequence Reordering

AAAI 2025technical

The language model (LM) approach based on acoustic and linguistic prompts, such as VALL-E, has achieved remarkable progress in the field of zero-shot audio generation. However, existing methods still have some limitations: 1) repetitions, transpositions, and omissions in the output synthesized speec…

2025

Enhancing Low-Resource ASR through Versatile TTS: Bridging the Data Gap

ICASSP 2025accepted

While automatic speech recognition (ASR) systems have achieved remarkable performance with large-scale datasets, their efficacy remains inadequate in low-resource settings, encompassing dialects, accents, minority languages, and long-tail hotwords, domains with significant practical relevance. With…

Cited by 0SourceScholar
2025

Enhancing Speech-to-Speech Dialogue Modeling with End-to-End Retrieval-Augmented Generation

EMNLP 2025

End-to-end speech-to-speech (S2S) dialogue systems have recently garnered increasing research attention for their lower latency and more natural integration of nonverbal cues such as emotion and speaker identity. However, these systems face key challenges, particularly in incorporating external know

2025

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

ACL 2025long

This paper introduces F5-TTS, a fully non-autoregressive text-to-speech system based on flow matching with Diffusion Transformer (DiT). Without requiring complex designs such as duration model, text encoder, and phoneme alignment, the text input is simply padded with filler tokens to the same length…

2025

GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement

ACL 2025long

The evolution of speech technology has been spurred by the rapid increase in dataset sizes. Traditional speech models generally depend on a large amount of labeled training data, which is scarce for low-resource languages. This paper presents GigaSpeech 2, a large-scale, multi-domain, multilingual s…

2025

Language Model Can Listen While Speaking

AAAI 2025technical

Dialogue serves as the most natural manner of human-computer interaction (HCI). Recent advancements in speech language models (SLM), have significantly enhanced speech-based conversational AI. However, these models are limited to turn-based conversation, lacking the ability to interact with humans i…

Cited by 2SourcePDFScholar
2025

MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix

NeurIPS 2025poster

We introduce MMAR, a new benchmark designed to evaluate the deep reasoning capabilities of Audio-Language Models (ALMs) across massive multi-disciplinary tasks. MMAR comprises 1,000 meticulously curated audio-question-answer triplets, collected from real-world internet videos and refined through ite…

Cited by 0SourcecodeScholar
2025

MUZO: Leveraging Multiple Queries and Momentum for Zeroth-Order Fine-Tuning of Large Language Models

EMNLP 2025

Fine-tuning pre-trained large language models (LLMs) on downstream tasks has achieved significant success across various domains. However, as model sizes grow, traditional first-order fine-tuning algorithms incur substantial memory overhead due to the need for activation storage for back-propagation

Cited by 0SourcePDFScholar
2025

Making LLMs Better Many-to-Many Speech-to-Text Translators with Curriculum Learning

ACL 2025long

Multimodal Large Language Models (MLLMs) have achieved significant success in Speech-to-Text Translation (S2TT) tasks. While most existing research has focused on English-centric translation directions, the exploration of many-to-many translation is still limited by the scarcity of parallel data. To…

2025

SLAM-AAC: Enhancing Audio Captioning with Paraphrasing Augmentation and CLAP-Refine through LLMs

ICASSP 2025accepted

Automated Audio Captioning (AAC) aims to generate natural textual descriptions for input audio signals. Recent progress in audio pre-trained models and large language models (LLMs) has significantly enhanced audio understanding and textual reasoning capabilities, making improvements in AAC possible.…

Cited by 0SourceScholar
2025

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

ACL 2025finding

Recent advancements highlight the potential of end-to-end real-time spoken dialogue systems, showcasing their low latency and high quality. In this paper, we introduce SLAM-Omni, a timbre-controllable, end-to-end voice interaction system with single-stage training. SLAM-Omni achieves zero-shot timbr…

2025

SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation

ACL 2025long

Simultaneous speech translation (SST) outputs translations in parallel with streaming speech input, balancing translation quality and latency. While large language models (LLMs) have been extended to handle the speech modality, streaming remains challenging as speech is pre-pended as a prompt for th…

2025

Speech Recognition Meets Large Language Model: Benchmarking, Models, and Exploration

AAAI 2025technical

In this paper, we focus on prompting one of the most important tasks in the field of speech processing, i.e., automatic speech recognition (ASR), with speech foundation encoders and large language models (LLM). Despite the growing body of research in this area, we find that many crucial design decis…

2025

Towards Reliable Large Audio Language Model

ACL 2025finding

Recent advancements in large audio language models (LALMs) have demonstrated impressive results and promising prospects in universal understanding and reasoning across speech, music, and general sound. However, these models still lack the ability to recognize their knowledge boundaries and refuse to…

2025

URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models

EMNLP 2025

Recent advances in large language models (LLMs) have driven significant progress in end-to-end spoken dialogue models (SDMs). In contrast to text-based LLMs, the evaluation framework for SDMs should encompass both cognitive dimensions (e.g., logical reasoning, knowledge) and speech-related aspects (

2025

VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

ICASSP 2025accepted

Recent TTS models with decoder-only Transformer architecture, such as SPEAR-TTS and VALL-E, achieve impressive naturalness and demonstrate the ability for zero-shot adaptation given a speech prompt. However, such decoder-only TTS models lack monotonic alignment constraints, sometimes leading to hall…

Cited by 0SourceScholar
2025

VQTalker: Towards Multilingual Talking Avatars Through Facial Motion Tokenization

AAAI 2025technical

We present VQTalker, a Vector Quantization-based framework for multilingual talking head generation that addresses the challenges of lip synchronization and natural motion across diverse languages. Our approach is grounded in the phonetic principle that human speech comprises a finite set of distinc…

Cited by 0SourcePDFScholar
2025

Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech Synthesis

NeurIPS 2025spotlight

While emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modelin…

Cited by 0SourceScholar
2024

BAT: Learning to Reason about Spatial Sounds with Large Language Models

ICML 2024poster

Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities…

Cited by 16SourcePDFScholar
2024

EAT: Self-Supervised Pre-Training with Efficient Audio Transformer

IJCAI 2024poster

Audio self-supervised learning (SSL) pre-training, which aims to learn good representations from unlabeled audio, has made remarkable progress. However, the extensive computational demands during pre-training pose a significant barrier to the potential application and optimization of audio SSL model…

2024

Leveraging Speech PTM, Text LLM, And Emotional TTS For Speech Emotion Recognition

ICASSP 2024accepted

In this paper, we explored how to boost speech emotion recognition (SER) with the state-of-the-art speech pre-trained model (PTM), data2vec, text generation technique, GPT-4, and speech synthesis technique, Azure TTS. First, we investigated the representation ability of different speech self-supervi…

Cited by 0SourceScholar
2024

SEF-VC: Speaker Embedding Free Zero-Shot Voice Conversion with Cross Attention

ICASSP 2024accepted

Zero-shot voice conversion (VC) aims to transfer the source speaker timbre to arbitrary unseen target speaker timbre, while keeping the linguistic content unchanged. Although the voice of generated speech can be controlled by providing the speaker embedding of the target speaker, the speaker similar…

Cited by 0SourceScholar
2024

StoryTTS: A Highly Expressive Text-to-Speech Dataset with Rich Textual Expressiveness Annotations

ICASSP 2024accepted

While acoustic expressiveness has long been studied in expressive text-to-speech (ETTS), the inherent expressiveness in text lacks sufficient attention, especially for ETTS of artistic works. In this paper, we introduce StoryTTS, a highly ETTS dataset that contains rich expressiveness both in acoust…

Cited by 0SourceScholar
2024

Towards Universal Speech Discrete Tokens: A Case Study for ASR and TTS

ICASSP 2024accepted

Self-supervised learning (SSL) proficiency in speech-related tasks has driven research into utilizing discrete tokens for speech tasks like recognition and translation, which offer lower storage requirements and great potential to employ natural language processing techniques. However, these studies…

Cited by 0SourceScholar
2024

UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and Vocoding

AAAI 2024technical

The utilization of discrete speech tokens, divided into semantic tokens and acoustic tokens, has been proven superior to traditional acoustic feature mel-spectrograms in terms of naturalness and robustness for text-to-speech (TTS) synthesis. Recent popular models, such as VALL-E and SPEAR-TTS, allow…

2024

VoiceFlow: Efficient Text-To-Speech with Rectified Flow Matching

ICASSP 2024accepted

Although diffusion models in text-to-speech have become a popular choice due to their strong generative ability, the intrinsic complexity of sampling from diffusion models harms their efficiency. Alternatively, we propose VoiceFlow, an acoustic model that utilizes a rectified flow matching algorithm…

Cited by 0SourceScholar
2024

emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation

ACL 2024findings

We propose emotion2vec, a universal speech emotion representation model. emotion2vec is pre-trained on open-source unlabeled emotion data through self-supervised online distillation, combining utterance-level loss and frame-level loss during pre-training. emotion2vec outperforms state-of-the-art pre…

2023

A Modular Biological Neural Network-Based Neuro-Robotic System via Local Chemical Stimulation and Calcium Imaging

RA-L 2023

Embodying in vitro biological neural networks (BNNs) with robots to explore the rise of intelligence in these simpler models and to endow robots with biological intelligence has been attracting increasing attention in the fields of neuroscience and robotics. However, current research suffers from un

Cited by 11SourceScholar
2023

An Adapter Based Multi-Label Pre-Training for Speech Separation and Enhancement

ICASSP 2023accepted

In recent years, self-supervised learning (SSL) has achieved tremendous success in various speech tasks due to its power to extract representations from massive unlabeled data. However, compared with tasks such as speech recognition (ASR), the improvements from SSL representation in speech separatio…

Cited by 0SourceScholar
2023

Emodiff: Intensity Controllable Emotional Text-to-Speech with Soft-Label Guidance

ICASSP 2023accepted

Although current neural text-to-speech (TTS) models are able to generate high-quality speech, intensity controllable emotional TTS is still a challenging task. Most existing methods need external optimizations for intensity calculation, leading to suboptimal results or degraded quality. In this pape…

Cited by 0SourceScholar
2023

Factorized AED: Factorized Attention-Based Encoder-Decoder for Text-Only Domain Adaptive ASR

ICASSP 2023accepted

End-to-end automatic speech recognition (ASR) systems have gained popularity given their simplified architecture and promising results. However, text-only domain adaptation remains a big challenge for E2E systems. Text-to-speech (TTS) based approaches fine-tune ASR models by synthesized speech with…

Cited by 0SourceScholar
2023

Front-End Adapter: Adapting Front-End Input of Speech Based Self-Supervised Learning for Speech Recognition

ICASSP 2023accepted

Recent years have witnessed a boom in self-supervised learning (SSL) in various areas including speech processing. Speech based SSL models present promising performance in a range of speech related tasks. However, the training of SSL models is computationally expensive and a common practice is to fi…

Cited by 0SourceScholar
2023

Improving Few-Shot Learning for Talking Face System with TTS Data Augmentation

ICASSP 2023accepted

Audio-driven talking face has attracted broad interest from academia and industry recently. However, data acquisition and labeling in audio-driven talking face are labor-intensive and costly. The lack of data resource results in poor synthesis effect. To alleviate this issue, we propose to use TTS (…

Cited by 0SourceScholar
2023

LongFNT: Long-Form Speech Recognition with Factorized Neural Transducer

ICASSP 2023accepted

Traditional automatic speech recognition (ASR) systems usually focus on individual utterances, without considering long-form speech with useful historical information, which is more practical in real scenarios. Simply attending longer transcription history for a vanilla neural transducer model shows…

Cited by 0SourceScholar
2022

Factorized Neural Transducer for Efficient Language Model Adaptation

ICASSP 2022accepted

In recent years, end-to-end (E2E) based automatic speech recognition (ASR) systems have achieved great success due to their simplicity and promising performance. Neural Transducer based models are increasingly popular in streaming E2E based ASR systems and have been reported to outperform the tradit…

Cited by 0SourceScholar
2021

Developing Real-Time Streaming Transformer Transducer for Speech Recognition on Large-Scale Dataset

ICASSP 2021accepted

Recently, Transformer based end-to-end models have achieved great success in many areas including speech recognition. However, compared to LSTM models, the heavy computational cost of the Transformer during inference is a key issue to prevent their applications. In this work, we explored the potenti…

Cited by 0SourceScholar
2021

Internal Language Model Training for Domain-Adaptive End-To-End Speech Recognition

ICASSP 2021accepted

The efficacy of external language model (LM) integration with existing end-to-end (E2E) automatic speech recognition (ASR) systems can be improved significantly using the internal language model estimation (ILME) method [1]. In this method, the internal LM score is subtracted from the score obtained…

Cited by 0SourceScholar
2021

Memory-Efficient Pipeline-Parallel DNN Training

ICML 2021spotlight

Many state-of-the-art ML results have been obtained by scaling up the number of parameters in existing models. However, parameters and activations for such large models often do not fit in the memory of a single accelerator device; this means that it is necessary to distribute training of large mode…

Cited by 274SourcePDFScholar
2021

Micro Robotic Manipulation System for the Force Stimulation of Muscle Fiber-like Cell Structure

ICRA 2021poster

Many previous works have facilitated muscle cell (C2C12) alignment to form fiber-like cell structures. However, there still remains a challenge how to induce C2C12 myoblasts in the cell structures to differentiate into matured myocytes to form a functional muscle tissue, while external mechanical st…

Cited by 2SourceScholar
2020

Low-bit Quantization of Recurrent Neural Network Language Models Using Alternating Direction Methods of Multipliers

ICASSP 2020accepted

The high memory consumption and computational costs of Recurrent neural network language models (RNNLMs) limit their wider application on resource constrained devices. In recent years, neural network quantization techniques that are capable of producing extremely low-bit compression, for example, bi…

Cited by 0SourceScholar
2019

Gaussian Process Lstm Recurrent Neural Network Language Models for Speech Recognition

ICASSP 2019accepted

Recurrent neural network language models (RNNLMs) have shown superior performance across a range of speech recognition tasks. At the heart of all RNNLMs, the activation functions play a vital role to control the information flows and tracking longer history contexts that are useful for predicting th…

Cited by 0SourceScholar
2019

Investigation of Sampling Techniques for Maximum Entropy Language Modeling Training

ICASSP 2019accepted

Maximum entropy language models (MaxEnt LMs) are log-linear models which are able to incorporate various hand-crafted features and non-linguistic information. Standard MaxEnt LMs are computationally heavy for tasks with a large vocabulary size due to the expensive normalization computation in the de…

Cited by 0SourceScholar
2019

Recurrent Neural Network Language Model Training Using Natural Gradient

ICASSP 2019accepted

Recurrent neural network language models (RNNLMs) have become an increasing popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. However, the SGD method doesn't consider…

Cited by 0SourceScholar
2018

Limited-Memory BFGS Optimization of Recurrent Neural Network Language Models for Speech Recognition

ICASSP 2018accepted

Recurrent neural network language models (RNNLM) have become an increasingly popular choice for state-of-the-art speech recognition systems. RNNLMs are normally trained by minimizing the cross entropy (CE) using the stochastic gradient descent (SGD) algorithm. The SGD method only uses first-order de…

Cited by 0SourceScholar
2018

Neural Network Language Modeling with Letter-Based Features and Importance Sampling

ICASSP 2018accepted

In this paper we describe an extension of the Kaldi software toolkit to support neural-based language modeling, intended for use in automatic speech recognition (ASR) and related tasks. We combine the use of subword features (letter n-grams) and one-hot encoding of frequent words so that the models…

Cited by 0SourceScholar
2018

Phonetic and Graphemic Systems for Multi-Genre Broadcast Transcription

ICASSP 2018accepted

State-of-the-art English automatic speech recognition systems typically use phonetic rather than graphemic lexicons. Graphemic systems are known to perform less well for English as the mapping from the written form to the spoken form is complicated. However, in recent years the representational powe…

Cited by 0SourceScholar
2017

Recurrent neural network language models for keyword search

ICASSP 2017accepted

Recurrent neural network language models (RNNLMs) have becoming increasingly popular in many applications such as automatic speech recognition (ASR). Significant performance improvements in both perplexity and word error rate over standard n-gram LMs have been widely reported on ASR tasks. In contra…

Cited by 0SourceScholar
2016

CUED-RNNLM - An open-source toolkit for efficient training and evaluation of recurrent neural network language models

ICASSP 2016accepted

In recent years, recurrent neural network language models (RNNLMs) have become increasingly popular for a range of applications including speech recognition. However, the training of RNNLMs is computationally expensive, which limits the quantity of data, and size of network, that can be used. In ord…

Cited by 0SourceScholar
2015

Improving the training and evaluation efficiency of recurrent neural network language models

ICASSP 2015accepted

Recurrent neural network language models (RNNLMs) are becoming increasingly popular for speech recognition. Previously, we have shown that RNNLMs with a full (non-classed) output layer (F-RNNLMs) can be trained efficiently using a GPU giving a large reduction in training time over conventional class…

Cited by 0SourceScholar
2015

Recurrent neural network language model training with noise contrastive estimation for speech recognition

ICASSP 2015accepted

In recent years recurrent neural network language models (RNNLMs) have been successfully applied to a range of tasks including speech recognition. However, an important issue that limits the quantity of data used, and their possible application areas, is the computational cost in training. A signi??…

Cited by 0SourceScholar
2015

Robust excitation-based features for Automatic Speech Recognition

ICASSP 2015accepted

In this paper we investigate the use of noise-robust features characterizing the speech excitation signal as complementary features to the usually considered vocal tract based features for Automatic Speech Recognition (ASR). The proposed Excitation-based Features (EBF) are tested in a state-of-the-a…

Cited by 0SourceScholar