← Search

Xueyuan Chen

10 accepted papers

2026

MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark

ICLR 2026poster

Speech inherently contains rich acoustic information that extends far beyond the textual language. In real-world spoken communication, effective interpretation often requires integrating semantic meaning (e.g., content), paralinguistic features (e.g., emotions, speed, pitch) and phonological charact…

Cited by 0SourcecodeScholar
2025

ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language Modeling

ICML 2025poster

Recent advancements in audio language models have underscored the pivotal role of audio tokenization, which converts audio signals into discrete tokens, thereby facilitating the application of language model architectures to the audio domain. In this study, we introduce ALMTokenizer, a novel low-bit…

Cited by 0SourcePDFScholar
2025

Integrating Potential Pronunciations for Enhanced Mispronunciation Detection and Diagnosis Ability in LLMs

ICASSP 2025accepted

Large Language Models (LLMs) have exhibited significant potentials across various tasks. However, how to leverage the power of LLMs in the mispronunciation detection and diagnosis (MDD) task is still under-explored. In this paper, we propose a PP-ATP model, which integrates potential pronunciations…

Cited by 0SourceScholar
2025

Speech Discrete Tokens or Continuous Features? A Comparative Analysis for Spoken Language Understanding in SpeechLLMs

EMNLP 2025

With the rise of Speech Large Language Models (SpeechLLMs), two dominant approaches have emerged for speech processing: discrete tokens and continuous features. Each approach has demonstrated strong capabilities in audio-related processing tasks. However, the performance gap between these two paradi

Cited by 0SourcePDFScholar
2024

Exploiting Audio-Visual Features with Pretrained AV-HuBERT for Multi-Modal Dysarthric Speech Reconstruction

ICASSP 2024accepted

Dysarthric speech reconstruction (DSR) aims to transform dysarthric speech into normal speech by improving the intelligibility and naturalness. This is a challenging task especially for patients with severe dysarthria and speaking in complex, noisy acoustic environments. To address these challenges,…

Cited by 0SourceScholar
2024

Stylespeech: Self-Supervised Style Enhancing with VQ-VAE-Based Pre-Training for Expressive Audiobook Speech Synthesis

ICASSP 2024accepted

The expressive quality of synthesized speech for audiobooks is limited by generalized model architecture and unbalanced style distribution in the training data. To address these issues, in this paper, we propose a self-supervised style enhancing method with VQ-VAE-based pre-training for expressive a…

Cited by 0SourceScholar
2023

SEGA: Structural Entropy Guided Anchor View for Graph Contrastive Learning

ICML 2023poster

In contrastive learning, the choice of "view" controls the information that the representation captures and influences the performance of the model. However, leading graph contrastive learning methods generally produce views via random corruption or learning, which could lead to the loss of essentia…

2022

A Character-Level Span-Based Model for Mandarin Prosodic Structure Prediction

ICASSP 2022accepted

The accuracy of prosodic structure prediction is crucial to the naturalness of synthesized speech in Mandarin text-to-speech system, but now is limited by widely-used sequence-to-sequence framework and error accumulation from previous word segmentation results. In this paper, we propose a span-based…

Cited by 0SourceScholar
2022

Unsupervised Multi-scale Expressive Speaking Style Modeling with Hierarchical Context Information for Audiobook Speech Synthesis

COLING 2022main

Naturalness and expressiveness are crucial for audiobook speech synthesis, but now are limited by the averaged global-scale speaking style representation. In this paper, we propose an unsupervised multi-scale context-sensitive text-to-speech model for audiobooks. A multi-scale hierarchical context e…

Cited by 9SourcePDFScholar