← Search

Xuesong Yang

9 accepted papers

2026

FlexiVideo: Variation-Aware Temporal Dynamics Modeling for Efficient Video Understanding

CVPR 2026

Natural videos exhibit heterogeneous temporal dynamics, with certain segments undergoing high-dynamic scene transitions and others dominated by low-dynamic visual changes. However, treating all frames identically, a common practice in most MLLMs, leads to redundant visual encoding, which results in

Cited by 0SourcecodeScholar
2026

LLaVA-UHD v2: Exploiting Hierarchical Vision Granularity in MLLMs via Inverse Semantic Pyramid

AAAI 2026technical

Vision transformers (ViTs) are widely employed in multimodal large language models (MLLMs) for visual encoding. However, they exhibit inferior performance on tasks regarding fine-grained visual perception. We attribute this to the inner limitations of ViTs in capturing diverse visual semantic level

Cited by 0SourcePDFScholar
2025

Koel-TTS: Enhancing LLM based Speech Generation with Preference Alignment and Classifier Free Guidance

EMNLP 2025

Autoregressive speech token generation models produce speech with remarkable variety and naturalness but often suffer from hallucinations and undesired vocalizations that do not conform to conditioning inputs. To address these challenges, we introduce Koel-TTS, an encoder-decoder transformer model f

2021

REDAT: Accent-Invariant Representation for End-To-End ASR by Domain Adversarial Training with Relabeling

ICASSP 2021accepted

Accents mismatching is a critical problem for end-to-end ASR. This paper aims to address this problem by building an accent-robust RNN-T system with domain adversarial training (DAT). We unveil the magic behind DAT and provide, for the first time, a theoretical guarantee that DAT learns accent-invar…

Cited by 0SourceScholar
2019

AutoVC: Zero-Shot Voice Style Transfer with Only Autoencoder Loss

ICML 2019oral

Despite the progress in voice conversion, many-to-many voice conversion trained on non-parallel data, as well as zero-shot voice conversion, remains under-explored. Deep style transfer algorithms, generative adversarial networks (GAN) in particular, are being applied as new solutions in this field.…

2019

When CTC Training Meets Acoustic Landmarks

ICASSP 2019accepted

Connectionist temporal classification (CTC) provides an end-to-end acoustic model (AM) training strategy. CTC learns accurate AMs without time-aligned phonetic transcription, but sometimes fails to converge, especially in resource-constrained scenarios. In this paper, the convergence properties of C…

Cited by 0SourceScholar
2018

Deep Learning Based Speech Beamforming

ICASSP 2018accepted

Multi-channel speech enhancement with ad-hoc sensors has been a challenging task. Speech model guided beamforming algorithms are able to recover natural sounding speech, but the speech models tend to be oversimplified or the inference would otherwise be too complicated. On the other hand, deep learn…

Cited by 0SourceScholar
2018

Joint Modeling of Accents and Acoustics for Multi-Accent Speech Recognition

ICASSP 2018accepted

The performance of automatic speech recognition systems degrades with increasing mismatch between the training and testing scenarios. Differences in speaker accents are a significant source of such mismatch. The traditional approach to deal with multiple accents involves pooling data from several ac…

Cited by 0SourceScholar
2017

End-to-end joint learning of natural language understanding and dialogue manager

ICASSP 2017accepted

Natural language understanding and dialogue policy learning are both essential in conversational systems that predict the next system actions in response to a current user utterance. Conventional approaches aggregate separate models of natural language understanding (NLU) and system action predictio…

Cited by 0SourceScholar