← Search

Xurong Xie

15 accepted papers

2026

LLM-Oriented Token-Adaptive Knowledge Distillation

AAAI 2026technical

Knowledge Distillation (KD) is a key technique for compressing Large-scale Language Models (LLMs), but prevailing logit-based methods employ static strategies misaligned with the student’s dynamic learning process. By treating all tokens indiscriminately with a fixed temperature, these methods resul

Cited by 0SourcePDFScholar
2026

MULTI-CHANNEL SPEECH ENHANCEMENT FOR COCKTAIL PARTY SPEECH EMOTION RECOGNITION

ICASSP 2026poster

This paper highlights the critical importance of multi-channel speech enhancement (MCSE) for speech emotion recognition (ER) in cocktail party scenarios. A multi-channel speech dereverberation and separation front-end integrating DNN-WPE and mask-based MVDR is used to extract the target speaker's sp…

Cited by 0SourcePDFScholar
2025

AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding

NeurIPS 2025poster

Multimodal Large Language Models (MLLMs) have demonstrated excellent performance in video understanding but suffer from degraded effectiveness when processing long videos due to fixed-length contexts and weaknesses in modeling long-term dependencies. Retrieval-Augmented Generation (RAG) technology c…

Cited by 0SourcecodeScholar
2025

CR-CLIP: Image-Text Contrastive Regression for Generalized Gaze Estimation

ICASSP 2025accepted

Gaze estimation methods typically encounter significant performance degradation in generalized tasks due to the domain mismatch between the source and target domains. Existing approaches attempt to utilize various domain generalization techniques. However, their generalization capabilities are limit…

Cited by 0SourceScholar
2025

Phone-purity Guided Discrete Tokens for Dysarthric Speech Recognition

ICASSP 2025accepted

Discrete tokens provide compact and domain-adaptable representations of speech features. However, their application to disordered speech, characterized by articulation imprecision and significant mismatch with normal voice, remains unexplored. To this end, this paper proposes novel phone-purity guid…

Cited by 0SourceScholar
2024

Towards Automatic Data Augmentation for Disordered Speech Recognition

ICASSP 2024accepted

Automatic recognition of disordered speech remains a highly challenging task to date due to data scarcity. This paper presents a reinforcement learning (RL) based on-the-fly data augmentation approach for training state-of-the-art PyChain TDNN and end-to-end Conformer ASR systems on such data. The h…

Cited by 11SourceScholar
2024

Towards High-Performance and Low-Latency Feature-Based Speaker Adaptation of Conformer Speech Recognition Systems

ICASSP 2024accepted

Practical application of model-based speaker adaptation techniques to end-to-end ASR systems is hindered by speaker-level data scarcity and latency in speaker-dependent (SD) parameters update. To this end, data-efficient and low-latency rapid feature-based speaker adaptation approaches are proposed…

Cited by 0SourceScholar
2023

Adversarial Data Augmentation Using VAE-GAN for Disordered Speech Recognition

ICASSP 2023accepted

Automatic recognition of disordered speech remains a highly challenging task to date. The underlying neuro-motor conditions, often compounded with co-occurring physical disabilities, lead to the difficulty in collecting large quantities of impaired speech required for ASR system development. This pa…

Cited by 0SourceScholar
2023

Exploring Self-Supervised Pre-Trained ASR Models for Dysarthric and Elderly Speech Recognition

ICASSP 2023accepted

Automatic recognition of disordered and elderly speech remains a highly challenging task to date due to the difficulty in collecting such data in large quantities. This paper explores a series of approaches to integrate domain adapted Self-Supervised Learning (SSL) pre-trained models into TDNN and C…

Cited by 0SourceScholar
2023

Unsupervised Model-Based Speaker Adaptation of End-To-End Lattice-Free MMI Model for Speech Recognition

ICASSP 2023accepted

Modeling the speaker variability is a key challenge for automatic speech recognition (ASR) systems. In this paper, the learning hidden unit contributions (LHUC) based adaptation techniques with compact speaker dependent (SD) parameters are used to facilitate both speaker adaptive training (SAT) and…

Cited by 0SourceScholar
2022

Exploiting Cross Domain Acoustic-to-Articulatory Inverted Features for Disordered Speech Recognition

ICASSP 2022accepted

Articulatory features are inherently invariant to acoustic signal distortion and have been successfully incorporated into automatic speech recognition (ASR) systems for normal speech. Their practical application to disordered speech recognition is often limited by the difficulty in collecting such s…

Cited by 0SourceScholar
2021

Development of the Cuhk Elderly Speech Recognition System for Neurocognitive Disorder Detection Using the Dementiabank Corpus

ICASSP 2021accepted

Early diagnosis of Neurocognitive Disorder (NCD) is crucial in facilitating preventive care and timely treatment to delay further progression. This paper presents the development of a state-of-the-art automatic speech recognition (ASR) system built on the Dementia-Bank Pitt corpus for automatic NCD…

Cited by 54SourceScholar
2021

Neural Architecture Search for LF-MMI Trained Time Delay Neural Networks

ICASSP 2021accepted

Deep neural networks (DNNs) based automatic speech recognition (ASR) systems are often designed using expert knowledge and empirical evaluation. In this paper, a range of neural architecture search (NAS) techniques are used to automatically learn two types of hyper-parameters of state-of-the-art fac…

Cited by 28SourceScholar
2019

BLHUC: Bayesian Learning of Hidden Unit Contributions for Deep Neural Network Speaker Adaptation

ICASSP 2019accepted

Speaker adaptation techniques play a key role in reducing the mismatch between speech recognition systems and target users. In order to robustly learn speaker-dependent adaptation parameters, model based DNN adaptation techniques often require a significant amount of data. For example, in the common…

Cited by 0SourceScholar
2019

Bayesian and Gaussian Process Neural Networks for Large Vocabulary Continuous Speech Recognition

ICASSP 2019accepted

The hidden activation functions inside deep neural networks (DNNs) play a vital role in learning high level discriminative features and controlling the information flows to track longer history. However, the fixed model parameters used in standard DNNs can lead to over-fitting and poor generalizatio…

Cited by 0SourceScholar