← Search

He Qu

3 accepted papers

2025

Boosting Code-Switching ASR with Mixture of Experts Enhanced Speech-Conditioned LLM

ICASSP 2025accepted

In this paper, we introduce a speech-conditioned Large Language Model (LLM) integrated with a Mixture of Experts (MoE) based connector to address the challenge of Code-Switching (CS) scenario in Automatic Speech Recognition (ASR). Specifically, we propose an Insertion and Deletion of Interruption To…

Cited by 0SourceScholar
2025

Dynamic Language Group-based MoE: Enhancing Code-Switching Speech Recognition with Hierarchical Routing

ICASSP 2025accepted

The Mixture of Experts (MoE) model is a promising approach for handling code-switching speech recognition (CS-ASR) tasks. However, the existing CS-ASR work on MoE has yet to leverage the advantages of MoE’s parameter scaling ability fully. This work proposes DLG-MoE, a Dynamic Language Group-based M…

Cited by 0SourceScholar
2024

Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic Coding

ICASSP 2024accepted

Recently, there has been a growing interest in text-to-speech (TTS) methods that can be trained with minimal supervision by combining two types of discrete speech representations and using two sequence-to-sequence tasks to decouple TTS. However, existing methods suffer from three problems: the high-…

Cited by 0SourceScholar