← Search

Qingyang Hong

17 accepted papers

2025

Dynamic Language Group-based MoE: Enhancing Code-Switching Speech Recognition with Hierarchical Routing

ICASSP 2025accepted

The Mixture of Experts (MoE) model is a promising approach for handling code-switching speech recognition (CS-ASR) tasks. However, the existing CS-ASR work on MoE has yet to leverage the advantages of MoE’s parameter scaling ability fully. This work proposes DLG-MoE, a Dynamic Language Group-based M…

Cited by 0SourceScholar
2025

SlimSpeech: Lightweight and Efficient Text-to-Speech with Slim Rectified Flow

ICASSP 2025accepted

Recently, flow matching based speech synthesis has significantly enhanced the quality of synthesized speech while reducing the number of inference steps. In this paper, we introduce SlimSpeech, a lightweight and efficient speech synthesis system based on rectified flow. We have built upon the existi…

Cited by 0SourceScholar
2024

Improving Multi-Speaker ASR With Overlap-Aware Encoding And Monotonic Attention

ICASSP 2024accepted

End-to-end (E2E) multi-speaker speech recognition with the serialized output training (SOT) strategy demonstrates good performance in modeling diverse speaker scenarios. However, the E2E architecture doesn’t explicitly address the modeling of overlapping speech areas, potentially limiting the model’…

Cited by 0SourceScholar
2024

MM-TTS: Multi-Modal Prompt Based Style Transfer for Expressive Text-to-Speech Synthesis

AAAI 2024technical

The style transfer task in Text-to-Speech (TTS) refers to the process of transferring style information into text content to generate corresponding speech with a specific style. However, most existing style transfer approaches are either based on fixed emotional labels or reference speech clips, whi…

2024

Reflow-TTS: A Rectified Flow Model for High-Fidelity Text-to-Speech

ICASSP 2024accepted

The diffusion models including Denoising Diffusion Probabilistic Models (DDPM) and score-based generative models have demonstrated excellent performance in speech synthesis tasks. However, its effectiveness comes at the cost of numerous sampling steps, resulting in prolonged sampling time required t…

Cited by 0SourceScholar
2024

SR-HuBERT : An Efficient Pre-Trained Model for Speaker Verification

ICASSP 2024accepted

Recently, pre-trained models (PTMs) have been extensively applied in speaker verification (SV) and greatly boosted system performance. However, mainstream PTMs currently concentrate on using frame-level universal representations. In this paper, we propose a novel pre-training framework that jointly…

Cited by 0SourceScholar
2023

Community Detection Graph Convolutional Network for Overlap-Aware Speaker Diarization

ICASSP 2023accepted

The clustering algorithm plays a crucial role in speaker diarization systems. However, traditional clustering algorithms suffer from the complex distribution of speaker embeddings and lack of digging potential relationships between speakers in a session. We propose a novel graph-based clustering app…

Cited by 0SourceScholar
2023

Meta Learning with Adaptive Loss Weight for Low-Resource Speech Recognition

ICASSP 2023accepted

Model Agnostic Meta-Learning (MAML) is an effective meta-learning algorithm for low-resource automatic speech recognition (ASR). It uses gradient descent to learn the initialization parameters of the model through various languages, making the model quickly adapt to unseen low-resource languages. Bu…

Cited by 0SourceScholar
2023

The XMU System for Audio-Visual Diarization and Recognition in MISP Challenge 2022

ICASSP 2023accepted

In this paper, we present our work in track 2 of the Multi-modal Information based Speech Processing (MISP) 2022 Challenge. We built a cascaded system and explored different acoustic front-ends and end-to-end speech recognition back-ends based on multimodal. To promote effective fusion between the d…

Cited by 0SourceScholar
2023

Unsupervised Speaker Verification Using Pre-Trained Model and Label Correction

ICASSP 2023accepted

Recently, the fine-tuning pre-trained model framework has emerged as a promising paradigm for speech-processing tasks. In this study, we present a novel strategy for unsupervised speaker verification using the Sub-structure of Pre-Trained Model (Sub-PTM), which consists of a CNN-based feature extrac…

Cited by 0SourceScholar
2022

Graph Convolutional Network Based Semi-Supervised Learning on Multi-Speaker Meeting Data

ICASSP 2022accepted

Unsupervised clustering on speakers is becoming increasingly important for its potential uses in semi-supervised learning. In reality, we are often presented with enormous amounts of unlabeled data from multi-party meetings and discussions. An effective unsupervised clustering approach would allow u…

Cited by 0SourceScholar
2021

ASV-SUBTOOLS: Open Source Toolkit for Automatic Speaker Verification

ICASSP 2021accepted

In this paper, we introduce a new open source toolkit for automatic speaker verification (ASV), named ASV-Subtools. Adopting PyTorch as main deep learning engine and Kaldi toolkit for data processing, ASV-Subtools allows users to develop modern speaker recognizers flexibly and efficiently. The toolk…

Cited by 0SourceScholar
2021

End-To-End Multi-Accent Speech Recognition with Unsupervised Accent Modelling

ICASSP 2021accepted

End-to-end speech recognition has achieved good recognition performance on standard English pronunciation datasets. However, one prominent problem with end-to-end speech recognition systems is that non-native English speakers tend to have complex and varied accents, which reduces the accuracy of Eng…

Cited by 0SourceScholar
2019

Training Multi-task Adversarial Network for Extracting Noise-robust Speaker Embedding

ICASSP 2019accepted

Under noisy environments, to achieve the robust performance of speaker recognition is still a challenging task. Motivated by the promising performance of multi-task training in a variety of image processing tasks, we explore the potential of multitask adversarial training for learning a noise-robust…

Cited by 0SourceScholar
2016

A transfer learning method for PLDA-based speaker verification

ICASSP 2016accepted

Currently, the state-of-the-art speaker verification system is based on i-vector and PLDA. However, PLDA requires tens of thousands of development data from many speakers. This makes it difficult to learn the PLDA parameters for a domain with scarce data. In this paper, we propose an effective trans…

Cited by 0SourceScholar