← Search

Dongxing Xu

5 accepted papers

2026

BRIDGING THE GAP: A COMPARATIVE EXPLORATION OF SPEECH-LLM AND END-TO-END ARCHITECTURE FOR MULTILINGUAL CONVERSATIONAL ASR

ICASSP 2026poster

The INTERSPEECH 2025 Challenge on Multilingual Conversational Speech Language Models (MLC-SLM) promotes multilingual conversational ASR with large language models (LLMs). Our previous SHNU-mASR system adopted a competitive parallel-speech-encoder architecture that integrated Whisper and mHuBERT with…

Cited by 0SourcePDFScholar
2024

Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech Recognition

ICASSP 2024accepted

How to effectively use limited supervised accent data to improve the accented ASR is of paramount importance. In this work, we propose an accent-specific quantization for joint unsupervised and supervised training (AQ-JUST) of end-to-end ASR models to address this issue. Specifically, two variants o…

Cited by 0SourceScholar
2024

Cross-Modal Parallel Training for Improving end-to-end Accented Speech Recognition

ICASSP 2024accepted

Multi-accent speech recognition is a key challenge in current speech recognition due to the pronunciation variations of different accents. In this study, we propose a Cross-modal Parallel Training (CPT) approach for improving the accent robustness of state-of-the-art Conformer-Transducer (Conformer-…

Cited by 0SourceScholar
2024

Score Calibration Based on Consistency Measure Factor for Speaker Verification

ICASSP 2024accepted

This paper proposes a new scoring calibration method named "Consistency-Aware Score Calibration", which introduces a Consistency Measure Factor (CMF) to measure the stability of audio voiceprints in similarity scores for speaker verification. The CMF is inspired by the limitations in segment scoring…

Cited by 0SourceScholar
2023

FEW-Shot Continual Learning with Weight Alignment and Positive Enhancement for Bioacoustic Event Detection

ICASSP 2023accepted

In this paper, we propose a new continual learning framework for few-shot bioacoustic event detection (BED). First, we modify the recently proposed dynamic few-shot learning (DFSL) and generalize it to the BED task. Then, we introduce a weight alignment loss to enhance the weight generator of modifi…

Cited by 0SourceScholar