← Search

Yanhua Long

10 accepted papers

2026

BRIDGING THE GAP: A COMPARATIVE EXPLORATION OF SPEECH-LLM AND END-TO-END ARCHITECTURE FOR MULTILINGUAL CONVERSATIONAL ASR

ICASSP 2026poster

The INTERSPEECH 2025 Challenge on Multilingual Conversational Speech Language Models (MLC-SLM) promotes multilingual conversational ASR with large language models (LLMs). Our previous SHNU-mASR system adopted a competitive parallel-speech-encoder architecture that integrated Whisper and mHuBERT with…

Cited by 0SourcePDFScholar
2025

Leveraging Out-of-Domain Noise for Unsupervised Domain Adaptation in Speech Enhancement

ICASSP 2025accepted

When there’s a mismatch between the training and test domains, supervised speech enhancement (SE) models trained on synthetic paired noisy-clean data often struggle in real-world scenarios, highlighting the industry’s strong demand for unsupervised training and domain adaptation methods. In this stu…

Cited by 0SourceScholar
2025

Personalized Speech Enhancement without User Enrollment for Real-World Audio Replay Scenarios

ICASSP 2025accepted

Many speech enhancement (SE) approaches have been proposed to deal with cocktail party problem. Personalized speech enhancement (PSE) approaches improve SE performance by utilizing user enrollment speech. However, PSE requires users to record additional clean audio for registration, which can be red…

Cited by 0SourceScholar
2025

SEF-PNet: Speaker Encoder-Free Personalized Speech Enhancement with Local and Global Contexts Aggregation

ICASSP 2025accepted

Personalized speech enhancement (PSE) methods typically rely on pre-trained speaker verification models or self-designed speaker encoders to extract target speaker clues, guiding the PSE model in isolating the desired speech. However, these approaches suffer from significant model complexity and oft…

Cited by 0SourceScholar
2024

Accent-Specific Vector Quantization for Joint Unsupervised and Supervised Training in Accent Robust Speech Recognition

ICASSP 2024accepted

How to effectively use limited supervised accent data to improve the accented ASR is of paramount importance. In this work, we propose an accent-specific quantization for joint unsupervised and supervised training (AQ-JUST) of end-to-end ASR models to address this issue. Specifically, two variants o…

Cited by 0SourceScholar
2024

Cross-Modal Parallel Training for Improving end-to-end Accented Speech Recognition

ICASSP 2024accepted

Multi-accent speech recognition is a key challenge in current speech recognition due to the pronunciation variations of different accents. In this study, we propose a Cross-modal Parallel Training (CPT) approach for improving the accent robustness of state-of-the-art Conformer-Transducer (Conformer-…

Cited by 0SourceScholar
2024

Score Calibration Based on Consistency Measure Factor for Speaker Verification

ICASSP 2024accepted

This paper proposes a new scoring calibration method named "Consistency-Aware Score Calibration", which introduces a Consistency Measure Factor (CMF) to measure the stability of audio voiceprints in similarity scores for speaker verification. The CMF is inspired by the limitations in segment scoring…

Cited by 0SourceScholar
2023

FEW-Shot Continual Learning with Weight Alignment and Positive Enhancement for Bioacoustic Event Detection

ICASSP 2023accepted

In this paper, we propose a new continual learning framework for few-shot bioacoustic event detection (BED). First, we modify the recently proposed dynamic few-shot learning (DFSL) and generalize it to the BED task. Then, we introduce a weight alignment loss to enhance the weight generator of modifi…

Cited by 0SourceScholar
2022

DPCCN: Densely-Connected Pyramid Complex Convolutional Network for Robust Speech Separation and Extraction

ICASSP 2022accepted

In recent years, a number of time-domain speech separation methods have been proposed. However, most of them are very sensitive to the environments and wide domain coverage tasks. In this paper, from the time-frequency domain perspective, we propose a densely-connected pyramid complex convolutional…

Cited by 0SourceScholar
2021

Multi-Channel Target Speech Extraction with Channel Decorrelation and Target Speaker Adaptation

ICASSP 2021accepted

The end-to-end approaches for single-channel target speech extraction have attracted widespread attention. However, the studies for end-to-end multi-channel target speech extraction are still relatively limited. In this work, we propose two methods for exploiting the multi-channel spatial informatio…

Cited by 0SourceScholar