← Search

Naijun Zheng

7 accepted papers

2025

CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition

ICASSP 2025accepted

Code-switching automatic speech recognition (ASR) aims to transcribe speech that contains two or more languages accurately. To better capture language-specific speech representations and address language confusion in code-switching ASR, the mixture-of-experts (MoE) architecture and an additional lan…

Cited by 0SourceScholar
2025

SCDiar: a streaming diarization system based on speaker change detection and speech recognition

ICASSP 2025accepted

In hours-long meeting scenarios, real-time speech stream often struggles with achieving accurate speaker diarization, commonly leading to speaker identification and speaker count errors. To address this challenge, we propose SCDiar, a system that operates on speech segments, split at the token level…

Cited by 0SourceScholar
2022

Multi-Channel Speaker Diarization Using Spatial Features for Meetings

ICASSP 2022accepted

Speaker identification for overlapped speech presents a great challenge for speaker diarization tasks in meeting scenarios. In order to overcome such challenges, several overlap-aware resegmentation methods based on deep learning have been integrated into speaker diarization systems. In this paper w…

Cited by 0SourceScholar
2022

Partially Fake Audio Detection by Self-Attention-Based Fake Span Discovery

ICASSP 2022accepted

The past few years have witnessed the significant advances of speech synthesis and voice conversion technologies. However, such technologies can undermine the robustness of broadly implemented biometric identification models and can be harnessed by in-the-wild attackers for illegal uses. The ASVspoo…

Cited by 0SourceScholar
2022

The CUHK-Tencent Speaker Diarization System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge

ICASSP 2022accepted

This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization and automatic speech recognition (ASR) tasks. In these meeting scenarios, the unce…

Cited by 0SourceScholar
2021

A Joint Training Framework of Multi-Look Separator and Speaker Embedding Extractor for Overlapped Speech

ICASSP 2021accepted

In multi-talker cases, overlapped speech degrades the speaker verification (SV) performance dramatically. To tackle this challenging problem, speech separation with multi-channel techniques can be adopted to extract each speaker’s signals to improve the SV performance. In this paper, a joint trainin…

Cited by 0SourceScholar
2017

LDPC code design for Gaussian multiple-access channels using dynamic EXIT chart analysis

ICASSP 2017accepted

We consider the degree distribution design of the low-density parity-check (LDPC) code ensembles for symmetric Gaussian multiple-access channels (GMAC). To characterize the probability density function (PDF) of the message passing in the process of joint decoding, we propose a new scheme to construc…

Cited by 0SourceScholar