← Search

Jiawen Kang

10 accepted papers

2026

Eliminate Distance Differences Induced by Backdoor Attacks: Layer-Selective Training and Clipping to Mask Backdoor Models

CVPR 2026

Federated learning (FL) enables a central server to collaboratively train a global model with multiple clients while preserving data privacy. However, the distributed nature of FL makes the paradigm vulnerable to backdoor attacks, as proved by numerous recent studies. Although existing studies impro

Cited by 0SourceScholar
2025

Disentangling Speakers in Multi-Talker Speech Recognition with Speaker-Aware CTC

ICASSP 2025accepted

Multi-talker speech recognition (MTASR) faces unique challenges in disentangling and transcribing overlapping speech. To address these challenges, this paper investigates the role of Connectionist Temporal Classification (CTC) in speaker disentanglement when incorporated with Serialized Output Train…

Cited by 0SourceScholar
2025

Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions

ICASSP 2025accepted

Recent advancements in large language models (LLMs) have revolutionized various domains, bringing significant progress and new opportunities. Despite progress in speech-related tasks, LLMs have not been sufficiently explored in multi-talker scenarios. In this work, we present a pioneering effort to…

Cited by 0SourceScholar
2024

Cross-Speaker Encoding Network for Multi-Talker Speech Recognition

ICASSP 2024accepted

End-to-end multi-talker speech recognition has garnered great interest as an effective approach to directly transcribe overlapped speech from multiple speakers. Current methods typically adopt either 1) single-input multiple-output (SIMO) models with a branched encoder, or 2) single-input single-out…

Cited by 0SourceScholar
2024

Generative Al-aided Joint Training-free Secure Semantic Communications via Multi-modal Prompts

ICASSP 2024accepted

Semantic communication (SemCom) holds promise for reducing network resource consumption while achieving the communications goal. However, the computational overheads in jointly training semantic encoders and decoders—and the subsequent deployment in network devices—are overlooked. Recent advances in…

Cited by 0SourceScholar
2023

A Sidecar Separator Can Convert A Single-Talker Speech Recognition System to A Multi-Talker One

ICASSP 2023accepted

Although automatic speech recognition (ASR) can perform well in common non-overlapping environments, sustaining performance in multi-talker overlapping speech recognition remains challenging. Recent research revealed that ASR model’s encoder captures different levels of information with different la…

Cited by 0SourceScholar
2022

The CUHK-Tencent Speaker Diarization System for the ICASSP 2022 Multi-Channel Multi-Party Meeting Transcription Challenge

ICASSP 2022accepted

This paper describes our speaker diarization system submitted to the Multi-channel Multi-party Meeting Transcription (M2MeT) challenge, where Mandarin meeting data were recorded in multi-channel format for diarization and automatic speech recognition (ASR) tasks. In these meeting scenarios, the unce…

Cited by 0SourceScholar
2021

Communication-efficient and Scalable Decentralized Federated Edge Learning

IJCAI 2021poster

Federated Edge Learning (FEL) is a distributed Machine Learning (ML) framework for collaborative training on edge devices. FEL improves data privacy over traditional centralized ML model training by keeping data on the devices and only sending local model updates to a central coordinator for aggrega…

Cited by 15SourcePDFScholar
2021

Squeezing Value of Cross-Domain Labels: A Decoupled Scoring Approach for Speaker Verification

ICASSP 2021accepted

Domain mismatch often occurs in real applications and causes serious performance reduction on speaker verification systems. The common wisdom is to collect cross-domain data and train a multi-domain PLDA model, with the hope to learn a domain-independent speaker subspace. In this paper, we firstly p…

Cited by 0SourceScholar
2020

CN-Celeb: A Challenging Chinese Speaker Recognition Dataset

ICASSP 2020accepted

Recently, researchers set an ambitious goal of conducting speaker recognition in unconstrained conditions where the variations on ambient, channel and emotion could be arbitrary. However, most publicly available datasets are collected under constrained environments, i.e., with little noise and limit…

Cited by 0SourceScholar