← Search

Huan Zhou

8 accepted papers

2025

CAMEL: Cross-Attention Enhanced Mixture-of-Experts and Language Bias for Code-Switching Speech Recognition

ICASSP 2025accepted

Code-switching automatic speech recognition (ASR) aims to transcribe speech that contains two or more languages accurately. To better capture language-specific speech representations and address language confusion in code-switching ASR, the mixture-of-experts (MoE) architecture and an additional lan…

Cited by 0SourceScholar
2025

SCDiar: a streaming diarization system based on speaker change detection and speech recognition

ICASSP 2025accepted

In hours-long meeting scenarios, real-time speech stream often struggles with achieving accurate speaker diarization, commonly leading to speaker identification and speaker count errors. To address this challenge, we propose SCDiar, a system that operates on speech segments, split at the token level…

Cited by 0SourceScholar
2023

X-SEPFORMER: End-To-End Speaker Extraction Network with Explicit Optimization on Speaker Confusion

ICASSP 2023accepted

Target speech extraction (TSE) systems are designed to extract target speech from a multi-talker mixture. The popular training objective for most prior TSE networks is to enhance reconstruction performance of extracted speech waveform. However, it has been reported that a TSE system delivers high re…

Cited by 0SourceScholar
2021

A Novel end-to-end Speech Emotion Recognition Network with Stacked Transformer Layers

ICASSP 2021accepted

Speech emotion recognition (SER) aims to automatically recognize emotional category for a given speech utterance. The performance of a SER system heavily relies on the effectiveness of global representation expressed at utterance level. To effectively extract such a global feature, the mainstream of…

Cited by 0SourceScholar
2021

Unimodal and Crossmodal Refinement Network for Multimodal Sequence Fusion

EMNLP 2021main

Effective unimodal representation and complementary crossmodal representation fusion are both important in multimodal representation learning. Prior works often modulate one modal feature to another straightforwardly and thus, underutilizing both unimodal and crossmodal representation refinements, w…

2017

A new noise annoyance measurement metric for urban noise sensing and evaluation

ICASSP 2017accepted

This paper investigates the problem of noise-induced annoyance level evaluation, and proposes a novel annoyance measurement metric for more efficient and accurate evaluation of annoyance level of different types of noises. Results from a large-scale subjective listening test using 90 different noise…

Cited by 0SourceScholar
2016

Enhanced vote count circuit based on nor flash memory for fast similarity search

ICASSP 2016accepted

A memory-based search circuit is introduced in this paper. In this circuit, the conventional memory structure is customized to provide equality comparison for each column of memory array, and a counting circuit is included at each column to record the degree of matches between query and reference da…

Cited by 0SourceScholar