← Search

Berlin Chen

23 accepted papers

2026

MALEFA: MULTI-GRANULARITY LEARNING AND EFFECTIVE FALSE ALARM SUPPRESSION FOR ZERO-SHOT KEYWORD SPOTTING

ICASSP 2026poster

User-defined keyword spotting (KWS) without resorting to domain-specific pre-labeled training data is of fundamental importance in building adaptable and personalized voice interfaces. However, such systems are still faced with arduous challenges, including constrained computational resources and li…

Cited by 0SourcePDFScholar
2026

Mamba-3: Improved Sequence Modeling using State Space Principles

ICLR 2026oral

The recent scaling of test-time compute for LLMs has restricted the practical deployment of models to those with strong capabilities that can generate high-quality outputs in an inference-efficient manner. While current Transformer-based models are the standard, their quadratic compute and linear me…

Cited by 0SourcecodeScholar
2025

Channel-Aware Domain-Adaptive Generative Adversarial Network for Robust Speech Recognition

ICASSP 2025accepted

While pre-trained automatic speech recognition (ASR) systems demonstrate impressive performance on matched domains, their performance often degrades when confronted with channel mismatch stemming from unseen recording environments and conditions. To mitigate this issue, we propose a novel channel-aw…

Cited by 0SourceScholar
2025

ConPCO: Preserving Phoneme Characteristics For Automatic Pronunciation Assessment Leveraging Contrastive Ordinal Regularization

ICASSP 2025accepted

Automatic pronunciation assessment (APA) manages to evaluate the pronunciation proficiency of a second language (L2) learner in a target language. Existing efforts typically draw on regression models for proficiency score prediction, wherein the models are trained to estimate target values without e…

Cited by 0SourceScholar
2025

Long-Context State-Space Video World Models

ICCV 2025poster

Video diffusion models have recently shown promise for world modeling through autoregressive frame prediction conditioned on actions. However, they struggle to maintain long-term memory due to the high computational cost associated with processing extended sequences in attention layers. To overcome…

Cited by 0SourcePDFScholar
2025

Towards Efficient and Multifaceted Computer-assisted Pronunciation Training Leveraging Hierarchical Selective State Space Model and Decoupled Cross-entropy Loss

NAACL 2025long

Prior efforts in building computer-assisted pronunciation training (CAPT) systems often treat automatic pronunciation assessment (APA) and mispronunciation detection and diagnosis (MDD) as separate fronts: the former aims to provide multiple pronunciation aspect scores across diverse linguistic leve…

2024

An Effective Automated Speaking Assessment Approach to Mitigating Data Scarcity and Imbalanced Distribution

NAACL 2024findings

Automated speaking assessment (ASA) typically involves automatic speech recognition (ASR) and hand-crafted feature extraction from the ASR transcript of a learner’s speech. Recently, self-supervised learning (SSL) has shown stellar performance compared to traditional methods. However, SSL-based ASA…

Cited by 4SourcePDFScholar
2024

An Effective Mixture-Of-Experts Approach For Code-Switching Speech Recognition Leveraging Encoder Disentanglement

ICASSP 2024accepted

With the massive developments of end-to-end (E2E) neural networks, recent years have witnessed unprecedented breakthroughs in automatic speech recognition (ASR). However, the code-switching phenomenon remains a major obstacle that hinders ASR from perfection, as the lack of labeled data and the vari…

Cited by 0SourceScholar
2024

An Effective Pronunciation Assessment Approach Leveraging Hierarchical Transformers and Pre-training Strategies

ACL 2024long

Automatic pronunciation assessment (APA) manages to quantify a second language (L2) learner’s pronunciation proficiency in a target language by providing fine-grained feedback with multiple pronunciation aspect scores at various linguistic levels. Most existing efforts on APA typically parallelize t…

2024

DANCER: Entity Description Augmented Named Entity Corrector for Automatic Speech Recognition

COLING 2024main

End-to-end automatic speech recognition (E2E ASR) systems often suffer from mistranscription of domain-specific phrases, such as named entities, sometimes leading to catastrophic failures in downstream tasks. A family of fast and lightweight named entity correction (NEC) models for ASR have recently…

2024

What Do Neural Networks Listen to? Exploring the Crucial Bands in Speech Enhancement Using SINC-Convolution

ICASSP 2024accepted

This study introduces a reformed Sinc-convolution (Sincconv) framework tailored for the encoder component of deep networks for speech enhancement (SE). The reformed Sinc-conv, based on parametrized sinc functions as band-pass filters, offers notable advantages in terms of training efficiency, filter…

Cited by 0SourceScholar
2023

Effective Graph-Based Modeling of Articulation Traits for Mispronunciation Detection and Diagnosis

ICASSP 2023accepted

Mispronunciation detection and diagnosis (MDD) manages to pinpoint phone-level erroneous pronunciation segmentations and provide instant and informative diagnostic feedback to L2 (second-language) learners. Among the various modeling paradigms for MDD, dictation-based neural methods have recently be…

Cited by 0SourceScholar
2022

Exploring Non-Autoregressive End-to-End Neural Modeling for English Mispronunciation Detection and Diagnosis

ICASSP 2022accepted

End-to-end (E2E) neural modeling has emerged as one predominant school of thought to develop computer-assisted pronunciation training (CAPT) systems, showing competitive performance to conventional pronunciation-scoring based methods. However, current E2E neural methods for CAPT are faced with at le…

Cited by 0SourceScholar
2020

Spoken Document Retrieval Leveraging Bert-Based Modeling and Query Reformulation

ICASSP 2020accepted

Spoken document retrieval (SDR) has long been deemed a fundamental and important step towards efficient organization of, and access to multimedia associated with spoken content. In this paper, we present a novel study of SDR leveraging the Bidirectional Encoder Representations from Transformers (BER…

Cited by 0SourceScholar
2019

What do you learn from context? Probing for sentence structure in contextualized word representations

ICLR 2019poster

Contextualized representation models such as ELMo (Peters et al., 2018a) and BERT (Devlin et al., 2018) have recently achieved state-of-the-art results on a diverse array of downstream NLP tasks. Building on recent token-level probing work, we introduce a novel edge probing task design and construct…

Cited by 1017SourcePDFScholar
2018

Automatic Music Transcription Leveraging Generalized Cepstral Features and Deep Learning

ICASSP 2018accepted

Spectral features are limited in modeling musical signals with multiple concurrent pitches due to the challenge to suppress the interference of the harmonic peaks from one pitch to another. In this paper, we show that using multiple features represented in both the frequency and time domains with de…

Cited by 0SourceScholar
2018

Essence Vector-Based Query Modeling for Spoken Document Retrieval

ICASSP 2018accepted

Spoken document retrieval (SDR) has become a prominently required application since unprecedented volumes of multimedia data along with speech have become available in our daily life. As far as we are aware, there has been relatively less work in launching unsupervised paragraph embedding methods an…

Cited by 0SourceScholar
2017

A locality-preserving essence vector modeling framework for spoken document retrieval

ICASSP 2017accepted

Because unprecedented volumes of multimedia data associated with spoken documents have been made available to the public, spoken document retrieval (SDR) has become an important research area in the past decades. Recently, representation learning has emerged as an active research topic in many machi…

Cited by 0SourceScholar
2017

Leveraging manifold learning for extractive broadcast news summarization

ICASSP 2017accepted

Extractive speech summarization is intended to produce a condensed version of the original spoken document by selecting a few salient sentences from the document and concatenate them together to form a summary. In this paper, we study a novel use of manifold learning techniques for extractive speech…

Cited by 0SourceScholar
2016

Improved spoken document summarization with coverage modeling techniques

ICASSP 2016accepted

Extractive summarization aims at selecting a set of indicative sentences from a source document as a summary that can express the major theme of the document. A general consensus on extractive summarization is that both relevance and coverage are critical issues to address. The existing methods desi…

Cited by 0SourceScholar