← Search

Honglie Chen

8 accepted papers

2025

Contextual Speech Extraction: Leveraging Textual History as an Implicit Cue for Target Speech Extraction

ICASSP 2025accepted

In this paper, we investigate a novel approach for Target Speech Extraction (TSE), which relies solely on textual context to extract the target speech. We refer to this task as Contextual Speech Extraction (CSE). Unlike traditional TSE methods that rely on pre-recorded enrollment utterances, video o…

Cited by 0SourceScholar
2025

Large Language Models are Strong Audio-Visual Speech Recognition Learners

ICASSP 2025accepted

Multimodal large language models (MLLMs) have recently become a focal point of research due to their formidable multimodal understanding capabilities. For example, in the audio and speech domains, an LLM can be equipped with (automatic) speech recognition (ASR) abilities by just concatenating the au…

Cited by 0SourceScholar
2025

MoME: Mixture of Matryoshka Experts for Audio-Visual Speech Recognition

NeurIPS 2025poster

Large language models (LLMs) have recently shown strong potential in audio-visual speech recognition (AVSR), but their high computational demands and sensitivity to token granularity limit their practicality in resource-constrained settings. Token compression methods can reduce inference cost, but t…

Cited by 0SourceScholar
2024

Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs

NeurIPS 2024poster

Research in auditory, visual, and audiovisual speech recognition (ASR, VSR, and AVSR, respectively) has traditionally been conducted independently. Even recent self-supervised studies addressing two or all three tasks simultaneously tend to yield separate models, leading to disjoint inference pipeli…

2023

Auto-AVSR: Audio-Visual Speech Recognition with Automatic Labels

ICASSP 2023accepted

Audio-visual speech recognition has received a lot of attention due to its robustness against acoustic noise. Recently, the performance of automatic, visual, and audio-visual speech recognition (ASR, VSR, and AV-ASR, respectively) has been substantially improved, mainly due to the use of larger mode…

Cited by 0SourceScholar
2023

SynthVSR: Scaling Up Visual Speech Recognition With Synthetic Supervision

CVPR 2023poster

Recently reported state-of-the-art results in visual speech recognition (VSR) often rely on increasingly large amounts of video data, while the publicly available transcribed video datasets are limited in size. In this paper, for the first time, we study the potential of leveraging synthetic visual…

Cited by 27SourcePDFScholar
2021

Localizing Visual Sounds the Hard Way

CVPR 2021poster

The objective of this work is to localize sound sources that are visible in a video without using manual annotations. Our key technical contribution is to show that, by training the network to explicitly discriminate challenging image fragments, even for images that do contain the object emitting th…

Cited by 234PDFScholar