← Search

Weiguang Chen

3 accepted papers

2025

Injecting Visual Features into Whisper for Parameter-Efficient Noise-Robust Audio-Visual Speech Recognition

ICASSP 2025accepted

Audio-visual speech recognition (AVSR) aims to enhance the robustness of an automatic speech recognition (ASR) systems by incorporating visual information from lip movements, especially in challenging noisy environments. Nevertheless, most current approaches either involve training from scratch or f…

Cited by 0SourceScholar
2024

Enhancing Low-Latency Speaker Diarization with Spatial Dictionary Learning

ICASSP 2024accepted

This study proposes a low-latency online speaker diarization framework. Specifically, we design a spatial dictionary learning module shared across different frequency bands, enabling spatial feature learning at each frequency bin. This contributes to reducing the latency constraints of the online di…

Cited by 0SourceScholar
2024

GAMMA: Graspability-Aware Mobile MAnipulation Policy Learning based on Online Grasping Pose Fusion

ICRA 2024poster

Mobile manipulation constitutes a fundamental task for robotic assistants and garners significant attention within the robotics community. A critical challenge inherent in mobile manipulation is the effective observation of the target while approaching it for grasping. In this work, we propose a gra…

Cited by 24SourcecodeScholar