← Search

Hongxu Zhu

3 accepted papers

2026

AV-SSAN: Audio-Visual Selective DOA Estimation Through Explicit Multi-Band Semantic-Spatial Alignment

AAAI 2026technical

Audio-visual sound source localization (AV-SSL) estimates the position of sound sources by fusing auditory and visual cues. Current AV-SSL methodologies typically require spatially-paired audio-visual data and cannot selectively localize specific target sources. To address these limitations, we intr

Cited by 0SourcePDFScholar
2024

An Empirical Study on the Impact of Positional Encoding in Transformer-Based Monaural Speech Enhancement

ICASSP 2024accepted

Transformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to enable Transformers to distinguish the order of elements in a sequence. However, it remains unclear how positional encod…

Cited by 0SourceScholar
2023

Ripple Sparse Self-Attention for Monaural Speech Enhancement

ICASSP 2023accepted

The use of Transformer represents a recent success in speech enhancement. However, as its core component, self-attention suffers from quadratic complexity, which is computationally prohibited for long speech recordings. Moreover, it allows each time frame to attend to all time frames, neglecting the…

Cited by 10SourceScholar