← Search

Jiadong Wang

7 accepted papers

2026

AV-SSAN: Audio-Visual Selective DOA Estimation Through Explicit Multi-Band Semantic-Spatial Alignment

AAAI 2026technical

Audio-visual sound source localization (AV-SSL) estimates the position of sound sources by fusing auditory and visual cues. Current AV-SSL methodologies typically require spatially-paired audio-visual data and cannot selectively localize specific target sources. To address these limitations, we intr

Cited by 0SourcePDFScholar
2024

Raindrop Clarity: A Dual-Focused Dataset for Day and Night Raindrop Removal

ECCV 2024poster

"Existing raindrop removal datasets have two shortcomings. First, they consist of images captured by cameras with a focus on the background, leading to the presence of blurry raindrops. To our knowledge, none of these datasets include images where the focus is specifically on raindrops, which result…

2024

Restoring Speaking Lips from Occlusion for Audio-Visual Speech Recognition

AAAI 2024technical

Prior studies on audio-visual speech recognition typically assume the visibility of speaking lips, ignoring the fact that visual occlusion occurs in real-world videos, thus adversely affecting recognition performance. To address this issue, we propose a framework that restores occluded lips in a vid…

Cited by 11SourcePDFScholar
2023

Seeing What You Said: Talking Face Generation Guided by a Lip Reading Expert

CVPR 2023poster

Talking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and visual quality. Despite much progress, they hardly focus on the content of lip move…

2022

A Hybrid Learning Framework for Deep Spiking Neural Networks with One-Spike Temporal Coding

ICASSP 2022accepted

Bio-inspired spiking neural networks (SNNs) are compelling candidates for spatio-temporal information processing on ultra-low power neuromorphic computing chips. However, the existing SNN training methods have not fully exploited the temporal information of spikes that plays a critical role in spars…

Cited by 0SourceScholar
2021

GCC-PHAT with Speech-oriented Attention for Robotic Sound Source Localization

ICRA 2021poster

Robotic audition is a basic sense that helps robots perceive the surroundings and interact with humans. Sound Source Localization (SSL) is an essential module for a robotic system. However, the performance of most sound source localization techniques degrades in noisy and reverberant environments du…

Cited by 19SourceScholar
2021

Multi-Target DoA Estimation with an Audio-Visual Fusion Mechanism

ICASSP 2021accepted

Most of the prior studies in the spatial Direction of Arrival (DoA) domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio and visual signals for multi-speaker local…

Cited by 0SourceScholar