← Search

Junwen Xiong

2 accepted papers

2024

DiffSal: Joint Audio and Video Learning for Diffusion Saliency Prediction

CVPR 2024poster

Audio-visual saliency prediction can draw support from diverse modality complements but further performance enhancement is still challenged by customized architectures as well as task-specific loss functions. In recent studies denoising diffusion models have shown more promising in unifying task fra…

Cited by 6SourcePDFScholar
2023

CASP-Net: Rethinking Video Saliency Prediction From an Audio-Visual Consistency Perceptual Perspective

CVPR 2023poster

Incorporating the audio stream enables Video Saliency Prediction (VSP) to imitate the selective attention mechanism of human brain. By focusing on the benefits of joint auditory and visual information, most VSP methods are capable of exploiting semantic correlation between vision and audio modalitie…