← Search

Hanyu Xuan

8 accepted papers

2024

Text-Video Completion Networks With Motion Compensation And Attention Aggregation

ICASSP 2024accepted

The purpose of video inpainting is to fill a specified area with reasonable content. However, in the case of multiple targets and complex textures, current methods struggle to distinguish between feature information of the targets, leading to confusing or fuzzy inpainting results. In this paper, we…

Cited by 0SourceScholar
2024

WaveFormer: Wavelet Transformer for Noise-Robust Video Inpainting

AAAI 2024technical

Video inpainting aims to fill in the missing regions of the video frames with plausible content. Benefiting from the outstanding long-range modeling capacity, the transformer-based models have achieved unprecedented performance regarding inpainting quality. Essentially, coherent contents from all th…

Cited by 18SourcePDFScholar
2023

Flow-Guided Deformable Alignment Network with Self-Supervision for Video Inpainting

ICASSP 2023accepted

Video inpainting aims to utilize plausible contents to fill missing regions in the video. State-of-the-art video inpainting methods typically generate the missing contents of the target frame (current frame) by aggregating the temporal information of reference frames (neighboring frames) aligned usi…

Cited by 0SourceScholar
2023

Semi-Supervised Video Inpainting With Cycle Consistency Constraints

CVPR 2023poster

Deep learning-based video inpainting has yielded promising results and gained increasing attention from researchers. Generally, these methods usually assume that the corrupted region masks of each frame are known and easily obtained. However, the annotation of these masks are labor-intensive and exp…

Cited by 18SourcePDFScholar
2022

A Proposal-Based Paradigm for Self-Supervised Sound Source Localization in Videos

CVPR 2022poster

Humans can easily recognize where and how the sound is produced via watching a scene and listening to corresponding audio cues. To achieve such cross-modal perception on machines, existing methods only use the maps generated by interpolation operations to localize the sound source. As semantic objec…

Cited by 22PDFScholar
2022

Active Contrastive Set Mining for Robust Audio-Visual Instance Discrimination

IJCAI 2022poster

The recent success of audio-visual representation learning can be largely attributed to their pervasive property of audio-visual synchronization, which can be used as self-annotated supervision. As a state-of-the-art solution, Audio-Visual Instance Discrimination (AVID) extends instance discriminati…

Cited by 1SourcePDFScholar