← Search

Zhiliang Wu

17 accepted papers

2026

DLVINet: Advancing Dual-Lens Video Inpainting Beyond Parallax Constraints

AAAI 2026technical

Dual-lens video inpainting aims to simultaneously restore missing or corrupted contents in videos captured by each lens of binocular systems. Although preliminary explorations have been conducted, existing methods still face two key challenges: limited exploitation of long-range reference informatio

Cited by 0SourcePDFScholar
2026

VAST: Video Ability-Stratified Taxonomy for Data-Efficient Video Reasoning

CVPR 2026

Reinforcement learning (RL) has emerged as an effective approach for improving video reasoning in multimodal large language models (MLLMs). However, existing methods remain inefficient for two reasons. First, training data are typically organized by task formats rather than underlying reasoning abil

Cited by 0SourcecodeScholar
2025

BVINet: Unlocking Blind Video Inpainting with Zero Annotations

ICCV 2025poster

Video inpainting aims to fill in corrupted regions of the video with plausible contents. Existing methods generally assume that the locations of corrupted regions are known, focusing primarily on the "how to inpaint". This reliance necessitates manual annotation of the corrupted regions using binary…

Cited by 0SourcePDFScholar
2025

Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs

EMNLP 2025

Sparse Mixture-of-Experts (SMoE) architectures are widely used in large language models (LLMs) due to their computational efficiency. However, though only a few experts are activated for each token, SMoE still requires loading all expert parameters, leading to high memory usage and challenges in dep

Cited by 0SourcePDFScholar
2025

MMAD: Multi-label Micro-Action Detection in Videos

ICCV 2025poster

Human body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, which are subtle, low-intensity body movements with promising applications in human emotion analysis. In real-world scenari…

2025

Prototypical Calibrating Ambiguous Samples for Micro-Action Recognition

AAAI 2025technical

Micro-Action Recognition (MAR) has gained increasing attention due to its crucial role as a form of non-verbal communication in social interactions, with promising potential for applications in human communication and emotion analysis. However, current approaches often overlook the inherent ambiguit…

2024

Text-Video Completion Networks With Motion Compensation And Attention Aggregation

ICASSP 2024accepted

The purpose of video inpainting is to fill a specified area with reasonable content. However, in the case of multiple targets and complex textures, current methods struggle to distinguish between feature information of the targets, leading to confusing or fuzzy inpainting results. In this paper, we…

Cited by 0SourceScholar
2024

WaveFormer: Wavelet Transformer for Noise-Robust Video Inpainting

AAAI 2024technical

Video inpainting aims to fill in the missing regions of the video frames with plausible content. Benefiting from the outstanding long-range modeling capacity, the transformer-based models have achieved unprecedented performance regarding inpainting quality. Essentially, coherent contents from all th…

Cited by 18SourcePDFScholar
2023

Flow-Guided Deformable Alignment Network with Self-Supervision for Video Inpainting

ICASSP 2023accepted

Video inpainting aims to utilize plausible contents to fill missing regions in the video. State-of-the-art video inpainting methods typically generate the missing contents of the target frame (current frame) by aggregating the temporal information of reference frames (neighboring frames) aligned usi…

Cited by 0SourceScholar
2023

Semi-Supervised Video Inpainting With Cycle Consistency Constraints

CVPR 2023poster

Deep learning-based video inpainting has yielded promising results and gained increasing attention from researchers. Generally, these methods usually assume that the corrupted region masks of each frame are known and easily obtained. However, the annotation of these masks are labor-intensive and exp…

Cited by 18SourcePDFScholar
2022

A Proposal-Based Paradigm for Self-Supervised Sound Source Localization in Videos

CVPR 2022poster

Humans can easily recognize where and how the sound is produced via watching a scene and listening to corresponding audio cues. To achieve such cross-modal perception on machines, existing methods only use the maps generated by interpolation operations to localize the sound source. As semantic objec…

Cited by 22PDFScholar
2022

Active Contrastive Set Mining for Robust Audio-Visual Instance Discrimination

IJCAI 2022poster

The recent success of audio-visual representation learning can be largely attributed to their pervasive property of audio-visual synchronization, which can be used as self-annotated supervision. As a state-of-the-art solution, Audio-Visual Instance Discrimination (AVID) extends instance discriminati…

Cited by 1SourcePDFScholar