← Search

Chunjie Zhang

11 accepted papers

2026

Divid: Disentangled Spatial-Temporal Modeling within LLMs for Temporally Grounded Video Understanding

ICLR 2026poster

Recent advances in Video LLMs have improved video understanding performance, but temporally grounded understanding in long-form videos remains challenging. Most models encode video frames into a flat sequence of visual tokens, which are then processed together with textual input by the LLM. While ef…

Cited by 0SourceScholar
2026

Leveraging Failed Samples: A Few-Shot and Training-Free Framework for Generalized Deepfake Detection

AAAI 2026technical

Recent deepfake detection studies often treat unseen sample detection as a ``zero-shot" task, training on images generated by known models but generalizing to unknown ones. A key real-world challenge arises when a model performs poorly on unknown samples, yet these samples remain available for analy

Cited by 0SourcePDFScholar
2025

Attend and Enrich: Enhanced Visual Prompt for Zero-Shot Learning

AAAI 2025technical

Zero-shot learning (ZSL) endeavors to transfer knowledge from the seen categories to recognize unseen categories, which mostly relies on the semantic-visual interactions between image and attribute tokens. Recently, the prompt learning has emerged in ZSL and demonstrated significant potential as it…

2025

Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

CVPR 2025poster

In recent years, numerous tasks have been proposed to encourage model to develop specified capability in understanding audio-visual scene, primarily categorized into temporal localization, spatial localization, spatio-temporal reasoning, and pixel-level understanding. Instead, human possesses a unif…

2025

Visual Relation Diffusion for Human-Object Interaction Detection

ICCV 2025poster

Human-object interaction (HOI) detection relies on fine-grained visual understanding to distinguish complex relationships between humans and objects. While recent generative diffusion models have demonstrated remarkable capability in learning detailed visual concepts through pixel-level generation,…

Cited by 0SourcePDFScholar
2023

An In-Depth Exploration of Person Re-Identification and Gait Recognition in Cloth-Changing Conditions

CVPR 2023poster

The target of person re-identification (ReID) and gait recognition is consistent, that is to match the target pedestrian under surveillance cameras. For the cloth-changing problem, video-based ReID is rarely studied due to the lack of a suitable cloth-changing benchmark, and gait recognition is ofte…

2023

CTP:Towards Vision-Language Continual Pretraining via Compatible Momentum Contrast and Topology Preservation

ICCV 2023poster

Vision-Language Pretraining (VLP) has shown impressive results on diverse downstream tasks by offline training on large-scale datasets. Regarding the growing nature of real-world data, such an offline training paradigm on ever-expanding data is unsustainable, because models lack the continual learni…

Cited by 33PDFcodeScholar
2023

Progressive Semantic-Visual Mutual Adaption for Generalized Zero-Shot Learning

CVPR 2023highlight

Generalized Zero-Shot Learning (GZSL) identifies unseen categories by knowledge transferred from the seen domain, relying on the intrinsic interactions between visual and semantic information. Prior works mainly localize regions corresponding to the sharing attributes. When various visual appearance…

2021

Progressively Complementary Network for Fisheye Image Rectification Using Appearance Flow

CVPR 2021poster

Distortion rectification is often required for fisheye images. The generation-based method is one mainstream solution due to its label-free property, but its naive skip-connection and overburdened decoder will cause blur and incomplete correction. First, the skip-connection directly transfers the im…

Cited by 57PDFcodeScholar
2021

Towards Fast and Accurate Real-World Depth Super-Resolution: Benchmark Dataset and Baseline

CVPR 2021poster

Depth maps obtained by commercial depth sensors are always in low-resolution, making it difficult to be used in various computer vision tasks. Thus, depth map super-resolution (SR) is a practical and valuable task, which upscales the depth map into high-resolution (HR) space. However, limited by the…

Cited by 100PDFScholar
2018

Two-Step Quantization for Low-Bit Neural Networks

CVPR 2018poster

Every bit matters in the hardware design of quantized neural networks. However, extremely-low-bit representation usually causes large accuracy drop. Thus, how to train extremely-low-bit neural networks with high accuracy is of central importance. Most existing network quantization approaches learn t…

Cited by 167SourcePDFScholar