← Search

Wei Zhai

44 accepted papers

2026

Diffusion Guided Chain-of-Vision for Large Autoregressive Vision Models

CVPR 2026

Chain-of-Thought (CoT) has recently shown encouraging progress in the vision language model. However, the pure-vision CoT (i.e., chain-of-vision) has been underexplored in visual in-context learning. In this paper, we introduce Diffusion Guided Chain-of-Vision, which integrates an explicit chain-of-

Cited by 0SourcecodeScholar
2026

E-MaT:Event-oriented Mamba for Egocentric Point Tracking

AAAI 2026technical

Egocentric point tracking aims to localize points on object surfaces from a first-person perspective and serves as a critical step toward embodied intelligence. Recent methods rely on video input, tracking query points through feature matching across consecutive frames. However, these methods strug

Cited by 0SourcePDFScholar
2026

Gloria: Consistent Character Video Generation via Content Anchors

CVPR 2026

Digital characters are central to modern media, yet generating character videos with long-duration, consistent multi-view appearance and expressive identity remains challenging. Existing approaches either provide insufficient context to preserve identity or leverage non-character-centric information

Cited by 0SourceScholar
2026

TOUCH: Text-guided Controllable Generation of Free-Form Hand-Object Interactions

ICLR 2026poster

Hand-object interaction (HOI) is fundamental for humans to express intent. Existing HOI generation research is predominantly confined to fixed grasping patterns, where control is tied to physical priors such as force closure or generic intent instructions, even when expressed through elaborate langu…

Cited by 0SourceScholar
2026

Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model

ICLR 2026poster

Autoregressive image generation aims to predict the next token based on previous ones. However, this process is challenged by the bidirectional dependencies inherent in conventional image tokenizations, which creates a fundamental misalignment with the unidirectional nature of autoregressive models.…

Cited by 0SourcecodeScholar
2026

UGround: Towards Unified Visual Grounding with Unrolled Transformers

ICML 2026poster

We present UGround, a **U**nified visual **Ground**ing paradigm that dynamically selects intermediate layers across **U**nrolled transformers as "mask as prompt'', diverging from the prevailing pipeline that leverages the fixed last hidden layer as "$\texttt{\}$ as prompt''. UGround addresses two pr…

Cited by 0SourceScholar
2026

Unbiased Gradient Estimation for Event Binning via Functional Backpropagation

ICLR 2026poster

Event-based vision encodes dynamic scenes as asynchronous spatio-temporal spikes called events. To leverage conventional image processing pipelines, events are typically binned into frames. However, binning functions are discontinuous, which truncates gradients at the frame level and forces most eve…

Cited by 0SourcecodeScholar
2026

WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens

CVPR 2026

Recent progress in multimodal large language models (MLLMs) has highlighted the challenge of efficiently bridging pre-trained Vision-Language Models (VLMs) with Diffusion Models. While methods using a fixed number of learnable query tokens offer computational efficiency, they suffer from task genera

Cited by 0SourceScholar
2025

Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning

CVPR 2025poster

Generating detailed captions comprehending text-rich visual content in images has received growing attention for Large Vision-Language Models (LVLMs). However, few studies have developed benchmarks specifically tailored for detailed captions to measure their accuracy and comprehensiveness. In this p…

2025

EF-3DGS: Event-Aided Free-Trajectory 3D Gaussian Splatting

NeurIPS 2025spotlight

Scene reconstruction from casually captured videos has wide real-world applications. Despite recent progress, existing methods relying on traditional cameras tend to fail in high-speed scenarios due to insufficient observations and inaccurate pose estimation. Event cameras, inspired by biological vi…

Cited by 0SourceScholar
2025

GREAT: Geometry-Intention Collaborative Inference for Open-Vocabulary 3D Object Affordance Grounding

CVPR 2025poster

Open-Vocabulary 3D object affordance grounding aims to anticipate "action possibilities" regions on 3D objects with arbitrary instructions, which is crucial for robots to generically perceive real scenarios and respond to operational changes. Existing methods focus on combining images or languages t…

2025

Generalizable Cross-Lingual Cognitive Distortion Detection with Standardized Annotations and Multi-Task Learning

ACL 2025finding

Cognitive distortion is a critical issue in psychology, with most existing studies based on Burns’ cognitive distortion theory. However, differences in annotation standards lead to variations in building analysis tools, resulting in inconsistent analyses and limiting the generalizability of findings…

2025

Improved Video VAE for Latent Video Diffusion Model

CVPR 2025poster

Variational Autoencoder (VAE) aims to compress pixel data into low-dimensional latent space, playing an important role in OpenAI's Sora and other latent video diffusion generation models. While most existing video VAEs inflate a pre-trained image VAE into the 3D causal structure for temporal-spatial…

2025

MATE: Motion-Augmented Temporal Consistency for Event-based Point Tracking

ICCV 2025poster

Tracking Any Point (TAP) plays a crucial role in motion analysis. Video-based approaches rely on iterative local matching for tracking, but they assume linear motion during the blind time between frames, which leads to point loss under large displacements or nonlinear motion. The high temporal resol…

Cited by 0SourcePDFScholar
2025

MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling

CVPR 2025poster

Recent advancements in multi-modal large language models have propelled the development of joint probabilistic models capable of both image understanding and generation. However, we have identified that recent methods suffer from loss of image information during understanding task, due to either ima…

Cited by 11SourcePDFScholar
2025

MentalGLM Series: Explainable Large Language Models for Mental Health Analysis on Chinese Social Media

EMNLP 2025

With the rise of mental health challenges, social media has become a key platform for emotional expression. Deep learning offers a promising solution for analyzing mental health but lacks flexibility and interpretability. Large language models (LLMs) introduce greater adaptability and can explain th

2025

PAID: Pairwise Angular-Invariant Decomposition for Continual Test-Time Adaptation

NeurIPS 2025poster

Continual Test-Time Adaptation (CTTA) aims to online adapt a pre-trained model to changing environments during inference. Most existing methods focus on exploiting target data, while overlooking another crucial source of information, the pre-trained weights, which encode underutilized domain-invaria…

Cited by 0SourcecodeScholar
2025

SIGMAN: Scaling 3D Human Gaussian Generation with Millions of Assets

ICCV 2025poster

3D human digitization has long been a highly pursued yet challenging task. Existing methods aim to generate high-quality 3D digital humans from single or multiple views, but remain primarily constrained by current paradigms and the scarcity of 3D human assets. Specifically, recent approaches fall in…

Cited by 0SourcePDFScholar
2025

ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models

NeurIPS 2025poster

Panoramic video generation aims to synthesize 360-degree immersive videos, holding significant importance in the fields of VR, world models, and spatial intelligence. Existing works fail to synthesize high-quality panoramic videos due to the inherent modality gap between panoramic data and perspecti…

Cited by 0SourceScholar
2024

Chinese MentalBERT: Domain-Adaptive Pre-training on Social Media for Chinese Mental Health Text Analysis

ACL 2024findings

In the current environment, psychological issues are prevalent and widespread, with social media serving as a key outlet for individuals to share their feelings. This results in the generation of vast quantities of data daily, where negative emotions have the potential to precipitate crisis situatio…

2024

EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views

NeurIPS 2024poster

Understanding egocentric human-object interaction (HOI) is a fundamental aspect of human-centric perception, facilitating applications like AR/VR and embodied AI. For the egocentric HOI, in addition to perceiving semantics e.g., ''what'' interaction is occurring, capturing ''where'' the interaction…

Cited by 6SourcePDFScholar
2024

Hypercorrelation Evolution for Video Class-Incremental Learning

AAAI 2024technical

Video class-incremental learning aims to recognize new actions while restricting the catastrophic forgetting of old ones, whose representative samples can only be saved in limited memory. Semantically variable subactions are susceptible to class confusion due to data imbalance. While existing method…

2024

LEMON: Learning 3D Human-Object Interaction Relation from 2D Images

CVPR 2024poster

Learning 3D human-object interaction relation is pivotal to embodied AI and interaction modeling. Most existing methods approach the goal by learning to predict isolated interaction elements e.g. human contact object affordance and human-object spatial relation primarily from the perspective of eith…

2023

Exploring Tuning Characteristics of Ventral Stream’s Neurons for Few-Shot Image Classification

AAAI 2023technical

Human has the remarkable ability of learning novel objects by browsing extremely few examples, which may be attributed to the generic and robust feature extracted in the ventral stream of our brain for representing visual objects. In this sense, the tuning characteristics of ventral stream's neurons…

Cited by 12SourcePDFScholar
2023

Grounding 3D Object Affordance from 2D Interactions in Images

ICCV 2023poster

Grounding 3D object affordance seeks to locate objects' "action possibilities" regions in the 3D space, which serves as a link between perception and operation for embodied agents. Existing studies primarily focus on connecting visual affordances with geometry structures, e.g., relying on annotation…

Cited by 34PDFcodeScholar
2023

Leverage Interactive Affinity for Affordance Learning

CVPR 2023poster

Perceiving potential "action possibilities" (i.e., affordance) regions of images and learning interactive functionalities of objects from human demonstration is a challenging task due to the diversity of human-object interactions. Prevailing affordance learning algorithms often adopt the label assig…

2023

Spatial-Aware Token for Weakly Supervised Object Localization

ICCV 2023poster

Weakly supervised object localization (WSOL) is a challenging task aiming to localize objects with only image-level supervision. Recent works apply visual transformer to WSOL and achieve significant success by exploiting the long-range feature dependency in self-attention mechanism. However, existin…

Cited by 13PDFcodeScholar
2023

Uncertainty-Aware Optimal Transport for Semantically Coherent Out-of-Distribution Detection

CVPR 2023poster

Semantically coherent out-of-distribution (SCOOD) detection aims to discern outliers from the intended data distribution with access to unlabeled extra set. The coexistence of in-distribution and out-of-distribution samples will exacerbate the model overfitting when no distinction is made. To addres…

2022

Exploring Figure-Ground Assignment Mechanism in Perceptual Organization

NeurIPS 2022accept

Perceptual organization is a challenging visual task that aims to perceive and group the individual visual element so that it is easy to understand the meaning of the scene as a whole. Most recent methods building upon advanced Convolutional Neural Network (CNN) come from learning discriminative rep…

Cited by 20SourcePDFScholar
2022

Self-Sustaining Representation Expansion for Non-Exemplar Class-Incremental Learning

CVPR 2022poster

Non-exemplar class-incremental learning is to recognize both the old and new classes when old class samples cannot be saved. It is a challenging task since representation optimization and feature retention can only be achieved under supervision from new classes. To address this problem, we propose a…

Cited by 207PDFScholar
2021

Self-Promoted Prototype Refinement for Few-Shot Class-Incremental Learning

CVPR 2021poster

Few-shot class-incremental learning is to recognize the new classes given few samples and not forget the old classes. It is a challenging task since representation optimization and prototype reorganization can only be achieved under little supervision. To address this problem, we propose a novel inc…

Cited by 203PDFcodeScholar
2018

A Generative Adversarial Network Based Framework for Unsupervised Visual Surface Inspection

ICASSP 2018accepted

Visual surface inspection is a challenging task due to the highly inconsistent appearance of the target surfaces and the abnormal regions. Most of the state-of-the-art methods are highly dependent on the labelled training samples, which are difficult to collect in practical industrial applications.…

Cited by 0SourceScholar