← Search

Hehe Fan

51 accepted papers

2026

4DPC$^2$hat: Towards Dynamic Point Cloud Understanding with Failure-Aware Bootstrapping

ICML 2026poster

Point clouds provide a compact and expressive representation of 3D objects, and have recently been integrated into multimodal large language models (MLLMs). However, existing methods primarily focus on static objects, while understanding dynamic point cloud sequences remains largely unexplored. This…

Cited by 0SourceScholar
2026

AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows

CVPR 2026

Training-free 3D editing aims to modify 3D shapes based on human instructions without model finetuning. It plays a crucial role in 3D content creation. However, existing approaches often struggle to produce strong or geometrically stable edits, largely due to inconsistent latent anchors introduced b

Cited by 0SourcecodeScholar
2026

DLVINet: Advancing Dual-Lens Video Inpainting Beyond Parallax Constraints

AAAI 2026technical

Dual-lens video inpainting aims to simultaneously restore missing or corrupted contents in videos captured by each lens of binocular systems. Although preliminary explorations have been conducted, existing methods still face two key challenges: limited exploitation of long-range reference informatio

Cited by 0SourcePDFScholar
2026

Endowing GPT-4 with a Humanoid Body: Building the Bridge Between Off-the-Shelf VLMs and the Physical World

ICLR 2026poster

In this paper, we explore how to empower general-purpose Vision-Language Models (VLMs) to control humanoid agents. General-purpose VLMs (e.g., GPT-4) exhibit strong open-world generalization, and remove the need for additional fine-tuning data. To build such an agent, two key components are required…

Cited by 0SourcecodeScholar
2026

Hi-Lo Prune: Look at What You'll Lose before Pruning with Hierarchical Token Selection

CVPR 2026

Multimodal Large Language Models (MLLMs) have achieved remarkable progress in vision-language understanding, yet processing long visual token sequences remains computationally expensive. Existing approaches mitigate this cost by reducing image tokens, either by discarding them after the visual encod

Cited by 0SourcecodeScholar
2026

Incentivizing Generative Zero-Shot Learning via Outcome-Reward Reinforcement Learning with Visual Cues

CVPR 2026

Recent advances in zero-shot learning (ZSL) have demonstrated the potential of generative models. Typically, generative ZSL synthesizes visual features conditioned on semantic prototypes to model the data distribution of unseen classes, followed by training a classifier on the synthesized data. Howe

Cited by 0SourceScholar
2026

PointThinker: Point-Incentivized Parallel Thinking for Multimodal Large Language Model

CVPR 2026

This paper explores parallel thinking for Multi-modal Large Language Models (MLLMs), aiming to improve Chain-of-Thought (CoT) through multiple diverse reasoning paths. We guide the model to list multiple visual key points and develop an independent reasoning path for each. Therefore, we term this me

Cited by 0SourceScholar
2026

SciEducator: Scientific Video Understanding and Educating via Deming-Cycle Multi-Agent System

CVPR 2026

Recent advancements in multimodal large language models (MLLMs) and video agent systems have significantly improved general video understanding. However, when applied to scientific video understanding and educating--a domain that demands external professional knowledge integration and rigorous step-

Cited by 0SourceScholar
2026

Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models

ICLR 2026poster

Rigged 3D assets are fundamental to 3D deformation and animation. However, existing 3D generation methods face challenges in generating animatable geometry, while rigging techniques lack fine-grained structural control over skeleton creation. To address these limitations, we introduce Stroke3D, a no…

Cited by 0SourceScholar
2026

Structured Reasoning for LLMs: A Unified Framework for Efficiency and Explainability

ICLR 2026poster

Recent Large Language Models (LLMs) have made remarkable progress, but they still struggle with complex reasoning tasks such as logical deduction and planning. This is partly because they rely primarily on token-level probability relationships, which limits their ability to reason effectively. In t…

Cited by 0SourcecodeScholar
2026

TransNormal: Dense Visual Semantics for Diffusion-based Transparent Object Normal Estimation

ICML 2026poster

Monocular normal estimation for transparent objects is critical for laboratory automation, yet it remains challenging due to complex light refraction and reflection. These optical properties often lead to catastrophic failures in conventional depth and normal sensors, hindering the deployment of emb…

Cited by 0SourceScholar
2026

UniF$^2$ace: A $\underline{Uni}$fied $\underline{F}$ine-grained $\underline{Face}$ Understanding and Generation Model

ICLR 2026poster

Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and generation. However, existing research in the face domain primarily faces two challenges: **(1) fragmentation development**…

Cited by 0SourcecodeScholar
2026

VAST: Video Ability-Stratified Taxonomy for Data-Efficient Video Reasoning

CVPR 2026

Reinforcement learning (RL) has emerged as an effective approach for improving video reasoning in multimodal large language models (MLLMs). However, existing methods remain inefficient for two reasons. First, training data are typically organized by task formats rather than underlying reasoning abil

Cited by 0SourcecodeScholar
2025

Adapting Text-to-Image Generation with Feature Difference Instruction for Generic Image Restoration

CVPR 2025poster

Diffusion-based Text-to-Image (T2I) models have demonstrated significant potential in image restoration. However, existing models continue to grapple with challenges such as complex training and prompt design. We introduce a new perspective for improving image restoration by injecting knowledge from…

Cited by 0SourcePDFScholar
2025

BVINet: Unlocking Blind Video Inpainting with Zero Annotations

ICCV 2025poster

Video inpainting aims to fill in corrupted regions of the video with plausible contents. Existing methods generally assume that the locations of corrupted regions are known, focusing primarily on the "how to inpaint". This reliance necessitates manual annotation of the corrupted regions using binary…

Cited by 0SourcePDFScholar
2025

DreamDPO: Aligning Text-to-3D Generation with Human Preferences via Direct Preference Optimization

ICML 2025poster

Text-to-3D generation automates 3D content creation from textual descriptions, which offers transformative potential across various fields. However, existing methods often struggle to align generated content with human preferences, limiting their applicability and flexibility. To address these limit…

Cited by 7SourcePDFScholar
2025

Dropping Experts, Recombining Neurons: Retraining-Free Pruning for Sparse Mixture-of-Experts LLMs

EMNLP 2025

Sparse Mixture-of-Experts (SMoE) architectures are widely used in large language models (LLMs) due to their computational efficiency. However, though only a few experts are activated for each token, SMoE still requires loading all expert parameters, leading to high memory usage and challenges in dep

Cited by 0SourcePDFScholar
2025

InfiniDreamer: Arbitrarily Long Human Motion Generation via Segment Score Distillation

ICCV 2025poster

We present InfiniDreamer, a novel framework for generating human motions of arbitrary length. Existing methods typically produce only short sequences, limited by the scarcity of long-range motion data. To address this, InfiniDreamer first generates short sub-motions for each textual description, the…

Cited by 0SourcePDFScholar
2025

MMAD: Multi-label Micro-Action Detection in Videos

ICCV 2025poster

Human body actions are an important form of non-verbal communication in social interactions. This paper specifically focuses on a subset of body actions known as micro-actions, which are subtle, low-intensity body movements with promising applications in human emotion analysis. In real-world scenari…

2025

OSDA Agent: Leveraging Large Language Models for De Novo Design of Organic Structure Directing Agents

ICLR 2025spotlight

Zeolites are crystalline porous materials that have been widely utilized in petrochemical industries as well as sustainable chemistry areas. Synthesis of zeolites often requires small molecules termed Organic Structure Directing Agents (OSDAs), which are critical in forming the porous structure. Mol…

Cited by 0SourcePDFScholar
2025

Prototypical Calibrating Ambiguous Samples for Micro-Action Recognition

AAAI 2025technical

Micro-Action Recognition (MAR) has gained increasing attention due to its crucial role as a form of non-verbal communication in social interactions, with promising potential for applications in human communication and emotion analysis. However, current approaches often overlook the inherent ambiguit…

2025

Reaction Graph: Towards Reaction-Level Modeling for Chemical Reactions with 3D Structures

ICML 2025poster

Accurately modeling chemical reactions using Artificial Intelligence (AI) can accelerate discovery and development, especially in fields like drug design and material science. Although AI has made remarkable advancements in single molecule recognition, such as predicting molecular properties, the st…

2025

VideoGrain: Modulating Space-Time Attention for Multi-Grained Video Editing

ICLR 2025poster

Recent advancements in diffusion models have significantly improved video generation and editing capabilities. However, multi-grained video editing, which encompasses class-level, instance-level, and part-level modifications, remains a formidable challenge. The major difficulties in multi-grained ed…

2025

Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion

CVPR 2025poster

Animatable head avatar generation typically requires extensive data for training. To reduce the data requirements, a natural solution is to leverage existing data-free static avatar generation methods, such as pre-trained diffusion models with score distillation sampling (SDS), which align avatars w…

2025

ZeroMamba: Exploring Visual State Space Model for Zero-Shot Learning

AAAI 2025technical

Zero-shot learning (ZSL) aims to recognize unseen classes by transferring semantic knowledge from seen classes to unseen ones, guided by semantic information. To this end, existing works have demonstrated remarkable performance by utilizing global visual features from Convolutional Neural Networks (…

2024

DocMSU: A Comprehensive Benchmark for Document-Level Multimodal Sarcasm Understanding

AAAI 2024technical

Multimodal Sarcasm Understanding (MSU) has a wide range of applications in the news field such as public opinion analysis and forgery detection. However, existing MSU benchmarks and approaches usually focus on sentence-level MSU. In document-level news, sarcasm clues are sparse or small and are of…

2024

Hand-Centric Motion Refinement for 3D Hand-Object Interaction via Hierarchical Spatial-Temporal Modeling

AAAI 2024technical

Hands are the main medium when people interact with the world. Generating proper 3D motion for hand-object interaction is vital for applications such as virtual reality and robotics. Although grasp tracking or object manipulation synthesis can produce coarse hand motion, this kind of motion is inevi…

2024

HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting

ECCV 2024poster

"Creating digital avatars from textual prompts has long been a desirable yet challenging task. Despite the promising results achieved with 2D diffusion priors, current methods struggle to create high-quality and consistent animated avatars efficiently. Previous animatable head models like FLAME have…

2024

Improving Context Understanding in Multimodal Large Language Models via Multimodal Composition Learning

ICML 2024poster

Previous efforts using frozen Large Language Models (LLMs) for visual understanding, via image captioning or image-text retrieval tasks, face challenges when dealing with complex multimodal scenarios. In order to enhance the capabilities of Multimodal Large Language Models (MLLM) in comprehending th…

2024

TOPA: Extending Large Language Models for Video Understanding via Text-Only Pre-Alignment

NeurIPS 2024spotlight

Recent advancements in image understanding have benefited from the extensive use of web image-text pairs. However, video understanding remains a challenge despite the availability of substantial web video-text data. This difficulty primarily arises from the inherent complexity of videos and the inef…

2024

VividDreamer: Invariant Score Distillation for Hyper-Realistic Text-to-3D Generation

ECCV 2024poster

"This paper presents Invariant Score Distillation (ISD), a novel method for high-fidelity text-to-3D generation. ISD aims to tackle the over-saturation and over-smoothing problems in Score Distillation Sampling (SDS). In this paper, SDS is decoupled into a weighted sum of two components: the reconst…

2023

Continuous-Discrete Convolution for Geometry-Sequence Modeling in Proteins

ICLR 2023poster

The structure of proteins involves 3D geometry of amino acid coordinates and 1D sequence of peptide chains. The 3D structure exhibits irregularity because amino acids are distributed unevenly in Euclidean space and their coordinates are continuous variables. In contrast, the 1D structure is regular…

Cited by 50SourcePDFScholar
2023

Masked Spatio-Temporal Structure Prediction for Self-supervised Learning on Point Cloud Videos

ICCV 2023poster

Recently, the community has made tremendous progress in developing effective methods for point cloud video understanding that learn from massive amounts of labeled data. However, annotating point cloud videos is usually notoriously expensive. Moreover, training via one or only a few traditional task…

Cited by 18PDFcodeScholar
2023

Point Contrastive Prediction with Semantic Clustering for Self-Supervised Learning on Point Cloud Videos

ICCV 2023poster

We propose a unified point cloud video self-supervised learning framework for object-centric and scene-centric data. Previous methods commonly conduct representation learning at the clip or frame level and cannot well capture fine-grained semantics. Instead of contrasting the representations of clip…

Cited by 24PDFScholar
2023

SEFormer: Structure Embedding Transformer for 3D Object Detection

AAAI 2023technical

Effectively preserving and encoding structure features from objects in irregular and sparse LiDAR points is a crucial challenge to 3D object detection on the point cloud. Recently, Transformer has demonstrated promising performance on many 2D and even 3D vision tasks. Compared with the fixed and ri…

2023

STPrivacy: Spatio-Temporal Privacy-Preserving Action Recognition

ICCV 2023poster

Existing methods of privacy-preserving action recognition (PPAR) mainly focus on frame-level (spatial) privacy removal through 2D CNNs. Unfortunately, they have two major drawbacks. First, they may compromise temporal dynamics in input videos, which are critical for accurate action recognition. Seco…

Cited by 24PDFScholar
2023

Text to Point Cloud Localization with Relation-Enhanced Transformer

AAAI 2023technical

Automatically localizing a position based on a few natural language instructions is essential for future robots to communicate and collaborate with humans. To approach this goal, we focus on a text-to-point-cloud cross-modal localization problem. Given a textual query, it aims to identify the descri…

2022

Point Cloud Domain Adaptation via Masked Local 3D Structure Prediction

ECCV 2022poster

"The superiority of deep learning based point cloud representations relies on large-scale labeled datasets, while the annotation of point clouds is notoriously expensive. One of the most effective solutions is to transfer the knowledge from existing labeled source data to unlabeled target data. Howe…

2022

Self-Supervised Global-Local Structure Modeling for Point Cloud Domain Adaptation With Reliable Voted Pseudo Labels

CVPR 2022poster

In this paper, we propose an unsupervised domain adaptation method for deep point cloud representation learning. To model the internal structures in target point clouds, we first propose to learn the global representations of unlabeled data by scaling up or down point clouds and then predicting the…

Cited by 68PDFScholar
2021

PSTNet: Point Spatio-Temporal Convolution on Point Cloud Sequences

ICLR 2021poster

Point cloud sequences are irregular and unordered in the spatial dimension while exhibiting regularities and order in the temporal dimension. Therefore, existing grid based convolutions for conventional video processing cannot be directly applied to spatio-temporal modeling of raw point cloud sequen…

2021

Point 4D Transformer Networks for Spatio-Temporal Modeling in Point Cloud Videos

CVPR 2021poster

Point cloud videos exhibit irregularities and lack of order along the spatial dimension where points emerge inconsistently across different frames. To capture the dynamics in point cloud videos, point tracking is usually employed. However, as points may flow in and out across frames, computing accur…

Cited by 214PDFcodeScholar
2017

Complex Event Detection by Identifying Reliable Shots From Untrimmed Videos

ICCV 2017poster

The goal of complex event detection is to automatically detect whether an event of interest happens in temporally untrimmed long videos which usually consist of multiple video shots. Observing some video shots in positive (resp. negative) videos are irrelevant (resp. relevant) to the given event cla…

Cited by 55PDFScholar