← Search

Lei Fan

26 accepted papers

2026

Chain of World: World Model Thinking in Latent Motion

CVPR 2026

Vision-Language-Action (VLA) models are promising for embodied intelligence, yet they often overlook the predictive and temporal-causal structure underlying visual dynamics. World-model VLAs address this by predicting future frames, but waste capacity reconstructing redundant backgrounds. To overcom

Cited by 0SourcecodeScholar
2026

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs

CVPR 2026

Multimodal Large Language Models (MLLMs) have been shown to be vulnerable to malicious queries that can elicit unsafe responses. Recent work uses prompt engineering, response classification, or finetuning to improve MLLM safety. Nevertheless, such approaches are often ineffective against evolving ma

Cited by 0SourcecodeScholar
2026

Localizing, Structuring, and Rendering: Bridging 3D and 2D Vision-Language-Action Models for Robotic Manipulation

CVPR 2026

Robotic manipulation in complex 3D environments requires unifying spatial reasoning with intuitive visual perception, which is a capability that current Vision-Language-Action paradigms address separately. While 3D VLAs excel in geometric and physical reasoning, they lack intuitive, image-level unde

Cited by 0SourcecodeScholar
2026

Sentinel-VLA: A Metacognitive VLA Model with Active Status Monitoring for Dynamic Reasoning and Error Recovery

ICML 2026poster

Vision-language-action (VLA) models have advanced the field of embodied manipulation by harnessing broad world knowledge and strong generalization. However, current VLA models still face several key challenges, including limited reasoning capability, lack of status monitoring, and difficulty in self…

Cited by 0SourceScholar
2025

DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation

ICCV 2025poster

Human-centric generative models are becoming increasingly popular, giving rise to various innovative tools and applications, such as talking face videos conditioned on text or audio prompts. The core of these capabilities lies in powerful pre-trained foundation models, trained on large-scale, high-q…

2025

Distributed Perception Aware Safe Leader Follower System via Control Barrier Methods

ICRA 2025

This paper addresses a distributed leader-follower formation control problem for a group of agents, each using a body-fixed camera with a limited field of view (FOV) for state estimation. The main challenge arises from the need to coordinate the agents' movements with their cameras' FOV to maintain

Cited by 1SourceScholar
2025

GPVK-VL: Geometry-Preserving Virtual Keyframes for Visual Localization under Large Viewpoint Changes

CVPR 2025poster

Visual localization, the task of determining the position and orientation of a camera, typically involves three core components: offline construction of a keyframe database, efficient online keyframes retrieval, and robust local feature matching. However, significant challenges arise when there are…

Cited by 0SourcePDFScholar
2025

GRPose: Learning Graph Relations for Human Image Generation with Pose Priors

AAAI 2025technical

Recent methods using diffusion models have made significant progress in human image generation with various control signals such as pose priors. However, existing efforts are still struggling to generate high-quality images with consistent pose alignment, resulting in unsatisfactory output. In this…

2025

Interpretable Image Classification via Non-parametric Part Prototype Learning

CVPR 2025poster

Classifying images with an interpretable decision-making process is a long-standing problem in computer vision. In recent years, Prototypical Part Networks has gained traction as an approach for self-explainable neural networks, due to their ability to mimic human visual reasoning by providing expla…

2025

MANTA: A Large-Scale Multi-View and Visual-Text Anomaly Detection Dataset for Tiny Objects

CVPR 2025poster

We present MANTA, a visual-text anomaly detection dataset for tiny objects. The visual component comprises over 137.3K images across 38 object categories spanning five typical domains, of which 8.6K images are labeled as anomalous with pixel-level annotations. Each image is captured from five distin…

Cited by 1SourcePDFScholar
2025

Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality Rebalance

IJCAI 2025

Multimodal pathology-genomic analysis has become increasingly prominent in cancer survival prediction. However, existing studies mainly utilize multi-instance learning to aggregate patch-level features, neglecting the information loss of contextual and hierarchical details within pathology images. F

2025

Prototype-Based Image Prompting for Weakly Supervised Histopathological Image Segmentation

CVPR 2025poster

Weakly supervised image segmentation with image-level labels has drawn attention due to the high cost of pixel-level annotations. Traditional methods using Class Activation Maps (CAMs) often highlight only the most discriminative regions, leading to incomplete masks. Recent approaches that introduce…

2025

Salvaging the Overlooked: Leveraging Class-Aware Contrastive Learning for Multi-Class Anomaly Detection

ICCV 2025poster

For anomaly detection (AD), early approaches often train separate models for individual classes, yielding high performance but posing challenges in scalability and resource management. Recent efforts have shifted toward training a single model capable of handling multiple classes. However, directly…

Cited by 0SourcePDFScholar
2024

AV4GAInsp: An Efficient Dual-Camera System for Identifying Defective Kernels of Cereal Grains

RA-L 2024

Grain Appearance Inspection (GAI) is a pre-requisite for grain quality determination, providing guidance for grain processing, storage, and trade. GAI is routinely performed by trained inspectors who are required to visually inspect cereal grains for each individual kernel. Since grain kernels (e.g.

Cited by 10SourceScholar
2024

Active Open-Vocabulary Recognition: Let Intelligent Moving Mitigate CLIP Limitations

CVPR 2024poster

Active recognition which allows intelligent agents to explore observations for better recognition performance serves as a prerequisite for various embodied AI tasks such as grasping navigation and room arrangements. Given the evolving environment and the multitude of object classes it is impractical…

Cited by 4SourcePDFScholar
2024

Evidential Active Recognition: Intelligent and Prudent Open-World Embodied Perception

CVPR 2024poster

Active recognition enables robots to intelligently explore novel observations thereby acquiring more information while circumventing undesired viewing conditions. Recent approaches favor learning policies from simulated or collected data wherein appropriate actions are more frequently selected when…

Cited by 6SourcePDFScholar
2024

Learning to Ask Denotative and Connotative Questions for Knowledge-based VQA

EMNLP 2024finding

Large language models (LLMs) have attracted increasing attention due to its prominent performance on various tasks. Recent works seek to leverage LLMs on knowledge-based visual question answering (VQA) tasks which require common sense knowledge to answer the question about an image, since LLMs have…

Cited by 0SourcePDFScholar
2023

Flexible Visual Recognition by Evidential Modeling of Confusion and Ignorance

ICCV 2023poster

In real-world scenarios, typical visual recognition systems could fail under two major causes, i.e., the misclassification between known classes and the excusable misbehavior on unknown-class images. To tackle these deficiencies, flexible visual recognition should dynamically predict multiple classe…

Cited by 5PDFScholar
2022

GrainSpace: A Large-Scale Dataset for Fine-Grained and Domain-Adaptive Recognition of Cereal Grains

CVPR 2022poster

Cereal grains are a vital part of human diets and are important commodities for people's livelihood and international trade. Grain Appearance Inspection (GAI) serves as one of the crucial steps for the determination of grain quality and grain stratification for proper circulation, storage and food p…

Cited by 25PDFcodeScholar
2017

RGB-T SLAM: A flexible SLAM framework by combining appearance and thermal information

ICRA 2017poster

Visual SLAM in low illumination scenes remains a considerably challenging task since the available amount of appearance information frequently stays insufficient. To tackle with this problem, we propose a novel SLAM framework by using both appearance information and thermal information, which posses…

Cited by 77SourceScholar