← Search

Xihan Wei

10 accepted papers

2026

Discriminative Visual Process Rewards for Scaling Thinking at Test-Time with Images

ICML 2026poster

The “thinking with images” paradigm has led multimodal large language models to generate intermediate visual steps—such as cropping, annotation, spatial localization, and sketches—to enhance high-resolution perception and complex reasoning. However, existing multimodal Process Reward Models (PRMs) e…

Cited by 0SourceScholar
2026

Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness

AAAI 2026technical

Facial expression captioning has found widespread application across various domains. Recently, the emergence of video Multimodal Large Language Models (MLLMs) has shown promise in general video understanding tasks. However, describing facial expressions within videos poses two major challenges for

Cited by 0SourcePDFScholar
2026

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding

CVPR 2026

We present SWIM (See What I Mean), a novel training strategy that aligns vision and language representations to enable fine-grained object understanding solely from textual prompts. Unlike existing approaches that require explicit visual prompts, such as masks or points, SWIM leverages mask supervis

Cited by 0SourcecodeScholar
2025

LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models

CVPR 2025highlight

Recent open-vocabulary detectors achieve promising performance with abundant region-level annotated data. In this work, we show that an open-vocabulary detector co-training with a large language model by generating image-level detailed captions for each image can further improve performance. To achi…

2025

Person De-reidentification: A Variation-guided Identity Shift Modeling

CVPR 2025poster

Person re-identification (ReID) is to associate images of individuals from different camera views against cross-view variations. Like other surveillance technologies, Re-ID faces serious privacy challenges, particularly the potential for unauthorized tracking. Although various tasks (e.g., face reco…

Cited by 0SourcePDFScholar
2025

ViSpeak: Visual Instruction Feedback in Streaming Videos

ICCV 2025poster

Recent advances in Large Multi-modal Models (LMMs) are primarily focused on offline video understanding. Instead, streaming video understanding poses great challenges to recent models due to its time-sensitive, omni-modal and interactive characteristics. In this work, we aim to extend the streaming…

2024

DreamView: Injecting View-specific Text Guidance into Text-to-3D Generation

ECCV 2024poster

"Text-to-3D generation, which synthesizes 3D assets according to an overall text description, has significantly progressed. However, a challenge arises when the specific appearances need customizing at designated viewpoints but referring solely to the overall description for generating 3D objects. F…

2024

Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models

NeurIPS 2024poster

Recent vision foundation models can extract universal representations and show impressive abilities in various tasks. However, their application on object detection is largely overlooked, especially without fine-tuning them. In this work, we show that frozen foundation models can be a versatile feat…

Cited by 4SourcePDFScholar
2022

Spatiotemporal Self-Attention Modeling with Temporal Patch Shift for Action Recognition

ECCV 2022poster

"Transformer-based methods have recently achieved great advancement on 2D image-based vision tasks. For 3D video-based tasks such as action recognition, however, directly applying spatiotemporal transformers on video data will bring heavy computation and memory burdens due to the largely increased n…

2021

Interactive Self-Training With Mean Teachers for Semi-Supervised Object Detection

CVPR 2021poster

The goal of semi-supervised object detection is to learn a detection model using only a few labeled data and large amounts of unlabeled data, thereby reducing the cost of data labeling. Although a few studies have proposed various self-training-based methods or consistency regularization-based metho…

Cited by 170PDFScholar