← Search

Xing Xi

8 accepted papers

2026

Awakening Visual Reasoning: Mitigating Post-Training Failure in Vision-Text Compression

ICML 2026poster

Vision-Text Compression (VTC) offers a scalable path for long-context multimodal modeling by rendering textual data into dense visual tokens. While recent Vision-Language Models (VLMs) demonstrate high decoding fidelity (OCR) on such inputs, they exhibit a severe reasoning gap: models that reason ro…

Cited by 0SourceScholar
2026

Let VLMs Grade Their Own Thoughts: A Self-Quantification Approach to Reasoning-Aware Reward Modeling

CVPR 2026

Chain-of-Thought (CoT) is a key technique for enhancing the reasoning capabilities of Vision Language Models (VLMs). Existing methods often employ Reinforcement Learning (RL) with external constraints to align the model's reasoning process with human cognitive patterns. However, we argue that the mo

Cited by 0SourceScholar
2026

OW-DAR: Dual-Granularity Adaptive Reconstruction-Error Modeling for Open-World Object Detection

AAAI 2026technical

Open-world object detection (OWOD) aims to detect known and unknown objects in dynamic environments. However, only known classes are labeled during training, making it challenging for detectors to recognize unknown objects during inference. Existing methods typically rely on supervision from known c

Cited by 0SourcePDFScholar
2026

Video-BCI: Bayesian Cognitive Integration of Self-Prior Hypotheses for Video Understanding

ICML 2026poster

Recent progress in vision-language models (VLMs) has driven significant advances in video understanding. However, existing methods often act as naive empiricists, mapping video input directly to output without any mechanism to introspect or challenge inherent bias. In this work, we challenge this pa…

Cited by 0SourceScholar
2025

OW-OVD: Unified Open World and Open Vocabulary Object Detection

CVPR 2025poster

Open world perception expands traditional closed-set frameworks, which assume a predefined set of known categories, to encompass dynamic real-world environments. Open World Object Detection (OWOD) and Open Vocabulary Object Detection (OVD) are two main research directions, each addressing unique cha…

2024

KTCN: Enhancing Open-World Object Detection with Knowledge Transfer and Class-Awareness Neutralization

IJCAI 2024poster

Open-World Object Detection (OWOD) has garnered widespread attention due to its ability to recall unannotated objects. Existing works generate pseudo-labels for the model using heuristic priors, which limits the model’s performance. In this paper, we leverage the knowledge of the large-scale visual…

2024

UMB: Understanding Model Behavior for Open-World Object Detection

NeurIPS 2024poster

Open-World Object Detection (OWOD) is a challenging task that requires the detector to identify unlabeled objects and continuously demands the detector to learn new knowledge based on existing ones. Existing methods primarily focus on recalling unknown objects, neglecting to explore the reasons behi…